Skip to main content

One-way ANOVA calculator

Paste two to six groups of raw values. You get the F-statistic, its exact p-value, η², and the per-group means — because an omnibus F tells you that something differs, and the means are where you start finding out what.

F-tests are always right-tailed — F is a ratio of between-group to within-group variance, and only a large ratio is evidence against the null. Why.

Paste each group to begin

You get the F-statistic, its p-value, η², and a per-group summary — because an omnibus F tells you that something differs, not what.

Why not just run t-tests?

Because the errors compound. Each test at α = 0.05 carries its own 5% false-positive risk, and comparing every pair of groups means running a lot of tests. Three groups give three comparisons and about a 14% chance of at least one false positive. Five groups give ten comparisons and about 40%. Six give fifteen and just over half.

ANOVA replaces all of them with one question — is there any difference among these groups? — answered at the error rate you actually chose. That is the whole reason it exists. It is not a more powerful t-test; it is a way of asking once instead of many times.

What F is measuring

F is a ratio of two estimates of variance. The numerator measures how far the group means spread out from the grand mean; the denominator measures how much the observations scatter within their own groups.

F = (between-group variance) / (within-group variance)

If the groups were all drawn from the same population, both numbers would estimate the same thing and their ratio would sit near 1. A large F means the means are further apart than the noise inside the groups can account for. That framing also explains why the test is always right-tailed: an F near zero means the groups look alike, which is exactly what the null hypothesis predicts.

The omnibus problem

A significant F establishes that at least one group differs from the others. It does not say which, it does not say how many, and it does not say by how much. This is a property of the test, not a shortcoming of any particular calculator — which is why this one shows no single effect estimate and no interval, and prints the group means instead.

The proper next step is a post-hoc procedure: Tukey's HSD for all pairwise comparisons, Dunnett's if you are comparing several treatments against one control, or Bonferroni- corrected t-tests if you only care about a handful of specific pairs decided in advance. What you must not do is run unadjusted pairwise t-tests after the ANOVA — that reinstates exactly the error inflation the ANOVA was protecting you from.

Report η² next to the p-value

Eta squared is the share of total variance that sits between the groups. It answers the question F cannot: not "is there a difference" but "how much of what I am measuring does group membership explain?"

The two numbers diverge fast with sample size. Six groups of 500 will detect a difference that accounts for 1% of the variance at p < 0.001, and that finding is real and almost certainly useless. Conventional benchmarks put η² of 0.01 as small, 0.06 as medium and 0.14 as large, though as with all such labels the honest answer depends on your field.

Assumptions, in order of how much they matter

  1. Independence. Each observation belongs to one group and does not depend on any other. Repeated measures on the same subjects break this and need a repeated-measures ANOVA instead. This is the assumption you cannot repair after the fact.
  2. Equal variances. Matters most when group sizes are unequal. The calculator warns when the largest group SD exceeds twice the smallest; Welch's ANOVA is the fix.
  3. Normality of residuals. The most-cited and least-critical of the three. With groups above about 30 the central limit theorem carries you. Strong skew or clear outliers in small groups are the real risk, and a Kruskal-Wallis test is the alternative.

If you already have an F-statistic from software and just need the p-value, the F-score to p-value calculator takes it directly. For two groups rather than three or more, use the t-test calculator — and note that they agree exactly, since F(1, df) = t(df)².

Frequently asked questions

When should I use ANOVA instead of t-tests?

Whenever you are comparing three or more group means at once. Running every pairwise t-test instead inflates your false-positive rate badly: with five groups there are ten comparisons, and at α = 0.05 the chance of at least one spurious "significant" result is about 40%. ANOVA asks a single question — is any group different? — at your chosen error rate.

My ANOVA is significant. Which groups differ?

The F-test cannot tell you, and that is a real limitation rather than a gap in this calculator. It is an omnibus test: it establishes only that at least one group mean is out of line with the rest. To find out where, run post-hoc comparisons such as Tukey's HSD, which control the error rate across the whole family of pairwise tests. The group summary table above is the place to start looking.

What is eta squared?

The proportion of total variance that lies between the groups rather than within them — the ANOVA equivalent of r². An η² of 0.15 means 15% of the variation in your outcome is accounted for by group membership and 85% is not. It is what tells you whether a significant F matters: with enough data, a group difference far too small to act on will still clear p < 0.05.

What assumptions does one-way ANOVA make?

Three. Observations are independent; the residuals are roughly normal; and the groups have similar variances. The first is the one you cannot fix afterwards. Normality matters less as group sizes grow, thanks to the central limit theorem. Unequal variances matter most when group sizes are also unequal — the calculator flags it when the largest group SD is more than twice the smallest.

What if my group variances are very different?

Use Welch's ANOVA, which drops the equal-variance assumption the same way Welch's t-test does, or a Kruskal-Wallis test if the data are also badly non-normal. Classic one-way ANOVA is reasonably robust when group sizes are equal and increasingly unreliable when they are not — unequal variances plus unequal n is the combination that produces genuinely misleading p-values.

Do my groups need the same number of observations?

No. Unbalanced designs are fine and one-way ANOVA handles them directly. Balance matters indirectly, though: with equal group sizes the test tolerates unequal variances well, and with unequal sizes it does not. If your groups differ a lot in size, check the variances before trusting the p-value.