ANOVA says at least one group differs. It never says which.
That is the whole result, and it is the half most people skip. This page runs a single-factor ANOVA end to end, prints every row of the report, then does the part the report leaves to you: finding which pair is responsible, and moving the threshold you judge it against.
Method page for Google Sheets Single factor, two factor, and the rank-based version free, and no plan buys a bigger one
What the F test actually asks
Group averages differ in every experiment ever run, including the ones where nothing is going on. ANOVA compares two kinds of variation and reports the ratio. The first is how far the group averages sit from the overall average. The second is how far the individual readings sit from their own group average. F is the first divided by the second. Near 1 and your groups look like one population that happens to have been split three ways. Well above 1 and the split is doing real work.
The reason a separate test exists at all is arithmetic you can do on the back of the page. Three groups make three pairs. Test each pair on its own at the usual 5% threshold and the chance of at least one false alarm is no longer 5%. Multiply the misses and 1 minus 0.95 cubed gives 14.3%, which is the quick count rather than the true rate: the three comparisons share the same three groups, so they are not three separate experiments, and the real figure comes in under it. Draw three groups of six from one bell-shaped population, put all three pairs through the t-test, and repeat half a million times: at least one pair clears 0.05 closer to 12% of the time. Four groups make six pairs, where the quick count says 26.5% and the measured rate is closer to 20%. Either way it is well past the 5% you thought you were spending, and it climbs with every group you add. ANOVA asks one question once, so it keeps the 5%.
The assumption worth reading twice is independence, and the tool prints it beside the result: “It assumes the groups are independent: each row belongs to one group and one only. If the same people appear in more than one column, or one person contributed several rows, the test does not apply and the p-value comes out too small.” If every subject appears in every column, that is a repeated-measures design and this is not the test for it. Sortia does not ship a repeated-measures ANOVA, and no arrangement of the ones it does ship substitutes for one.
One run, printed in full
Eighteen exam scores, three revision methods, six students in each. It is the worked model called Which Study Method Actually Works?, and every figure below came out of the tool itself, called through the repository’s own test harness so nothing here is a mock-up of a report.
| Flashcards (exam score) | Rereading (exam score) | Practice tests (exam score) |
|---|---|---|
| 78 | 72 | 86 |
| 85 | 75 | 91 |
| 82 | 78 | 88 |
| 88 | 74 | 93 |
| 80 | 80 | 85 |
| 84 | 77 | 90 |
Point the tool at that block, tick Labels in first row/column, leave Grouped by on columns, and the run writes a tab into your file. Two blocks go down the left of it, and they are the whole of the report proper.
| Groups | Count | Sum | Average | Variance |
|---|---|---|---|---|
| Flashcards (exam score) | 6 | 497 | 82.833 | 12.967 |
| Rereading (exam score) | 6 | 456 | 76.000 | 8.400 |
| Practice tests (exam score) | 6 | 533 | 88.833 | 9.367 |
| Source of Variation | SS | df | MS | F | P-value | F crit |
|---|---|---|---|---|---|---|
| Between Groups | 494.778 | 2 | 247.389 | 24.149 | 2.04E-05 | 3.682 |
| Within Groups | 153.667 | 15 | 10.244 | |||
| Total | 648.444 | 17 |
- F statistic
- 24.149
- F critical, 2 and 15 df
- 3.682
- P-value
- 2.04E-05
- Widest gap between averages
- 12.83 points
The sentence the tool writes underneath, word for word: “This run: F = 24.15. The widest gap is Practice tests (exam score) at 88.83 against Rereading (exam score) at 76, a gap of 12.83. p is well under 0.001, below your alpha of 0.05, so the difference between groups is more than chance would ordinarily produce.” Notice what it does not say. It does not name a winner.
Nothing here is seeded, because none of these tests has a random step. The same eighteen numbers give the same table on any machine, which is also why you can rebuild the whole thing with a calculator:
- Overall average of all eighteen scores: 1486 / 18 = 82.556.
- SS between groups. For each group, its size times the square of the distance from its average to the overall average, added up: 6(82.833 − 82.556)² + 6(76.000 − 82.556)² + 6(88.833 − 82.556)² = 494.778.
- SS within groups. For every reading, the square of its distance from its own group average, added up: 153.667. The two add to the total, 648.444, which is the only cross-check the table gives you and is worth using.
- Degrees of freedom. Groups minus one, so 2. Readings minus groups, so 18 − 3 = 15.
- Mean squares. 494.778 / 2 = 247.389. 153.667 / 15 = 10.244.
- F. 247.389 / 10.244 = 24.149. The p-value is the share of the F distribution with 2 and 15 degrees of freedom lying above that, which is 0.0000204.
F crit is the other side of the same coin: 3.682 is the point on that distribution with 5% of the area above it. F past the critical value and p below alpha always agree, so the two columns are never a second opinion.
A significant F is the start of the answer
The report above contains no post-hoc test. Neither does the product: there is no Tukey, no Scheffé, no Games-Howell anywhere in it, and this page is not going to pretend otherwise. What there is, is the second run, and it is three t-tests and one division.
Run a two-sample t-test on each pair of columns in turn. These are Welch’s version, the unequal-variances one, which is the safer default. With six readings in every group the pooled version returns the same t statistic on this data and differs only in its degrees of freedom, and here both round to 10.
| Pair | t Stat | df | P two-tail | Verdict at 0.0167 |
|---|---|---|---|---|
| Flashcards against rereading | 3.621 | 10 | 0.0047 | clears |
| Flashcards against practice tests | −3.110 | 10 | 0.0111 | clears |
| Rereading against practice tests | −7.458 | 10 | 0.0000217 | clears |
0.0167 is 0.05 divided by 3, one share of the budget per comparison. The threshold moves because you are now asking three questions instead of one, and the 12% at the top of this page is what happens if you leave it where it was. Dividing alpha by the number of comparisons is the plainest correction there is and a deliberately strict one: it costs power, and a real difference between small groups can fail to clear it. It is not the only correction anyone uses. It is the one you can do in your head, against a number the report has already given you.
All three clear it here, so the ordering is real and the finding is a full ranking: practice tests, then flashcards, then rereading. That is a stronger result than the F test alone was entitled to state.
Now look at the middle row. Flashcards against practice tests comes back at 0.0111. Against a bare 0.05 that is comfortable. Against 0.0167 it clears with a third of the room to spare. Had the study carried a fourth revision method, there would have been six pairs rather than three, the line would have sat at 0.05 / 6 = 0.0083, and that same 0.0111 would not have cleared it. Same students, same scores, same p-value, opposite verdict, because the question got bigger. That is not a flaw in the arithmetic. It is the arithmetic telling you that a wider search finds more by accident.
Two things not to do with this. Do not run the pairwise tests after an F that did not clear alpha: the F test is the thing protecting you from exactly that, and going round it is how a difference gets found in noise. And do not report the pairwise p-values without saying which threshold you judged them against, because the threshold is the reason they mean anything.
Which of the three t-tests matches your design, and what the report prints beside the p-value: t-tests in Google Sheets.
Two runs where the second one changes the story
Three suppliers, eight goods-in lots each, defects per thousand units. Averages of 8.463 for Northgate, 9.075 for Belmar and 12.563 for Calder. F comes back 42.113 against an F critical of 3.467, with a p-value of 4.48E-08, which settles that the three are not interchangeable. But the argument in the room is almost never anyone against Calder. It is Northgate against Belmar, the two good ones, and a t-test on just those two columns returns t of −1.284 on 14 degrees of freedom with a two-tail p of 0.2198. That 0.61 gap is noise, and moving volume between them on this evidence would be moving it for no reason. Calder is the finding. The ranking of the other two is not a finding at all. Worked in full at Are These Three Suppliers Really Different?
Three crews, eight days each, square meters laid a day. Averages of 61.500, 67.750 and 58.500, so the best crew is 15.8% ahead of the worst and everybody on site already believes it. F comes back 2.421 against the same F critical of 3.467, with a p-value of 0.1132, so this evidence does not establish that the crews differ at all. The reason is in the SUMMARY block rather than the ANOVA block: the standard deviation inside a single crew runs from 7.91 to 8.97 square meters a day, and the entire gap between the best and worst crew averages is 9.25. One crew’s own day-to-day swing is about the size of the thing you are trying to measure, and eight days is not enough to see through it. The fix is not a better test, it is more days. Worked in full at Do the Three Crews Work at the Same Rate?
Both runs print the same two blocks. What separates them is which block you read first.
Two factors, and the interaction between them
When you changed two things at once, a one-factor test cannot tell you which one worked, and averaging over the other one can hide the answer completely. Twelve loaves: two ovens, two recipes, three bakes of every combination, rise measured in centimeters. That is Oven, Recipe, or Both?, and it is the smallest honest example of the design.
| Source of Variation | SS | df | MS | F | P-value | F crit |
|---|---|---|---|---|---|---|
| Rows (Deck oven, Convection) | 0.367 | 1 | 0.367 | 36.750 | 3.02E-04 | 5.318 |
| Columns (Sourdough (rise, cm), Rye (rise, cm)) | 0.908 | 1 | 0.908 | 90.750 | 1.22E-05 | 5.318 |
| Interaction | 0.187 | 1 | 0.187 | 18.750 | 2.51E-03 | 5.318 |
| Within | 0.080 | 8 | 0.010 | |||
| Total | 1.543 | 11 |
All three rows clear 0.05, and the third one is the reason the design exists. The interaction says the answer to which oven depends on which recipe you are baking. The four cell averages say it in centimeters: deck oven 4.300 on sourdough and 4.000 on rye, convection 4.900 and 4.100. Convection lifts sourdough by 0.600 cm and rye by 0.100 cm, and the interaction is the 0.500 cm difference between those two lifts. With it live, neither factor can be quoted on its own, which is what the tool tells you in its reading.
Where those four averages come from, and where they do not go. The panel draws them while you are looking at it: a two-by-two grid of cell averages, rows and columns named as you typed them, under the heading “The factors interact: one’s effect depends on the other (p = 0.003)” and the line “Each cell is that combination’s average. Rows changing pattern as you read across the columns is the interaction the test scores.” The tab written into your file is the other half, and it is one ANOVA table: Rows, Columns, Interaction, Within, Total, with no summary block and no cell averages anywhere on it. So the grid is on the screen and not in the thing you send anyone. Copy the four numbers across, or rebuild them with AVERAGE over each combination, because a significant interaction with no cell averages beside it is a result nobody in the room can act on.
What the second factor buys is visible if you take it away. Put the same twelve loaves through a single-factor ANOVA on the two recipe columns alone and F comes back 14.291, not 90.750, on 1 and 10 degrees of freedom with a p-value of 0.0036 against an F critical of 4.965. Six times weaker on identical data, because the oven difference that the two-factor test holds apart is dumped into the noise term instead.
Two limits worth knowing before you design the sheet. The first is how many repeats each combination needs. Two is the minimum, not three: point the tool at a grid with two values per cell and it runs and prints a full Interaction row, and the worked example in the panel’s own guide is exactly that, two fertilizers by two crops with two measurements each. What a third repeat buys is error degrees of freedom, which come to the number of cells times one less than the repeats: eight for this two-by-two grid at three bakes, four at two. All three F-ratios are divided by that same error term, so the whole table rests on it. Drop the third bake of each combination here and the interaction that came back F 18.750 at p 0.0025 comes back F 7.000 at p 0.0572, the same pattern in the same direction, no longer clearing 0.05. Two will run. Three is the fewest worth designing around. And the two-factor test assumes every combination has the same number of repeats. If you have one reading per combination rather than several, that is the third tool, Two-Factor Without Replication, whose table has no Interaction row by construction: with one observation per cell there is nothing left to separate an interaction from error, so it is folded in. Worked at Layout and Shift: Which One Moves Output? and Shift or Line: What Moves Output?
When the groups are not bell-shaped
ANOVA compares averages, and an average is the wrong summary of a skewed column: one enormous customer, one 289-unit day, one claim ten times the size of the others, and the group average is describing the outlier rather than the group. Kruskal-Wallis is the same question asked of the rank order instead. It is in the same panel, it is free, and it is what to reach for when the shape is wrong rather than when the p-value is disappointing.
Put the eighteen exam scores through it and you can price the swap exactly. H comes back 13.189 on 2 degrees of freedom with a p-value of 0.001368, against ANOVA’s 0.0000204 on the same numbers. Same verdict, a p-value about 67 times larger. That is the cost of throwing away the distances between the scores and keeping only their order, and it is the right trade when you do not trust the distances and a bad one when you do.
It inherits the same limit, and the tool says so in its own reading: “P below your alpha says at least one group differs. It does not say which.” The medians table it prints beside the result is where you look next.
Deciding between the two is not a matter of running a normality test first, and Sortia does not ship one. Is my data normally distributed? walks through what to look at instead, and which statistical test should I use takes the question one step further back, to how many groups you have and whether the rows are paired.
Worked: Which Channel Brings the Big Spenders? and Do the Regions Rate Us Differently?
Getting the range right
- One group per column, the name in the top cell. Select the block including that header row and tick Labels in first row/column, and the group names in the report are yours rather than “Column 1”. If your groups run across rows instead, switch Grouped by to rows and leave the data where it is.
- Groups do not have to be the same size. Blanks are dropped column by column, so a short group is simply a smaller group. Drop the last rereading score, the 77, from the run above and F becomes 21.271 on 2 and 14 degrees of freedom, p 5.71E-05, F critical 3.739. The test carries on; the degrees of freedom move with it.
- A stray word in the range stops the run rather than guessing. The refusal names the cell by row and column and never quotes what is in it, and if it is your header it tells you which box to tick.
- Select every group. If only one column of numbers comes out of the range, the run refuses with “Only one group came out of the input range, and ANOVA compares the averages of two or more. Widen the range to cover every group, or switch Grouped by to Rows if your groups run across rows rather than down columns.” That message exists because selecting one column is the commonest way to get this wrong.
- Two-factor with replication needs its grid to divide. Set Rows per sample to the number of repeats per combination, and make the height of the block an exact multiple of it. Twelve loaves in two blocks of three rows by two columns is 6 by 2 with rows per sample 3.
What the report contains, and what it does not
Single factor writes two blocks. SUMMARY: Groups, Count, Sum, Average, Variance. ANOVA: Source of Variation, SS, df, MS, F, P-value, F crit, over rows for Between Groups, Within Groups and Total. Both two-factor tools write the ANOVA block only, with a row per effect. To the right of the widest table, every one of the four opens a panel: the tool’s name, the range it read, what the tool is for, then the plain-language reading quoted higher up this page, written in your file from the numbers the run just produced, then the short notes on how to read the result, which is where the two quotations higher up came from. Under the table goes one stamp naming the tool, the source tab and the build, and on the free plan one more line saying Made with Sortia.
Each of these four runs also draws one chart in the panel, and none of the four writes that chart into your file: single factor a bar per group average, both two-factor tools a grid of the cell values, Kruskal-Wallis a bar per group median. Every one of those headings states the verdict rather than the variable, which is why the two-factor grid quoted higher up leads with the interaction. It is the surface to read the result on and the tab is the surface to send, so anything you want a colleague to see has to be in the tab.
What is not in any of them, stated plainly so nobody plans around it: no post-hoc test of any kind, no effect size, no confidence interval on a group average, and no test of the equal-variance assumption. The two pairwise sections above are the honest workaround for the first of those, and they are work you do rather than a button you press.
The one figure people usually want next is at least derivable from the table you have. The share of the total variation that sits between the groups rather than inside them is SS between divided by SS total, 494.778 / 648.444 = 0.763 on this run. It is a description of these eighteen students and it runs high in small samples, so treat it as a way of reading your own table rather than as a number to publish.
All three ANOVA tools and Kruskal-Wallis are free, and no plan buys a bigger one: the only ceiling is the 500,000-cell selection gate every analysis tool reads through, which is an engineering limit rather than a priced one and is measured on the performance page at both sizes. Nothing on this page needs a paid feature. What Pro is for is the five paid engines, which are a different job entirely.
Next question
- Three averages differ. Is that more than chance?Which Study Method Actually Works?
- You changed two things and it worked. Which one did it?Oven, Recipe, or Both?
- Crew B is 16% ahead. Is that the crew, or the fortnight?Do the Three Crews Work at the Same Rate?
The others in the library that run one of these four tools:
- Is your worst supplier really worse than the rest?Are These Three Suppliers Really Different?
- Two layouts, three shifts, one output number.Layout and Shift: Which One Moves Output?
- Does the night shift really produce less?Shift or Line: What Moves Output?
- Is it the venue, or is it the month?Venue and Month: What Drives Attendance?
- Is that channel better, or just one big customer?Which Channel Brings the Big Spenders?
- Three regions share a median. Do they share an answer?Do the Regions Rate Us Differently?
- Two crews, almost identical averages. Same crew?Four Shifts, Ranked Output, One Question
Every statistics tool in the sidebar and what each one is for: Statistics in Google Sheets.