Two questions pick your test. The rest is a tiebreak.

How many groups are you comparing, and is each row in one group tied to one particular row in the other. Answer those two and the shortlist is usually down to one test, sometimes two.

Method guide Three worked comparisons, every figure computed the data is printed on this page

The two questions

Nearly every argument about which test to use is really an argument about the design of the study, and the design was settled before anyone opened the file. So start there rather than with the numbers.

  1. How many groups am I comparing? One column against a target you already have in mind; two columns; or three and more. Two groups and three groups are different tests, not the same test run twice, and that is the single most common way a comparison goes wrong.
  2. Is each row in one group tied to one particular row in the other? The same person weighed twice. The same weekday before and after. The same store under two layouts. If it is, the rows are paired, and a paired test is not merely allowed, it is the reason a small study can settle anything at all: it takes the differences between people out of the comparison and leaves only the change.

Question two has a test you can apply without knowing any statistics. Shuffle one of the two columns. If the shuffle destroys information, the rows were paired and you must use a paired test. If it changes nothing, they were independent and you must not.

After those two comes a third question that is about the data rather than the design, and it is a tiebreak between two tests you have already narrowed to, never the first question you ask: would one or two extreme values move the average enough to change the verdict? If yes, the rank-based version of the test you picked is the safer read. The load-time example below is that situation happening, on eight page loads a side, where the two answers come out as 0.0104 and 0.7253.

Not the same question as “is my data normally distributed”, which is a worse question than it looks: why the shape question is usually the wrong one.

The chooser

Read down the first column until a row describes the thing on your screen. The middle column is the name as it is written in the tool list, so you can search for it. The last is a worked example with real numbers, readable whether or not you ever run anything.

What you haveThe testWorked on real data
One group, against a number you already have
One column of amounts, and a target you have to clearDescriptive StatisticsOne Table That Describes Your Numbers
Two groups, rows paired
Two columns, row for row the same subject, amountst-Test: Paired Two Sample for MeansDoes Cutting Caffeine Buy You Sleep?
The same design, but the numbers are skewed or one pair is extremeWilcoxon Signed-Rank (nonparametric)Did the New Process Cut Ticket Times?
Two groups, rows independent
Two separate groups of amounts, spreads about the samet-Test: Two-Sample Assuming Equal VariancesAre These Two Stores Really Different?
Two separate groups of amounts, one group much wilder than the othert-Test: Two-Sample Assuming Unequal VariancesManual vs Automated, Tested Fairly
Two separate groups, skewed or carrying an outlierMann-Whitney U (nonparametric)Faster Site, Proven Without the Bell Curve
Two separate groups, and the spread of the process is already documentedz-Test: Two Sample for MeansAre Two Fill Lines Filling the Same?
Two separate groups, and the question is consistency rather than averageF-Test Two-Sample for VariancesIs the New Machine More Consistent?
Two variants, and every row is a yes or a no rather than an amount. Total them first: this is the one tool here that takes four typed counts rather than a rangeA/B Test (two proportions)Did the B Version Win?
Three groups or more
Three or more groups, one thing changing between them, amountsANOVA: Single FactorWhich Study Method Actually Works?
Three or more groups, skewed or carrying outliersKruskal-Wallis (nonparametric)Which Channel Brings the Big Spenders?
Two things changing at once, several readings of each combinationANOVA: Two-Factor With ReplicationOven, Recipe, or Both?
Two things changing at once, one reading of each combinationANOVA: Two-Factor Without ReplicationVenue and Month: What Drives Attendance?
Counts, rather than amounts
Counts in a table, and the question is whether the two ways of slicing are relatedChi-Square: IndependenceDoes the Court Change the Settlement Rate?
Counts against a split you expected in advanceChi-Square: Goodness of FitIs the Product Mix What We Planned For?
Not a comparison at all
Two columns, and the question is whether they move togetherCorrelationSpend, Visits, Sales: What Moves Together?
Several columns, and the question is which of them carries the outcomeRegressionWhat Actually Drives Your Sales?

The middle column is the tool list of the Sortia sidebar, verbatim, down to the three names that carry (nonparametric) in the picker itself. That is the list telling you the same thing this page does: those three make no assumption about the shape of your numbers. Two of them have a floor on how little data they will accept: the signed-rank test refuses fewer than five pairs that actually changed, and Mann-Whitney refuses fewer than three values on either side.

Two honest gaps in that table. There is no one-sample t-test on this bench. For a single column against a fixed target, the half-width that Descriptive Statistics prints on its Confidence Level row answers the same question, once you do the last step yourself: add it to the mean and subtract it, and if your target sits outside that pair of numbers, the gap is larger than samples that size usually produce by chance at that level. And there is no repeated-measures ANOVA and no mixed model anywhere, so a design with three or more measurements on the same subjects is past what this list covers. Reducing it to one difference per subject and running a paired test answers a narrower question honestly, which is better than running a wider test that is not there.

Two defensible tests, one set of nine days

A support team added a triage step and measured average resolution hours on nine matched weekdays either side of the change. The rows are paired, so the shortlist is two tests, and both of them are defensible on this data.

Before, hours per ticket18, 22, 15, 26, 20, 17, 24, 19, 21
After, hours per ticket14, 16, 13, 19, 15, 16, 17, 15, 18
The nine differences4, 6, 2, 7, 5, 1, 7, 4, 3

Run those two columns twice, once through each test:

What the report gives youWilcoxon Signed-Rankt-Test: Paired
The statisticW+ 45t Stat 6.128
z2.611not printed
dfnot printed8
P, two-tail0.00900.000281
P computed bynormal approximationnot printed
The size of the changemedian difference 4 hmeans 20.222 and 15.889
What the pairing bought9 non-zero pairsPearson correlation 0.842

Both columns are printed above, so this is reproducible in any tool. In the add-on they are the model named Did the New Process Cut Ticket Times?, and no seed is involved: these tests are deterministic, so the same rows always return the same figures. Computed 6 September 2026 by calling the shipped functions directly through the repository’s test harness.

The two p-values are about thirty-two times apart, on identical rows, and neither of them is a mistake. That gap is the whole subject of this page, so it is worth being exact about where it comes from.

The rank test hits a ceiling, and you can watch it happen

W+ adds up the ranks of the improvements, so on nine pairs the largest value it can possibly take is 1 + 2 + … + 9, which is 45. This data reached it: every one of the nine days improved. Once that happens the statistic has nowhere left to go, and neither does the p-value.

Here is the same nine days with a fixed number of hours added to every improvement. Adding a constant leaves the order of the differences alone, and the rank test only ever looks at the order:

The same nine daysMedian cutWilcoxon W+Wilcoxon PPaired t StatPaired t, P
as measured4 h450.00906.1280.000281
20 hours added to every improvement24 h450.009034.4130.00000000056

Add a hundred hours to every improvement instead and the rank test still returns W+ 45 and p 0.0090, while the paired t reaches roughly 0.0000000000000050. The rank test is not being stubborn. It was never told how big the improvements were, only that all nine of them pointed the same way, and nine out of nine is as convincing as nine pairs get.

Why 0.0090 rather than something smaller. The nine differences are 4, 6, 2, 7, 5, 1, 7, 4, 3, and two of those sizes occur twice. Ties among the sizes rule out counting every possible arrangement of signs, so the tool falls back on the normal approximation and prints a row saying so. Hand it nine differences with no ties and it counts all 29 = 512 sign patterns exactly instead, and the floor on nine pairs becomes 2 ÷ 512 = 0.0039. Either way there is a floor, and either way it is a fact about how many pairs you collected rather than about how much better things got.

This is the cost of not assuming a bell curve, priced on one dataset. The rank test cannot be embarrassed by a wild value because it never looks at the size of one. The same blindness is why it cannot tell a four-hour improvement from a hundred-hour one. That is a trade, not a safety feature, and anyone who tells you the rank test is simply the safer choice is quoting one half of it.

One thing both tests are silent about: nine weekdays from one team is a before and after, not a controlled comparison. Neither p-value says the triage step caused the improvement. The 4-hour median is the size of the effect; the p-value is only the odds of a run like this turning up by chance.

When the disagreement is total

Eight page loads served by an old content network and eight by a new one. Different pages, not the same eight measured twice, so the rows are independent and the shortlist is the two-sample tests. Each column happens to carry one multi-second load, which is what real latency data looks like.

Old network, milliseconds420, 380, 460, 2900, 510, 440, 395, 470
New network, milliseconds210, 180, 240, 195, 260, 225, 205, 3100

Two tests, both legitimate for this design, on those sixteen numbers:

What the report gives youMann-Whitney Ut-Test: Unequal Variances
The statisticU 8t Stat 0.359
P, two-tail0.01040.7253
P computed byexactnot printed
The middle of each groupmedians 450 and 217.5 msmeans 746.875 and 576.875 ms
Spreadnot usedvariances 758,635 and 1,040,014
Observations8 and 88 and 8, df 14

The equal-variances t-test returns exactly the same figures here, t 0.359 and p 0.7253, so this gap is not about which t-test you pick. Both columns are above; in the add-on they are the model named Faster Site, Proven Without the Bell Curve. Computed 6 September 2026.

One test says 0.0104 and the other says 0.7253. That is not a rounding disagreement. One of them says the new network is faster and the other says there is nothing here to see.

The two slow loads are doing all of it. A t-test measures the gap between the averages against the spread inside the groups, and 2,900 ms and 3,100 ms push those variances to roughly 760,000 and 1,040,000. Against a spread that size a 170 ms gap between the means is nothing, which is exactly what a p of 0.73 is reporting. The rank test never sees the number 3,100 at all. It records only that this was the slowest of the sixteen loads, worth one rank like every other value, and then asks whether the new network’s loads tend to sit lower in the pooled order. Compare every old load against every new one, all 64 pairings, and the new network is the faster of the two in 56 of them. The 8 pairings that go the other way are the U of 8 in the table above.

So report the medians. 450 ms against 217.5 ms, roughly twice as fast for a typical visitor. The means, 746.9 and 576.9, describe neither network’s behavior on any real page: no load in either column is anywhere near them. Once you have chosen a rank test because the extremes were distorting the averages, quoting those same averages in the summary undoes the choice.

And notice the row that says how the p was computed. Here it says exact: no two values in the pooled sixteen are tied and the groups are small enough to enumerate, so every possible arrangement was counted rather than approximated. In the ticket example above, the same row said normal approximation, because ties among the differences made counting impossible. The same tool, the same row, two different answers about its own arithmetic. Any tool that prints a p-value without telling you how it got there is asking for more trust than it has earned.

Choosing a rank test is not the same as throwing the outliers away. A 2.9-second load is a real load and a real person waiting. The rank test keeps that row in the data and refuses to let it set the verdict, which is a very different act from deleting it. If the slow tail is the thing you actually care about, then the extremes are the finding rather than the noise, and the instruments are a look at the outliers and a histogram with the cumulative column, not a test of the middle.

The same statistic, two different answers

Eight invoices processed by hand against eight through a new automated flow. Independent rows again, amounts again, and no outliers this time. The shortlist is the two ordinary t-tests, and most people pick between them by feel.

By hand, minutes per invoice24, 31, 18, 38, 22, 12, 29, 26
Automated, minutes per invoice9.5, 10.1, 8.9, 9.8, 10.4, 9.2, 10.9, 8.5

Run those two columns through both of the ordinary t-tests:

What the report gives youEqual VariancesUnequal Variances
Mean25.000 and 9.66225.000 and 9.662
Variance64.286 and 0.63764.286 and 0.637
Pooled Variance32.461not printed
df147
df (exact)not printed7.139
t Stat5.3845.384
P, two-tail0.0000960.001026
t Critical two-tail2.1452.365

The two tests return the identical statistic. That is not a coincidence and not a bug: when the two groups are the same size, the pooled formula and the unpooled one reduce to the same arithmetic. Everything that separates these two columns happens after the statistic, in the degrees of freedom, and that alone moves the p-value by a factor of about ten and a half.

What the degrees of freedom are buying. The variances are 64.286 and 0.637, a ratio of about 101 to 1. That is not a judgment call: run the same two columns through F-Test Two-Sample for Variances and it returns F 100.925 against a one-tail critical value of 3.787, with a one-tail p of 0.0000017. The equal-variances test answers by averaging those two numbers into a pooled variance of 32.461, which describes neither column, and then claims 14 degrees of freedom as though it had sixteen comparable readings. The unequal-variances test refuses to pool, and pays for the refusal: 7.139 degrees of freedom, rounded to 7 for the printed p-values, with the unrounded figure on a row of its own so you can see the price.

Choosing between these two is not a choice about the statistic. It is a choice about how much evidence you are entitled to claim from it. Unequal variances is the safer default because when the spreads genuinely do match it costs you almost nothing, and when they do not it is the only one of the two telling the truth. That is not the same as saying the other one is wrong.

And here the disagreement does not matter. 0.000096 and 0.001026 are both far under any threshold anyone works to, and both tests say the same thing: automation saves about 15.3 minutes an invoice. That is worth holding on to, because it is the honest counterweight to the load-time example. Two tests disagreeing is information about your data, not automatically a crisis. The question to ask is whether the disagreement moves the decision. Here it does not. There it moved the answer from act to there is nothing here, which is the whole distance.

The variance rows also carried a finding neither p-value mentions. Automation did not only move the average from 25 minutes to 9.7. It moved the variance from 64.286 to 0.637, which is what removing the 38-minute disasters looks like in numbers. Whichever test you run, the rows beside the p-value are usually where the useful part is.

What a result has to carry

A p-value on its own is not a result. Nobody can check it, nobody can tell whether the effect is worth acting on, and nobody can tell which test produced it. Whatever tool you use, these are the things to insist on beside it, and every one of them appears in the tables above:

  • The size of the effect, in the units of your problem. A median cut of 4 hours per ticket. 15.3 minutes an invoice. 450 milliseconds against 217.5. That is the number the decision is made on; the p-value only says whether it is worth believing.
  • The middle of each group, computed the way the test computes it. Means for a t-test, medians for a rank test. Quoting the means beside a rank test undoes the reason you chose it.
  • The spread, and how many rows. Two groups with the same averages and hundredfold different variances are not the same finding, and the variance rows are where that shows up.
  • The degrees of freedom. The invoice example is two tests with the same statistic separated by nothing else, and it moved the p-value more than tenfold.
  • How the p-value was computed. The two rank tests above print it: exact when every arrangement was counted, normal approximation when ties or size made that impossible. Kruskal-Wallis prints no such row, and nothing else here does either. It is a rare row and it deserves to be common.

Sortia writes all of those into a tab of your own spreadsheet, as a table, with a panel beside it saying in plain words what the figures mean, so the artifact you forward is checkable by whoever receives it rather than a screenshot of one number. One display note worth knowing: in the report tab the t-test, z-test and F-test tables format their value columns to three decimals, so a p-value below 0.0005 reads as 0.000 in the cell. The full value is still there and the formula bar shows it when you click the cell. The reading beside the table will not: that sentence writes anything under 0.001 as “well under 0.001” rather than printing the figure. The panel itself goes to the right of the table on a fresh report tab; append a run into a tab of your own and the numbers land without it.

What this page does not settle

  • Whether your data is normally distributed. Nothing in the tool list returns a verdict on that. There is no Shapiro-Wilk and no Levene. Distribution Fitting comes closest, and what it does is narrower than a verdict. It ranks candidate shapes against each other by Kolmogorov-Smirnov distance, the normal curve among them, and prints an Anderson-Darling statistic beside each one. It does carry two p-value columns beside the candidates, one from a chi-square and one from the K-S distance, and both headings call the figure an approximation, because the parameters were fitted from the same data those p-values are then judged against. The Anderson-Darling column deliberately carries no p-value at all, for the same reason taken to its conclusion: once the parameters were fitted from the same data no honest generic one exists. The report’s own note says to read a small p as evidence against a family and not to read a large one as proof of it, which is a ranking with a caveat on it rather than an answer to whether your data is normal. The absence is real, and it is less serious than it sounds: the shape question matters far less for comparing two averages than most people have been told, and far more for a tail percentile. Why the shape question is usually the wrong one takes that apart.
  • Which group differs, once three or more of them do. A significant F says at least one group is unlike the others and never which, and there is no post-hoc test here. ANOVA in Google Sheets does the follow-up arithmetic out loud, including why the threshold has to move when you run the pairs.
  • Which of the three t-tests fits your design, in more detail than one row of the chooser. t-Test in Google Sheets works through each of them, and through the built-in formula beside them.
  • Anything with repeated measures or mixed effects. Three or more measurements on the same subjects, or subjects nested inside groups, is past what this bench covers. Saying so is more useful than bending a test that is here into a shape it does not fit.

And it does not settle the biggest question of all, because nothing can: there is no single correct test for a situation. Three of the sections above run two defensible tests on one dataset and get different answers back, and in one of the three the difference changes the verdict. What exists instead is a test that matches the design you actually ran, a set of assumptions you have to be willing to say out loud, and an effect size that has to be worth acting on before the p-value is even interesting.

If you ran two and they disagreed, the useful move is not to quietly keep the smaller p. It is to say which one you chose before you looked, and why, and to report the other one beside it.