Google Sheets has three t-tests. It runs whichever one you name.

The formula is =T.TEST(range1, range2, tails, type). Three of those arguments are easy. The fourth decides which test you actually performed, and on one real pair of columns below it moves the answer by a factor of about 400.

Method guide for Google Sheets Six datasets, every figure computed all three t-tests are free

The formula that is already in your file

You do not need anything installed to run a t-test in Google Sheets. Type this into any empty cell:

=T.TEST(A2:A15, B2:B15, 2, 1)

Four arguments, and the first three are the obvious ones.

range1, range2the two columns you are comparing
tails1 or 2
type1, 2 or 3

tails is 2 unless you decided before you looked at the data that only one direction of change counts. If you are choosing it after seeing which way the numbers went, it is 2.

type is the argument this whole page is about. 1 is the paired test, 2 assumes the two groups spread about the same amount, 3 does not. Those are three different tests, they answer three different questions, and the formula runs whichever one you name without ever looking at your data to see whether it fits.

What comes back is a probability and nothing else. No means, no degrees of freedom, no variances, no sample sizes. One number in one cell.

The fourth argument, priced

Fourteen students sat an assessment on the way into a two-day course and again on the way out. Two columns, one row per student, same order. Here is the identical pair of ranges through all three types:

Formulap (two-tail)What it assumes
T.TEST(…, 2, 1)0.00049the rows are pairs
T.TEST(…, 2, 2)0.1949two groups, equal spread
T.TEST(…, 2, 3)0.1957two groups, unequal spread

Pre-course and post-course scores, 14 students. Computed 6 September 2026; the run is in Did the Course Move the Scores?

Nothing about the data changed between those three lines. One of them reports a result you could take to whoever paid for the course. Two of them report nothing at all. The second p-value is about 400 times the first.

Type 1 is the honest one here, because row 3 on the left and row 3 on the right are the same student. The other two throw that away and treat 28 unrelated scores as two piles. These students range from 35 to 88 on the way in, so the pile-to-pile spread swamps a 6.8-point course effect and the unpaired forms never see past it. Pairing is what makes fourteen people enough.

Now the same mistake pointing the other way. Eight invoices handled by hand, eight through a new automated flow, different invoices on each side. Run type 1 on those two columns and it returns 0.00081, a confident-looking number that means nothing, because there is no sense in which the third manual invoice and the third automated one are the same invoice. The formula cannot know that. It has two ranges of eight numbers and a type you typed.

type is not a setting. It is you telling the formula what you did, and it will answer whichever thing you tell it, in the same confident tone, to five decimal places.

Question one: is row three the same thing on both sides?

If yes, it is the paired test. Before and after on the same people, the same machines, the same accounts, the same weeks. If no, it is one of the two-sample tests, and question two below picks which.

That is the whole decision, and it is a fact about how you collected the data rather than anything you can read off the numbers.

Pairing is worth understanding because of what it buys. Eight people logged their nightly sleep, dropped the afternoon coffee for a fortnight, and logged again: 6.03 hours on average before and 6.56 after, a gain of about 32 minutes, and all eight improved.

Cost of ignoring the pairing
123×
Paired, p (two-tail)
0.00003
Same two columns, unpaired
0.0037

Eight people, sleep before and after. Computed 6 September 2026; the run is in Does Cutting Caffeine Buy You Sleep?

Some of those people are natural five-and-a-half-hour sleepers and some are closer to six and a half. An unpaired test has to see past all of that before it can see a half-hour change. A paired test never looks at it: it works only on each row's difference, so the between-person spread is gone before the arithmetic starts.

One practical difference. The paired test needs the two ranges to be the same length, and both the formula and the tool refuse rather than guess. The tool says so in words: “A paired t-test compares each subject against itself, so both ranges need the same number of rows: extend the shorter one, or trim the longer.” The two-sample tests are happy with columns of different lengths, which is the other reason a blank row in the middle of your data changes the answer to a question you did not think you were asking.

Question two: do the two columns spread the same?

This one only comes up if the rows are not pairs. Type 2 assumes the two groups vary about the same amount and pools them. Type 3, Welch, does not assume it.

Back to the invoices, which are the case Welch exists for. Manual handling swings from 12 to 38 minutes. The automated flow sits between 8.5 and 10.9. Same eight and eight, run both ways:

type 2, pooledtype 3, Welch
t5.38405.3840
df147.139
p (two-tail)0.0000960.000962

Manual variance 64.29, automated 0.64. Computed 6 September 2026; the run is in Manual vs Automated, Tested Fairly.

Look at the first row. The statistic did not move at all. Both tests compute t = 5.3840 from exactly the same means and variances. Every bit of the ten-fold difference in the p-value comes from the second row, the degrees of freedom, which is the test's own statement about how much evidence it thinks it is holding.

That is the honest way to describe Welch. It does not disbelieve your data. It notices that one column carries a hundred times the variance of the other, decides that eight wild numbers and eight steady ones are not sixteen equally informative numbers, and cuts the degrees of freedom from 14 to 7.139 accordingly.

The pre-check, and what it costs to skip it

You do not have to guess whether the spreads match. The F-test compares two variances directly, and running it first is the standard move before a two-group t-test. On these two datasets it says opposite things:

DataFp (one-tail)F crit
Invoices, manual vs automated100.930.00000173.787
Takings, two stores1.870.1572.818

Computed 6 September 2026. The store data is Are These Two Stores Really Different?, twelve trading days each.

The invoices fail the check comfortably, so type 3 is the one to report. The stores pass it, and there the choice barely matters: type 2 gives p 0.0000428 on 22 degrees of freedom, type 3 gives p 0.0000552 on 20.145, and t is 5.0864 either way. Same decision, same meeting.

That is the whole argument for making Welch your default, stated as a price rather than a rule. When the spreads really do match, using Welch anyway costs you about 29% on the p-value and nothing on the conclusion. When they do not match and you pooled anyway, you publish a number ten times smaller than the data supports.

One trap worth naming while you are here. The built-in F.TEST formula returns the two-tailed probability, and the F-Test tool in the panel reports the one-tail figure beside its F critical value. On the invoices those are 0.0000034 and 0.0000017. Same conclusion, one is twice the other, and you should know which one you are about to quote in a document.

When one number in one cell is not enough

T.TEST is genuinely fine for a lot of work. If you know your design, you know which type you need, and you only want the probability, the formula is free, it is already there, and nothing on this page beats it.

It stops being enough the moment somebody has to check you. A p-value on its own cannot be audited: nobody reading it can see how many rows went in, what the two averages were, whether a blank thinned one column, or which of the three tests produced it. Here is what the same Welch run writes into a tab of your spreadsheet instead, every row of it, and the value each cell holds:

Mean25 · 9.6625
Variance64.2857 · 0.63696
Observations8 · 8
Hypothesized Mean Difference0
df7
df (exact)7.1387
t Stat5.38395
P(T<=t) one-tail0.000513
t Critical one-tail1.89458
P(T<=t) two-tail0.001026
t Critical two-tail2.36462

The t-Test: Two-Sample Assuming Unequal Variances report, on the invoice data.

One thing to know about that block before you screenshot it. The figures column is displayed at three decimals, so on the sheet both p rows read 0.001, even though one of them is exactly twice the other. The cells hold the full values above, and so does anything you point a formula at; only the display is rounded. The written reading beside it keeps the distinction in words, which is why it says “p is 0.001” here and “well under 0.001” on a sharper result.

Beside the figures on the same tab the tool writes one sentence about this run, from the numbers it just computed:

“This run: t = 5.38. The two groups average 25 and 9.66, a gap of 15.34. p is 0.001, below your alpha of 0.05, so the difference is more than chance would ordinarily produce. (two-tail)”

Why there are two df rows

Welch's degrees of freedom are fractional. Here they are 7.139. There are two conventions for what to do next: use the fraction as it stands, or round it to the nearest whole number, which here is 7. They give slightly different p-values, 0.000962 and 0.001026 on this data.

The formula uses the fraction. The report rounds, and then prints the exact figure on the very next line so you can see that it did. Neither number changes any decision on the invoices, which is exactly why it is worth showing: on a result sitting near your threshold it would, and you would want to know which convention produced the number you are quoting. A report that shows only the rounded df is hiding a choice it made on your behalf.

Three things the report does not contain

Said plainly, because a page that only lists what a tool has is an advertisement.

No confidence interval on the difference. The panel puts a range around the gap while you are looking at the result, and the tab does not: it gives you both means and both variances, and no interval around the gap between them. So report the gap in your own units beside the p-value, always: 15.3 minutes an invoice, 6.79 points on the assessment, $342.50 a day of takings. The p-value is the odds of a run like this happening by chance. It has never once told anybody how big anything is.

No effect size on a t-test. There is no Cohen's d anywhere in this product, and no page here is going to pretend otherwise. The gap in your own units does the same job for most audiences and needs no explaining.

No one-sample t-test. Neither the formula nor the panel has one, so “is my column different from 10?” is not a t-test you can run here. What you can do is run Descriptive Statistics on the single column and read its confidence interval. On the eight automated invoices the mean is 9.66 minutes and the 95% half-width is 0.667, so the interval runs from 8.995 to 10.330. A ten-minute target sits inside it. Those eight invoices settle comfortably that automation beat the manual flow, and they do not settle that it beats ten minutes. Two different questions, and only one of them is answered.

And when a t-test is the wrong family entirely

You already know both variances. Then it is a z-test, not a t-test. Two bottling lines with variances of 0.25 and 0.64 documented from months of process history: Line 1 averages 500.23 g, Line 2 sits at 499.03 g, z is 3.60 and p is 0.0003. The numbers being known rather than estimated from today's eight bottles is the entire reason that test applies. If you do not have documented variances, do not invent them. The worked run is here.

The question is about spread, not average. Two machines cutting to 250 mm with nearly identical averages are not the same machine if one wanders 2.40 mm and the other 0.24 mm. A t-test finds nothing there and should not. The F-test does.

The differences are badly skewed, or the values are ranks. Then the rank-based tests ask the same question without assuming a bell curve, and they cost you something measurable for the privilege. How much the shape actually matters is a question with a real answer, and it is not always “a lot”.

How to check every number on this page

Every dataset on this page is small on purpose, and every one of them is a free model you can load: the two columns land on a tab in your own file, and the formula goes in the cell beside them. It takes about a minute per figure, and if your sheet disagrees with anything here, tell us: we would rather hear about it than not.

Every figure above was computed on 6 September 2026 by the engines that ship in the add-on. Two of the three T.TEST values in the first table, the paired one and the equal-variance one, were computed a second time by the spreadsheet-formula engine in the same repository, which is separate code with its own arithmetic, so those two came back the same from two code paths rather than from one printed twice. The third one does not agree, and it is the disagreement this page has already described: 0.1957 is the formula reading the table at Welch’s fractional degrees of freedom, 24.402 on these fourteen students, while the add-on’s own Welch test returns 0.1959 because it rounds that to 24 first. Same convention split as the invoices above, on a pair of columns where it changes nothing at all. The t-tests themselves, pooled and Welch, including the degrees-of-freedom rounding, are checked on every commit against published critical values and against fixtures generated by separate statistical software. What else is checked, and how.

Running one on your own data

The three t-tests are three of the 38 statistics tools in the panel, under the names the report uses: t-Test: Paired Two Sample for Means, t-Test: Two-Sample Assuming Equal Variances and t-Test: Two-Sample Assuming Unequal Variances. All three are free, at any number of rows.

  1. Open Sortia in Google Sheets and pick the test. If you are not sure which, two questions decide it.
  2. Point Variable 1 range and Variable 2 range at your two columns. Tick Labels in first row/column if the top cell is a heading, and the report will use your own column names instead of “Variable 1”.
  3. Leave Hypothesized mean difference at 0 unless you are testing against a specific change rather than against no change, and leave Alpha (significance level) at 0.05 unless you have a reason.
  4. Press Run analysis. The report lands on its own tab in your file, with the reading written beside the figures.

Never used Google Sheets? Start here goes the whole way, in seven steps, and assumes nothing.