Statistics, in the sheet the numbers already live in.

38 statistics tools sit in the sidebar. Point one at a column, press Run, and a report tab appears in your file with the table, a chart where the tool draws one, and a sentence beside it saying what it found.

Method guide for Google Sheets The data toolbox, 38 tools free, and no plan buys a bigger one

Regression in Google Sheets

Regression is the one most people arrive looking for, so it goes first. Name the column you are trying to explain and the columns you think move it, and the report comes back with the coefficient on each input, its standard error, its t statistic and its p-value, a confidence interval either side of it, Multiple R, R Square, Adjusted R Square and the ANOVA table with the F test underneath.

It is written into a tab of your own spreadsheet, not into a panel you have to screenshot. Beside the table, on a fresh report tab, there is a short reading: a sentence giving this run’s R Square and its Significance F, which says whether the model as a whole beats a flat line against the alpha you set, and then a line on how to use the rest, that each coefficient is the change in the outcome per one-unit change in that input with the others held still, and that its P-value says whether that input is pulling any weight. Which input carries the outcome is a call you make from that table. The tool sets the evidence out for it rather than naming a winner.

Two more tools sit either side of it. Correlation gives you the whole matrix of which pairs move together before you decide what belongs in the model, and Logistic Regression takes over when the outcome is a yes or a no rather than an amount. That one is on the machine learning page.

All 38, and what each one is for

Nine groups, the way the panel itself is arranged. Every one of them is free, and every one is deterministic: the same data and the same seed give the same answer, on any machine.

Describe

  • Descriptive Statistics mean, median, spread and the confidence half-width around the mean, for a column.
  • Histogram bins, counts and the cumulative share, with the chart drawn. How to make one, and why the bins are an argument.
  • Rank & Percentile tie-aware ranks, and where each row sits in the pack.
  • Outlier Detection IQR fences, and the rows that fall outside them.
  • Distribution Fitting fits candidate distributions and ranks them by goodness of fit.

Compare two or more groups

  • ANOVA: Single Factor one factor, three groups or more. How to run it, and what a significant F still leaves you to do.
  • ANOVA: Two-Factor With Replication two factors and the interaction between them.
  • ANOVA: Two-Factor Without Replication two factors, one reading per cell.
  • t-Test: Paired the same subjects measured twice, before and after.
  • t-Test: Equal Variances two groups that vary about the same amount.
  • t-Test: Unequal Variances Welch’s version, and the safer default.
  • z-Test: Two Sample two means when the variance is already known.
  • F-Test: Two-Sample whether two processes differ in spread rather than in average.
  • Chi-Square: Independence whether two categories move together at all.
  • Chi-Square: Goodness of Fit observed counts against the split you expected.
  • Mann-Whitney U two groups, with no bell curve assumed.
  • Wilcoxon Signed-Rank paired readings, with no bell curve assumed.
  • Kruskal-Wallis three groups or more, with no bell curve assumed.

Relate

  • Correlation the matrix of how every pair of columns moves together.
  • Covariance the same pairs, kept in the units of your data.
  • Regression coefficients, p-values, R Square and the ANOVA table.

Match

  • Fuzzy Lookup lines up rows that two systems spell differently.

Forecast

  • Moving Average the smoothed line, and its own error.
  • Exponential Smoothing recent weeks weigh more than old ones.
  • Holt-Winters trend and season carried together.
  • Autocorrelation (ACF) finds the repeat length hiding in the history.
  • ETS & ARIMA Forecast compares a family of models, names the winner, and bands the forecast.

Machine learning

  • Neural Network learns a curved pattern and reports a held-out score.
  • Logistic Regression the odds of a yes or no, with odds ratios beside them.
  • Classification Tree plain if-then rules learned from your rows, scored on rows it never saw.
  • K-Means Clustering groups the rows that resemble each other.
  • Association Rules which items show up together, with support, confidence and lift.

A/B testing

  • Two-Proportion z-Test whether B really beat A, or the gap is noise.
  • A/B Test Sample Size how many visitors you need before you look.
  • Sample Size (General) the sample a survey or a study has to reach.

Generate

  • Random Number Generation seeded draws from a distribution you name.
  • Sampling a random or periodic sample pulled from your rows.

Transform

  • Fourier Analysis pulls the repeating frequencies out of a signal.

Which of them do I actually need?

Two questions settle it, and both are about the design of the study rather than the numbers in it: how many groups are you comparing, and is each row in one group tied to one particular row in the other. Answer those and the shortlist above is usually down to one test, sometimes two.

Which statistical test should I use? is the chooser, one row per situation, and then the part that matters more: three cases where two defensible tests run on one dataset disagree. Nine matched weekdays where a rank test returns 0.0090 and a paired t-test returns 0.000281 on the same rows. Sixteen page loads where one test says 0.0104 and the other says 0.7253. And two t-tests that return the identical statistic and differ only in the degrees of freedom they will admit to.

What Google Sheets already does on its own

Some of these tools are a formula away, and a page that did not say so would be selling rather than explaining. T.TEST and F.TEST each return a probability, CORREL gives you a correlation, LINEST fits a regression and reports the standard error on every coefficient, and FREQUENCY, QUARTILE and RAND get you most of the way to a histogram, a set of outlier fences and a column of random draws.

Where is data analysis in Google Sheets? is the whole list, one row per analysis, with the built-in function named wherever there is one. Eight of them come back from a single call, twelve you can assemble out of built-ins and your own arithmetic, and eighteen have nothing behind them at all, including every ANOVA and every test that works on ranks instead of on a bell curve. That column was checked against Google’s published function list on 6 September 2026, and it carries the date because Google can move it.

Which t-test, and is the formula enough?

Three of the tools above are t-tests, and Google Sheets already has one of its own as a formula. Both make you choose between paired, equal-variance and Welch, and neither looks at your data to check that the choice fits. On one real pair of columns that choice moves the p-value by a factor of about 400.

Which of the three t-tests you need works through it on six datasets: what the built-in formula returns for each type, why Welch changes the degrees of freedom rather than the statistic, and the three things the report does not contain.

Three groups or more, and what a significant F leaves you

Four of the tools above take three groups or more: the three ANOVAs and Kruskal-Wallis, which asks the same question of the rank order instead of the averages. All four answer one narrow thing, whether the groups differ at all, and none of them says which pair is responsible. Eighteen exam scores across three revision methods come back at F 24.15 with a p-value of 0.00002, which settles that the methods are not interchangeable and names no winner, because naming one is not what the test does.

ANOVA in Google Sheets prints that run in full, both blocks and every column, and then does the part the report leaves to you: three pairwise t-tests judged against 0.05 divided by 3, so 0.0167 rather than 0.05, and why the threshold has to move once you ask three questions instead of one. It also runs the two-factor case, where twelve loaves separate what the recipe did from what the oven did and from the 0.5 cm interaction between them, and the case nobody writes up, where F comes back 2.42 on three crews and the honest answer is more days rather than a better test.

Fuzzy matching, for two lists that nearly agree

Fuzzy Lookup is in the same panel and is the one people are most surprised to find. Give it two columns from two exports and it scores every candidate pair, so “Globex Inc.” and “Globex Incorporated” end up on the same row instead of in two different totals.

It writes the matched pairs and their scores into a tab, which means you can read the near-misses and fix them by hand rather than trusting a join you cannot see.

Every result explains itself

A test that only prints a p-value is half a tool. Every run here ends with a card called What this means: the figure that matters, what it implies, and the caveat that goes with it. On the free plan that sentence is written by the add-on from the numbers the tool just computed, in your file, with nothing sent anywhere.

Pro adds an AI reading of the same result, written by Google’s Gemini API from the shape of the result alone, with a See what was sent panel showing the exact payload. Ratios, counts and cell references go; cell values, labels and formulas never do.

Free or Pro

All 38 are free, and no plan buys a bigger one. No run counter, no card, and no number a paid plan would unlock. A statistics tool on a hundred thousand rows costs the same as one on twelve. There is one ceiling and it sits on both plans alike: a run reads at most 500,000 cells, because the whole block is copied out of the sheet before anything starts. Two tools carry a limit of their own as well: Fourier Analysis reads at most 4,096 points, and k-means fits at most 50 clusters.

Pro covers five simulation engines: Monte Carlo risk simulation, decision trees, schedule risk, critical chain and optimization under uncertainty, plus the AI readings and a clean report footer. Every install gets five full-quality runs before it asks, shared across all five rather than five for each. Pro is $199/year.

Worked examples to start from

Each one loads into your sheet with real numbers already in it, and every figure on its page came out of the tool itself.

The other 80 models in the library that use one of these:

Try it in your own sheet

  1. Open Sortia in Google Sheets and choose Start from a template.
  2. Pick one of the models above, and it loads with the inputs filled in.
  3. Change the assumptions to fit your situation and press Run.

Never used Google Sheets? Start here goes the whole way, in seven steps, and assumes nothing.

Other methods: Monte Carlo  Decision trees  Schedule risk  Critical chain  Optimization under uncertainty  Machine learning  Forecasting  Optimization  What-if analysis