Is my data normally distributed?
Computed 6 September 2026. Every figure below comes from running the shipped engine functions on the data that ships inside two free templates, through the repository's own test harness (test/harness.js), at the parameters printed beside each table. The simulations carry a seed, so they reproduce to the digit rather than approximately, and the last section is a step-by-step recipe for getting every number back in your own spreadsheet.
Almost certainly not. Real measurements are hardly ever exactly normal, and a large enough sample will always prove it. The useful question is narrower, and it decides everything: does the number you are about to produce come out of the middle of your data, or out of its tail? In the middle, the shape is mostly forgiven. In the tail, the shape is most of your answer. This page measures that difference on forty real task durations, and then shows the case where getting it wrong reverses a decision.
There is no normality test here, and none in the sheet either
If you came looking for Shapiro-Wilk, you will not find it. Sortia ships no normality test at all: no Shapiro-Wilk, no Jarque-Bera, no Levene, no Bartlett test of variances, anywhere in the add-on. Google Sheets has no built-in function that takes a column and returns a verdict on the bell curve either. That gap is not an oversight to work around, and the rest of this page is why.
What does ship is Distribution Fitting, which is a different thing, and its own help text is careful to say so:
It ranks the candidates against each other. It cannot tell you the best of them is any good, and under a few dozen points the ranking is mostly noise. A good fit to last year is not a promise about next year. Distribution Fitting, "What it cannot tell you", in the tool guide
That is the honest shape of the problem. You can measure how far your data sits from a family of curves, and you can compare one family against another. You cannot get a certificate saying your data belongs to one.
A pass-or-fail normality test also answers a question nobody asked. Its power grows with your sample size, so on forty rows it waves almost anything through, and on forty thousand rows it rejects data that is normal enough for every purpose you had in mind. The verdict turns on how much data you happened to collect rather than on how much the departure would cost you. That is the wrong axis to be measuring.
Forty task durations, looked at properly
The worked example is the free template What Shape Are Our Task Durations?, which ships with forty finished tasks and the days each one took from start to done. Nothing here is invented for the page: it is one column of forty numbers, run through the tools as they ship.
First, the picture
Descriptive Statistics on that column, one table:
| Statistic | Value |
|---|---|
| Count | 40 |
| Mean | 7.45 days |
| Median | 6.25 days |
| Standard deviation | 4.27 |
| Skewness | 0.94 |
| Kurtosis | 0.13 |
| Minimum / maximum | 1.5 / 18 |
The mean sits a full day above the median. That gap, and a positive skewness, is what a long right tail does to an average, and the two numbers have already told you the answer before any test runs.
Then Histogram with no bin limits supplied, so it picks its own. The rule it uses is worth knowing rather than trusting blindly: k equal-width bins from the minimum to the maximum, where k is the square root of the row count rounded up, and the top edge is snapped to the maximum. Forty values gives seven bins. That is the square-root rule, which is a convention, not an optimum.
| Bin, up to | Tasks | Cumulative |
|---|---|---|
| 3.86 days | 7 | 17.5% |
| 6.21 days | 13 | 50.0% |
| 8.57 days | 8 | 70.0% |
| 10.93 days | 4 | 80.0% |
| 13.29 days | 3 | 87.5% |
| 15.64 days | 2 | 92.5% |
| 18.00 days | 3 | 100.0% |
Half the tasks are finished inside 6.2 days. The other half runs all the way out to 18. That is the diagnosis, and it took one chart and no p-value.
One more reading, because it shows how a symmetric habit behaves on a lopsided column. Outlier Detection on the same forty values, with the usual 1.5 times the interquartile range, puts its fences at minus 3.50 and 17.50 days and flags exactly one value, the 18-day task. A rule that goes looking for suspiciously short tasks below minus three and a half days is not calibrated to this data. It is calibrated to a bell curve.
Then the ranking
Distribution Fitting ranks six candidate shapes by Kolmogorov-Smirnov distance, smallest first. This is the whole ranking, in the number formats the report itself uses.
| Rank | Distribution | KS | Anderson-Darling | AIC | BIC | Chi-square | p, chi-sq (approx.) | p, KS (approx.) |
|---|---|---|---|---|---|---|---|---|
| 1 | lognormal | 0.0449 | 0.120 | 222.3 | 225.7 | 0.29 | 0.8640 | 1.0000 |
| 2 | triangular | 0.0835 | 0.554 | 226.3 | 231.4 | 1.54 | 0.2148 | 0.9327 |
| 3 | normal | 0.1420 | 1.132 | 232.7 | 236.0 | 5.18 | 0.0749 | 0.3693 |
| 4 | PERT | 0.2518 | 10.195 | 273.7 | 278.8 | 41.01 | 0.0000 | 0.0100 |
| 5 | exponential | 0.2565 | 3.710 | 242.7 | 244.3 | 13.09 | 0.0044 | 0.0082 |
| 6 | uniform | 0.2785 | 4.971 | 230.2 | 233.6 | 15.15 | 0.0005 | 0.0031 |
Lognormal wins on every column that has an opinion: smallest KS distance, smallest Anderson-Darling,
best AIC and best BIC. The normal comes third, three times further from the data than the winner. The
report also hands over the parameters, which is what makes the next section possible:
Lognormal(mean 7.5505, sd 4.8977) and
Normal(mean 7.4500, sd 4.2710).
The number that would have fooled you
Look at the normal row again, last column. The approximate KS p-value against the bell curve is 0.3693.
Read that as a normality test and it says: no evidence against the normal, carry on. It is the answer most people arrive at this question hoping for. Here it would be wrong three times over.
- It is not a comparison. The same run, on the same forty numbers, put lognormal three times closer. A p-value that fails to reject a family is not evidence that the family is the best one available. It is evidence that forty points cannot rule it out, which is a much smaller claim.
- It is optimistic by construction. The parameters of that normal curve were estimated from the very data being tested against it, which drags the curve toward the sample and shrinks the measured distance. The engine names this in its own source and prints "(approx.)" in both p-value column headers for exactly that reason. The true p is smaller than 0.3693, by an amount no generic formula can quote.
- The tool tells you not to read it that way. Printed with every fit report: "Read a small p as evidence against a family; do not read a large p as proof of it."
There is a second decision in that same report, and it is the same lesson from the other side. The Anderson-Darling column ships with no p-value at all, deliberately: under estimated parameters its null distribution is specific to each family, and no honest generic approximation exists, so the report gives you the statistic and refuses to give you a verdict. That is what an honest goodness-of-fit answer looks like. A distance you can compare, not a pass mark.
The question that actually matters: middle, or tail?
So the ranking says lognormal and the p-value says the normal is fine. Which of those you should care about depends entirely on what you are about to do with the answer, and that dependency can be measured rather than argued.
Both shapes below are the fitter's own output at the parameters it printed. Each was put into a Risk Analysis input and run at 10,000 trials with Latin hypercube sampling, which are the panel's own defaults, and the seed set to 12345. The seed box is blank by default, which means a fresh random draw on every run; typing a number into it is what makes a result somebody else can reproduce. The ladder below is each run's percentile table put side by side, with three more readings from the same two reports under it.
One task
| Percentile | Normal(7.4500, 4.2710) | Lognormal(7.5505, 4.8977) |
|---|---|---|
| 1% | -2.48 | 1.60 |
| 5% | 0.43 | 2.39 |
| 10% | 1.98 | 2.96 |
| 25% | 4.57 | 4.25 |
| 50% | 7.45 | 6.33 |
| 75% | 10.33 | 9.45 |
| 90% | 12.92 | 13.54 |
| 95% | 14.47 | 16.78 |
| 99% | 17.37 | 25.10 |
| Simulated skewness | -0.00 | 2.16 |
| Worst trial of 10,000 | 23.51 | 58.81 |
| Trials below zero | 405 | 0 |
Read the middle of that table and the two shapes barely disagree. Read the ends and they disagree completely.
The median moves about a day, and it moves in the direction nobody expects: assuming the bell curve makes your typical task look worse, 7.45 days against 6.33. Meanwhile the one-task-in-a-hundred case moves from 17.37 days to 25.10, more than a week apart on the same forty rows. At the 95th percentile, which is where most plans actually commit, it is 14.47 against 16.78, sixteen percent apart. Choosing a shape does not add a safety margin evenly, it moves mass from the middle into the tail. If you are quoting a date you intend to hit 99 times in 100, the shape you assumed is most of your answer.
The bell curve also says something impossible on the way. 405 of its 10,000 trials came out negative, and the lowest of all ten thousand was minus 9.21 days: tasks that finish before they start. That is not a rounding artefact, it is the shape being wrong in a way you can see. Sortia's own note on the normal distribution says so before you ever pick it: "It has no floor. A normal can go negative, so it is the wrong shape for a price, a duration or a headcount unless you add a lower limit."
The same ten tasks, added up
Now ten independent tasks of that kind, summed, under each shape. Same trial count, same seed.
| Percentile | Normal, sum of 10 | Lognormal, sum of 10 | Gap |
|---|---|---|---|
| 5% | 52.31 | 53.16 | 1.6% |
| 25% | 65.24 | 64.44 | -1.2% |
| 50% | 74.34 | 73.68 | -0.9% |
| 75% | 83.49 | 84.60 | 1.3% |
| 90% | 92.20 | 95.80 | 3.9% |
| 95% | 97.15 | 104.35 | 7.4% |
| 99% | 106.17 | 120.04 | 13.1% |
| Simulated skewness | 0.05 | 0.69 |
At the median the two shapes now differ by under one percent. At the 95th percentile the gap has fallen from 16 percent to 7. And the simulated total's skewness drops from 2.16 to 0.69: the sum of ten lopsided things is markedly less lopsided than any one of them.
That is why the textbook answer to "does my data have to be normal" is so often "not really". Most of the numbers people compute are sums or averages, and sums and averages drift toward a bell curve whatever the pieces look like. The shape gets averaged away. What never gets averaged away is a single extreme value, and a far tail is a question about single extreme values by definition.
So the rule the whole page reduces to:
- If your answer is a middle, a mean, a total, a difference between two averages, and you have plenty of rows, the shape of the raw data is largely forgiven. Stop worrying about it.
- If your answer is a tail, a P95, a worst case, the chance of breaching a limit, the shape is your answer. Choose it on purpose and write down which one you chose, so the next person can disagree with the assumption rather than with the number.
When the shape decides the verdict outright
"Plenty of rows" is doing real work in that rule. Here is the case where it fails, and it fails hard enough to reverse a decision rather than soften one.
The free template Faster Site, Proven Without the Bell Curve ships eight page-load times from an old CDN and eight from a new one, in milliseconds. One bad measurement on each side, which is exactly what real telemetry looks like:
| Old CDN | 420 | 380 | 460 | 2900 | 510 | 440 | 395 | 470 |
|---|---|---|---|---|---|---|---|---|
| New CDN | 210 | 180 | 240 | 195 | 260 | 225 | 205 | 3100 |
Both of the tools below legitimately apply. Run them on the identical sixteen numbers:
| Welch t-test, on means | Mann-Whitney U, on ranks | |
|---|---|---|
| What it compares | 746.9 ms against 576.9 ms | median 450 ms against 217.5 ms |
| Statistic | t 0.3585 | U 8 |
| P (two-tail) | 0.7253 | 0.0104 |
| P computed by | t distribution, df 14 (exact 13.67) | exact, every arrangement counted |
| Verdict at 0.05 | no difference detected | the new CDN is faster |
The t-test does not merely lose a little power here. At the 0.05 gate it sends the decision the other way, and, as with the other large p-value on this page, what it returns is a failure to tell the two apart rather than evidence that they match. Two values out of sixteen push the two sample variances to 758,635 and 1,040,014, and against spreads that wide a 170 ms difference between the means is invisible. The rank test never sees the size of those two values, only their position, so it reports what the other fourteen rows plainly show.
Notice what did the damage, because it is not quite what this page has been about. It was not skewness in the abstract. It was that at eight rows a single value can move a mean on its own. A normality verdict on either column would not have helped. Both are lopsided, skewed at about 2.8 each, and knowing that does not tell you which way the error runs or what to do. The question that catches it in four seconds, with no test at all, is: can any single row here move my answer? Sort the column and look at both ends.
One detail worth keeping. That 0.0104 is not an approximation. With no tied values and eight against eight, the report says "P computed by: exact", because every possible arrangement of the sixteen numbers was counted rather than modeled.
What to do instead, in order
- Draw the histogram, and draw it twice. Once with the automatic bins, once with bin limits that mean something in your world. If the two pictures tell different stories, that disagreement is the finding, and bin width is an argument you are making rather than a setting.
- Read the mean against the median, and read the skewness. Descriptive Statistics gives you both in one table. Mean above median with positive skewness is a right tail, and you do not need a test to confirm what those two numbers have already said.
- Rank the families, and take the ranking as a ranking. Distribution Fitting will tell you which of up to six candidate shapes sits closest and hand you its parameters. It will not tell you the winner is any good, and on a few dozen points the order itself can move.
- Then ask the only question that pays. Middle or tail? A middle with plenty of rows: proceed. A tail: pick the shape deliberately and state it. Two groups where one row could move a mean: use the rank-based test and report the medians, not the means the outliers have already bent.
Nowhere in that list is there a step called "test for normality", and that is the point. Every step is either a picture you can read or a decision you have to make anyway.
Reproducing every figure on this page
All of it comes from two free templates and seven tools. Six of the seven are free, with no run counter and no cap a plan can raise. The seventh, Risk Analysis, is a paid engine, and the four runs this page uses fit inside the five full-quality runs of the paid engines that every free install gets, with one to spare.
- Open Extensions, then Sortia, then Open Sortia, click Start from a template and load What Shape Are Our Task Durations? It arrives on its own tab with the forty durations already in it.
- Run Descriptive Statistics, Histogram with the bin range left empty and Cumulative percentage ticked, and Outlier Detection on that column for the three readings in the first section.
- Run Distribution Fitting on the same column for the six-row ranking. The parameters in its last column are the ones used below.
- For the one-task ladder, the smallest model there is: put a number in A1 and
=A1in B1. In Risk Analysis, make A1 the input withNormal, mean 7.45, sd 4.2710and B1 the output, leave trials at 10,000 and Latin hypercube on, type 12345 into the seed box, and run. Then change the input toLognormal, mean 7.5505, sd 4.8977and run again. The skewness and the worst trial are the Skewness and Maximum rows of the Simulation Statistics table beside the ladder, the minus 9.21 is its Minimum, and the trials below zero come from setting the chance-of-clearing box to At or below and typing 0. For the ten-task ladder, ten input cells with the same shape and=SUM(A1:J1)as the output. Worth knowing before you type: the lognormal here is parameterised by the distribution's own mean and standard deviation, not by log-space mu and sigma, which is what lets you paste the fitted numbers straight in. - For the two-group comparison, load Faster Site, Proven Without the Bell Curve and run Mann-Whitney U and then t-Test: Unequal Variances (Welch) on the same two columns.
Because the simulations carry a seed, they reproduce exactly rather than closely: the same model at 10,000 trials with Latin hypercube sampling and seed 12345 returns this table, not one like it. If any figure here does not come back, that is a defect and we would like to know: support@sortia.io.
Where to go next
- Statistics in Google Sheets, the full bench and what each tool is for.
- Which statistical test should I use?, if the answer you needed was a test rather than a shape.
- What Shape Are Our Task Durations?, the forty durations used above, and the three-point estimate they produce.
- What Distribution Do Claim Sizes Follow?, the same method on thirty-six settled insurance claims, where the tail is the entire business.
- What Shape Are Your Wait Times?, twenty-four support calls and the bin limits that turn a shape into a promise you can check.
- Monte Carlo simulation, for what to do once you have picked a shape on purpose.
Sortia is a Google Sheets add-on for risk analysis (estimates in, odds out), optimization and decision analysis. The six analysis tools this page used are free on every plan, with no run counter and no cap a plan can raise, and the simulation engine is the paid half. If you have barely used Google Sheets before, Start here takes you from an empty browser tab to your first answer. You can install it free (opens the Google Workspace Marketplace in a new tab) or read how the numbers are checked first.