Version B is up 18%. Is that a real win or just noise?
Thirty six extra sign-ups is the entire case for Version B. Spread across 4,000 visitors per version it reads as an 18% lift, which is enough to get a variant shipped in most companies on most Mondays. Run the actual test and the gap comes back at p = 0.076, meaning a difference this big turns up roughly 8 times in 100 when the two versions are genuinely identical.
Words on this sheet
- Chi-square: A test on a table of counts: are two categories linked, or do the counts match the pattern you expected?
Analytics Intermediate Statistics free
After you install, this is the model to open.
Did the A/B Test Actually Win?
- In your spreadsheet, click the Sortia icon in the strip of icons down the right-hand edge. No strip? Click the arrow at the bottom-right to open it. You can also use Extensions, then Sortia, then Open Sortia.
- Click Start from a template and put that name in the search box.
- Pick the card with that name and click Load this template. It arrives on a new tab with real numbers already in it.
The answer
If the two versions really were identical, the expected count in every cell is the same for both rows: 218 conversions and 3,782 non-conversions. Version B is 18 conversions above that line and Version A is 18 below it, which sounds decisive until you weigh it against 8,000 visitors.
- Relative lift
- 18% 5.0% vs 5.9%
- p-value
- 0.076 above the 0.05 bar
- Chi-square, 1 df
- 3.14 cutoff is 3.84
- More traffic needed
- 888 per version, on top of 4,000
An 18% lift built on 36 extra conversions is not a win. The chi-square statistic comes to 3.14 against a cutoff of 3.84, giving p = 0.076, so this result is still inside the range that pure chance produces. Keep the same rates and run 888 more visitors per version, about 4,900 per arm instead of 4,000, and the identical gap would clear the bar. The lesson is that the size of a lift tells you almost nothing on its own, and calling the winner two weeks early is how teams end up shipping the losing variant with full confidence.
The model
The sheet holds one 2x2 table of raw counts, which is all a chi-square independence test needs. Rates and lift are computed below it so you can see the tempting number and the honest number side by side.
| Visitors per version | 4,000 each (8,000 total) |
| Version A conversions | 200 of 4,000 (5.0%) |
| Version B conversions | 236 of 4,000 (5.9%) |
| Extra conversions from B | 36 |
| Relative lift | 18% |
| Test | Chi-square independence, 1 degree of freedom |
| Bar for calling a winner | p below 0.05 |
Once it is in your sheet
- The model arrives with real numbers in it and runs as it stands, so you can press the button first and understand it second.
- Change the numbers to yours. The sheet marks which cells are inputs and which hold formulas, and most labels carry a note explaining the row.
- Press the run button at the bottom of the panel. It is labeled for the tool you are in, and the result lands on its own tab, with a written reading of it beside the figures.
Never used Google Sheets? Start here goes the whole way, in seven steps, and assumes nothing.
Next question
- Why are most of your flags false alarms?Bayes Flip: P(A|B) vs P(B|A)
- Your dashboard says stores win. Does the data agree?Simpson's Paradox Detector
- How long will the next batch take?Learning Curve: How Long Will the Next Batch Take?
- What is your average hiding?One Table That Describes Your Numbers
- You guessed min, likely, max. What does history say?Which Distribution Fits Your Data?
- What season length should your forecast use?Does Your Revenue Have a Season?
Every model like this one, and the method behind them: Statistics in Google Sheets.