Your traffic is fixed. What lift can it even see?
The usual question runs the other way: you have 3,800 signups per arm and want to know what that buys. Reading the table backwards says a 25% lift, which is far bigger than this team has ever shipped.
SaaS Intermediate A/B Test Sample Size free
After you install, this is the model to open.
What Lift Could We Even Detect?
- In your spreadsheet, click the Sortia icon in the strip of icons down the right-hand edge. No strip? Click the arrow at the bottom-right to open it. You can also use Extensions, then Sortia, then Open Sortia.
- Click Start from a template and put that name in the search box.
- Pick the card with that name and click Load this template. It arrives on a new tab with real numbers already in it.
The answer
- Your traffic
- 3,800 signups per arm in four weeks
- A 20% lift needs
- 4,321 per arm; it does not fit
- First lift that fits
- 25% 2,821 per arm, with room to spare
- Ever shipped
- 11% the best lift this team has produced
This template runs the sample-size question backwards, which is how it actually arrives. You do not get to choose your traffic: 1,900 signups a week, four weeks, split two ways, is 3,800 per arm and that is the whole budget. The question is what that buys. Click Run at the prefilled minimum detectable effect of 20%: the answer is 4,321 per arm.
You have 3,800, so a 20% lift is out of reach. Change the effect to 25 and run again: 2,821 per arm, which fits with room to spare. Read down the Fits in our traffic column and the first yes is at 25%, so 3,800 signups per arm can reliably detect a lift of about a quarter and nothing smaller. Now look at the three onboarding changes listed at the foot of the sheet, which is why this is worth doing before the test rather than after.
They produced lifts of 6%, 11% and 4%. Nothing this team has ever shipped has moved trial-to-paid by 25%. So a four-week test of the next change will almost certainly end inconclusive, not because the change does nothing but because the test cannot see effects of the size this team actually produces. That is a finding, and it has three honest responses: run for twelve weeks instead of four, test something with a bigger expected effect, or ship on judgment and measure the cohort afterwards.
Running the four-week test anyway and calling the null result evidence of no effect is the one response that is wrong, and it is the most common. Note the shape of the Needed per arm column. Going from 25% to 5% multiplies the requirement by twenty-three, from 2,821 to 64,914, because the sample needed grows with the square of the effect you are chasing.
Small effects are not slightly harder to prove, they are categorically harder. Two runs worth doing. Set Power to 90 and rerun at 25: the requirement rises by a third to 3,776, which still just fits inside the 3,800 you have, so you can cut your chance of missing a real effect from 20 in 100 to 10 in 100 for nothing. And set the baseline to 4.5 to see what a lower-converting product costs: 5,961 per arm at the same 25% lift, which is roughly double, everywhere.
To use your own product, change the conversion rate, the weekly signups and the number of weeks at the top of the sheet, and put your own baseline into the panel.
The model
It arrives on a tab called Template: What Lift Could We Detect, carrying these columns:
- Trial-to-paid conversion today
- 0.09
with the model computed beside the data:
| Signups per arm available | 3,800 |
| Smallest lift our traffic can detect (%) | 0.25 |
Once it is in your sheet
- The model arrives with real numbers in it and runs as it stands, so you can press the button first and understand it second.
- Change the numbers to yours. The sheet marks which cells are inputs and which hold formulas, and most labels carry a note explaining the row.
- Press the run button at the bottom of the panel. It is labeled for the tool you are in, and the result lands on its own tab, with a written reading of it beside the figures.
Never used Google Sheets? Start here goes the whole way, in seven steps, and assumes nothing.
Next question
- Do the service credits or the churn cost you more?What Would an Outage Cost Us?
- Commit to the servers or rent them by the hour?Which Instance Mix Holds the Load for the Least Money?
- Is the migration in trouble, or just early?Buffer the Migration, Not the Tasks
- Is it cheaper to staff the night or to buy it in?Cover the Support Hours for the Least Money
- How late will the release really be?Can We Ship on That Date?
- Does your roadmap leave revenue on the floor?Which Features Fit the Quarter?
Every model like this one, and the method behind them: Statistics in Google Sheets.