Did the new format really rate better?
Twelve people in each of two cohorts, rating the same session one to seven. Averaging a rating scale is something everybody does and nobody can defend, so this test compares the orderings instead and still returns a verdict.
Words on this sheet
- Median: The middle value: half the readings sit above it and half below.
Analytics Starter Statistics free
After you install, this is the model to open.
Two Groups, Ranked Feedback, One Answer
- In your spreadsheet, click the Sortia icon in the strip of icons down the right-hand edge. No strip? Click the arrow at the bottom-right to open it. You can also use Extensions, then Sortia, then Open Sortia.
- Click Start from a template and put that name in the search box.
- Pick the card with that name and click Load this template. It arrives on a new tab with real numbers already in it.
The answer
- The verdict
- p = 0.0081 the redesigned session genuinely rated higher
- Medians
- 5.5 vs 4 new format against old, on the same 1-to-7 scale
- The honest size
- 9 of 12 new-format ratings sit above the old median
- The statistic
- U = 26.5 z 2.65 across all twenty-four pooled ratings
Twelve attendees rated the redesigned session and twelve different attendees rated the original, both on the same one-to-seven scale. The scale key sitting beside the ratings is the reason this is not a t-test: a seven is better than a six, but nothing anywhere says it is better by the same amount that a four is better than a three, so an average of these numbers is arithmetic performed on labels.
Mann-Whitney sidesteps that entirely. It pools all twenty-four responses, ranks them, and asks whether one group's ratings sit higher in the order than chance would put them. Click Run: U comes back 26.5 with a z of 2.65 and a two-tail p-value of 0.0081, and the medians are 5.5 for the new format against 4 for the old. Two identical sessions would produce a separation this clean about eight times in a thousand, so the redesign genuinely rated better.
What the test gives you is a direction and a confidence, not a size. It will never tell you the new format is one and a half points better, because on this scale that sentence has no meaning. If you want a size, report the two medians and the share of new-format ratings that beat the old-format median, which here is nine of twelve and is on the sheet.
Swap the two ranges over and rerun: U stays at 26.5 and the p-value stays at 0.0081, because the report always shows the smaller of the two U values. Only the sign of z flips, and that sign is the report telling you which group sits higher. If your two groups are the same people rated twice, this is the wrong test and Wilcoxon Signed-Rank is the right one, because pairing throws away the differences between people and this test cannot.
What no test can fix is a feedback form that only the enthusiastic filled in: a rank test is robust to the shape of the ratings, not to who chose to give them. To use your own survey, put one group under each rating heading, keep the two groups independent, and widen both ranges to match.
The model
It arrives on a tab called Template: Ranked Feedback, Two Groups, carrying these columns:
- Attendee
- Rating, new format (1-7)
- Attendee (old format)
- Rating, old format (1-7)
with the model computed beside the data:
| Median, new format | 5.5 |
| Median, old format | 4 |
| New ratings above the old median | 9 |
Once it is in your sheet
- The model arrives with real numbers in it and runs as it stands, so you can press the button first and understand it second.
- Change the numbers to yours. The sheet marks which cells are inputs and which hold formulas, and most labels carry a note explaining the row.
- Press the run button at the bottom of the panel. It is labeled for the tool you are in, and the result lands on its own tab, with a written reading of it beside the figures.
Never used Google Sheets? Start here goes the whole way, in seven steps, and assumes nothing.
Next question
- How do you know a forecast method actually works?Make a Demand Series to Test a Model
- Do ads move your sales, or is it price?What Actually Drives Your Sales?
- Which subscribers are about to cancel?Churn Early-Warning Scorecard
- How good is each team, really?Power Ratings & Spread Predictor
- How many survey responses do you actually need?Survey Sample Size Planner
- Version B is up 18%. Is that a real win or just noise?Did the A/B Test Actually Win?
Every model like this one, and the method behind them: Statistics in Google Sheets.