Did the workshop actually move anyone?
Twelve assessors rated before and after, on a one-to-ten competency scale. Two scored the same both times and the test drops them, which is the rule everybody gets wrong and the reason the sample is smaller than the headcount.
School Intermediate Statistics free
After you install, this is the model to open.
Did the Training Change Each Person's Score?
- In your spreadsheet, click the Sortia icon in the strip of icons down the right-hand edge. No strip? Click the arrow at the bottom-right to open it. You can also use Extensions, then Sortia, then Open Sortia.
- Click Start from a template and put that name in the search box.
- Pick the card with that name and click Load this template. It arrives on a new tab with real numbers already in it.
The answer
- The verdict
- p = 0.0069 the workshop genuinely moved the ratings
- Typical move
- +2 points median shift on a ten-point observer scale
- Pairs the test can use
- 10 of 12 two assessors scored identically and are dropped
- The statistic
- W = 1.5 z of -2.70, by the normal approximation
Twelve assessors rated on a competency scale before a calibration workshop and again after it. Click Run: the statistic comes back at 1.5 with a z of -2.70 and a two-tail p-value of 0.0069, and the median difference is minus 2 points, which is the before rating sitting two below the after one. So the workshop moved people, and on a ten-point observer scale a two-point median shift is a large move.
The number to look at, and the one that trips people up, is the count of non-zero pairs: 10, not 12. Wilcoxon works on the differences, and two assessors scored identically both times, so their differences are zero and there is no direction to rank. The test drops them. That is the correct behavior and it has a consequence worth understanding: a study where most people do not change is a study with a much smaller effective sample than its headcount suggests, and a run where eight of twelve are unchanged can fail to find an effect that is plainly there in the four who moved.
Count your zeroes before you interpret your p-value, and the Pairs the test actually uses line beside the data does it for you. Notice also that the report says the p-value came from the normal approximation rather than from an exact count. That is because several of the differences are the same size, and ties rule out enumerating every arrangement of signs.
At this size the approximation is reliable, and the report tells you which one it used rather than leaving you to assume. This is a rank test on purpose. The gap between a 7 and an 8 on an agreed observer scale is not the same quantity as the gap between a 3 and a 4, so summing those differences and dividing would be arithmetic on labels.
Run the same two ranges through t-Test: Paired if you want to see what that assumption buys you: it returns a p-value of 0.0015, comfortably smaller, on an assumption this scale does not support. Smaller is not better when it was bought by pretending your scale is a ruler. What the test cannot tell you is whether the ratings after the workshop are more accurate or merely more generous, which is a question about moderation rather than about statistics.
To use your own cohort, put the before ratings in one column and the after ratings in the next, one row per person, and widen both ranges.
The model
It arrives on a tab called Template: Did the Training Move People, carrying these columns:
- Assessor
- Competency rating before (1-10)
- Competency rating after (1-10)
- Change (rating points)
- =COUNTIF(D2:D13,">0")
with the model computed beside the data:
| Unchanged | 2 |
| Went backwards | 1 |
| Pairs the test actually uses | 10 |
Once it is in your sheet
- The model arrives with real numbers in it and runs as it stands, so you can press the button first and understand it second.
- Change the numbers to yours. The sheet marks which cells are inputs and which hold formulas, and most labels carry a note explaining the row.
- Press the run button at the bottom of the panel. It is labeled for the tool you are in, and the result lands on its own tab, with a written reading of it beside the figures.
Never used Google Sheets? Start here goes the whole way, in seven steps, and assumes nothing.
Next question
- How many teachers will next summer actually need?How Many Students Next Term?
- Build the course yourself or license one?Is the New Course Worth Building?
- Does a small cohort need fewer replies?How Many Responses for a Course Survey?
- How often does the cohort miss its budget?What If the Cohort Does Not Fill?
- Two tasks both have slack. Can you spend it twice?When Will the Course Be Ready?
- What are the odds you land an A?What Grade Will You Get?
Every model like this one, and the method behind them: Statistics in Google Sheets.