Regression to the Mean: Why Your Best Week Is Followed by a Worse One
You praised the top performers and they got worse. You coached the worst and they improved. Before concluding anything about praise or coaching, meet regression to the mean.
A sales manager pilots a coaching program with the bottom 10% of reps. Next quarter, every one of them improves. The program is declared a triumph and rolled out company-wide — where it does nothing. What happened? The bottom 10% were selected at an extreme, and extremes fade on their own. That fade has a name: regression to the mean.
Any metric with a luck component behaves this way: an unusually good week is usually followed by a merely good one, and an unusually bad week by a merely bad one. No cause, no intervention, no lesson — just the noise component of the metric returning to its average contribution of zero.
Why It Must Happen
Think of each observation as skill + luck. When you select the top performers, you select people with high skill and people who drew good luck. Next period, skill persists but luck redraws around zero — so the group's average falls even though nobody got worse. Selecting the bottom works in reverse: bad luck does not repeat itself, so the group improves without any intervention.
The strength of the pull depends on the noise share. A metric that is 90% stable skill barely regresses; a metric that is 90% weekly noise snaps back almost fully. This is why regression ambushes exactly the metrics managers watch most closely: short-window numbers (weekly sales, daily conversions) are noise-dominated, while the stable annual numbers nobody checks would barely move.
The Three Classic Traps
- Fake treatment effects. Any intervention applied to an extreme-selected group "works" by construction — the control group you forgot to include would have improved too. Before/after comparisons on extremes are evidence of nothing.
- Punishing variance. A rep who swings between great and terrible weeks looks worse than a steady mediocre rep if you react to each extreme. The swings are the metric's noise, not the rep's character.
- The sophomore slump. Rookie of the year underperforms in year two; a record quarter is followed by a "disappointing" one. Some of the slump is real (adjustment, pressure), but much of it is pure regression wearing a narrative costume.
How to Control for It
The gold standard is a control group: select extremes, then treat only half. The untreated half shows you the regression baseline, and only improvement beyond that baseline counts as a treatment effect. Without a control, a weaker but useful habit: select on one period, measure on a later one. Pick the bottom 10% by Q1, but evaluate the program on Q3 vs Q2 — skipping the selection window removes the mechanical bounce (most of it, anyway).
And when a spike demands explanation, check the anomaly playbook before the narrative one: is this a level shift or a spike? Spikes regress; shifts persist. Treating a spike as a shift is how teams "fix" problems that already fixed themselves — and how forecasts built on an extreme week (see Holt's method with a reactive level) overshoot the month after.