A/B Testing on Small Traffic: Sample Size, Peeking, and When Not to Test
You ran an A/B test for a week, B won by 12%, you shipped it — and revenue did not move. Small-traffic testing fails in predictable ways. Here is how to survive it.
The most common A/B testing story on small sites goes like this: run a test for a week, variant B leads by 12%, declare victory, ship it — and watch revenue do nothing. The test did not lie; it was never capable of telling the truth. Low traffic does not forbid experimentation. It forbids casual experimentation, where every shortcut compounds into a false positive.
The Sample-Size Math Nobody Wants to Hear
Detecting a 10% relative lift on a 5% baseline conversion rate at standard power needs roughly 30,000 visitors per variant. At 1,000 visitors a week, that is a 60-week test. The math is brutal because small effects hide inside binomial noise: with 500 visitors per side, a "12% lift" is well within the range of pure chance.
Three honest responses exist. Test bigger changes — redesigns and offer changes move conversion by 30–100%, which small samples can detect. Test further down the funnel — checkout completion has a higher baseline rate than homepage signup, so the same visitor count buys more power. Or don't test: ship the obviously-better variant and spend the six months on the next improvement. Running an underpowered test and believing it is worse than running none, because it manufactures false certainty.
Peeking: The Silent p-Value Killer
Checking results daily and stopping when significance appears — "peeking" — inflates the false-positive rate from 5% to 20–30%. Each peek is another lottery ticket: with enough looks, noise eventually crosses the threshold. The fix is procedural, not mathematical: fix the sample size (or runtime) in advance and do not look until it is reached. If you must monitor, monitor only for breakage (errors, zero conversions), never for winners.
Sequential methods exist that allow peeking with corrected thresholds, but they still need the same order of data — there is no statistical trick that conjures power from 200 visitors. When someone promises "valid peeking," check what the method costs in extra runtime; it is never free.
When Randomization Is Off the Table
Some changes cannot be A/B tested at all: pricing overhauls, rebrands, features with network effects. The fallback ladder, from strongest to weakest evidence:
- Before/after with a holdout. Change one region or segment, keep another as control. Imperfect, but it separates your change from seasonality and campaigns.
- Interrupted time series. Model the baseline trend (a simple trend model works), then measure the break at launch. Credible only with a long, stable baseline.
- Qualitative plus guardrails. Ship with session recordings and a rollback metric. "Users complete the flow without new errors and support tickets don't spike" is weak evidence of improvement but strong evidence of safety.
And whatever you measure, measure it like a forecaster: define the success metric and its error tolerance before launch. A test whose metric was chosen after seeing the data is not an experiment — it is a story with numbers in it.