p50, p95, p99: Percentiles for Monitoring, SLAs and Honest Latency Talk
Average latency is 120ms — and 5% of users wait over 3 seconds. Averages hide the suffering tail; percentiles price it. How to compute, plot, and SLA them.
Average latency is 120ms. Sounds fine — until you learn that 5% of users wait over 3 seconds. The average did not lie; it averaged away exactly the users who are suffering. Latency, like revenue and deal size, is heavily skewed, which is why the mean-vs-median warning applies doubly to performance: the mean describes the infrastructure's total work, but percentiles describe what users actually experience.
Reading p50, p95, p99
The p95 latency is the threshold under which 95% of requests fall — equivalently, 1 request in 20 is slower. Each percentile answers a different stakeholder's question:
- p50 (the median): the typical experience. Track it for trends — a rising p50 means the common case is degrading, which usually indicates systemic load, not edge cases.
- p95: the "almost everyone" experience. The workhorse SLA percentile: sensitive enough to catch real regressions, stable enough to alert on without constant noise.
- p99 (and p99.9): the tail. Dominated by garbage collection, retries, cold starts, and cross-AZ hops — infrastructure realities, not application logic. High-volume services watch p99.9 because at scale the tail is a large population.
Report them as a set, never alone: "p50 80ms / p95 400ms / p99 1.8s" tells the full story in one line. A p99 far above p95 (here 4.5×) points at a distinct tail cause worth hunting separately from median performance.
Computing Percentiles on Streams
Exact percentiles need all the data (sort it, pick the rank), which is fine for offline analysis but impossible for high-volume live monitoring. The standard answer is a sketch: HDR Histogram and DDSketch compress a distribution into kilobytes with bounded relative error, merge across servers, and answer any percentile query. If your metrics pipeline offers histogram-type metrics (Prometheus histograms, StatsD timers with backends that keep distributions), use them — pre-aggregated averages from agents cannot be recombined into correct global percentiles later.
Two computation traps: averaging percentiles ("average p95 across servers") is mathematically meaningless — merge the underlying distributions, then take the percentile. And sampling distorts tails first: a 1% sample estimates p50 fine but may never observe the p99.9 events at all. Sample the body if you must, but capture tail events exhaustively.
Writing SLAs Percentiles Can Verify
"Average response under 200ms" is an SLA the tail can violate for years undetected. A percentile SLA — "p95 under 500ms measured over 5-minute windows, 99.9% of windows in a month" — names the experience, the window, and the compliance fraction, and each part is checkable. Pair the SLA with burn-rate alerting (alert on the rate of budget consumption, fast-burn and slow-burn windows) rather than on raw threshold crossings, and the on-call rotation stops drowning in tail noise.
Percentiles also discipline how you communicate performance upward: leadership summaries built on p50/p95/p99 (see writing summaries people read) survive contact with reality far better than average-based ones. The tail is where your worst user experiences and your most interesting engineering problems both live — measure it, plot it, and put it in the SLA.