Systems & Backend

p99 Latency

also: tail latency · percentile latency

The response time your worst 1% of requests exceed — the tail users actually feel.

p99 latency is the 99th-percentile response time: 99% of requests are faster, 1% slower. Averages hide this tail, so SLOs are written against p95/p99. Because users touch many services per action, one service’s tail becomes the common experience.

Worked example: a page that makes 10 backend calls, each with a 1% chance of hitting its p99 slow path, has a 1 − 0.99¹⁰ ≈ 9.6% chance that at least one call is slow — so a per-service p99 becomes roughly a per-page p90. Tail latency amplifies under fan-out. Gotcha: you cannot average percentiles across services (p99 of a sum ≠ sum of p99s), and a healthy-looking mean can hide an ugly p99 — always optimize the tail users feel, not the average.