Percentiles describe a distribution: P50 is the median, P95 and P99 are the slow (or expensive) tail. LLM latency is long-tailed, so averages hide the requests users complain about. Monitor P95/P99 latency and cost, not means, to catch real regressions.
Why it matters
Averages hide the requests users complain about. LLM latency is long-tailed, so the slow tail is where the experience actually degrades.
How it works
Percentiles describe the distribution: P50 is the median, P95 and P99 the slow tail. Monitor and alert on P95/P99 latency and cost rather than means so regressions surface before they spread.