Percentiles describe a distribution: P50 is the median, P95 and P99 are the slow (or expensive) tail. LLM latency is long-tailed, so averages hide the requests users complain about. Monitor P95/P99 latency and cost, not means, to catch real regressions.

Why it matters

Averages hide the requests users complain about. LLM latency is long-tailed, so the slow tail is where the experience actually degrades.

How it works

Percentiles describe the distribution: P50 is the median, P95 and P99 the slow tail. Monitor and alert on P95/P99 latency and cost rather than means so regressions surface before they spread.

Example

If P50 is 700ms but P99 is 9s, one in a hundred users waits several seconds - invisible in the average.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features