Latency is the total time a request takes from send to complete response. It combines network time, provider queueing, and generation time, which grows with output length. Track it as percentiles (P50/P95/P99), not averages, because tail latency is what users actually feel.

Why it matters

Latency is a product feature. Long, variable response times erode trust even when answers are correct, and tail latency is what users actually remember.

How it works

Total latency is network time plus provider queueing plus generation time, which grows with output length. Look at percentiles rather than averages, and use streaming to improve the perceived result.

Example

A model with 800ms median but 6s P99 feels slow in production even though the average looks fine.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features