Latency is the total time a request takes from send to complete response. It combines network time, provider queueing, and generation time, which grows with output length. Track it as percentiles (P50/P95/P99), not averages, because tail latency is what users actually feel.
Why it matters
Latency is a product feature. Long, variable response times erode trust even when answers are correct, and tail latency is what users actually remember.
How it works
Total latency is network time plus provider queueing plus generation time, which grows with output length. Look at percentiles rather than averages, and use streaming to improve the perceived result.