Time-to-first-token measures how long until the first streamed token arrives. It is dominated by queueing and prompt processing, so long prompts and busy providers push it up. TTFT and total latency together describe the streaming experience better than either alone.

Why it matters

In a streaming UI, the wait before the first token is the wait users notice most. TTFT is often what makes an app feel fast or sluggish.

How it works

TTFT is dominated by queueing and prompt processing, so long prompts and busy providers push it up. Reducing prompt size and routing to fast models are the primary ways to lower it.

Example

Trimming a 12,000-token prompt to the relevant 2,000 tokens can cut time-to-first-token substantially on the same model.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features