Streaming sends a model's response back token by token as it is produced, so users see output immediately rather than waiting for the whole answer. It improves perceived latency without changing the total cost, and is the default for chat interfaces.

Why it matters

Waiting for a full answer feels slow even when total latency is unchanged. Streaming makes the same request feel far more responsive.

How it works

The provider sends tokens as they are generated, usually as server-sent events. Time-to-first-token becomes the latency users perceive, while total latency and cost stay the same.

Example

A chat reply appears word by word; the first token arrives in a few hundred milliseconds instead of after the full multi-second response.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features