Streaming sends a model's response back token by token as it is produced, so users see output immediately rather than waiting for the whole answer. It improves perceived latency without changing the total cost, and is the default for chat interfaces.
Why it matters
Waiting for a full answer feels slow even when total latency is unchanged. Streaming makes the same request feel far more responsive.
How it works
The provider sends tokens as they are generated, usually as server-sent events. Time-to-first-token becomes the latency users perceive, while total latency and cost stay the same.