The context window is the maximum number of tokens a model can consider in one request, counting the prompt and the response together. Current models range from a few thousand to over a million tokens. Prompts that exceed the window must be trimmed, compressed, or split - and some providers raise rates once a prompt crosses a context threshold.

Why it matters

The window caps how much a model can "see" at once, so it decides whether a task is even possible in a single call. It also sets a hard cost boundary: every input token in the window is billed, even if the model ignores it.

How it works

Input tokens plus the requested output tokens must fit inside the window. When they do not, you trim history, compress context, or split the work. Some providers also raise the per-token rate once a prompt crosses a threshold - for example 200K or 272K tokens.

Example

A 1M-token model can read a 750,000-word book in one request, but you pay input rates for all of it - so trimming to the relevant chapters is often cheaper and more accurate than sending the whole thing.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features