The context window is the maximum number of tokens a model can consider in one request, counting the prompt and the response together. Current models range from a few thousand to over a million tokens. Prompts that exceed the window must be trimmed, compressed, or split - and some providers raise rates once a prompt crosses a context threshold.
Why it matters
The window caps how much a model can "see" at once, so it decides whether a task is even possible in a single call. It also sets a hard cost boundary: every input token in the window is billed, even if the model ignores it.
How it works
Input tokens plus the requested output tokens must fit inside the window. When they do not, you trim history, compress context, or split the work. Some providers also raise the per-token rate once a prompt crosses a threshold - for example 200K or 272K tokens.