A token is the unit of text an LLM reads and writes. It is roughly three-quarters of an English word, so a million tokens is about 750,000 words. Providers bill per token, split into input (your prompt) and output (the response), which is why token counts - not requests - are the real unit of LLM cost.

Why it matters

Tokens are the unit everything is priced in. Input and output tokens are billed separately, output usually costs several times more, and every prompt, cached prefix, and generated word adds to the count.

How it works

Text is broken into tokens by the model's tokenizer - common words are one token, rare words several, punctuation and spaces their own. Providers report input, output, and cached token counts per request, and Infere charges against your balance using those counts.

Example

The word "unbelievable" is typically 2-3 tokens, and a 1,000-word document is roughly 1,300 tokens. When a rate reads $5/1M input, that is about $5 per 750,000 words of prompt.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features