Prompt caching lets a provider store a repeated prefix - a system prompt, tool definitions, or a document - and charge a cache read at roughly one-tenth of the normal input rate. Put stable content at the front of the request and measure the hit rate, because a silent cache miss re-bills the full input price.
Why it matters
Repeated context - system prompts, tool definitions, long documents - is often the bulk of input cost. Caching turns that repeated cost into a fraction of the base rate.
How it works
The provider stores a prefix and charges a cache read at roughly one-tenth of normal input on later requests. Put stable content first, keep prompts deterministic, and watch the hit rate - a silent cache miss re-bills the full input price.