Prompt caching lets a provider store a repeated prefix - a system prompt, tool definitions, or a document - and charge a cache read at roughly one-tenth of the normal input rate. Put stable content at the front of the request and measure the hit rate, because a silent cache miss re-bills the full input price.

Why it matters

Repeated context - system prompts, tool definitions, long documents - is often the bulk of input cost. Caching turns that repeated cost into a fraction of the base rate.

How it works

The provider stores a prefix and charges a cache read at roughly one-tenth of normal input on later requests. Put stable content first, keep prompts deterministic, and watch the hit rate - a silent cache miss re-bills the full input price.

Example

A 20,000-token system prompt cached once can cost about 90% less on every subsequent request that reuses it.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features