Context compression reduces input tokens by trimming or summarizing context before it reaches the model, directly cutting input cost and latency. Infere offers strategies including head/tail trimming, middle-out, and LLM-based semantic compression. Validate each strategy with evaluators so compression does not quietly degrade quality.

Why it matters

Long context is often mostly redundant. Compressing it cuts input cost and latency at the same time, unlike many optimizations that trade one for the other.

In Infere

Infere offers strategies including head/tail trimming, middle-out, and LLM-based semantic compression. Because compression can affect answer quality, validate each strategy with evaluators rather than assuming it is free.

Example

Summarize a long chat history into a compact running summary before sending it, cutting input tokens while keeping the important facts.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features