A frontier model is the newest and most capable model a provider offers. It usually carries the highest per-token price and the strongest reasoning, but most production traffic does not need it. Routing easier requests to a cheaper mid-tier model and reserving the frontier tier for hard cases is the single largest cost lever.

Why it matters

Frontier models often cost 10-50x more per token than entry-tier models. Using one for every request is the most common and most expensive mistake in production LLM applications.

How it works

Capability tiers are stable even as model names churn: an entry model handles classification and extraction, a mid model handles most generation, and a frontier model is reserved for hard reasoning. A router sends each request to the lowest tier that clears your quality bar.

Example

Route 90% of traffic to a $0.20/1M model and 10% to a $5/1M model, and the blended input rate lands near $0.68/1M instead of $5/1M.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features