A frontier model is the newest and most capable model a provider offers. It usually carries the highest per-token price and the strongest reasoning, but most production traffic does not need it. Routing easier requests to a cheaper mid-tier model and reserving the frontier tier for hard cases is the single largest cost lever.
Why it matters
Frontier models often cost 10-50x more per token than entry-tier models. Using one for every request is the most common and most expensive mistake in production LLM applications.
How it works
Capability tiers are stable even as model names churn: an entry model handles classification and extraction, a mid model handles most generation, and a frontier model is reserved for hard reasoning. A router sends each request to the lowest tier that clears your quality bar.