A model is a specific trained artifact identified by an ID, such as claude-sonnet-5 or gpt-5.6-terra. Each model has its own rates, context window, capabilities (tools, vision, JSON mode), and speed/quality profile. The Infere Model Registry lists every available model with live pricing so you can pick the cheapest model that meets your quality bar.

Why it matters

Each model has its own price, context window, capabilities, and quality profile. The model ID you send is effectively the single most important cost lever in a request.

In Infere

The Model Registry lists every available model with live per-1M-token pricing, context window, max output, and capability tags. You select a model by ID in the request, or let the Auto-Router pick one per request.

Example

Switching a summarization task from a premium model to a mid-tier model can cut a request's cost roughly tenfold, with no code change beyond the model ID.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features