A model is a specific trained artifact identified by an ID, such as claude-sonnet-5 or gpt-5.6-terra. Each model has its own rates, context window, capabilities (tools, vision, JSON mode), and speed/quality profile. The Infere Model Registry lists every available model with live pricing so you can pick the cheapest model that meets your quality bar.
Why it matters
Each model has its own price, context window, capabilities, and quality profile. The model ID you send is effectively the single most important cost lever in a request.
In Infere
The Model Registry lists every available model with live per-1M-token pricing, context window, max output, and capability tags. You select a model by ID in the request, or let the Auto-Router pick one per request.