Models Registry
The Model Registry is a live, read-only catalog of every AI model available through Infere - a nightly sync job pulls model data directly from provider APIs, so the catalog is never stale. Each model card shows capabilities, context window, max output tokens, and pricing per 1M tokens.
- Multi-provider - one OpenAI-compatible API across providers including OpenAI, Anthropic, Fireworks AI, Together AI, and Infercom.
- Capability tags - streaming, tool/function calling, vision, JSON mode, embeddings, image generation, audio, plus quality and speed ratings.
- Live pricing - input, output, and cached-input rates per 1M tokens; nothing is hardcoded.
- Auto-Router - enable it per token to select a model per request based on prompt analysis, required capabilities, routing mode, availability, and provider preferences.
- Fallback chains - if the primary model fails (5xx, rate limit, timeout), the gateway retries with the next model in your chain.
Open the registry from the command palette (Cmd/Ctrl+K → "Go to Models") or directly
at /models. Model IDs from the registry are used directly in API requests and PML
frontmatter.