API Reference
The Infere API is REST-based and supports all major AI operations. It is fully compatible with the OpenAI API specification, making migration seamless.
Base URL
https://api.infere.com/v1
Authentication
All requests require a Bearer token in the Authorization header. Tokens carry their
own configuration - model routing, budgets, rate limits, IP allowlists, and optional LLM
security scanning - and can be created and managed from the Tokens page.
Endpoints
| Method | Endpoint | Description |
|---|---|---|
POST |
/v1/chat/completions |
Chat completions with intelligent routing (streaming supported) |
POST |
/v1/embeddings |
Text embeddings generation |
POST |
/v1/multimodalembeddings |
Multimodal embeddings |
POST |
/v1/audio/transcriptions |
Speech-to-text (multipart) |
POST |
/v1/audio/speech |
Text-to-speech generation |
GET |
/v1/models |
List available models (add ?mode=all for non-chat models) |
Error Responses
Errors use the OpenAI error envelope with an Infere code in the code field:
{ "error": { "message": "...", "type": "authentication_error", "code": "INFERE_ERR_1005_TOKEN_NOT_FOUND" } }
401- missing, unknown, or inactive token402- no credit remaining or token budget exceeded403- capability not enabled, model excluded, or IP not allowed429- rate or concurrency limit hit (respect theRetry-Afterheader)
Token status, credit status, budgets, IP allowlists, rate limits, and allowed models are all enforced at the edge before a request ever reaches a model provider.