The Infere API is REST-based and supports all major AI operations. It is fully compatible with the OpenAI API specification, making migration seamless.

Base URL

https://api.infere.com/v1

Authentication

All requests require a Bearer token in the Authorization header. Tokens carry their own configuration - model routing, budgets, rate limits, IP allowlists, and optional LLM security scanning - and can be created and managed from the Tokens page.

Endpoints

Method Endpoint Description
POST /v1/chat/completions Chat completions with intelligent routing (streaming supported)
POST /v1/embeddings Text embeddings generation
POST /v1/multimodalembeddings Multimodal embeddings
POST /v1/audio/transcriptions Speech-to-text (multipart)
POST /v1/audio/speech Text-to-speech generation
GET /v1/models List available models (add ?mode=all for non-chat models)

Error Responses

Errors use the OpenAI error envelope with an Infere code in the code field:

{ "error": { "message": "...", "type": "authentication_error", "code": "INFERE_ERR_1005_TOKEN_NOT_FOUND" } }
  • 401 - missing, unknown, or inactive token
  • 402 - no credit remaining or token budget exceeded
  • 403 - capability not enabled, model excluded, or IP not allowed
  • 429 - rate or concurrency limit hit (respect the Retry-After header)

Token status, credit status, budgets, IP allowlists, rate limits, and allowed models are all enforced at the edge before a request ever reaches a model provider.