An API token is the credential your application uses to call Infere. Each token carries its own configuration - allowed models, routing and auto-router settings, budgets, rate limits, and IP allowlists - so a single workspace can have many tokens with different behavior. Tokens are shown once at creation and stored hashed.

Why it matters

Tokens are how you grant access with least privilege. Putting routing, budgets, and limits on the token - not just the account - is what makes per-app control possible.

How it works

Each token carries its own configuration: allowed models, Auto-Router and routing mode, budgets, rate limits, and IP allowlists. The secret is shown once at creation and stored hashed.

Example

A frontend token gets a tight budget and no access to expensive models; a backend token gets production models and a higher cap.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features