Getting Started
Infere is an intelligent AI routing and management platform: one OpenAI-compatible API across leading AI providers, with automatic model selection, Git-native prompt management, AI evaluation, observability, and enterprise controls. It's built to solve cost, reliability, and security together - reducing AI spend via AI-powered routing.
1. Create an Account
Sign up at app.infere.com/signup to get your account. Registration is confirmed with a 6-digit code sent by email, and a personal workspace is created automatically. Infere uses a prepaid balance with no Infere markup - add balance when you're ready to make requests.
2. Generate an API Token
Navigate to Tokens in the sidebar and create a new token - either a Quick Dev Token for instant local testing or the Advanced Wizard to configure routing, budgets, rate limits, and security scanning. Your token secret is displayed once - copy it and store it securely.
# Your token follows this format:
sk-toj-{shard}-{signature}
3. Make Your First Request
Use any HTTP client to send a request. Infere exposes an OpenAI-compatible API at
https://api.infere.com/v1, so switching from OpenAI requires just a URL and token
change.
curl -X POST https://api.infere.com/v1/chat/completions \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "Hello, world!"}
]
}'
Name a specific model like "gpt-4o" or "claude-sonnet-4-5", or enable
Auto-Routing on your token and let Infere select the model per request based on
prompt complexity and your routing mode (balanced, cost-optimized, or quality-optimized).
4. OpenAI SDK Compatibility
import OpenAI from 'openai';
const client = new OpenAI({
apiKey: INFERE_API_KEY,
baseURL: 'https://api.infere.com/v1',
});
const response = await client.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Summarize this document.' }],
});
5. Monitor Your Request
Open Requests in the sidebar (Observe & Improve) to inspect the logged call - latency, token usage, cost, and the full request/response. Costs are settled against your credit balance in real time, request by request.