Infere is an intelligent AI routing and management platform: one OpenAI-compatible API across leading AI providers, with automatic model selection, Git-native prompt management, AI evaluation, observability, and enterprise controls. It's built to solve cost, reliability, and security together - reducing AI spend via AI-powered routing.

Welcome! You can start sending AI requests through Infere in under 5 minutes. Follow this quickstart guide to get your first response.

1. Create an Account

Sign up at app.infere.com/signup to get your account. Registration is confirmed with a 6-digit code sent by email, and a personal workspace is created automatically. Infere uses a prepaid balance with no Infere markup - add balance when you're ready to make requests.

2. Generate an API Token

Navigate to Tokens in the sidebar and create a new token - either a Quick Dev Token for instant local testing or the Advanced Wizard to configure routing, budgets, rate limits, and security scanning. Your token secret is displayed once - copy it and store it securely.

# Your token follows this format:

sk-toj-{shard}-{signature}
Security Note: Only a hash of the token secret is stored - the plaintext is only shown at creation. If you lose your token, you must generate a new one. Set an expiration and rotate tokens regularly.

3. Make Your First Request

Use any HTTP client to send a request. Infere exposes an OpenAI-compatible API at https://api.infere.com/v1, so switching from OpenAI requires just a URL and token change.

curl -X POST https://api.infere.com/v1/chat/completions \

  -H "Authorization: Bearer ***" \

  -H "Content-Type: application/json" \

  -d '{

    "model": "gpt-4o",

    "messages": [

      {"role": "user", "content": "Hello, world!"}

    ]

  }'

Name a specific model like "gpt-4o" or "claude-sonnet-4-5", or enable Auto-Routing on your token and let Infere select the model per request based on prompt complexity and your routing mode (balanced, cost-optimized, or quality-optimized).

4. OpenAI SDK Compatibility

import OpenAI from 'openai';



const client = new OpenAI({

  apiKey: INFERE_API_KEY,

  baseURL: 'https://api.infere.com/v1',

});



const response = await client.chat.completions.create({

  model: 'gpt-4o',

  messages: [{ role: 'user', content: 'Summarize this document.' }],

});

5. Monitor Your Request

Open Requests in the sidebar (Observe & Improve) to inspect the logged call - latency, token usage, cost, and the full request/response. Costs are settled against your credit balance in real time, request by request.