A large language model is a neural network trained on large text corpora to predict the next token. That single objective is what lets it generate, summarize, classify, translate, and reason over language. Infere routes requests to LLMs from OpenAI, Anthropic, Fireworks AI, Together AI, and Infercom through one OpenAI-compatible API, so you can switch or combine models without rewriting code.

Why it matters

Choosing between frontier, mid-tier, and open-weight models is the core cost-and-quality decision in an AI product. A single request can cost fractions of a cent on a small model or tens of cents on a frontier model, and most traffic does not need the most expensive tier.

How it works

An LLM converts your prompt into tokens, predicts the most likely next token over and over, and returns the sequence. Capability broadly scales with parameter count and training, but so do price and latency - which is why routing and model choice matter more than raw model quality alone.

Example

A support assistant can classify a ticket with a cheap small model, draft a reply with a mid-tier model, and escalate only hard cases to a frontier model - all through the same OpenAI-compatible call.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features