A large language model is a neural network trained on large text corpora to predict the next token. That single objective is what lets it generate, summarize, classify, translate, and reason over language. Infere routes requests to LLMs from OpenAI, Anthropic, Fireworks AI, Together AI, and Infercom through one OpenAI-compatible API, so you can switch or combine models without rewriting code.
Why it matters
Choosing between frontier, mid-tier, and open-weight models is the core cost-and-quality decision in an AI product. A single request can cost fractions of a cent on a small model or tens of cents on a frontier model, and most traffic does not need the most expensive tier.
How it works
An LLM converts your prompt into tokens, predicts the most likely next token over and over, and returns the sequence. Capability broadly scales with parameter count and training, but so do price and latency - which is why routing and model choice matter more than raw model quality alone.