The AI glossary
Clear definitions of the terms behind modern AI systems and the Infere platform - from tokens and context windows to routing, evaluation, observability, and cost control.
80 terms · 9 categories
A
Annotation
The act of humans labeling model outputs - pass/fail, quality score, or category - for evaluation data.
Read → PlatformAPI Token
The credential your application uses to call Infere, carrying its own routing, budgets, and security settings.
Read → Cost & BillingAuto-Recharge
An automatic top-up that buys credits when the balance falls below a set threshold.
Read → Gateway & RoutingAuto-Router
Infere's AI-powered router that analyzes each prompt and picks the optimal model for cost, speed, and quality.
Read →C
Chain-of-Thought (CoT)
A prompting technique that asks the model to reason step by step before giving a final answer.
Read → RAG & DataChunking
Splitting source documents into smaller passages so retrieval can fetch just the relevant parts.
Read → Gateway & RoutingCircuit Breaker
A guard that stops sending traffic to a failing provider for a cooldown period instead of retrying endlessly.
Read → Safety & SecurityContent Moderation
Filtering or flagging harmful, abusive, or disallowed content in inputs and outputs.
Read → Cost & BillingContext Compression
Reducing the number of input tokens sent to a model by trimming or summarizing context.
Read → FundamentalsContext Window
The maximum number of tokens (input plus output) a model can consider in a single request.
Read → ObservabilityCost Attribution
Mapping LLM spend to the team, feature, user, or model that caused it.
Read → Cost & BillingCost Optimization
The tactics that lower LLM spend: routing, caching, compression, output limits, batching, and budgets.
Read → Cost & BillingCredit Balance
The remaining prepaid funds available for requests, tracked to high precision.
Read → Cost & BillingCredits
The prepaid balance Infere draws down as you use models; credits never expire.
Read →D
Data Retention
How long request and prompt data is stored, and how it is deleted after that.
Read → EvaluationDataset
A curated collection of inputs (and often expected outputs) used to run evaluations.
Read → ObservabilityDrift
A gradual change in model behavior, quality, or cost over time, even when the code has not changed.
Read →E
Embedding
A numeric vector that represents the meaning of text, used for search, retrieval, and clustering.
Read → ObservabilityError Rate
The share of requests that fail, by provider, model, or error type.
Read → EvaluationEvaluator
An automated check that scores a model or prompt output against defined criteria.
Read →F
Fallback Chain
An ordered list of backup models the gateway retries with if the primary model fails.
Read → PromptingFew-Shot & Zero-Shot Prompting
Guiding a model with a few worked examples (few-shot) or none (zero-shot).
Read → FundamentalsFine-Tuning
Further training a base model on your own examples to specialize its behavior and output style.
Read → FundamentalsFrontier Model
The newest, most capable - and usually most expensive - model from a provider, reserved for the hardest reasoning tasks.
Read →G
Golden Dataset
A trusted, human-verified dataset used as the reference for evaluating changes.
Read → EvaluationGround Truth
The correct or expected answer for a task, used as the reference when scoring model output.
Read → RAG & DataGroundedness
How well a model's answer is supported by the context it was given, rather than invented.
Read → Safety & SecurityGuardrails
Rules and checks that constrain what goes into and comes out of a model.
Read →H
L
Large Language Model (LLM)
A neural network trained on large text corpora to predict the next token and, in turn, generate, summarize, translate, and reason over language.
Read → Gateway & RoutingLatency
The total time a request takes from send to complete response.
Read → Gateway & RoutingLLM Gateway
A proxy that sits in front of multiple model providers and exposes one API, one key, and one bill.
Read → EvaluationLLM-as-a-Judge
Using one model to grade the output of another against a rubric, at scale and at low cost.
Read →M
Model
A specific trained artifact identified by an ID, such as claude-sonnet-5 or gpt-5.6-terra, that you select per request.
Read → PlatformModel Registry
Infere's live, read-only catalog of every available model with capabilities, context, and per-token pricing.
Read → Gateway & RoutingModel Router
Logic that selects which model handles each request based on cost, capability, and quality.
Read → ObservabilityMonitoring
Continuously watching production metrics - latency, errors, cost, quality - and alerting when they move.
Read → PlatformMulti-Provider Fallback
Serving one request across several providers so another picks up when one fails.
Read → FundamentalsMultimodal
A model that can accept or generate more than text - typically images, and sometimes audio or video.
Read →O
Observability
Understanding what a system did and why, from its logs, traces, and metrics - in production.
Read → EvaluationOffline vs Online Evaluation
Offline evaluation runs before deploy on a fixed dataset; online evaluation scores real production traffic.
Read → FundamentalsOpen-Weight Model
A model whose weights are publicly downloadable, so it can be self-hosted or served by third parties.
Read → Gateway & RoutingOpenAI-Compatible API
An API that matches the OpenAI request/response shape, so existing SDKs work with only a base URL change.
Read → PlatformOrganization
A shared account above workspaces, with members, roles, and a common credit balance.
Read →P
Pass-Through Pricing
Charging exactly the provider's per-token rate with no gateway markup.
Read → EvaluationPass@k
The share of tasks solved at least once across k attempts - a measure of how often a model can succeed.
Read → ObservabilityPercentiles (P50 / P95 / P99)
Latency or cost at given percentiles of traffic; tail percentiles reveal the experience users actually hit.
Read → Safety & SecurityPII (Personally Identifiable Information)
Data that identifies a person - names, emails, IDs - which may need redaction before reaching a model.
Read → PromptingPML (Prompt Markup Language)
Infere's file format for prompts and evaluators, with model config and variables in frontmatter.
Read → Cost & BillingPrepaid Billing
Paying for usage in advance by funding a balance, instead of being invoiced after the fact.
Read → PromptingPrompt
The input you send to a model - system instructions, user messages, and any context or examples.
Read → PromptingPrompt Caching
A provider feature that reuses a repeated prompt prefix and bills cache hits at a fraction of the input rate.
Read → PromptingPrompt Enhancement
Infere feature that uses AI to rewrite a prompt for a specific target model to improve output quality.
Read → Safety & SecurityPrompt Injection
An attack where untrusted input overrides a model's instructions or leaks its system prompt.
Read → PromptingPrompt Repository
A Git-backed repository that stores prompts and PML files as versioned source code.
Read → PromptingPrompt Template
A reusable prompt with placeholders that are filled in at runtime for each request.
Read → ObservabilityProperty
A custom tag attached to a request (for example user, feature, or environment) used to slice analytics.
Read → FundamentalsProvider
The vendor that serves a model. Infere routes across OpenAI, Anthropic, Fireworks AI, Together AI, and Infercom.
Read →R
RAG (Retrieval-Augmented Generation)
Feeding a model relevant retrieved context so it answers from your data instead of its memory alone.
Read → Gateway & RoutingRate Limit
A cap on how many requests or tokens can be sent in a time window.
Read → ObservabilityRequest Ledger
A searchable record of every request that passes through the gateway, with model, tokens, cost, and latency.
Read → RAG & DataRetrieval
Fetching the passages most relevant to a query, usually by vector similarity, before generation.
Read → EvaluationRubric
The written criteria and scale an evaluator - human or LLM - uses to score output.
Read →S
Semantic Search
Search ranked by meaning (via embeddings) rather than exact keyword matches.
Read → ObservabilitySpan
A single unit of work within a trace, such as one model call or one tool invocation.
Read → Cost & BillingSpend Cap
A hard upper bound on spending over a period, which blocks further requests once reached.
Read → Gateway & RoutingStreaming
Delivering a model's response token by token as it is generated, instead of waiting for the full output.
Read → PromptingSystem Prompt
The high-level instructions that set a model's role, rules, and tone before the user's messages.
Read →T
Temperature
A parameter that controls randomness: lower is more deterministic, higher is more varied and creative.
Read → Gateway & RoutingTime-to-First-Token (TTFT)
How long until the first token arrives; the latency users perceive most in streaming chat.
Read → FundamentalsToken
The unit of text a model reads and writes; roughly three-quarters of an English word. Usage and cost are measured in tokens.
Read → Cost & BillingToken Budget
A spend or usage cap attached to an API token, with optional daily, weekly, or monthly windows.
Read → ObservabilityTrace
A group of related requests that together form one workflow, such as an agent loop or RAG pipeline.
Read →Put the platform behind the terms
Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.