An LLM gateway is a proxy between your application and multiple model providers. It normalizes provider formats behind one OpenAI-compatible API, centralizes keys and billing, and adds routing, fallback, caching, budgets, and logging. An API gateway routes service traffic; an LLM gateway understands models, tokens, and cost.

Why it matters

Without a gateway, every new provider means another integration, another key, another bill, and another place reliability can break. A gateway centralizes all of it behind one endpoint.

How it works

Requests hit the gateway, which normalizes them, applies routing, caching, fallback, and budgets, forwards to a provider, then logs the result. Because it sits in the path of every call, it is also the natural place for observability and cost control.

Example

Point your existing OpenAI SDK at the Infere base URL and keep the same code; the gateway handles provider translation, routing, and logging behind it.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features