A rate limit caps how many requests or tokens a key may send in a given window. Providers apply their own limits, and a gateway can apply yours per token or workspace. Rate limits protect availability and, with token budgets, prevent a runaway loop from spending without bound.

Why it matters

Rate limits protect availability and stop a runaway loop or a buggy client from exhausting capacity or budget. Providers enforce theirs; a gateway lets you enforce yours per token.

How it works

Limits cap requests or tokens per time window, per key or token. Exceeding the limit returns an error the client can back off on. Budgets and rate limits together bound both cost and load.

Example

Set a token to 100 requests per minute so a retry storm cannot flood a provider or your balance.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features