Observability is the ability to understand what a system did and why, using its logs, traces, and metrics. For LLM applications that means capturing prompts, model, tokens, cost, latency, and quality for every request - and connecting them to the prompt version and evaluator that produced the result.

Why it matters

LLM behavior is probabilistic and multi-step, so traditional uptime monitoring is not enough. You need to see the prompt, model, tokens, cost, latency, and quality for every request.

How it works

Requests are logged to a ledger, grouped into traces, and tagged with properties so you can slice them. Because the gateway sees every call, this data is captured without instrumenting each service.

Example

When a user reports a bad answer, you open the request, see the exact prompt and model, the token cost, and the evaluator score that produced it.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features