Monitoring is continuously watching production - latency, error rate, cost, throughput, and quality - with alerts when something moves. For LLMs it also means watching quality drift and spend, not just uptime, since a model can be "up" and still producing worse or pricier answers.

Why it matters

Model quality and cost change without any code change, because providers update models and input traffic shifts. Continuous monitoring is how you notice.

How it works

Production metrics - latency, errors, cost, throughput, and quality scores - are watched continuously with alerts when they move. For LLMs this includes spend and quality, not just uptime.

Example

Alert when daily spend rises 30% above its baseline or when judge scores on sampled traffic trend downward.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features