Error rate is the share of requests that fail - spiking on provider outages, rate limits, or bad requests. Track it by provider and error type so you can tell an upstream outage from a client bug, and pair it with circuit breakers and fallback chains to contain the damage.
Why it matters
For LLM apps, an error might be a failed request or a confidently wrong answer. Tracking failures by provider and type separates client bugs from upstream outages.
How it works
Error rate is the share of requests that fail, sliced by provider, model, and status. Combined with circuit breakers and fallbacks, a spike in one provider's errors is contained instead of becoming a user-visible outage.