Offline evaluation runs before deploy against a frozen dataset, so results are reproducible and can gate a release. Online evaluation scores sampled production traffic, so it catches drift and real-world edge cases a test set misses. Mature teams run both: offline to gate, online to monitor.

Why it matters

Offline evaluation catches regressions before users do; online evaluation catches problems your test set never imagined. Neither alone is sufficient.

How it works

Offline runs before deploy on a frozen dataset for reproducible, gating results. Online scores sampled production traffic to detect drift and real-world edge cases. Mature teams run both and treat online failures as new offline test cases.

Example

Gate the deploy with an offline golden-set run, then sample 5% of live traffic through judges to catch what the test set missed.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features