Offline evaluation runs before deploy against a frozen dataset, so results are reproducible and can gate a release. Online evaluation scores sampled production traffic, so it catches drift and real-world edge cases a test set misses. Mature teams run both: offline to gate, online to monitor.
Why it matters
Offline evaluation catches regressions before users do; online evaluation catches problems your test set never imagined. Neither alone is sufficient.
How it works
Offline runs before deploy on a frozen dataset for reproducible, gating results. Online scores sampled production traffic to detect drift and real-world edge cases. Mature teams run both and treat online failures as new offline test cases.