An evaluator scores model output automatically against defined criteria. Evaluators can be LLM judges, deterministic code checks, or a mix. Running a set of evaluators over a dataset turns "it looks fine" into a measurable pass rate you can gate releases on.

Why it matters

Evaluation turns subjective "it looks fine" into a measurable number you can compare across runs and gate releases on. Without it, quality changes ship silently.

How it works

An evaluator takes an input and the model output and returns a score or pass/fail. Evaluators can be LLM judges, deterministic code checks, or a mix, and are run over a dataset to produce comparable results.

Example

Combine an exact-match check for JSON validity with an LLM judge for tone, and report a single pass rate per run.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features