An evaluator scores model output automatically against defined criteria. Evaluators can be LLM judges, deterministic code checks, or a mix. Running a set of evaluators over a dataset turns "it looks fine" into a measurable pass rate you can gate releases on.
Why it matters
Evaluation turns subjective "it looks fine" into a measurable number you can compare across runs and gate releases on. Without it, quality changes ship silently.
How it works
An evaluator takes an input and the model output and returns a score or pass/fail. Evaluators can be LLM judges, deterministic code checks, or a mix, and are run over a dataset to produce comparable results.