Annotation is humans labeling model outputs, typically as pass/fail, a quality score, or a category. Annotations feed golden datasets, calibrate LLM judges, and measure inter-annotator agreement. They are the ground-truth substrate that makes automated evaluation trustworthy.

Why it matters

Annotations are the raw material of trustworthy evaluation. Judges and gates are only as reliable as the human labels they were calibrated against.

How it works

Humans label outputs - pass/fail, a quality score, or a category - often against a shared rubric. Measuring agreement between annotators (IAA) tells you whether the task itself is well-defined.

Example

Two reviewers independently label 50 replies pass/fail; their agreement rate reveals whether the rubric is clear enough to trust.

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features