Annotation is humans labeling model outputs, typically as pass/fail, a quality score, or a category. Annotations feed golden datasets, calibrate LLM judges, and measure inter-annotator agreement. They are the ground-truth substrate that makes automated evaluation trustworthy.
Why it matters
Annotations are the raw material of trustworthy evaluation. Judges and gates are only as reliable as the human labels they were calibrated against.
How it works
Humans label outputs - pass/fail, a quality score, or a category - often against a shared rubric. Measuring agreement between annotators (IAA) tells you whether the task itself is well-defined.