A rubric is the set of criteria and the scale used to score an output, shared by human reviewers and LLM judges. A precise rubric is what makes scores comparable across runs and models; a vague rubric produces noisy scores no one can act on.
Why it matters
Scores are only comparable if everyone - human or model - is applying the same criteria. The rubric is what makes evaluation results actionable rather than noisy.
How it works
A rubric defines the dimensions to score and the scale for each. It is shared between human reviewers and LLM judges so their outputs can be compared and the judges calibrated against people.