A rubric is the set of criteria and the scale used to score an output, shared by human reviewers and LLM judges. A precise rubric is what makes scores comparable across runs and models; a vague rubric produces noisy scores no one can act on.

Why it matters

Scores are only comparable if everyone - human or model - is applying the same criteria. The rubric is what makes evaluation results actionable rather than noisy.

How it works

A rubric defines the dimensions to score and the scale for each. It is shared between human reviewers and LLM judges so their outputs can be compared and the judges calibrated against people.

Example

A 1-5 scale where 5 is "fully correct and well-cited", 3 is "correct but unsupported", and 1 is "hallucinated".

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features