LLM-as-a-judge uses a model to score another model's output against a rubric. It scales human-style judgment across large datasets cheaply, but judges have known biases (verbosity, position, self-preference). Calibrate judges against human labels and watch for drift before trusting their scores as gates.

Why it matters

Human review does not scale to thousands of examples, but automated judges can. Judges make continuous, cheap quality measurement possible.

How it works

One model is prompted to score another model's output against a rubric. Judges have known biases - preferring longer answers, favoring the first option, or rewarding their own style - so calibrate them against human labels and watch for drift.

Example

A judge prompt asks: "Score this answer 1-5 for factual grounding against the context, and explain the score."

← Back to the full glossary

Put the platform behind the terms

Route, evaluate, and monitor every AI request from one OpenAI-compatible platform.

Start Free → Explore the Features