LLM-as-a-judge uses a model to score another model's output against a rubric. It scales human-style judgment across large datasets cheaply, but judges have known biases (verbosity, position, self-preference). Calibrate judges against human labels and watch for drift before trusting their scores as gates.
Why it matters
Human review does not scale to thousands of examples, but automated judges can. Judges make continuous, cheap quality measurement possible.
How it works
One model is prompted to score another model's output against a rubric. Judges have known biases - preferring longer answers, favoring the first option, or rewarding their own style - so calibrate them against human labels and watch for drift.