LLM-as-judge means using a capable language model to evaluate the outputs of an AI system, in place of (or alongside) human reviewers. The judge receives the input, the system's output, and optionally a reference answer and a rubric. It returns a verdict: pass/fail, a score, or a preference between two candidates.
It offers a practical middle ground. Human review is valuable but slow, while code assertions cannot judge qualities such as faithfulness, helpfulness, or tone. A model judge can score a larger sample at lower cost, though the exact speed and price depend on the chosen model.
The critical caveat: a judge is itself a model with biases (position, verbosity, self-preference) and error rates. It must be validated against human labels before its scores are trusted. An uncalibrated judge is just an opinion generator.
Treat the judge prompt as production code: version it, test it, and re-check its agreement with humans over time.
Judge input: user question + system answer + rubric
Judge output: {"verdict": "fail", "reason": "cites a price not present in context"}
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓