LLM-as-a-judge
LLM-as-a-judge means using one strong model to grade another model's output. The judge scores each answer against a rubric or compares it to a reference, so instead of a human reading a thousand responses, the model reads them and returns a score or a verdict. The catch: your judge has its own biases. A grader you never checked is just a second opinion you decided to trust blindly.
Why it matters
Human grading does not scale. You cannot pay someone to read every answer on every prompt change, so most teams grade nothing and ship on gut feel. A judge scores hundreds of open-ended answers in minutes, which is often the only way to catch a regression before users do. The danger is a lazy judge that rubber-stamps everything, so test the judge against a small set of human-graded cases before you trust it.
How it works
Write a clear rubric. Then prompt the judge with the question, the answer, and the grading criteria, and ask for a score plus a short reason. The reason matters: it shows you when the judge is grading on the wrong thing. Judges drift with tiny changes in wording, and they favor longer answers written in their own style, so pin the format and calibrate against human labels. Asking "is answer A or B better" tends to be more reliable than asking for an absolute number.
A support bot answers "how do I cancel my plan" a hundred different ways across a prompt change. You hand each answer, plus the correct cancellation steps, to a judge with a rubric: does it name the right menu, and does it avoid inventing a fee? The judge flags eleven answers that quietly made up a cancellation charge, and you catch the bug before it reaches a single customer.