Rubric-based grading
Rubric-based grading scores an answer against a written checklist of what a good answer must contain, instead of asking someone to eyeball it and pick a number. You spell out the criteria in advance: does it cite the source, does it name the right refund window, does it stay on topic. The hard part is writing criteria specific enough that two graders land on the same score.
Why it matters
A bare "rate this 1 to 10" score drifts. The same answer gets a 7 on Monday and a 9 on Friday, and you can never explain why one version beat another. A rubric makes the grade legible: you can point at the exact criterion that failed and fix that one thing. It also lets an LLM judge grade consistently, because you have told it what to look for instead of leaving it to taste.
How it works
Break "good" into a handful of concrete checks, each with a clear pass condition and a weight. Run every answer past every check. Do it by hand for small sets, or use an LLM as a judge to scale. Sum the points into a score you can track over time. Make checks binary where you can ("mentions the 14-day limit: yes/no"), because a yes/no question is much harder to grade inconsistently than a vague 1-to-5 feel.
A support bot answers refund questions. Your rubric has four checks: names the correct 14-day window, tells the user how to start the refund, stays polite, and adds nothing the policy does not say. One reply is friendly and clear but invents a "store credit" option, so it passes three checks and fails the last one, scoring 3 of 4. Now you know exactly what to fix.