Why Your LLM Eval Scores Do Not Match Production
Your eval says 94 percent and your users are unhappy. Both can be true, and the reason matters.
Your eval says 94 percent and your users are unhappy. Both can be true. The number measures something real, just not the thing your users feel. Eval scores diverge from production for reasons you can find, and once you know where to look, the gap stops being a mystery and turns into a to-do list.
#Your test set is not your traffic
Most eval sets get built once, early, from cases someone could think of at a desk. Production traffic gets written by real users who are tired, terse, and weird in ways you did not predict. The distribution drifts the day you launch and keeps drifting.
Here is the tell. Your eval is full of clean, well-formed questions. Production is full of typos, half-sentences, pasted logs, and people asking three things at once. A model that scores 94 percent on the tidy set can crater on the messy one, because the messy one is a different test. If you have never sampled a few hundred real production inputs and diffed them against your eval set, you do not know what you are measuring.
#The average hides the failures that matter
A single aggregate score hides the shape of the errors. 94 percent correct sounds great until you notice the failing 6 percent is not random. It clusters on your highest-value customers, or your newest feature, or the intent that drives revenue.
Users do not experience your average. They experience their own requests. One confidently wrong answer to a paying customer does more damage than a hundred correct ones they took for granted. Slice the score by intent, by user segment, by input length, by whatever your revenue depends on. The average lies by leaving things out. The slices show you where the pain lives.
#You are grading answers, not journeys
Offline evals score a single turn: one input, one output, one grade. Production is multi-turn and stateful. A user rephrases, the context window fills, a tool call fails silently, an earlier answer poisons the next one. None of that shows up when you grade isolated pairs.
Then there is correct versus useful. An answer can be factually right and still fail the user because it buries the point, ignores a constraint they stated two turns ago, or does not fit the workflow they are actually in. Your rubric rewards correctness because correctness is easy to grade. Users are rating whether you solved their problem, which is harder to grade and the only thing that counts.
#The bottom line
Do not throw out the eval, and do not trust it past what it measures. Close the gap three ways: feed real production traffic back into your test set, read the slices instead of the average, and grade whole interactions instead of isolated pairs. The 94 percent is not a lie. It just answers a smaller question than the one your users are asking.