Elo rating
Elo rating scores models by who wins head-to-head matchups, the same math chess uses to rank players. You show people two models' answers to the same prompt, let them pick the better one, and every win nudges the winner's number up and the loser's down. The number means nothing on its own. It only makes sense relative to the other models in the pool.
Why it matters
Some qualities are easy to feel but hard to score in the abstract. Ask a grader "how good is this answer, 0 to 100?" and you get noisy, drifting numbers. Ask "which of these two is better?" and people are far more reliable. Elo turns a pile of those cheap pairwise judgments into one clean ranking, which is how you settle "did the new model actually get better, or does it just feel that way?"
How it works
Collect many head-to-head comparisons on the same prompts, from real users or an LLM judge. Each match has a winner, and Elo updates both scores based on the result and how surprising it was: beating a much higher-rated model earns a big jump, beating a weaker one barely moves it. Run thousands of these and the ratings settle into a stable order. This is exactly how public leaderboards like Chatbot Arena rank models.
You are choosing a model for a support bot. Instead of scoring answers in isolation, you feed the same 500 customer questions to two candidates and have judges pick the better reply each time. Model A wins 62% of the matchups, its Elo climbs above Model B's, and now you have a defensible reason to ship it instead of a hunch.