JTjason.teixeira() Glossary
Home / Learn / Glossary / Pairwise comparison
Evaluation

Pairwise comparison

Pairwise comparison is asking a judge, human or model, "which of these two answers is better?" instead of scoring each answer alone. You pick the winner. People and models are far more reliable at ranking two things side by side than at putting an absolute number on one thing in isolation.

Why it matters

Absolute scores drift. Ask a model to rate an answer 1 to 10 and you get a 7 today and an 8 tomorrow, with no idea what the gap means. That wobble makes it nearly useless for telling whether your new prompt actually beat the old one. Pairwise comparison sidesteps the problem: put version A next to version B on the same input, ask who won, and the noisy scale stops mattering.

How it works

Run both versions on the same set of inputs. Show each input's two outputs to a judge and record the winner. Tally the wins into a rate, like "B beat A on 63 percent of cases," and offer a tie option so forced choices do not add noise. To kill position bias, swap which answer appears first and average the two runs. With many competing versions, feed the results into an Elo or Bradley-Terry model to get a full ranking from the head-to-head matches.

In practice

You tweak the prompt on a support bot and want to know if it improved. Instead of scoring 200 replies one by one, you show a judge the old reply and the new reply for each question and ask which better answers the customer. The new prompt wins 130 of 200. That is a clear call to ship, and you never had to argue about what a 7 out of 10 means.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy