JTjason.teixeira() Glossary
Home / Learn / Glossary / Eval harness
Evaluation

Eval harness

An eval harness is the plumbing that runs your test cases through a model and scores the answers for you. You feed it inputs, it collects the model's outputs, applies your grading rules, and hands back a number. The hard part is not running the model. It is deciding how each answer gets scored, because that choice quietly decides what good means.

Why it matters

Without a harness, you test by hand. You check three cases, get bored, and ship. That falls apart the first time you tweak a prompt, so most regressions slip through. A harness runs 200 cases in a minute every time you change something, so "I think it's better" turns into a score you can compare. It also keeps you honest, because the grading is written down instead of living in your gut.

How it works

At its core it loops over a golden set, sends each input to the model, and passes the answer to a scorer. The scorer might be exact match for closed answers, semantic similarity or F1 for fuzzy ones, or an LLM judge with a rubric for open-ended text. It logs every result so you can see which cases moved, then rolls them into an aggregate like pass rate or average score. Good harnesses also save the raw outputs, so a suspicious score is easy to inspect.

In practice

Say you are building a refund-policy bot. Your harness holds 80 real questions with known-good answers, runs each through the new prompt, and grades them with a judge that checks the answer against the policy. One run tells you the new version scored 88, down from 92, and points at the four cases it broke before a single customer sees them.

Want this checked on your own AI feature?
Get a free mini-eval — real findings on your live feature, no call required.
Free mini-eval →
© 2026 Jason Teixeira · Sage Ideas LLC · Glossary · Learn · privacy