Bias evaluation
Bias evaluation checks whether a model treats similar people differently based on things that should not matter, like a name, gender, or ZIP code. You feed it cases that are identical except for one such detail and watch whether the answer changes. The catch: the bias rarely shows in any single answer. It only appears when you compare across groups.
Why it matters
A biased model can quietly hurt real people and put you on the wrong end of a lawsuit. If your loan helper is stricter with applicants named Jamal than ones named Jake, no single response looks wrong, but the pattern is. You will not catch this by reading a few outputs and nodding. It hides in the aggregate, so if you never measure it, you never know it is there.
How it works
The common trick is counterfactual testing: take one input, swap only the sensitive attribute, and check if the output or its tone changes. Run this across many paired cases and compare outcome rates between groups, like how often each group gets approved or the average sentiment of replies. A gap past a threshold you set in advance is a flag. This sits alongside red teaming and adversarial testing as a way to probe behavior you cannot see in normal use.
A resume screener rates two identical resumes, changing only the name from "Sarah" to "Mohammed." Sarah gets "strong fit" and Mohammed gets "consider with reservations." Neither answer looks broken on its own, but the paired comparison exposes the bias before it ever touches a real candidate.