Semantic similarity
Semantic similarity is a score for how close two pieces of text are in meaning, ignoring the exact words. "Refunds take about two weeks" and "you get your money back in 14 days" share almost no vocabulary but mean nearly the same thing, and a good similarity score sees that. The catch: it scores closeness of meaning, so two wrong answers can still look nearly identical even when both are false.
Why it matters
Exact-match grading breaks the moment your model phrases a right answer differently, which is most of the time. You either fail correct answers on wording or loosen the check until it waves junk through. Semantic similarity gives you a middle path: credit an answer that says the right thing in its own words. Without it, your eval scores punish the model for being fluent, and you end up chasing wording instead of meaning.
How it works
You turn each text into an embedding, a long list of numbers that captures its meaning, then measure the angle between the two vectors with cosine similarity. The result runs from about 0 (unrelated) to 1 (nearly identical), and you pick a threshold where "close enough" starts. It is cheap and fast, which is why golden sets lean on it for open-ended answers. Watch the failure mode: opposite statements often land close, because "the API is down" and "the API is up" live in nearly the same neighborhood.
A doc Q&A bot is asked how to reset a password. The reference answer says "go to Settings, click Security, then Reset." The bot replies "open your account security settings and choose reset password." Word overlap is thin, but the similarity score is high, so the case passes instead of getting flagged as wrong on phrasing alone.