F1 score
The F1 score is a single number from 0 to 1 that combines two things you care about: how much of the real stuff you caught, and how often you were right when you raised your hand. It rewards a system only when both are good. A lopsided model can still look fine on plain accuracy, and F1 is what exposes that.
Why it matters
Plain accuracy lies when one answer is rare. A fraud detector that says "not fraud" every time can score 99% accurate and catch zero fraud. F1 refuses to reward that, because catching nothing tanks one half of the math. If your classifier deals with a rare positive, F1 is the number that tells you the truth.
How it works
F1 balances precision (of the things you flagged, how many were right) and recall (of the things that mattered, how many you caught). It is the harmonic mean of the two, which is a fancy way of saying it stays low unless both are decent. 90% precision with 10% recall gives an F1 near 0.18, nowhere near 0.50. You compute it on a labeled test set, then watch it move as you tune your threshold.
A support bot flags messages that need a human. It only escalates when very sure, so precision is high. But it quietly ignores half the angry customers, so recall is bad. Accuracy looks great. F1 comes back at 0.4, and that one number tells you the bot is calm and useless.