Test Coverage Metrics That Actually Matter
One hundred percent coverage can still miss every bug that matters. Here is what to measure instead of a number.
Coverage is the one testing number everyone knows, which is why it gets abused. A team hits 100 percent, ships, and still gets paged at 2am for a null that a test walked right past. The problem: coverage measures whether a line of code ran during your tests, not whether you checked what it did. The gap between those two questions is where bugs live.
#What coverage actually measures
Line coverage answers one thing. Did this line execute while the suite ran? A test can call a function, run every line inside it, and assert nothing at all. The line still lights up green. The bug it contains stays invisible.
Here is the trap. You have function applyDiscount(price, pct) { return price - price * pct; }. A test calls it with (100, 0.1) and asserts the result is 90. That is 100 percent line coverage. Now pass a negative percentage, a pct above 1, or a null price. The function returns garbage and no test notices, because coverage never asked what the function should reject. It tells you the line ran. It says nothing about whether you tested the cases that matter.
#Mutation testing: the metric coverage wishes it were
Mutation testing asks the question coverage cannot. It takes your passing code, makes one small deliberate break (flips a > to >=, changes a + to a -, deletes a line), then reruns your tests. If a test fails, your suite caught the change. If every test still passes, that mutation survived, which means you have code whose behavior no assertion pins down.
The number it gives you is the percentage of mutations your tests kill, and it is worth far more than a coverage percentage. It measures whether your assertions have teeth. A suite at 95 percent line coverage and 40 percent mutation score is mostly theater. It runs the code but does not check it. Tools exist for most languages: Stryker for JS, mutmut for Python, PIT for Java. It is slow, so point it at your core logic and skip the rest of the repo.
#Measure the paths that would hurt
Not all uncovered code is equal, and a global percentage hides that. Ninety percent coverage that skips your payment retry logic is worse than 70 percent that nails it. The number that matters is coverage of the paths where a failure is expensive or hard to reverse.
So split the question. Which code, if it breaks silently, causes data loss, a wrong charge, or a security hole? Those branches need tests that assert real outcomes, including the errors and edge cases. Let the getters and the log lines stay lightly tested. A coverage report scores a logging statement and a funds transfer the same. Your attention should not.
#The bottom line
Coverage is not useless. A sharp drop in it is a real signal, and zero coverage on a module is worth knowing. But treat it as a floor and never let a percentage stand in for the harder question of whether your tests assert anything. Run mutation testing on the code that would hurt if it failed. Spend your test-writing time where a silent break is expensive. That is the difference between a suite that looks safe and one that is.