Test metrics that matter, and vanity metrics that do not
Teams love to measure testing — number of test cases, bugs found, percentage coverage, tests passing. But most testing metrics are misleading, and some are actively harmful: they get gamed, they reward the wrong behaviour, and they create a false sense of safety. A good QA understands which numbers mean something and which are theatre. This lesson is the metrics that matter, the vanity metrics that do not, and why measuring testing well is harder than it looks.
Why metrics mislead
The core problem: testing metrics measure activity or proxies, not the thing you actually care about, which is confidence that the software works. And the moment a metric becomes a target, people optimise the metric rather than the goal (this is Goodhart's law: "when a measure becomes a target, it ceases to be a good measure"). So a number that looked useful becomes gamed and meaningless. Almost every bad testing metric fails this way — it is easy to move the number without improving quality.
The vanity metrics
Be sceptical of these, which are commonly reported and commonly misleading:
- Number of test cases. More tests is not better testing — a thousand shallow, redundant tests are worse than fifty sharp ones (they cost more to maintain and catch no more). Counting test cases rewards volume, not value. It is trivially gamed by splitting one test into ten.
- Number of bugs found. Tempting, but perverse. A high bug count might mean thorough testing — or buggy code, or trivial nitpicks padding the number. And it pits QA against developers (more bugs "for" QA looks like more bugs "against" them — the adversarial trap). Bugs found says little on its own.
- Test pass rate / "98% passing". Passing tests feel reassuring but measure almost nothing about quality: they could be shallow tests, or tests of the wrong things, and a green suite over poor tests is false comfort. And it pressures people to make tests pass rather than to test well.
- Code coverage percentage. The most abused of all. Coverage tells you which lines ran during tests — not whether they were checked. You can have 100% coverage with assertions that verify nothing (the test executes the line and asserts nothing meaningful). High coverage does not mean well-tested; chasing a coverage target produces tests written to touch lines, not to catch bugs.
None of these is useless, but each is dangerous as a target, because each is easy to move without improving quality — and once it is a target, it will be moved.
What is actually worth measuring
Better signals focus on outcomes and health, not activity:
- Escaped bugs (defects found in production). The most honest testing metric: bugs that reached users are the ones your testing missed. Tracking these — how many, how severe, where they came from — tells you where testing is weak and improving. Low and falling escaped-defect counts are real evidence testing is working.
- Where bugs are found (which stage). Bugs caught early (in design, in review, in test) cost less than bugs caught in production (the cost-of-bugs lesson); a healthy trend is bugs found earlier over time.
- Time to detect and fix. How fast a regression is caught (fast if CI runs on every change) and fixed — a measure of the feedback loop's health.
- Flaky-test rate. How many tests fail intermittently — directly measures the suite's trustworthiness (the whole automation part's theme). This one is worth watching closely, because it predicts whether the suite will be believed.
- Coverage as a diagnostic, not a target. Use coverage to find untested important areas ("the payment module has no tests — that's a risk"), not as a score to maximise. As a map it is useful; as a target it is harmful.
The pattern: measure outcomes (did bugs reach users? are they caught earlier? is the suite trusted?), not activity (how many tests, how many bugs, what percentage). Outcomes are harder to game and closer to what you actually care about.
The honest limit of measuring testing
A mature position to hold: the value of testing is partly unmeasurable. The bugs you prevented by testing early, the disaster you avoided, the confidence that let the team ship — these do not show up as a number, and the best testing often produces quiet results (nothing went wrong). So resist the pressure to reduce testing to a dashboard. Use metrics as signals to investigate — a rising escaped-bug count or flaky rate is a prompt to look closer — not as targets to hit or a verdict on a QA's worth. A QA who games metrics to look productive is less valuable than one who tests the right things and reports risk honestly, even when that is harder to put on a chart. Judgement, not numbers, is the job.
Check your work
Why metrics mislead: they measure activity/proxies, not confidence-the-software-works; and by Goodhart's law, a measure that becomes a target gets gamed. Most bad testing metrics move easily without improving quality.
Vanity metrics (dangerous as targets): number of test cases (rewards volume, not value; trivially gamed), bugs found (high could mean thorough or buggy or nitpicks; fuels the adversarial trap), pass rate ("98% passing" over shallow tests is false comfort), and code coverage % (measures lines run, not checked — 100% with empty assertions; chasing it writes tests to touch lines).
Worth measuring (outcomes): escaped bugs (the honest one — what testing missed), where/when bugs are caught (earlier is healthier), time to detect/fix (feedback-loop health), flaky-test rate (suite trustworthiness), and coverage as a diagnostic to find untested important areas, never a target.
Honest limit: much of testing's value is unmeasurable (prevented bugs, avoided disasters, confidence); use metrics as signals to investigate, not targets or a verdict. Judgement over numbers.
Practice
- Explain Goodhart's law with a testing example (e.g. a code-coverage target being gamed).
- For each vanity metric (test count, bugs found, pass rate, coverage %), give one way it misleads and one way it is gamed.
- Explain why "100% code coverage" does not mean "well tested", with a concrete example.
- Choose three outcome metrics you would actually track and justify each.
- Explain how you would use code coverage as a diagnostic without making it a target.
- Argue why a QA who reports honest remaining risk is more valuable than one who maximises a metric.
Official documentation
- Martin Fowler — Test coverage — Why coverage is a useful diagnostic but a bad target.
- Google Testing Blog — Code coverage best practices — Using coverage well, and its limits.
- Ministry of Testing — Metrics — Practitioner views on measuring testing.
Next: growing as a QA — from manual to automation to SDET.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship