RizTech Academy logo
RizTech Academy
Automation: When, Why, and the PyramidLesson 3 of 430 min

Flaky tests destroy trust — the most important lesson

If you remember one lesson from this whole course about automation, make it this one. A flaky test is a test that sometimes passes and sometimes fails without anything real changing — same code, same app, run it twice, get two answers. Flaky tests are the single biggest reason automated suites fail in practice, not because they miss bugs but because they destroy trust, and a suite nobody trusts is worse than no suite at all. This lesson is what flakiness is, why it is so corrosive, where it comes from, and how to fight it.

Why flakiness is the worst problem

A flaky test does something subtle and fatal: it teaches people to ignore failures.

Think it through. The suite runs, and a test goes red. Is it a real bug, or just flakiness again? If a tester has learnt that failures are usually flakiness, the rational response is to shrug, re-run it, and move on when it passes. But now the suite has trained everyone to treat red as noise — so the day a red is a real bug, it gets the same shrug-and-re-run, and the bug ships. The suite still runs, still costs money and time, and catches nothing anyone acts on. It has become worse than nothing, because it consumes effort and provides false comfort.

That is why flakiness matters more than coverage, more than cleverness, more than how many tests you have. A small suite of tests that are always right when they go red is worth more than a huge suite that cries wolf. The whole value of a test is that a failure means something. Flakiness takes that away.

Where flakiness comes from

Flaky tests are not random bad luck; they have causes, and almost all of them live at the UI/end-to-end level (another reason the pyramid keeps that layer thin). The common ones:

  • Timing / race conditions — the test checks for something before the app has finished producing it. The page is still loading, the animation still running, the API response not yet back, and the test looks too early. This is the most common cause.
  • Hard-coded waits — sleep(2) to "let the page load". Sometimes two seconds is too short (flaky failure); sometimes it is wasteful. Fixed sleeps are both a flakiness source and a speed killer (the actions-and-assertions lesson shows the fix: wait for the condition, not the clock).
  • Test order and shared state — tests that depend on each other, or share data, so one test's leftovers break another. Run them in a different order and they fail.
  • Test data — a test that assumes a record exists, or that it is the only thing touching the data, fails when that is not true.
  • The environment — a slow or overloaded CI machine, a flaky network, an external service that is sometimes down. The test is at the mercy of things outside it.
  • Non-determinism in the app — dates/times, random values, ordering that is not guaranteed — the test asserts something that is not actually stable.

Notice how many are timing. Most flakiness is a test not waiting correctly for an asynchronous app — which is why modern tools build in smart waiting, and why you must never paper over it with a sleep.

Fighting flakiness

You fight flakiness with discipline, and it starts the moment a test goes flaky:

  • Treat a flaky test as a bug — a serious one. Do not ignore it, do not just add a re-run. A flaky test is actively damaging the suite's trust; fixing it is higher priority than writing new tests.
  • Wait for conditions, never the clock. Replace every sleep with "wait until this element/state is ready". Good tools (Playwright's web-first assertions, the next module) do this automatically — lean on it.
  • Make each test independent. It sets up its own data, does not depend on other tests, and cleans up after itself; it must pass run alone or in any order (the fixtures-and-data lesson).
  • Control the data. A test should own the data it needs — create it, use it, not assume it — so it does not depend on a shared, changing state.
  • Isolate from flaky externals. Where a test depends on an unreliable external service, stub or mock it so the test checks your app, not the internet.
  • Quarantine, do not tolerate. If a test is flaky and cannot be fixed immediately, move it out of the trusted suite (quarantine it) so it stops poisoning the signal — then fix it. Never leave a known-flaky test failing intermittently in the main suite.

The mindset: a test that is not reliable is not done. A test that passes when it should but also fails when it should not has negative value. Reliability is not a nice-to-have on top of an automated test; it is the whole point.

The rule to carry forward

Every automation module after this builds on one principle from here: guard the trust. Keep the suite small enough to be reliable, wait for conditions not clocks, make tests independent and data-owning, and treat every flake as a fire to put out. Do that, and a red result means "stop, there is a bug" — which is the only reason to have the suite. Fail at it, and you have built an expensive machine everyone has learnt to ignore.

Check your work

What flakiness is. A test that passes sometimes and fails sometimes with nothing real changed. It is the biggest reason automated suites fail — not by missing bugs, but by destroying trust.

Why it is the worst. It trains people to treat red as noise (shrug and re-run), so a real bug gets the same shrug and ships. The suite then costs effort and catches nothing acted on — worse than nothing. The whole value of a test is that a failure means something; flakiness removes that.

Where it comes from. Mostly UI-level and mostly timing: race conditions (checking too early — the top cause), hard-coded sleeps, test order / shared state, assumed test data, flaky environment/network/ externals, and non-determinism (dates, random, ordering).

Fighting it. Treat a flake as a serious bug (fix before new tests); wait for conditions not the clock; make each test independent and self-cleaning; control/own the data; isolate flaky externals with stubs; quarantine what you cannot fix now. A test that is not reliable is not done.

Practice

  1. Explain, step by step, how a single flaky test can lead to a real bug shipping.
  2. List the six common causes of flakiness and, for each, one way to remove it.
  3. Take sleep(2) in a test and rewrite the intent as "wait until <condition>"; say why the sleep was both flaky and slow.
  4. Describe how you would make two tests that currently share data independent of each other.
  5. Argue why a small always-reliable suite is worth more than a large flaky one, using the "cries wolf" idea.
  6. You find a flaky test the day before a release and cannot fix it in time — what do you do with it, and why not just leave it?

Official documentation

Next: choosing what to automate first.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship