Flaky tests: causes and a triage method
A flaky test fails and then passes with nothing changed in between. Nearly every one has one of five causes: waiting on the wrong thing, shared state, time, a real intermittent bug, or an environment that differs between runs. The triage method is to make the tool tell you which, by retrying once and recording both results, then fixing the cause rather than the symptom.
Why flakiness matters more than it looks
One flaky test in a suite of a hundred means every run has a chance of being red for no reason. People learn to re-run, then to ignore, and the suite stops protecting anything. The cost is not the one test; it is the trust in the other ninety-nine.
The five causes
1. Waiting on the wrong thing. The test clicks "Save" and immediately asserts the banner. Usually the banner is there in time; under load it is not. The fix is to wait for the outcome, not for a duration: the banner's text, the row in the table, the URL. Good tools wait on outcomes by default and never sleep for a fixed time.
2. Shared state. Two tests use the same account and one leaves it in a state the other does not expect. Or a test depends on the previous one having created something. The fix is a fresh session per test and data the test creates for itself or reads from a fixture that nobody mutates.
3. Time. Assertions on "today", on a relative date ("2 minutes ago"), on something that expires. The fix is to assert on the stable part of the text, or on a value the test controls.
4. A real intermittent bug. A race in the product that loses the save one time in fifty. This is not a flaky test; it is a found bug, and the test is doing its job. The tell is that the failure is the same every time and reproduces, eventually, by hand.
5. Environment. A different browser build, a slower worker, a font that changes a line break and moves a button. Visual checks suffer most; the fix is to compare only against baselines from the same rendering profile, never from a different machine.
The triage method
- Retry once, automatically, and record both results. A fail-then-pass is marked flaky, not passed and not failed. Treating it as a pass hides it; treating it as a failure trains people to ignore red.
- Look at the two runs side by side. The screenshots at the failing step, the timing of each step, the error text. A wait problem shows a page that was still loading. A state problem shows data that should not be there.
- Name the cause before touching the test. Write it on the test: "flaky: waits on banner". A suite where every flaky test has a named cause is a suite someone is in charge of.
- Fix the cause. Change the wait, isolate the data, stabilise the assertion, or file the bug.
- Quarantine what you cannot fix this week. A flaky test that blocks merges does more harm than a missing test. Move it out of the gate, keep it running, keep its cause visible.
A worked example
A composed example: the team, the timings and the counts show the shape of the work and are not measurements from one customer.
A suite of sixty tests shows two flaky verdicts a day. Side by side, the first failing run shows a spinner where the second shows the table: cause one. The test is changed to wait for the table's first row, and the verdict stops. The other test fails on an assertion that the order count is "3"; the failing run shows "4". Another test, run in parallel, creates an order against the same account: cause two. Each test gets its own account from a pool, and the suite goes quiet. Two causes, two fixes, no re-runs.
FlowQA explores your site, writes the tests, proves them against a broken page, and keeps proving them after every change.
See FlowQA