Flaky Tests — Isolation, Time, Concurrency & Quarantine
Root causes: time, randomness, shared state, async races, order dependence, externals. Quarantine vs fix; retries hide bugs; freeze clocks and isolate state.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
The same test failed on retry, then passed
Prefer
Make it deterministic
Inject the clock, seed the random source, give the test a fresh store, and await a condition instead of sleeping.
- A failure you can replay is a bug you can fix.
- Order randomization finds shared mutable state.
- The default suite stays something developers believe.
Alternative
Retry until green
Three automatic retries turn an intermittent logic bug into a slow pass.
- The flake rate disappears from the dashboard and stays in production.
- CI time grows because every red test runs again.
- New flakes hide inside the retry budget.
From symptom to a deterministic fix
The source diagram fans out from one failure. This column is the order to check the causes.
- 1
Time
Inject a clock. Freeze it. Stop calling now in five places. - 2
Randomness and shared state
Seed the generator. Give each test a fresh database or map. - 3
Async, order, and externals
Await a condition. Forbid cross-test writes. Fake the network boundary.
Overview
A flaky test sometimes passes and sometimes fails without a relevant code change. Flakes destroy trust faster than a missing test, because the team learns that red CI is noise. The usual causes are time, randomness, shared state, async races, order dependence, and external services.
The cure is determinism and isolation. It is not an infinite retry loop. Quarantine is temporary, with an owner. Retries hide bugs.
Higher layers of the pyramid flake more often. A unit test of a pure function has little room to flake. An end-to-end journey has clocks, animations, networks, and shared users. Shrink that mass before you invest in a smarter retry.
The six causes
Flow
- 1
1. Flaky failure, no code change
- next2. Time: inject a clock
- 2
2. Time: inject a clock
- next3. Random: seed the RNG
- 3
3. Random: seed the RNG
- next4. Shared state: fresh store
- 4
4. Shared state: fresh store
- next5. Async: await a condition
- 5
5. Async: await a condition
- next6. Order: no cross-test writes
- 6
6. Order: no cross-test writes
- next7. External: fake the boundary
- 7
7. External: fake the boundary
Lesson map
Flaky Tests — Isolation, Time, Concurrency & Quarantine
Root causes: time, randomness, shared state, async races, order dependence, externals. Quarantine vs fix; retries hide bugs; freeze clocks and isolate state.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Flaky failure, no code change"] b["2. Time: inject a clock"] c["3. Random: seed the RNG"] d["4. Shared state: fresh store"] a -->|1. Flaky failure, no code change| b b -->|2. Time: inject a clock| c c -->|3. Random: seed the RNG| d
Time. A trial expiry that calls the system clock fails when the suite runs near midnight or across a timezone. Inject a Clock. Freeze it in tests. Vitest does this with vi.useFakeTimers() and vi.setSystemTime. The sandbox below uses an explicit clock so it stays deterministic without a test runner.
Randomness. An unordered assertion, or a generator with no seed, fails on one worker. Seed the source. Sort before you compare. Property tests that flake belong back on the property page with a printed seed.
Shared state. Two tests write the same row, the same temp directory, or a module singleton. Each test gets a fresh database, or a unique key, and the singleton resets in setup. "Passes alone, fails in the suite" is this bug until proven otherwise.
Async races. A fixed sleep is shorter than the machine is slow. Await a condition: the row exists, the locator is visible. Playwright's web-first assertions retry the condition. waitForTimeout sleeps and hopes. Disable animations and use stable test ids so the condition is about state, not a transition frame.
Order dependence. Test A leaks a mutation that test B reads. Randomize order in CI. The leak shows up as a flake. The fix is isolation, not pinning the order forever.
External services. A live third-party API in the pull-request job will fail when they do. Fake at the boundary. A sandbox smoke can run outside the merge gate if you truly need the vendor. Contracts cover shape. Doubles cover the unit and integration job.
Quarantine, fix, or retry
Fix when you can reproduce. Most time and shared-state flakes reproduce once the clock is frozen or the order is randomized.
Quarantine when you cannot reproduce yet. The test leaves the default suite so it stops blocking merges, and it keeps an owner, a ticket, an expiry, and a budget. Quarantine is not a graveyard. A quarantine with no expiry becomes the new permanent suite, and the flake rate moves rather than falls.
Retry only for a known infrastructure blip: a runner that lost the network for a second. Never as the default for logic tests. A retry that converts an intermittent assertion into a green build has deleted the signal. Track how often the retry saved the build. That number is the flake rate you are choosing not to fix.
A frozen clock and an isolated store
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The Python date is 28 September 2026 at noon UTC, plus 14 days, which is 12 October. The same instant tomorrow is a different answer if the test calls the real clock. That is the flake.
Interview Q&A
What is a flaky test?
Answer
A test whose pass or fail changes without a relevant code change. Rerun it and the result flips. The defect may still be in the product, but the test cannot point at a commit. Until it is deterministic, the team cannot tell a regression from noise.
Why are retries dangerous?
Answer
They convert intermittent failures into silent passes. The bug remains. CI gets slower. Dashboards show a healthy suite because the third attempt went green. Retry a known runner blip if you must, and count the saves. Do not retry assertion failures in logic tests.
How do you test time-dependent code?
Answer
Inject a clock. Production passes a system clock. Tests pass a fixed instant. Avoid now scattered through helpers, because each call can observe a different second. Fake timers in Vitest are one way to freeze the process clock when you cannot thread a clock through old code. Prefer an explicit parameter when you can change the design.
It passes alone and fails in the suite. What now?
Answer
Assume shared state or order dependence. Run the suite in random order. Look for module singletons, leftover rows, and files in a shared temp directory. Give the test a fresh store. Pinning the order hides the leak until someone adds a test in the middle.
What about third-party APIs?
Answer
Fake them at the boundary for the pull-request suite. Add a contract if the risk is response shape. An optional sandbox smoke can live outside the merge gate. A unit test that needs the vendor to be up is a flake you scheduled on purpose.
What quarantine policy is acceptable?
Answer
An owner, a ticket, an expiry, and a budget for how many tests may sit there. Quarantine is for a flake you cannot reproduce yet, so the default suite stays trustworthy. It is not acceptable as a permanent label on a test everyone is afraid to fix. When the expiry hits, the test returns to the suite or it is deleted.
How do you fix an end-to-end animation flake?
Answer
Use stable test ids. Disable animations in the test build. Assert with a web-first condition that retries until the state is true, or times out with the last observation. Do not waitForTimeout for a round number of milliseconds. If the journey is not critical, delete the test and cover the rule lower in the pyramid. Higher layers flake more. Less end-to-end mass means fewer of these bugs.
How do flakes interact with the pyramid?
Answer
The tip flakes more than the base, because it includes time, networks, and shared environments. A strategy that piles tests into end-to-end also piles flakes. Push assertions down. The flake budget you spend should be on the few journeys that need a browser, not on every field validation.
When is quarantine the wrong call?
Answer
When you can already reproduce the failure. A frozen clock, a seed, or a fresh database is a fix, and shipping the test to quarantine instead leaves a known hole. Quarantine is for the unreproducible remainder, with a person attached. A reproducible logic bug is a mandatory fix.
What do you say in the first minute?
Answer
A flake passes and fails without a code change, and it destroys trust. Causes are time, randomness, shared state, async races, order, and external services. Inject a clock, isolate state, await conditions. Quarantine with an owner is temporary. Retries hide bugs. Shrink the end-to-end tip, because that is where flakes concentrate.
Pitfalls
- Adding retries to the default logic suite and calling the flake solved.
- Sleeping for a fixed delay instead of awaiting a condition.
- Sharing one database across tests that each assume an empty table.
- Leaving quarantine entries with no owner and no expiry.
- Calling the system clock inside a property test.
- Debugging a pixel animation before asking whether the test should exist.
One test fails only near midnight. One fails only after the user-admin test. One fails when the payments sandbox is down. Name the cause for each, the fix, and which one is allowed to be quarantined overnight.