EN·ES·PT
Test automation · flake protocol

A flaky suite is worse than no suite at all.

When a red run is just noise, engineers stop reading it — and the one time red meant a real regression, it sails straight through to production. Flakiness isn't a nuisance; it's the failure of the whole gate. This is the protocol I use to take a suite from red-as-noise to red-as-signal: diagnose the sources, quarantine and set a retry policy, fix the isolation — and keep the tests you already wrote.

Watch · test automation & CI in ~45s
10% → <1%
HighStrike flake rate, live trading workflows
500+
tests stabilized on real market flows
37 / 37
specs green · 0 flakes · 15.3s (public suite)
4h → 75m
Home Depot regression, 2,300+ stores

Every figure here is from real work. The public ones link to the receipt — the Playwright suite is on GitHub and this site's own QA runs live at eval.html; the HighStrike & Home Depot numbers are from prior full-time roles and are self-reported. No rounded-up demo numbers.

The real cost · why a buyer should care

Noise doesn't just waste time. It hides the one failure that mattered.

The trust collapse

A 10% flake rate on a 500-test suite means dozens of false reds per run. Engineers learn to click "re-run" on reflex — and the habit that clears noise also clears the one real regression hiding in it.

The velocity tax

Every flaky red is a context switch: investigate, dismiss, retry, wait. Multiply that across a team and a day, and the suite you built to move faster is the thing slowing every merge down.

The gate that isn't

A CI gate only protects you if red is believed. Once it's noise, the gate is decorative — you're shipping on hope, with a green-ish dashboard that gives false confidence.

The shift · what the protocol changes

From red-as-noise to red-as-signal.

The goal isn't a suite that's always green. It's a suite where a red run means something — so the team reads it, trusts it, and stops when it fires.

Flaky pipeline versus stable pipeline A flaky pipeline emits random red runs that are treated as noise and re-run. The flake protocol — diagnose, quarantine, set retry policy, fix isolation — turns it into a stable pipeline where a red run is a trusted signal that blocks the release. BEFORE — RED IS NOISE Random reds → "just re-run it" → real regression slips through FLAKE PROTOCOL 1 · Diagnose sources 2 · Quarantine + retry 3 · Fix isolation AFTER — RED IS SIGNAL real bug CI GATE red = blocked, believed SHIP Green is trusted. The suite you kept still runs. BLOCKED The real regression is caught.
passing run failing run the gate
The protocol doesn't delete your tests — it makes their verdict trustworthy again.
The diagnosis · flake source → fix

Most flakiness comes from four places.

Flakiness is rarely random — it's a symptom with a source. The protocol starts by naming which one you have, because the fix is different for each. This is the map I work from.

Flake source → tell → fix
SourceHow it shows upThe fix
Timing & races Passes locally, fails in CI; fails under load; fixed by adding a sleep(). Replace fixed waits with state-based waits (web-first assertions, poll-until-condition). Never wait on the clock — wait on the app.
Shared state Fails only when run with other tests; a "poisoned" record from a prior test leaks in. Isolate every test: fresh fixtures, unique data per run, teardown that actually cleans up. No test should depend on another having run.
Order dependence Green in one order, red when the runner shuffles or parallelizes. Make tests order-independent and parallel-safe: no reliance on execution sequence, no global singletons mutated across specs.
Environment & data Depends on a live third party, a seeded row, a time-of-day, or network weather. Control the boundary: deterministic seeds, mocked or contract-tested externals, frozen clocks where behavior is time-sensitive.

This is the same diagnosis that took the HighStrike suite from a 10% flake rate to under 1% — on 500+ tests running against live trading workflows, where a false red is expensive and a missed real one is worse.

The protocol · how the work runs

Quarantine first. Fix second. Never delete the suite.

1 · Quarantine lane
  • Known-flaky specs move to a separate, non-blocking lane — CI goes green-trustworthy immediately
  • Nothing is deleted; every quarantined test is a tracked debt with an owner
  • The blocking suite only contains tests you can believe today
2 · Retry policy
  • A bounded, measured retry — not a blanket retry that hides real failures
  • Retries are counted: a spec that only passes on retry is flagged, not celebrated
  • Flake rate becomes a metric on the dashboard, so it can't silently creep back
3 · Isolation fixes
  • Work the quarantine lane down source-by-source using the matrix above
  • Each fixed spec graduates back into the blocking suite, proven stable
  • You end with the same coverage you started with — now trustworthy
▸ how I build test automation & CI ISTQB CT-AI + Test Automation Engineer certified
The proof · not a claim, a link

A public suite that runs green, fast, and flake-free.

The public regression suite

An open Playwright SDET suite you can read and run: 37 of 37 specs passing, 0 flakes, 15.3-second full run. Web-first assertions, isolated fixtures, parallel-safe — the protocol on this page, in code you can inspect.

37/37 specs0 flakes15.3sPlaywright
This site QAs itself

The site you're reading runs its own QA — 100+ checks, axe-clean accessibility — and hosts a live eval demo. Proof over claims is the whole thesis, so the proof is one click away, not a case-study PDF.

100+ checksaxe-cleanlive demo
Related · more on testing AI

Have a suite the team stopped trusting?

Book a call and we'll diagnose the flake sources on your actual suite — then quarantine, fix, and hand it back as a gate red means something again.

Book a call → Test automation & CI service →
frequently asked

Questions, answered.

What actually causes flaky tests?
Almost always shared state, timing and races, or test-ordering dependencies — not bad luck. I diagnose the source instead of blanket-retrying.
Do you just add retries?
Retries are a stopgap that hides the signal. I install a quarantine lane and retry policy while fixing the root cause — isolation, waits, fixtures.
Will you rebuild our whole suite?
Rarely. Keeping your existing suite and making it trustworthy is faster and cheaper; a rebuild is the last resort, not the first move.
What result can I expect?
A fast, stable suite where red means red. In a prior full-time role I cut a live suite’s flake rate from ~10% to under 1% and runtime from 45 to 8 minutes (self-reported).
How do you measure our real flake rate before starting?
I re-run your suite across many CI passes on unchanged code and count the tests that flip between pass and fail. That gives a real flake rate and a ranked list of the worst offenders before I touch anything.
What is the quarantine policy — do flaky tests get disabled?
Flaky tests move to a quarantine lane, not the trash. They still run and report, but they stop blocking the main pipeline while I fix the root cause, then graduate back once they hold.
How long does it take to get under one percent?
It depends on suite size and how deep the flake sources run. I quarantine early so your pipeline becomes trustworthy fast, then work the root-cause fixes down in priority order.
Does this work with our existing suite or require a rewrite?
It works with your existing suite. Whether you run Playwright, Pytest, or another framework, I stabilize what you already have rather than starting over — a rebuild is the last resort.
How do you find the root cause of a flaky test?
I isolate the failing test, run it in and out of order, and inspect shared state, timing, and fixtures until the failure reproduces on demand. A flake you can reproduce is a bug you can fix.
What do we have at the end that keeps flake from returning?
A stabilized suite plus the guardrails that hold it there — a quarantine lane, a retry policy, and isolation patterns your team applies to new tests. The signal stays trustworthy after I leave.