DEV Community

Cover image for I built a tool that finds the minimal cause of a browser bug by replaying it
Zayd Mulani
Zayd Mulani

Posted on

I built a tool that finds the minimal cause of a browser bug by replaying it


Some bugs only happen with one combination of API data, timing and local state. The failing run and a working run differ in dozens of ways, and figuring out which difference matters usually means commenting things out until something changes.

Causality Cage automates that. You give it a failing session and a passing session. It replays the flow in Chromium with the differences switched on and off (API response fields, latency, response order, localStorage, cookies, environment), uses hierarchical delta debugging to shrink the set, then re-checks the answer by removing each piece and replaying.

What it looks like

demo

The report is a single offline HTML file with a causal graph. Click a factor to neutralize it and the failure node flips from red to green.

A real bug

I ran it on a closed issue in the RealWorld React app: gothinkster/react-redux-realworld-example-app#187. The app crashes with "Cannot read property 'tags' of undefined", and the reporter traced it to a stale token in localStorage. The original API host is dead, so I pointed the app at a public RealWorld API.

On the issue's commit, cage returned the full chain in 19 experiments (25.1 s): the stale token AND the API rejecting the two requests it triggers.

My first version said "the token" alone. Its minimal configuration crashed with a different error, and it counted any crash as the bug. So it now fingerprints the original failure and only counts runs that fail the same way. --loose-oracle restores the old behavior, which names the token alone in 10 experiments. The same run also exposed a redaction leak in my own report, which is fixed. The raw output is in launch/real-world.md.

What it does not claim

A result is necessary and sufficient only within the factors it models, under your flow and your oracle, at the measured repeat rates. It is 1-minimal, not guaranteed globally smallest. It does not model server-side state, WebSockets, IndexedDB or static assets. When the cause is out there, it says "unexplained" instead of naming something. Every report prints that scope box.

On a benchmark of planted bugs in my own demo app it matched ground truth on 10 of 10 rows, with caveats:

  • The OR row (two causes that crash with different messages) needs --loose-oracle. In strict mode it finds only one.
  • The flaky-oracle row failed after I added failure matching. It passes because of a rule I added afterwards: runs that crash before reaching the target failure are excluded from the fail rate.
  • The latency race flip point was 808 ms against an expected 600 to 900 ms band.

The app is mine, so treat that as a self-check, not evidence it works on yours.

Try it

npm install -g causality-cage
npx playwright install chromium
cage doctor
Enter fullscreen mode Exit fullscreen mode

Everything runs locally. There are no LLM calls (the verdict text is a template), and it can export a standalone Playwright test that fails until the bug is fixed.

Repo: https://github.com/zaydmulani09/causality-cage (MIT)

I'd most like to hear about bugs where it gives a wrong or useless answer.

Top comments (0)