DEV Community

Scott Mallinson
Scott Mallinson

Posted on Originally published at scottmallinson.com on

A passing test that never saw the page

The transcript in a desktop app I'm building is supposed to scroll to the newest message as it arrives. It didn't. The test suite covering that behaviour was green, and had been green the whole time.

The tests ran in jsdom. jsdom implements the DOM but not layout, so every element reports a height of zero. That isn't an error and not a warning. It's zero, returned in exactly the same shape a real browser would return a real number. Any code asking whether the content is taller than its container gets told no, every time. An assertion that the transcript has scrolled to the bottom passes, because in a world where nothing has height, everything is already at the bottom.

The problem was never thin coverage. There were tests for the scrolling. They passed. They were interrogating an environment that had quietly agreed with them.

The same failure in a stylesheet

I hit the same shape of failure on a dashboard in the same period, in a completely different stack.

A tile was parked behind a feature flag and marked hidden. It stayed on screen anyway. The site's own stylesheet contained .tile { display: flex }, and that beats the browser's built-in [hidden] { display: none } on specificity. So the attribute was set, the built-in rule existed, and "hidden" resolved to "visible". Nothing reported a problem, because from the browser's point of view nothing had gone wrong. It applied the rules in the order the cascade demands.

Both failures look like bugs in my code, and in the narrow sense they are. But they share something more specific than that. In each case a system was asked a question it wasn't equipped to answer, and it responded with a confident, well-formed, wrong answer instead of declining.

That's a much harder failure to defend against than a crash. A crash would have shown me where to look. A plausible zero gave me no reason to look at the tests at all.

What actually fixed it

The scrolling fix wasn't more tests. It was tests in an environment that has the property under test: real browsers driven by Playwright, against a stubbed IPC layer so the app's backend doesn't need to be running. The stub keeps the tests fast and deterministic. The real browser supplies the one thing jsdom structurally can't, which is layout.

A second decision inside that fix went the other way. The obvious call for the scrolling itself is scrollIntoView. I didn't use it. It walks every scrollable ancestor of the target, including the document, so asking it to bring one message into view can move parts of the page you never mentioned. Setting scrollTop directly on the transcript element is the narrower operation, and narrower is what you want when the thing you're fixing is unwanted movement.

The tile fix was a higher-specificity rule, .tile[hidden] { display: none }. This is a patch rather than a principle. The transferable part is smaller and duller: in any codebase that sets display at the component level, [hidden] is a suggestion. If you rely on it to hide something that matters, say so explicitly at the same specificity as the rule that shows it.

The question to ask of a test environment

The instinct after something like this is to write more tests. That instinct is wrong here, or at least insufficient, because the tests that existed were reasonable tests of the right behaviour.

The better question is which of the properties you care about your test environment actually implements, and what it returns for the ones it doesn't. jsdom isn't defective. It's a DOM implementation without a layout engine, doing exactly what it says on the label. The defect was mine: I asked it a layout question and believed the answer.

Framed that way, the audit is short and specific. Layout, fonts, timing, animation, real network behaviour, actual pixel rendering: which of these does the environment model, and which does it stub with a default that happens to be a valid-looking value? Anything in the second group is a place where a test can pass while the behaviour is broken, and no amount of additional assertions in that environment will help.

I stopped short of the fullest version of this. Testing against the real WebView the app ships with was on the table and I decided against it for now, on the grounds that a Chromium substitute catches most of what matters and the extra weight in CI wasn't justified by a single incident. That's a judgement rather than a principle, and if a second layout bug slips through a Playwright suite it's the first decision I'll revisit.

Why this is worse than a gap in coverage

I've written before about how a demo surfaces things a test suite structurally can't, which is about the limits of what you thought to check. This is narrower and more uncomfortable.

A gap in coverage is a known unknown. You can look at a test file and notice nothing exercises a path. What happened here is different: the test existed, it named the right behaviour, it ran, and it came back green while the behaviour was broken. It didn't fail to catch the bug. It actively asserted the bug was absent.

The green suite here was claiming the transcript scrolled, and the environment underneath it had no way to know.

What I don't have is a general way to spot this before it bites. Every environment stubs something, and the stubs that hurt are the ones returning valid-looking values rather than throwing. I found both of these by chasing a visible bug backwards into a passing test, which is a terrible detection strategy and the only one I've got. If you have a better one, I'd like to hear it.

Top comments (0)