DEV Community

Vitalii Holben
Vitalii Holben

Posted on

I Caught a Production CSS Bug Before Users Did


title: I Caught a Production CSS Bug Before Users Did — Here's the Visual Regression Setup published: true tags: testing, webdev, css, devops

We shipped a "minor" font-size change last month. Updated the base size from 16px to 15px in our design tokens. Seemed harmless — the designer wanted tighter spacing.

It broke the checkout flow.

The payment form had a hardcoded min-height based on content size. With the smaller font, the content shifted up, and the submit button overlapped the terms checkbox on mobile viewports. Users on iPhone SE literally couldn't complete a purchase.

We didn't catch it in code review because the CSS diff looked fine. We didn't catch it in unit tests because we don't render real layouts in Jest. We caught it 4 hours after deploy when support tickets started coming in.

What visual regression testing actually looks like

The concept: screenshot key pages before deploy, screenshot them after, diff the images. If pixels changed where they shouldn't have, block the deploy.

In practice it's messier than that.

# pseudo pipeline
1. Take "baseline" screenshots of production
2. Deploy to staging
3. Take "candidate" screenshots of staging
4. Compare each pair, flag differences above threshold
5. Human reviews flagged diffs
6. Approve or reject deploy

Step 4 is where most tools fall apart. Pixel-perfect comparison produces way too many false positives — browser antialiasing, font rendering differences, dynamic content like timestamps, ads, user avatars.

What I actually use

After trying a few setups (BackstopJS, Percy, custom Puppeteer scripts), I landed on a simpler approach:

  1. Monitor key pages periodically with SnapshotArchive — it takes scheduled screenshots and stores the history
  2. When we deploy, I compare the latest pre-deploy screenshot against the staging version
  3. Significant visual changes get flagged in Slack

The advantage over CI-integrated tools: I get visual monitoring even when there's NO deploy. So if a third-party script changes, or a CDN serves different assets, or a CMS editor modifies content — I still see it.

The false positive problem

Every visual testing tool struggles with this. Here's what reduced noise for me:

Mask dynamic regions — hide elements that change on every load (timestamps, personalized content, ads). Most tools support CSS selectors for this.

Perceptual diffing over pixel diffing — a 1px antialiasing difference isn't a real change. Tools that use perceptual comparison (SSIM or similar) produce way fewer false alarms.

Separate critical from cosmetic — the checkout page gets a 0.1% threshold. The blog page gets 5%. Not everything needs the same sensitivity.

My current setup in numbers

  • 12 pages monitored (homepage, pricing, checkout flow, key landing pages)
  • Screenshots every 6 hours outside of deploys
  • ~2 real alerts per week after tuning thresholds
  • 0 CSS incidents reaching production since setting this up

The checkout font-size bug would've shown up as a visual diff on staging before we merged. Instead it cost us 4 hours of lost conversions and a hotfix.

The irony: the monitoring setup took an afternoon. The incident cost way more.

What's your visual testing setup look like? Or are you still yolo-deploying CSS changes like we used to?

Top comments (0)