I run a nightly audit across five static sites. One night it came back with 17 failures: eight from the site verifier, three unreachable assets, six index checks returning no HTTP code at all.
I spent the first hour preparing to fix 17 problems. The real number was zero. Every one of those failures was produced by my checking code, not by my sites — and the reason they were convincing is that each one had a specific, reproducible mechanism behind it.
Here are the four traps, in the order I hit them.
Trap 1: a JSON-LD grep that can't see spaces
After deploying a fix, I grepped the live HTML to confirm the new WebSite block was there:
grep -o '"@type":"WebSite"' page.html | wc -l
# 0
Zero matches. The deploy had to have failed, right? Except the byte count on the server matched my local file exactly, which is the hardest evidence I have that a deploy landed.
The block was pretty-printed across multiple lines, so the actual text was "@type": "WebSite" — with a space. My pattern had none. It would return 0 forever, on a page that was completely correct.
grep -o '"@type": *"WebSite"' page.html | wc -l
# 1
The rule: when a negative result contradicts a positive one, check your pattern before you check your deploy.
Trap 2: a 404 probe that had been poisoned
My verifier confirms a site returns 404 for pages that don't exist, using a fixed probe URL. It reported a failure, and it kept reporting it — three retries, all returning HTTP 000.
A consistent failure looks like a real failure. That's exactly what fooled me. The fixed URL had become a reliable timeout through my local proxy, so the probe was measuring my network, not my site.
Swapping in a random path fixed the diagnosis immediately:
curl -s -o /dev/null -w '%{http_code}' https://example.com/nope-$RANDOM.html
# 404
Same site, same minute, correct answer. When a probe fails repeatedly at one spot, change the input and see whether the failure follows you. If it doesn't, the probe is the broken part.
Trap 3: a truncated response that opens with <!DOCTYPE html>
I deploy by comparing the live file against my local copy byte for byte. One deploy looked wrong: the server returned 19,139 bytes, then 1,360, against a local file of 49,901.
Both of those short responses began with <!DOCTYPE html>, so they looked like real pages that happened to be different. They weren't different — they were cut off. Fetching four more times gave me two complete responses that matched byte for byte.
The lesson I'd keep: cmp disagreement means re-fetch before it means re-deploy. A truncated response is indistinguishable from a real one if you only look at the first line.
Trap 4: the CDN serving you yesterday's page
The same night, another site compared unequal — 110,058 bytes live against 110,293 locally. Same trap, different cause: the edge was still holding the previous version. Five fetches later, every one returned 110,293.
Verifying immediately after a deploy is the one moment you're most likely to read a stale copy.
The one that looked like a real diff and wasn't
A byte comparison came back 156 bytes short across four lines. The difference turned out to be Cloudflare's email protection rewriting every mailto: link into an obfuscated /cdn-cgi/l/email-protection# path. Normalising for that made the files identical.
I nearly "fixed" a deploy that had already succeeded perfectly.
What I changed
The bug wasn't any single one of these. It was that my checker treated its own result as evidence about the site. Now it separates them:
- Byte-identical beats a green checkmark. If the live file matches my local file byte for byte, the deploy worked — regardless of what a grep says.
- Network failures get retried with different inputs, not the same input. Three identical failures on one URL is one data point about that URL, not three.
- Re-fetch before concluding. Truncation and cache staleness both disappear on a second or third fetch.
- When a checker contradicts itself, fix the checker first. All 17 failures that night were instrumentation. I would have found that in ten minutes if I'd started there.
I also hit a separate crash writing the scanner itself: subprocess with text=True throws UnicodeDecodeError: 0xff the moment it reads a binary asset like a .webp or .ico. errors="ignore" fixes it.
The general shape: a tool that reports problems is also a thing that can have problems. Before you spend an hour fixing what it found, spend ten minutes proving it can find things correctly.
If your monitoring, deploy checks or CI have started crying wolf and you can't tell which alarms are real, that's the kind of thing I do: https://yongrui-services.pages.dev/
Top comments (0)