I was cleaning up templated text across 385 pages. Partway through, someone asked how many were left. I said 208. Then later I said 188. Then 56. Then 20.
Only the last one was measured. The rest were numbers I had said out loud once and then re-used as if they were still true.
That is the part worth writing down, because the failure was not arithmetic. It was that I treated a previous statement as a data source.
How it actually happened
Two mistakes, and they compound.
1. I counted from a snapshot I had already changed.
The scan takes a few seconds and reads from disk. I had just rewritten dozens of files. When I quoted a count without re-running it, I was reporting a state that no longer existed on disk.
Worse: the number I kept quoting had come from before the first batch. So every time I repeated it, the gap between the number and reality grew.
2. I used two different definitions without saying which one.
I had two ways to call a page "still templated":
- EXACT — the sentence is byte-identical to another page's
- SKELETON — same structure after replacing product names, numbers and prices with placeholders
The second is much stricter. A page can be perfectly unique word-for-word and still match something else structurally, because there are only so many natural ways to write "X is better than Y when Z."
I quoted whichever number suited the moment without saying which definition produced it. Both numbers were internally correct. They just were not the same measurement.
The fix that made it stop
A script with the definitions hard-coded into it, run fresh every time:
EXACT = clean(text) # strip tags, unescape, collapse whitespace
SKELETON = placeholders(EXACT) # product names / numbers / prices -> X
dup_pages = [p for p in pages if exact_counter[text_of(p)] >= 2]
Two properties matter more than the code:
- It reads from disk on every run. No cached counts, no remembering what the number was last time.
- The definition is in the file, not in my head. EXACT and SKELETON are written down, so "56" and "173" can never be confused for each other again.
The rule I now follow: before quoting how much is left, re-run the scan. It is the difference between reporting and remembering.
The bit that saved the most time
Running the scan also told me when to stop.
Template boilerplate across all five sites measured 3.1%–10.7% of body text, all inside the acceptable band. And the repeated sentences were almost entirely affiliate disclosures and pricing-disclaimer lines — site-wide boilerplate that Google does not treat as templated content.
So the correct action was not "rewrite 120 more pages." It was "close this and work on something that moves money." A negative result, measured early, is worth more than a positive one discovered late.
Built while running five comparison sites as a one-person operation. If you are doing the same, more notes at https://yongrui-services.pages.dev/
Top comments (0)