DEV Community

Cover image for My regex inserted 116 links. Six of them landed inside the title tag.
孙永瑞
孙永瑞

Posted on

My regex inserted 116 links. Six of them landed inside the title tag.

I wrote a script to add 116 vendor pricing links across 48 HTML pages. It ran clean. No exceptions, no warnings, and the diff looked exactly like what I asked for.

Then my title-length check failed on six pages.

What the pages actually looked like

Here is a real one, reproduced from the broken output:

<title>Airtable vs<a href="https://asana.com/pricing" target="_blank"
rel="noopener noreferrer"> Asana: 2026 C</a>omparison</title>
Enter fullscreen mode Exit fullscreen mode

The word "Comparison" got cut in half. An anchor tag opened in the middle of it, closed, and the rest of the word trailed after. On a sixth page the <title> tag itself was no longer well-formed at all.

Every one of those six pages still rendered perfectly in a browser.

That is the part I want to emphasize, because it is why this survived as long as it did. <head> is not painted. Nothing I could click through would have shown me this. The pages looked fine, the links worked, and the only thing that noticed was an unrelated gate that measured title length and complained that six titles were now the wrong size.

The bug

I was locating table cells with nested re.finditer: one pass over the document to find tables, then an inner pass over each table to find cells.

for tb in re.finditer(TABLE_RE, doc):
    for cell in re.finditer(CELL_RE, tb.group(1)):
        pos = cell.start(1)          # <-- wrong
        insert_link(doc, pos)
Enter fullscreen mode Exit fullscreen mode

.start(1) on the inner match is an offset relative to tb.group(1), not relative to the whole document. I knew that on some level and still forgot to add the outer table's own start:

        pos = tb.start(1) + cell.start(1)   # correct
Enter fullscreen mode Exit fullscreen mode

Without that term, every position resolved as if each table began at character zero. So the offsets pointed near the top of the file — which, in an HTML document, is <head>. The script then did precisely what it was told and inserted links there. On six pages, "there" happened to be inside <title>.

Why I consider this a near miss and not a funny story

The failure mode is invisible by construction. A broken link, a missing element, a 500 — those surface immediately. An edit that lands in a region the renderer ignores produces a document that is silently wrong until something unrelated measures it.

I already had a rule from an earlier incident: strip <head> before evaluating visible text. That rule was about reading. This was the same trap on the writing side, and I had not carried the lesson across.

So I stopped fixing the arithmetic and added a guard instead:

def guard_outside_head(doc, pos):
    head = re.search(r"<head[^>]*>.*?</head>", doc, re.S | re.I)
    if head and head.start() <= pos < head.end():
        log_warning("INSERT_REJECTED_IN_HEAD", pos)
        return False
    return True
Enter fullscreen mode Exit fullscreen mode

Any insertion landing inside <head> is now discarded and logged. The offset fix makes the guard redundant on this particular code path; I kept both, because the offset bug is one specific way to compute a bad position and the guard covers all of them.

The rollback that made it worse

Before fixing anything I tried to restore the 48 pages from a backup I had made:

cp -r _backups/x/. pmcompared/guides/
Enter fullscreen mode Exit fullscreen mode

That trailing /. copies the contents of x, and x already contained a guides directory. So I got guides/guides/. Now I had two problems instead of one.

The cleaner restore was to stop treating my backup as authoritative and use the thing that actually tracks state:

git checkout -- guides/
Enter fullscreen mode Exit fullscreen mode

One command, exact restoration, no nesting. Git is the source of truth and should have been the first move, not the fallback.

Verifying the fix

Rather than check that the six known-bad pages looked right, I compared every page against the committed baseline:

  • Titles changed vs. HEAD: 0. Not "six fixed" — zero drift anywhere.
  • <a href inside <head>: 1 remaining. I pulled the committed version with git show HEAD:<path> and confirmed that one was pre-existing in the original commit, not mine.

That second check mattered more than it sounds. "One left" could mean my guard missed one. Checking it against HEAD is what turned it from a possible bug into a known, pre-existing artifact — and I noted it as a separate item rather than quietly deleting it.

What I'd take from it

  1. Nested regex offsets are relative to the enclosing match. If you compute a position with re.finditer inside another re.finditer, add the outer start every time. Better: stop computing document offsets from nested matches and slice once at the end.
  2. An edit that lands in a non-rendered region is invisible. <head>, HTML comments, display:none containers, JSON-LD blocks — if your tooling writes into documents, add an explicit guard rejecting those regions rather than trusting the arithmetic.
  3. A check that catches your bug by accident is still a check worth having. The title-length gate was not written for this. It paid for itself anyway.
  4. Restore from version control first. My own backup copy was the thing that introduced a second failure.

The uncomfortable part is not the bug. It is that I only found it because an unrelated assertion happened to measure the right thing. Everything else I looked at said the run was clean.


If you've got batch edits running against markup that nobody looks at directly, I fix those: https://yongrui-services.pages.dev/

Top comments (0)