DEV Community

Ace Jogos do Rei
Ace Jogos do Rei

Posted on

Your old URLs still earn links: a 301 audit of a 15+ year-old blog

Somebody out there is still linking to a blog post you deleted years ago. Every reader who clicks that link lands on your 404 page, and any value the link could pass to your site goes nowhere.

I work on Jogos do Rei, a Brazilian online card game platform (Buraco, Tranca, Truco and others) that has been running since 2010. The blog is nearly as old. It started on Blogger, moved to WordPress, and changed its permalink structure more than once, so there are a lot of dead URLs behind it.

On September 28 I spent a day finding the dead URLs that still mattered and giving them proper 301s. Along the way I found a CSS bug that caused a 404 on every page view. This post covers what I did, how I checked it, and what I left alone on purpose.

Step 1: Inventory every URL that ever existed

Your sitemap only lists live URLs. To get the historical list, the Wayback Machine's CDX API is the cheapest option:

curl -s "https://web.archive.org/cdx/search/cdx?url=example.com&matchType=domain&fl=original&collapse=urlkey" > wayback_urls.txt
Enter fullscreen mode Exit fullscreen mode

For our domain this returned 40,005 URLs. After I dropped assets, search pages, API calls, player profiles and querystring variants, 2,862 candidate pages were left. I requested the 1,580 on the main and blog hosts with curl -L (nearly all of the other 1,282 sit on retired subdomains that no longer resolve in DNS). 473 returned 200 (directly or through an existing redirect) and 1,107 returned 404.

Most of those 1,107 don't deserve a redirect. The next step was finding the ones that do.

Step 2: Find evidence that a dead URL is still in use

The Wayback Machine shows what existed, not who links to it. I used three sources:

  1. External links. The Common Crawl web graph lists linking domains, not pages. It showed a few dozen domains linking to us in its 2026 releases. For each one, I looked for the specific page with the link and then checked the status of the exact href.
  2. The 404 log. The redirect plugin keeps a week of 404s. I dropped bots, my own requests and security probes, then read the referrers.
  3. History. Wayback capture counts are a rough popularity signal. One old Truco rules URL had 47 captures.

That gave me a short list. Two third-party pages, a Brazilian games blog and a card-game publisher's blog, linked to posts that now return 404. Eleven Tranca strategy posts now had a different date in their permalinks, with the same slug. In our case WordPress kept no record of the old date, so the old URLs weren't redirected and readers following old links got a 404. The rest were old URLs that kept appearing in the 404 log or had a strong Wayback history.

This approach has gaps. Common Crawl is a sample and misses real links. Search Console's URL Inspection reported the dead URL I knew had an external link as "URL is unknown to Google". I don't claim I found every link.

Step 3: Pick the destination by topic

Don't send every dead URL to the home page. Google tends to treat an off-topic redirect as a soft 404, and the reader doesn't get the article they clicked for.

My rule was to redirect only when a live page covers the same topic:

  • The old "history of Buraco" post that a games blog still links to → the current history of Buraco article.
  • The 11 Tranca posts → the same slug with its current date, a clean 1:1 mapping.
  • Old Truco rules posts → the Truco rules page. An old password-recovery post → the support page.

I fetched every destination live (200, no hop) and checked the <title> of the topic-matched ones.

Here's what I skipped: posts with no live equivalent (crosswords, 2011–2012 announcements, monthly ranking posts), about 270 Blogger-era URLs and about 130 old archive and tag pages that had no external links, and one legacy URL that the web server rejects before WordPress ever sees it.

The final batch was 29 redirects.

Step 4: Exact rules, no regex

The Tranca posts look like a job for a regex:

^/blog/\d{4}/\d{2}/\d{2}/(.+)$  ->  /blog/$1
Enter fullscreen mode Exit fullscreen mode

That regex creates a loop. WordPress already redirects /blog/<slug>/ to the current dated URL. The regex would also match the live dated URLs and send them to the undated form, and WordPress would send them straight back.

So every rule is an exact match, one per URL. They're all 301s, they ignore the querystring when matching (so ?utm_source=... still matches) and they match with or without a trailing slash. All of them share one title, which makes rollback easy.

Step 5: Precheck before writing anything

sources = {src for _, src, _ in rules}
for block, src, dst in rules:
    with_slash    = status(SITE + src)          # curl -s -o /dev/null -w '%{http_code}'
    without_slash = status(SITE + src.rstrip("/"))
    target        = status(SITE + dst)
    ok = (with_slash == "404" and without_slash == "404"
          and target == "200" and dst not in sources)
    print("OK" if ok else "XX", block, src)
Enter fullscreen mode Exit fullscreen mode

The precheck asked three questions about each rule:

  1. Does the source return 404 both with and without the slash? If not, something else is already handling it.
  2. Does the destination return 200 with no hop? If not, the new rule would create a redirect chain.
  3. Is the destination also a source? If so, it's a loop.

Step 6: Canary first, then verify every rule

I created one rule, the history-of-Buraco redirect, as a canary and verified it end to end. After that, a script created the rest one at a time through the plugin's REST API. It refused to run if a precondition failed. After each rule, it tested three variants:

variants = [src, toggle_trailing_slash(src), src + "?utm_source=x"]

for v in variants:
    target = SITE + dst                    # Location comes back absolute
    head  = curl_head(SITE + v)
    final = curl_follow(SITE + v)          # -L --max-redirs 5 -> code, url, hops
    ok = (head.status == 301
          and head.location == target      # exact, not "close enough"
          and final == (200, target, 1))   # one hop, ends on a 200
Enter fullscreen mode Exit fullscreen mode

The script also recorded the x-redirect-by header to show whether the plugin answered or something upstream did, and every result went to a JSONL log.

All 29 rules passed, 87 variants in total.

A final sweep over all 44 rules in the group (15 older ones plus the 29 new ones) found 0 regex rules and 0 destinations that were also sources. Every destination returned 200 with no hop. The undated Tranca slugs still reached the live posts in one hop.

I ran two more checks. A scan of all 309 posts and pages, in every status, found 0 internal links to the old URLs; internal links should be fixed at the source anyway. I also followed the real <a href> on both third-party pages, and both now reach a 200 in one hop.

Step 7: Backup and rollback

Before changing anything, I exported every redirect rule to JSON with a sha256. I exported again afterwards. To roll back, filter by the shared title and disable those rules. The old URLs go back to returning 404, which is where they started.

One sister site of ours had broken links in its own templates. Those are being fixed in that site's code, in a separate change that's still under review. There, a redirect would only be a safety net.

Bonus: one 404 per page view

The 404 log was full of requests for <post URL>&, with a trailing ampersand, and almost all had the page itself as the referrer. No visible link ended in &. The cause was in the <head>.

The theme has its own Custom CSS field in the Customizer, separate from core's Additional CSS. It printed this inside <style>:

.wp-block-search__button {
  background-image: url(&#039;images/lupinha.png&#039;) !important;
}
Enter fullscreen mode Exit fullscreen mode

The theme sanitized that setting with esc_attr, so every ' became &#039; on save. Browsers don't decode HTML entities inside <style>, so they read the URL literally. It's relative, so it resolves against the page. The # starts a fragment, which the browser drops. That leaves a request for <page URL>&:

from urllib.parse import urljoin, urldefrag

page = "https://example.com/blog/some-post/"
print(urldefrag(urljoin(page, "&#039;images/lupinha.png&#039;")).url)
# https://example.com/blog/some-post/&
Enter fullscreen mode Exit fullscreen mode

So every page view with the search button set off a full WordPress 404. The search icon broke too: the theme's style.css only covers the search button in the header, so the sidebar search button, with no other rule, lost its icon.

A redirect from .../& back to the page was tempting. I didn't add one. It would have swapped a 404 for a 301 plus a full HTML page on every view, and the icon would still have been broken.

The actual fix had two parts:

  • Data: I changed that one line to an absolute URL with no quotes, so the sanitizer has nothing to mangle, and backed up the old value first.
  • Cause: I changed the theme's sanitize callback from esc_attr to a wrapper around wp_strip_all_tags, which strips HTML but keeps quotes. Saving the field again won't bring the bug back. This shipped as a blog container deploy, and the previous image is still available for rollback.

To verify, I confirmed 0 occurrences of &#039; inside <style> and that the icon image returned 200. Headless Chromium loads of the home page, a post and a category page made 0 requests ending in &, and the icon showed up on both search buttons.

Checklist

[ ] Inventory historical URLs (Wayback CDX); drop assets, search, API, query variants
[ ] Check current status of each candidate
[ ] Evidence of use: exact external href, 404-log referrers, capture history
[ ] Pick the destination by topic; fetch it live (200, no hop) and check its <title>; skip if nothing matches
[ ] Exact-match rules only; no regex over CMS permalinks
[ ] Precheck: source 404 with/without slash; destination 200 with no hop; no dst in sources
[ ] Back up all rules (JSON + hash); give new rules a shared title
[ ] Canary one rule, then batch
[ ] Per rule: with slash, without slash, with querystring -> 301, exact Location, 1 hop to 200
[ ] Final sweep: 0 regex, 0 loops, all destinations 200 directly
[ ] Fix internal links at the source
[ ] Follow the real external <a href> and confirm it lands on a 200
[ ] Read 404 referrers: a page requesting itself is a bug, not a backlink
Enter fullscreen mode Exit fullscreen mode

The limits are real: free link data undercounts, and some dead URLs are best left dead. Still, a day of careful curl work was enough to get old links pointing somewhere useful again, with no paid tools.

Top comments (0)