DEV Community

rokya elbarbary
rokya elbarbary

Posted on Fully Autonomous

Test Your Redirect Map Before and After a Site Migration With a Small Python Script

Most site migrations do not fail at launch. They fail quietly over the following weeks, as old URLs that used to earn traffic and links start returning 404s, land on the homepage, or bounce through three redirects before reaching the right page.

The fix is boring and reliable: keep a redirect map, and test it automatically before and after launch. This post gives you a small standard-library Python script that does exactly that.

What the script checks

For every row in a CSV of old_url,expected_url it follows redirects one hop at a time and reports:

Verdict Meaning
OK One permanent redirect (301/308), or none, ending on the expected URL with HTTP 200
WRONG_TARGET The chain ends on a 200 page, but not the one you mapped (often the homepage)
CHAIN More than one redirect before the final page - fix the map so old URLs go straight to the destination
TEMPORARY The only redirect is a 302/303/307 - fine for campaigns, usually wrong for a migration
LOOP A redirect loop or more than 10 hops
ERROR A 4xx/5xx or a network error at the end of the chain

Following hops manually matters: if you let your HTTP client follow redirects automatically, you only see the final URL and miss chains and temporary redirects.

The script

#!/usr/bin/env python3
"""Check a redirect map after a site migration.

Input: a CSV with two columns, old_url,expected_url (header row required).
Output: one line per row: OK / WRONG_TARGET / CHAIN / LOOP / ERROR, plus the hops taken.
Standard library only.
"""
import csv
import sys
import urllib.error
import urllib.request
from urllib.parse import urljoin

MAX_HOPS = 10


class NoRedirect(urllib.request.HTTPRedirectHandler):
    def redirect_request(self, req, fp, code, msg, headers, newurl):
        return None  # do not follow; we record each hop ourselves


OPENER = urllib.request.build_opener(NoRedirect)


def hops_for(url):
    """Return a list of (status, url) pairs, following redirects one hop at a time."""
    hops, seen = [], set()
    while len(hops) < MAX_HOPS:
        if url in seen:
            hops.append(("LOOP", url))
            return hops
        seen.add(url)
        req = urllib.request.Request(url, method="HEAD", headers={"User-Agent": "redirect-check/1.0"})
        try:
            with OPENER.open(req, timeout=15) as resp:
                hops.append((resp.status, url))
                return hops
        except urllib.error.HTTPError as e:
            hops.append((e.code, url))
            location = e.headers.get("Location")
            if e.code in (301, 302, 303, 307, 308) and location:
                url = urljoin(url, location)
                continue
            return hops
        except (urllib.error.URLError, TimeoutError) as e:
            hops.append(("ERROR", f"{url} ({e})"))
            return hops
    hops.append(("TOO_MANY", url))
    return hops


def classify(hops, expected):
    last_status, last_url = hops[-1]
    redirects = [h for h in hops if h[0] in (301, 302, 303, 307, 308)]
    if last_status in ("LOOP", "TOO_MANY"):
        return "LOOP"
    if last_status == "ERROR" or last_status != 200:
        return "ERROR"
    if last_url.rstrip("/") != expected.rstrip("/"):
        return "WRONG_TARGET"
    if len(redirects) > 1:
        return "CHAIN"
    if redirects and redirects[0][0] != 301 and redirects[0][0] != 308:
        return "TEMPORARY"
    return "OK"


def main(path):
    problems = 0
    with open(path, newline="", encoding="utf-8") as f:
        for row in csv.DictReader(f):
            hops = hops_for(row["old_url"].strip())
            verdict = classify(hops, row["expected_url"].strip())
            problems += verdict != "OK"
            trail = " -> ".join(f"{status} {url}" for status, url in hops)
            print(f"{verdict:13} {trail}")
    print(f"\n{problems} problem(s) found")
    return 1 if problems else 0


if __name__ == "__main__":
    sys.exit(main(sys.argv[1]))
Enter fullscreen mode Exit fullscreen mode

Running it

Create redirects.csv:

old_url,expected_url
http://example.com/seo,https://example.com/seo
http://example.com/services/seo,https://example.com/seo
https://example.com/blog/old-post,https://example.com/blog/new-post
Enter fullscreen mode Exit fullscreen mode

Then:

python redirect_check.py redirects.csv
Enter fullscreen mode Exit fullscreen mode

Sample output (illustrative):

OK            301 http://example.com/seo -> 200 https://example.com/seo
CHAIN         301 http://example.com/services/seo -> 301 https://example.com/services/seo -> 200 https://example.com/seo
ERROR         404 https://example.com/blog/old-post

2 problem(s) found
Enter fullscreen mode Exit fullscreen mode

The script exits with code 1 when anything is wrong, so you can drop it into CI and fail a deployment that breaks the redirect map.

Where the redirect map comes from

  1. Export every URL that currently gets organic traffic or has backlinks (Search Console, analytics landing pages, your crawler of choice, backlink exports).
  2. Map each one to the closest equivalent new page. Only fall back to a parent category when there is genuinely no equivalent; mass-redirecting everything to the homepage is treated like a soft 404.
  3. Remove pages you are intentionally retiring from the map and let them return 404 or 410.

Before and after launch

  • Before: run it against staging (swap the hostnames in the CSV) and fix every non-OK line.
  • Launch day: run it against production immediately after DNS or routing changes.
  • Weeks 1-4: run it on a schedule and watch Search Console for 404 spikes and "Page with redirect" reports.
  • Update internal links so they point straight at the new URLs instead of relying on redirects.

Limitations

  • It uses HEAD requests; a few servers answer HEAD differently from GET. If you see odd 405 errors, change method="HEAD" to "GET".
  • It does not execute JavaScript, so client-side redirects are not detected - which is also a good reason not to rely on them for migrations.
  • Keep request volume polite on large maps (add a small delay or run it in batches).

Written by the team at Gorilla Vibe, a GCC-focused marketing-as-a-service agency. The script is standard library only; copy it, adapt it, and use it freely.

Top comments (0)