DEV Community

Cover image for I Crawled 4 More Dev Sites. 3 Had Every Tracked Redirect Point Back To Itself
Hermis
Hermis

Posted on

I Crawled 4 More Dev Sites. 3 Had Every Tracked Redirect Point Back To Itself

Continuing the site-crawl series, I ran the free RankForge crawler against four more sites developers touch daily — python.org, redis.io, rust-lang.org, and elastic.co — 10 pages each, starting from the homepage. This time the interesting finding wasn't orphan pages, it was redirects: on 3 of the 4 sites, nearly every tracked internal redirect target resolved back to the exact same URL it started from.

What I ran

No account needed, so this is reproducible:

  • POST https://rankforge.cc/api/analyzer with {"seed_url": "..."}, then poll GET /api/analyzer/<id> until status is completed.
  • The crawler follows internal links breadth-first from the seed, caps at 10 pages for the anonymous tier, and separately audits every internal link that returns a redirect status, following it to its final destination.
  • Raw JSON for all four runs: content/analyzer_20260924_<host>.json.

The self-redirect pattern: python.org and rust-lang.org

python.org: 9 internal links flagged as redirects, and 8 of the 9 resolved to a URL identical to the one that was linked — /about/apps redirects (301) to /about/apps, /community/irc redirects (301) to /community/irc, and so on. Average internal authority for the sample came back at a perfect 100 — the crawl itself is healthy, cross-linked, zero orphans — the anomaly is purely in how the redirect layer reports itself.

rust-lang.org was the cleanest case: all 9 of 9 tracked redirects resolved to their own identical URL — /tools/install → /tools/install, /learn → /learn, /governance → /governance. 100% of this site's tracked redirect surface is a same-URL 301.

This is almost certainly a trailing-slash or scheme-normalization artifact in each site's redirect middleware: the crawler requests a bare path, the server 301s to what it considers the canonical form, and once both ends of the analyzer's own URL-normalization strip the differing bit (a trailing slash, a scheme, a case difference), the before and after look identical in the report. It's not a real infinite loop — the browser lands on a working page in one hop — but it does mean a same-URL redirect chain fired at all on pages that gain nothing from it, which is 8-9 extra round-trips per crawl for no navigational benefit.

redis.io: the same pattern, plus one real 3-hop chain

redis.io had 8 tracked redirects, 7 of them self-redirects matching the same pattern as python.org and rust-lang.org. But one was different: /search3 redirects through 3 hops (308) before landing on /search — a genuine multi-step chain, not a same-URL artifact. Average internal authority for the sample was 28.17, pulled down by a smaller strongly-linked core (only /docs/latest hit the maximum score).

Three hops for a single internal link is the kind of thing that's invisible in a browser (it still resolves, instantly) but adds latency and crawl-budget waste at scale — every hop is a full extra request before a crawler or a client actually reaches content.

elastic.co: the control case

elastic.co reported zero redirects in the tracked sample — every internal link it crawled returned a direct 200. Average internal authority 43.95, with authority concentrated in a few top pages (homepage, /customers, and a top gated report) and a long tail of thinner pages, but no orphans and nothing broken in the redirect layer at all. Useful as a baseline: the self-redirect pattern on the other three isn't just "how redirect audits always look."

Why this keeps showing up

Across three installments of this series now (twelve sites total), two structural patterns keep recurring: internally orphaned content clusters (the first two rounds), and now same-URL self-redirects on the majority of a site's tracked redirect surface (this round, 3 of 4 sites). Neither is catastrophic on its own. Both are the kind of thing that's completely invisible from a normal page-by-page walkthrough and only shows up once something actually crawls the link graph and follows every redirect to its real destination.

Try it on your own site

Same checks (orphan pages, near-orphans, internal authority distribution, redirect-chain hops) run in one pass, no signup: rankforge.cc/audit.

Top comments (0)