DEV Community

Cover image for 4 More Big Tech Sites Crawled. Only 1 Had Self-Redirects, But All 4 Had Zero Nofollowed Outbound Links
Hermis
Hermis

Posted on

4 More Big Tech Sites Crawled. Only 1 Had Self-Redirects, But All 4 Had Zero Nofollowed Outbound Links

Continuing the site-crawl series, I ran the free RankForge crawler against four large tech platforms — stripe.com, vercel.com, cloudflare.com, and github.com — 10 pages each, starting from the homepage. This round the same-URL self-redirect pattern from the last installment only showed up on one site, not three, but a second, unrelated finding was consistent across all four: not a single outbound external link on any of the 40 crawled pages was marked nofollow.

What I ran

No account needed, so this is reproducible:

  • POST https://rankforge.cc/api/analyzer with {"seed_url": "..."}, then poll GET /api/analyzer/<id> until status is completed.
  • The crawler follows internal links breadth-first from the seed, caps at 10 pages for the anonymous tier, audits every internal redirect to its final destination, and separately classifies every outbound external link as followed or nofollowed.
  • Raw JSON for all four runs: content/analyzer_20260925_<host>.json.

cloudflare.com: 10 of 11 tracked redirects point back to themselves

Same pattern documented in the last installment (python.org, rust-lang.org, redis.io): 10 of 11 internal redirects the crawler followed on cloudflare.com resolved to a URL identical to the one that was linked — /plans 301s to /plans, /careers 301s to /careers, /resource-hub 301s to /resource-hub, and seven more. The one exception was a genuine cross-subdomain hop: /blog redirects (2 hops, 301) to blog.cloudflare.com/ — a real destination change, not a same-URL artifact.

stripe.com, vercel.com, and github.com reported zero tracked redirects in the sampled crawl — a useful reminder that this isn't a universal artifact of how the redirect audit works, it's specific to how individual sites' routing layers are configured.

All 4 sites: 0% of external links are nofollowed

The more consistent finding this round was in the external-link audit, not the redirect audit. Across all four sites and all 40 crawled pages combined (565 total outbound external links: 0 stripe, 162 vercel, 147 cloudflare, 256 github), every single one was dofollow — total_nofollowed_external came back 0 on every site.

That includes links handed to genuinely separate properties, not just self-references: vercel.com sent 27 dofollow links to github.com, cloudflare.com sent 36 dofollow links to its own dash.cloudflare.com login portal, and github.com sent 60 dofollow links to docs.github.com. None of it was nofollowed, even where a stricter policy (say, nofollowing a login/dashboard CTA that isn't meant to pass ranking signal) would be a defensible choice.

It's a small sample — 4 sites, 10 pages each — but it lines up with a pattern worth checking on your own site: rel="nofollow" on external links is usually applied deliberately (UGC, sponsored, low-trust destinations), and its complete absence here just means these four teams either don't have a nofollow policy for external links or don't need one at their content mix. Either way, it's not something you'd catch without a crawler that actually reads the rel attribute on every outbound anchor.

Why this keeps showing up

Four installments into this series now (sixteen sites total), the recurring theme isn't one specific bug — it's that structural link-graph properties (orphaned pages, self-redirecting 301s, nofollow policy) are invisible from a normal page-by-page walkthrough and only surface once something actually crawls the link graph and reads every link's real destination and attributes.

Try it on your own site

Same checks (orphan pages, near-orphans, internal authority distribution, redirect-chain hops, external nofollow audit) run in one pass, no signup: rankforge.cc/audit.

Top comments (0)