DEV Community

SwiftKit
SwiftKit

Posted on

A small SEO crawler in Node: the 20 checks worth running on every page (and a score that's easy to explain)

Most technical SEO problems are boring and findable: a template that drops the meta description, a staging noindex that shipped to production, 40 pages with the same title, links to pages that were deleted last year. A crawler finds all of them in minutes.

Here's the one I built, what it checks, and what it found on a real docs site.

Crawl politely

  • Start at the home page, follow links on the same site (subdomains included), skip files (.pdf, .jpg, .zip…).
  • Read robots.txt first and skip disallowed paths (longest matching rule wins).
  • A few requests at a time, an honest User-Agent, and a hard page limit.

Parse with linkedom. It's a fraction of jsdom's weight and plenty for reading tags:

import { parseHTML } from 'linkedom';

const t0 = Date.now();
const res = await fetch(url, { headers: { 'User-Agent': 'my-seo-audit/1.0' }, redirect: 'follow' });
const responseTimeMs = Date.now() - t0;
const html = await res.text();
const { document } = parseHTML(html);

const meta = (k) => document.querySelector(`meta[name="${k}" i], meta[property="${k}" i]`)?.getAttribute('content')?.trim() || null;
const page = {
  title: document.querySelector('title')?.textContent.replace(/\s+/g, ' ').trim() || null,
  description: meta('description'),
  h1: [...document.querySelectorAll('h1')].map((h) => h.textContent.trim()).filter(Boolean),
  canonical: document.querySelector('link[rel="canonical" i]')?.getAttribute('href') || null,
  robots: (meta('robots') || '') + ' ' + (res.headers.get('x-robots-tag') || ''),
  imagesWithoutAlt: [...document.querySelectorAll('img')].filter((i) => !i.hasAttribute('alt')).length,
  viewport: !!meta('viewport'),
  lang: document.documentElement.getAttribute('lang'),
};
Enter fullscreen mode Exit fullscreen mode

Note the X-Robots-Tag header: a noindex there doesn't show up in the HTML at all, and it's a classic leftover from staging.

The checks

Errors (the page is hurt):

  • missing <title>
  • noindex in meta robots or the X-Robots-Tag header
  • links to pages that return 4xx/5xx

Warnings (fix soon):

  • title over 60 characters or under 15
  • missing meta description
  • no <h1>
  • the same title on several pages
  • images without alt
  • under 200 words
  • server response over 1.5 s
  • no mobile viewport tag
  • plain HTTP

Notices (nice to fix):

  • meta description over 160 or under 50 characters
  • more than one H1
  • duplicate meta descriptions
  • canonical pointing elsewhere, or missing
  • no lang
  • no Open Graph tags
  • HTML over 1 MB
  • redirects

Broken links without hammering the site

Every page you crawl already gives you a status code. For links you didn't crawl (because of the page limit, or because they point elsewhere), check them once at the end: HEAD first, and GET if the server answers 405/403 to HEAD (many do). Cap the count, and keep external links optional.

A score people understand

I tried weighted formulas. Clients didn't trust them. What works is a checklist score: every page starts at 100, and each error costs 15, each warning 6 and each notice 2. Anyone can recompute it from the issue list, and fixing an issue visibly moves the number. It's a hygiene score, not a guess at Google's ranking, and I say so.

What it found on a real site

On 30 pages of a well-maintained open-source docs site:

  • 27 titles over 60 characters. The template appends a long tagline to every page title.
  • 24 pages sharing one meta description. It's the site-wide default.
  • 10 pages with the same title, all versioned copies of one docs page.
  • 1 page without an H1, 2 thin pages, no broken links.
  • Average score: 90.

Nothing dramatic, but each is a one-line template fix. That's the typical result: a handful of template issues multiplied across hundreds of pages.

The packaged version

I put this on Apify as SEO Audit Crawler. You get one row per page with the score and issues, plus a free site summary: issues ranked by pages affected, every broken link with the page it's on, duplicate groups, non-indexable pages and the slowest pages. 40 pages take about 15 seconds.

What's the most common SEO bug you keep finding on client sites?


Written with AI assistance; the code and numbers were tested before publishing.

Top comments (0)