DEV Community

Cover image for I Built a TypeScript SEO Content Checker in ~100 Lines
Mitu Das
Mitu Das

Posted on

I Built a TypeScript SEO Content Checker in ~100 Lines

I shipped a blog post with a 94-character title, two H1s, and no meta description. Nobody caught it, not me, not code review, not the linter. Google caught it for me in the form of a truncated, ugly search snippet.

SEO mistakes are just bugs, and bugs can be tested. In this post we'll build a small TypeScript SEO content checker that validates metadata, headings, keyword density, and readability, then run it in CI so bad content fails the build. Everything is copy-paste ready.

Step 1: Check the Metadata

The cheapest wins are title and description length. Search engines truncate both, so they're easy to validate with plain functions. Start with a shared type so every check returns the same shape:

// seo-check.ts
export interface Issue {
  level: 'error' | 'warn';
  rule: string;
  message: string;
}

export function checkMetadata(title: string, description: string): Issue[] {
  const issues: Issue[] = [];

  if (!title) {
    issues.push({ level: 'error', rule: 'title-missing', message: 'Title is missing' });
  } else if (title.length < 30 || title.length > 60) {
    issues.push({
      level: 'warn',
      rule: 'title-length',
      message: `Title is ${title.length} chars (aim for 30–60)`,
    });
  }

  if (!description) {
    issues.push({ level: 'error', rule: 'description-missing', message: 'Meta description is missing' });
  } else if (description.length < 70 || description.length > 160) {
    issues.push({
      level: 'warn',
      rule: 'description-length',
      message: `Description is ${description.length} chars (aim for 70–160)`,
    });
  }

  return issues;
}
Enter fullscreen mode Exit fullscreen mode

Returning plain data instead of throwing keeps the checker composable. You can print the issues, render them in a UI, or fail a build on them.

Step 2: Analyze the Heading Structure

Headings are where content quietly breaks. Common problems are multiple H1s, a missing H1, and skipped levels (an H2 jumping straight to an H4). A regex is fine for a quick checker. For production, use a real HTML parser like node-html-parser.

export function checkHeadings(html: string): Issue[] {
  const issues: Issue[] = [];
  const levels = [...html.matchAll(/<h([1-6])[\s>]/gi)].map((m) => Number(m[1]));

  const h1Count = levels.filter((l) => l === 1).length;
  if (h1Count === 0) {
    issues.push({ level: 'error', rule: 'h1-missing', message: 'Page has no H1' });
  } else if (h1Count > 1) {
    issues.push({ level: 'warn', rule: 'h1-multiple', message: `Page has ${h1Count} H1 tags` });
  }

  for (let i = 1; i < levels.length; i++) {
    if (levels[i] - levels[i - 1] > 1) {
      issues.push({
        level: 'warn',
        rule: 'heading-skip',
        message: `Heading jumps from H${levels[i - 1]} to H${levels[i]}`,
      });
    }
  }

  return issues;
}
Enter fullscreen mode Exit fullscreen mode

Run it on <h1>A</h1><h3>B</h3> and you get one warning: Heading jumps from H1 to H3.

Step 3: Keyword Density and Readability

Two content-level checks next. Keyword density tells you whether your focus keyword shows up at all (or way too much). Flesch Reading Ease gives a rough readability score.

const stripHtml = (html: string): string =>
  html.replace(/<[^>]*>/g, ' ').replace(/\s+/g, ' ').trim();

function countSyllables(word: string): number {
  const w = word.toLowerCase().replace(/[^a-z]/g, '');
  if (!w) return 0;
  if (w.length <= 3) return 1;
  const groups = w
    .replace(/(?:[^laeiouy]es|ed|[^laeiouy]e)$/, '')
    .replace(/^y/, '')
    .match(/[aeiouy]{1,2}/g);
  return groups ? groups.length : 1;
}

export function keywordDensity(text: string, keyword: string): number {
  const words = text.toLowerCase().split(/\s+/).filter(Boolean);
  if (words.length === 0) return 0;
  const phrase = keyword.toLowerCase().split(/\s+/);
  let hits = 0;
  for (let i = 0; i <= words.length - phrase.length; i++) {
    if (phrase.every((p, j) => words[i + j] === p)) hits++;
  }
  return (hits * phrase.length / words.length) * 100;
}

export function fleschScore(text: string): number {
  const sentences = Math.max(1, (text.match(/[.!?]+/g) ?? []).length);
  const words = text.split(/\s+/).filter(Boolean);
  if (words.length === 0) return 0;
  const syllables = words.reduce((sum, w) => sum + countSyllables(w), 0);
  return 206.835 - 1.015 * (words.length / sentences) - 84.6 * (syllables / words.length);
}

export function checkContent(html: string, keyword: string): Issue[] {
  const issues: Issue[] = [];
  const text = stripHtml(html);
  const wordCount = text.split(/\s+/).filter(Boolean).length;

  if (wordCount < 300) {
    issues.push({ level: 'warn', rule: 'thin-content', message: `Only ${wordCount} words (aim for 300+)` });
  }

  const density = keywordDensity(text, keyword);
  if (density === 0) {
    issues.push({ level: 'error', rule: 'keyword-missing', message: `"${keyword}" never appears` });
  } else if (density > 3) {
    issues.push({ level: 'warn', rule: 'keyword-stuffing', message: `Density is ${density.toFixed(1)}% (keep under 3%)` });
  }

  const flesch = fleschScore(text);
  if (flesch < 50) {
    issues.push({ level: 'warn', rule: 'hard-to-read', message: `Flesch score ${flesch.toFixed(0)} (60+ is easy to read)` });
  }

  return issues;
}
Enter fullscreen mode Exit fullscreen mode

The 3% ceiling and the Flesch cutoffs aren't laws of nature, just sane defaults. Treat them as tunable config, not gospel.

Step 4: Run It in CI

Now combine everything into a script that exits non-zero on errors. Warnings print but don't block the build.

// run-check.ts
import { readFileSync } from 'node:fs';
import { checkMetadata, checkHeadings, checkContent, Issue } from './seo-check';

const [file, keyword, title = '', description = ''] = process.argv.slice(2);
const html = readFileSync(file, 'utf8');

const issues: Issue[] = [
  ...checkMetadata(title, description),
  ...checkHeadings(html),
  ...checkContent(html, keyword),
];

for (const i of issues) {
  console.log(`${i.level === 'error' ? '✗' : '!'} [${i.rule}] ${i.message}`);
}

const errors = issues.filter((i) => i.level === 'error').length;
console.log(`\n${errors} error(s), ${issues.length - errors} warning(s)`);
process.exit(errors > 0 ? 1 : 0);
Enter fullscreen mode Exit fullscreen mode

Run it with npx tsx run-check.ts post.html "typescript seo" "My Title" "My description". Add that line to a GitHub Actions step and SEO regressions become failing builds instead of mystery traffic drops.

When to Stop Hand-Rolling

My checker is intentionally simple. Once you want Yoast-style scoring, such as keyphrase in the first paragraph, subheading distribution, and link checks, you're maintaining a rules engine. That's where I switched to an existing library for the scoring part. I used @power-seo/content-analysis because it's TypeScript-native:

import { analyzeContent } from '@power-seo/content-analysis';

const result = analyzeContent({
  title: 'TypeScript SEO Content Checker',
  content: '<h1>TypeScript SEO Content Checker</h1><p>Your article HTML...</p>',
  focusKeyphrase: 'typescript seo content checker',
});

console.log(result.score);           // numeric score
console.log(result.recommendations); // string[] of fixes
Enter fullscreen mode Exit fullscreen mode

That gave me a score and a list of recommendations in a few lines. I kept my own metadata and heading checks for the project-specific rules (like a house style on title length). If you want to see how this kind of checker fits into a broader workflow, this write-up on ccbd.dev goes deeper. Check the package docs for the current API, since options may change between versions.

What I Learned

  • Treat SEO as testable output: If a rule can be expressed as a function returning Issue[], it can run in CI.
  • Errors vs. warnings matter: Fail builds only on things that are objectively broken (no H1, no description). Everything else should be advisory, or the team will disable the check.
  • Regex is fine until it isn't: It works for headings and word counts, but switch to a real HTML parser before you trust it on messy CMS output.
  • Don't over-engineer scoring: Hand-roll the checks specific to your site, and use a library for generic content scoring.

If you want to try this approach, here's the repo: https://github.com/CyberCraftBD/power-seo

Your Turn

What's the most embarrassing SEO bug you've shipped to production, and would an automated check have caught it? Drop it in the comments. I'm also curious whether you'd run something like this at build time or as a pre-commit hook.

Top comments (0)