DEV Community

Chen Tao
Chen Tao

Posted on

Why We Treat Static HTML as a Graph: 3 Insidious SSG Traps ESLint Can't Catch

Every production frontend pipeline has the usual suspects: TypeScript in strict mode, ESLint, Prettier, maybe a Playwright smoke test, and a Lighthouse run to ensure Core Web Vitals are green.

If the types check out and the bundle compiles, CI turns green and we ship.

Yet, after deploying hundreds of static pages, we kept running into a class of subtle, catastrophic failures that no type system, bundler, or unit test could ever detect:

  • Crawlers were silently dropping detail pages because they had become "orphan islands" despite rendering cleanly.
  • The alphabetically first dirty file in our repo had its "Last Updated" timestamp frozen in time for months due to a single-character string slicing bug.
  • Well-meaning contributors were accidentally leaking defensive, self-deprecating copy ("Why we don't know X yet") and raw backend enums (confidence: "unverified") directly into production headers.

When building Valor Mortis Guide—a high-density, fully static reference wiki and boss database for an upcoming first-person Soulslike—we stopped treating static pages as isolated React components. Instead, we started treating our build output as a directed graph of content invariants.

Here are three real-world, non-obvious engineering traps we hit in Next.js Static Site Generation (SSG), and how we built post-build graph linters to eradicate them.


Trap 1: The Invisible Git Porcelain Whitespace Bug (And the "Date Churn" Penalty)

In content-heavy static sites, maintaining an accurate "Last updated" date is critical for both user trust and search engine freshness.

Most developers do one of two things:

  1. The Lazy Way: new Date().toISOString(). This is fatal in production. Every time your CI runs a routine build, every single static HTML file gets stamped with today’s date. Search crawlers notice 100% of your URLs changing timestamps without a single byte of editorial diff, triggering freshness spam penalties.
  2. The Naive Git Way: Running git log -1 --format=%cs path/to/page.tsx.

The Git approach sounds clean until you deploy locally or in preview environments. A developer edits copy, runs next build, and previews the output—but because the change isn't committed yet, Git reports the file was modified three weeks ago.

So we wrote a prebuild script to derive modification timestamps using a simple rule: Working tree changes win over commit history. If a file is dirty in the working tree, use fs.statSync(file).mtime; otherwise, fall back to git log.

To detect dirty files, we ran:

git status --porcelain
Enter fullscreen mode Exit fullscreen mode

And that’s where the trap sprung.

The Slicing Glitch That Affected Exactly One Page

In Git's porcelain format, an unstaged modification is padded with a leading space:

 M src/app/about/page.tsx
 M src/app/bosses/page.tsx
Enter fullscreen mode Exit fullscreen mode

Notice the first character: a space (" M ").

A developer wrote a standard parser:

// ❌ The silent bug:
const stdout = execSync('git status --porcelain', { encoding: 'utf8' });
const lines = stdout.trim().split('\n');

for (const line of lines) {
  // Attempting to strip the 3-character status prefix (' M ')
  const filePath = line.slice(3);
  // ...
}
Enter fullscreen mode Exit fullscreen mode

Look closely at what stdout.trim() does.

stdout.trim() strips whitespace from the beginning of the entire output string! That means the leading space on line 1 is erased:

  • Line 1 becomes: "M src/app/about/page.tsx" (length 24)
  • Line 2 remains: " M src/app/bosses/page.tsx" (length 26)

When the code ran line.slice(3):

  • Line 2 was properly sliced into "src/app/bosses/page.tsx".
  • But Line 1 was sliced by 3 characters from "M s", resulting in "rc/app/about/page.tsx"!

Because "rc/app/..." didn't match any real file on disk, the script silently ignored it and fell back to the hardcoded repository launch date. The alphabetically first modified file was permanently stuck in the past.

The Fix: Multi-File Dependency Graphs

We also realized a page's freshness rarely depends on page.tsx alone. A page like our Valor Mortis Guide homepage or our boss database and combat mechanics imports external JSON data: game.config.json, bosses.json, combat.json.

If an editor updates a boss's parry timing in bosses.json, page.tsx is completely untouched by Git.

We replaced naive file checking with a mapped dependency graph:

// scripts/gen-content-dates.mjs
const ROUTE_DEPENDENCIES = {
  '/': ['src/app/page.tsx', 'src/data/game.config.json', 'src/data/steam.json'],
  '/bosses/': ['src/app/bosses/page.tsx', 'src/data/bosses.json'],
  '/combat/': ['src/app/combat/page.tsx', 'src/data/combat.json'],
};

// Process each line individually WITHOUT trimming the global buffer:
const statusLines = execSync('git status --porcelain', { encoding: 'utf8' })
  .split('\n')
  .filter(Boolean);

const dirtyPaths = new Set(
  statusLines.map(line => {
    // Preserve the exact 3-character porcelain column structure:
    const status = line.slice(0, 2);
    const path = line.slice(3).trim();
    return path;
  })
);
Enter fullscreen mode Exit fullscreen mode

Now, the timestamp is calculated from Math.max() across all dependencies of a route, and the working tree is respected without silent string corruption.


Trap 2: The "Island Page" & Dynamic Template BFS Crawl Trap

When using Next.js App Router with output: 'export', dynamic routes like src/app/bosses/[slug]/page.tsx require generateStaticParams().

If your data array contains 20 bosses, Next.js will cleanly export 20 static .html files into your /out directory. Your type checker passes. Your build finishes in seconds.

However, generating an HTML file does not mean anyone can reach it.

Consider this real-world UI scenario:
You have a hub page /bosses/. To make the interface slick on mobile, you introduce a client-side filter or a custom <select> dropdown to switch between encounters. When a user clicks, Next.js does a client-side router.push('/bosses/the-hollow-knight/').

What happens to a search crawler, a privacy browser with JavaScript disabled, or an RSS archiver?

  1. When they request the static HTML of /bosses/index.html, there is no literal <a href="/bosses/the-hollow-knight/"> tag in the raw server-rendered markup.
  2. Even though out/bosses/the-hollow-knight/index.html exists on the server, it has 0 inbound links in the static HTML topology.
  3. The page becomes an absolute orphan. Crawlers will never index it, and its PageRank remains exactly zero.

The Solution: A Post-Build Graph Crawler (audit-orphans.mjs)

Instead of hoping developers remember to include fallback semantic links, we created a post-build invariant audit that parses the exported /out directory as a directed graph:

// scripts/audit-orphans.mjs (runs on postbuild)
import { readdirSync, readFileSync } from 'node:fs';
import { parse } from 'node-html-parser';

// 1. Discover all exported HTML files
const pages = getAllHtmlFiles('out');
const inlinkMap = new Map(); // url -> Set of referring urls

// 2. Extract every literal <a href> from the static HTML
for (const page of pages) {
  const html = readFileSync(page.path, 'utf8');
  const root = parse(html);
  const anchors = root.querySelectorAll('a[href]');

  for (const a of anchors) {
    const target = normalizeUrl(a.getAttribute('href'));
    if (isInternalLink(target)) {
      if (!inlinkMap.has(target)) inlinkMap.set(target, new Set());
      inlinkMap.get(target).add(page.url);
    }
  }
}

// 3. Enforce Invariant Rules
for (const page of pages) {
  const inlinks = inlinkMap.get(page.url)?.size || 0;

  // RULE A: Absolute Orphans
  if (inlinks === 0 && page.url !== '/') {
    throw new Error(`[ABSOLUTE_ORPHAN] ${page.url} has 0 inbound links in exported HTML!`);
  }

  // RULE B: Island Pages (must have at least 2 distinct inbound paths)
  if (inlinks < 2 && page.url !== '/') {
    throw new Error(`[ISLAND_PAGE] ${page.url} only has ${inlinks} inlink. Minimum 2 required.`);
  }
}

// 4. BFS Reachability Check from Root ("/")
const visited = new Set(['/']);
const queue = ['/'];
while (queue.length > 0) {
  const current = queue.shift();
  for (const neighbor of getOutlinks(current)) {
    if (!visited.has(neighbor)) {
      visited.add(neighbor);
      queue.push(neighbor);
    }
  }
}

const unreachable = pages.filter(p => !visited.has(p.url));
if (unreachable.length > 0) {
  throw new Error(`[UNREACHABLE] ${unreachable.map(u => u.url).join(', ')} cannot be reached from homepage!`);
}
Enter fullscreen mode Exit fullscreen mode

If a developer builds a shiny client-side dropdown that omits crawlable <a href="..."> anchors from the hub HTML, CI immediately crashes with:
ENTITY_HUB_MISSING_LITERAL_LINK: /bosses/ requires raw HTML links to all 10 child encounters.


Trap 3: The "Trust Copy" Linter (Why We Ban Defensive Headings)

In technical documentation and game guides, engineers and writers share a subconscious defensive reflex: opening with what is missing.

We noticed drafts frequently contained sections like:

  • <h2>Why there is no multiplayer confirmed yet</h2>
  • <h3>What we do not know about weapon scaling</h3>
  • <p>This page might be out of date as the demo closed.</p>

Worse, data engineers frequently passed raw backend schema values directly to React children:

// ❌ Leaking internal schema doubts into UI:
<span>Confidence: {item.confidence}</span> 
// Renders: "Confidence: unverified" or "Confidence: community-corroborated"
Enter fullscreen mode Exit fullscreen mode

Why is this an architectural bug?

When a visitor lands on your page from a search engine, their eye tracks the <h1> and <h2> elements first. If your headings lead with self-deprecation ("Why we don't know...", "What is missing..."), the user immediately assumes the resource is abandoned or useless.

We realized this wasn't an editorial problem; it was an engineering boundary problem.

The 4-Part Truth Formula

We codified a rule for how uncertain facts must be presented:

  1. Answer First: State what is known before anything else.
  2. Primary Source + Verification Date: State who said it and when.
  3. Targeted Uncertainty Only: Restrict doubt to specific parameters ("The studio has not announced cross-play"), never the whole page.
  4. Actionable Next Step: Tell the reader where to look or what to expect.

To enforce this, we built scripts/audit-trust-copy.mjs into our build gate:

// scripts/audit-trust-copy.mjs
const NEGATIVE_HEADING_REGEX = /^(?:why (?:there (?:is|are) no|there's no)|what we (?:do not|don't|cannot|can't)|we do not publish)\b/i;
const UNFILTERED_ENUM_REGEX = /\bunverified\b|\bcommunity-corroborated\b/i;
const STALE_PAGE_REGEX = /\b(?:this|that|the) page\b[^.]{0,80}?\bout[- ]of[- ]date\b/i;

// Recursively scans src/app and flags infractions:
for (const file of sourceFiles) {
  const content = readFileSync(file, 'utf8');

  if (NEGATIVE_HEADING_REGEX.test(content)) {
    console.error(`[TRUST_VIOLATION] Found negative heading pattern in ${file}`);
    process.exit(1);
  }

  if (UNFILTERED_ENUM_REGEX.test(content)) {
    console.error(`[ENUM_LEAK] Raw internal status enum rendered in ${file}`);
    process.exit(1);
  }
}
Enter fullscreen mode Exit fullscreen mode

Now, instead of a heading reading:

"Why there is no official release time announced"

The writer is forced by CI to structure it as:

"Launch Schedule & Storefront Verification: Confirmed for October 13, 2026 across Steam and Xbox Game Pass; specific unlock hours are pending developer confirmation."


Moving Beyond Component-Level CI

Static site generation in 2026 is often treated as simple templating: throw some React components together, generate static HTML, and push to Cloudflare or Vercel.

But when you build high-density reference sites with hundreds of entity relationships, the real bugs don't happen in JavaScript execution. They happen in the graph topology:

  1. Date Invariants: Slicing strings from Git porcelain output can quietly freeze your freshness metadata.
  2. Graph Reachability: Generating static files is pointless if your static HTML link graph doesn't connect them via literal BFS paths.
  3. Editorial Invariants: Defensive copywriting and leaked database enums erode user trust before a single paragraph is read.

If you're maintaining a static documentation hub, reference wiki, or product catalogue, stop stopping at TypeScript. Write post-build graph linters that treat your entire exported directory as an immutable, verifiable state machine.


Curious to see what this looks like in production? You can explore the live architecture and reference implementation over at Valor Mortis Guide, or inspect how our boss and combat matrices are structured with zero orphan links.

Top comments (0)