DEV Community

Daniel Pertu
Daniel Pertu

Posted on

69 pages in one array, and 29 tests that fail if any of them is an orphan

Our site has 72 URLs in its sitemap. Sixty-nine of them come from one array.

curl -s https://pub-trivia.app/sitemap.xml | grep -c '<loc>'
# 72
Enter fullscreen mode Exit fullscreen mode

Before that array existed, this project had already taught me the lesson the hard way, at a scale of eight pages. Three lists answered three questions independently: which routes are public, which are in the sitemap, which are disallowed to crawlers. They disagreed. /about shipped in the sitemap while the auth gate bounced every crawler that followed the link to /login.

Nothing broke. No build failed, no page 404'd, no error appeared anywhere. The page simply did not rank, for a reason invisible in review.

Then the content plan committed to sixty more pages, each of which needs a title, a description, a canonical, a sitemap entry, a breadcrumb trail, a place in the navigation and links to its neighbours. Eight hand-kept lists instead of three, and the same class of disagreement at eight times the scale.

One node per page, and everything else is a derivation

export type ContentNode = {
  /** Root-relative, with a leading slash. The homepage is '/'. */
  path: string
  /** The SHORT title. The root layout's template appends the site name. */
  title: string
  /** The page's <h1>, which is allowed to be longer than the title. */
  heading: string
  /** What to call this page in a breadcrumb or a footer column. */
  shortLabel?: string
  /** The meta description, and the blurb wherever another page links here. */
  description: string
  cluster?: ClusterKey
  /** The date the content last changed, as YYYY-MM-DD. */
  updated: string
  /** Sideways links: siblings, plus at least one page in another cluster. */
  related: readonly string[]
}
Enter fullscreen mode Exit fullscreen mode

Adding a page is one node and one page.tsx. The sitemap entry, the canonical, the <title>, the breadcrumb trail, the hub index listing, the footer column, the related links block and the link preview card all follow from the node.

Two fields are worth dwelling on.

description is deliberately one string used in two places: the meta description, and the line under this page's name wherever another page links to it. The pages that predate the registry had three copies of that sentence and two of them had already drifted, which means the site was describing a page one way to Google and another way to a reader.

shortLabel exists because of exactly one problem. "Frequently Asked Questions" is the right <title>, since it is the phrase people search for, and the wrong footer link, where it is three times the width of everything around it. Only the handful of pages with that tension carry one, and the rest default to their title.

Declared is not published

A cluster gets declared when the plan commits to it, which is usually a phase or two before anything is written. What decides whether it appears in the header, the footer and the sitemap is whether its hub page exists:

export function publishedClusters(): Cluster[] {
  return CLUSTERS.filter((cluster) => BY_PATH.has(cluster.path))
}
Enter fullscreen mode Exit fullscreen mode

So a half-built cluster cannot leak into the navigation as a dead link, and finishing the hub publishes it everywhere at once. The corresponding house rule is blunter: the registry describes pages that exist. A node with no page behind it is a 404 in your sitemap and a dead link in your footer, so nodes land in the same commit as their page.

The graph, and the twenty-nine assertions on it

The derived link graph is one function, and it returns what the components actually render, because the components take their links from it:

export function outboundLinksFor(path: string): string[] {
  // breadcrumb ancestors, a hub's spokes, and the node's own related links
}
Enter fullscreen mode Exit fullscreen mode

That is what the tests run against. Here is the useful half of them, in plain English:

  • No duplicate paths, and no duplicate titles, because two pages with the same title are two pages competing for the same search result.
  • Every description within 160 characters, where a search result truncates. Past that, the end of the sentence is written for nobody, and since the description is also the blurb in the footer and on the hub, a long one is a ragged wrapped line in three places rather than one.
  • Every rendered title within 60 characters. The thing a crawler sees is what the layout's %s | PubTrivia template produces, so the budget has to be measured after the template is applied, not against the short title on the node.
  • Every internal link points at a page that exists, and no page links to itself.
  • Every spoke links up to its hub, every hub links down to all of its spokes, and every spoke's path sits underneath its hub's path.
  • Breadcrumbs for a spoke are exactly Home, hub, page, and the homepage has no trail at all.
  • Every content page is in the sitemap, and every content page is reachable without a session. That last one is the original bug, encoded as an assertion.
  • The sign-in pages are crawlable but not in the sitemap. They stay fetchable, because a crawler following a link to them should get the page rather than a redirect, but a sitemap is a statement about which pages we would like ranked, and a login form is not one of them.
Test Files  2 passed (2)
     Tests  29 passed (29)
  Duration  261ms
Enter fullscreen mode Exit fullscreen mode

A quarter of a second, no network, no browser, no build. These are the cheapest tests in the repo and they catch the bugs that are hardest to see.

The orphan test is the one with a subtlety in it

it('leaves no page orphaned', () => {
  // The header and footer are on every page, so what they link to counts.
  const inbound = new Set<string>(globalNavPaths())
  for (const node of CONTENT) {
    for (const path of outboundLinksFor(node.path)) inbound.add(path)
  }
  for (const node of CONTENT) {
    expect(inbound.has(node.path), `nothing links to ${node.path}`).toBe(true)
  }
})
Enter fullscreen mode Exit fullscreen mode

Seeding the inbound set with the global navigation is not a loophole, it is the point. A hub with no editorial link pointing at it is still reachable from every page on the site, and that is precisely why the hubs are in the footer. Leave the nav out and the test demands fake editorial links to satisfy it, which is how a test starts producing worse pages than no test.

The complementary assertion is that every page links somewhere. A page with no outgoing links is a dead end, and a reader who reaches one leaves the site rather than going deeper.

What the tests cannot catch, and what that implies

Nothing stops a page rendering links the registry does not know about. You cannot assert the absence of a hardcoded anchor tag from a test that only reads data.

What the tests do catch is a page rendering fewer links than the registry promises, which is the failure worth catching, and the reason the components are built to read from the registry rather than to hold their own lists. The test is only meaningful because of that discipline. It verifies the data, and the architecture is what makes the data a description of the pages.

That is the general shape of this kind of testing, and it is worth being clear-eyed about: these assertions do not prove the site is good. They prove it is coherent, which is the part humans are worst at maintaining across sixty pages and six months.

Have a look at the output

  • pub-trivia.app/sitemap.xml is the array, serialised. Every entry is a page that exists and is reachable without a session.
  • pub-trivia.app/guides is a hub whose index, structured data and breadcrumbs are all derived from the same query. Click into any guide and the breadcrumb is Home, Guides, page, with no hand-written trail anywhere.
  • Pick any page and check its footer. The columns are generated, and a column that would hold one link is not emitted, because "Resources" over a single link reads as a bug.
  • pub-trivia.app/features is what all those pages are selling, if you want the product rather than the plumbing. The first quiz session is free and needs no card.

Top comments (0)