pub-trivia.app publishes 72 URLs. Sixty-nine of them are content pages described by a single array of nodes, and three are legal documents that are not.
That split is the whole post, but it took measuring the thing twice to see it.
One field does the sideways linking
Each node in the content registry carries a related array:
{
path: '/features/team-scoring',
title: 'Team and Individual Scoring',
shortLabel: 'Team scoring',
description: 'Team mode gives a table one shared score and ranks tables, ...',
cluster: 'features',
updated: '2026-09-14',
related: [
'/features/live-leaderboard',
'/features/question-timer',
'/solutions/social-clubs',
],
}
The convention is siblings in the same cluster plus at least one page in a different one, so a reader going deep into features is always one click from the venue pages or the guides. The link up to the hub is deliberately not in there, because it is structural: it comes from the breadcrumb trail instead.
The block at the foot of the page takes paths and looks the titles and descriptions up rather than being handed link text. That is the point of it being a path list: the words describing a page are written once, on that page's node, so every link to it says the same thing, and a path with no node renders nothing instead of a link to a 404.
The graph the tests run on
One function assembles the full set, and it has to match what the components actually render:
export function outboundLinksFor(path: string): string[] {
const node = nodeFor(path)
if (!node) return []
const links = new Set<string>()
for (const crumb of breadcrumbFor(path)) {
if (crumb.path !== path) links.add(crumb.path)
}
if (node.cluster && isHub(node)) {
for (const spoke of spokesOf(node.cluster)) links.add(spoke.path)
}
for (const related of node.related) links.add(related)
links.delete(path)
return [...links]
}
Breadcrumb ancestors, a hub's spokes, and the page's own related list. Run it over all 69 nodes and the numbers are these:
- 405 edges, 5.9 outbound links per page on average
-
219 of those come from
related, and 98 of the 219 are reciprocal, which is 44.7 per cent - the most linked-to pages are the hubs, at 17 for features and 16 for solutions
-
/has 68 inbound, one from every other page, because the first breadcrumb is always Home - 9 pages have exactly 1 inbound link in this graph: eight spokes whose only pointer is their own hub's list of pages, plus the about page, which one FAQ entry relates to
The reciprocity number is the one worth not optimising. A page about scoring systems should link to the guide on writing questions; the guide on writing questions has three better things to link to than scoring systems. Pushing that 44.7 per cent towards 100 would mean filling the related arrays with links chosen to balance a matrix rather than to be followed, which is the opposite of what the field is for.
What matters more is the floor. Two assertions in the registry test cover it:
it('leaves no page orphaned', () => {
// The header and footer are on every page, so what they link to counts.
const inbound = new Set<string>(globalNavPaths())
for (const node of CONTENT) {
for (const path of outboundLinksFor(node.path)) inbound.add(path)
}
for (const node of CONTENT) {
expect(inbound.has(node.path), `nothing links to ${node.path}`).toBe(true)
}
})
it('gives every page somewhere to go', () => {
for (const node of CONTENT) {
expect(outboundLinksFor(node.path).length).toBeGreaterThan(0)
}
})
Counting the global navigation as inbound is a judgement, not a loophole: a hub with no editorial link pointing at it is still one click from every page on the site, and being in the footer is exactly why. Twelve pages have no related entry anywhere pointing at them. Four of those are the homepage and three hubs, which the navigation carries, and the other eight are spokes whose only pointer is the list of pages on their own hub, which is why the test counts a hub's output as well as its related array.
Both assertions pass. So the generated half of the site is sound, and I believed that was the end of it.
Then measure the deployed HTML
The registry is a model of the site. The thing search engines read is the HTML. So: fetch all 72 URLs from the sitemap, pull every root-relative href, and see which links really appear where.
curl -s https://pub-trivia.app/sitemap.xml \
| grep -o '<loc>[^<]*' | cut -c6- > urls.txt
while read -r url; do
curl -s "$url" | grep -o 'href="/[a-z/-]*"' | sort -u > "$(basename "$url").links"
done < urls.txt
Two things came out of that which the registry cannot tell you.
The first is a count of the chrome. Eleven href values appear on all 72 pages, and once you drop the favicon, the manifest and the two bundles, the ones that are links to pages are: the homepage, the about page, and the three legal documents. Everything else that looks site-wide is not quite.
Because the second finding is that thirteen content paths appear on 69 of the 72 pages, at 68 inbound links each. Those are the six cluster hubs, pricing, the FAQ, and five of the eleven venue pages, which the footer lists by name rather than folding them behind their hub, because "is there a page for my kind of venue" is the question that sends a landlord looking. Sixty-nine, not seventy-two.
The three pages
The missing three are the legal documents. The cookie policy is the quickest one to open, and if you scroll to the bottom you will find a different footer from the one on the rest of the site: a copyright line and four company links. No hubs, no venue columns, no breadcrumb, because there is no registry node to build one from.
Its complete set of root-relative href values, in the deployed HTML, is five: the homepage, the about page, the other two legal documents, and itself.
So three URLs we submit in the sitemap sit at the edge of the site with a single link into the content: the wordmark in the header, which goes home. Everything the registry generates is one click from everything else, and the pages the registry does not describe are not.
This is survivable, and partly on purpose. Nobody is trying to rank a terms of service, and the three documents were given a narrow light-theme layout so they read like documents instead of like marketing pages. But the two places where that decision shows up are not the same size. One is "the privacy policy has a plainer footer". The other is "the privacy policy is where a crawler's path through the site ends", and nothing in the test suite can object, because the assertions all iterate CONTENT, and these three pages are appended to the sitemap list by hand, immediately after it:
export const INDEXABLE_ROUTES: readonly string[] = [
...INDEXABLE_CONTENT_PATHS,
'/legal/privacy',
'/legal/terms',
'/legal/cookies',
]
The derivation stops at exactly that line, and so does the graph. That is the general shape of the lesson, and it is not really about legal pages: the invariants you generate hold over the set you generated, and the exceptions you append by hand are outside every test you wrote about it.
Worth doing on your own site
The registry measurement took one script and the live measurement took one more, and only the second one found anything. If you generate your internal links from a content model, run the crawl anyway, then diff the two. The interesting rows are the pages in your sitemap that your model has never heard of.
If you want to see the generated half working, the guides hub lists fourteen pages and marks up the same fourteen in its ItemList, and every one of them carries a "Keep reading" block at the bottom built from its own related array. The solutions hub is the long one, eleven venue types, five of which are the names in the footer.
Top comments (0)