DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our brand voice rules are unit tests, including the one that bans the em dash

Nakodo (nakodo.app) has 63 hand-written search pages: guides, four free calculators, per-niche pages for YouTube, TikTok and Instagram, audience pages, blog posts and competitor comparisons. They are not CMS entries. They are TypeScript objects in src/content, listed in one registry.

That part is not unusual. What has turned out to matter is the file next to the registry: a test suite that reads every page and refuses the ones that are thin, duplicated, mis-titled, wrongly linked or off-voice. It has rejected more of my own writing than anyone else's.

Everything derives from the registry

// Every public page, in one place. The sitemap, llms.txt, llms-full.txt, the
// hubs, the footer and "read next" links all derive from these lists, so a new
// page can't be left out of any of them: a page missing from the sitemap still
// works, it just never gets found.
export const ENTRIES: Entry[] = [
  ...entries("guide", GUIDES),
  ...entries("tool", TOOLS),
  ...entries("niche", NICHES),
  ...entries("tiktok-niche", TIKTOK_NICHES),
  ...entries("instagram-niche", INSTAGRAM_NICHES),
  ...entries("audience", AUDIENCES),
  ...entries("post", POSTS),
  ...entries("alternative", ALTERNATIVES),
];
Enter fullscreen mode Exit fullscreen mode

The comment is the entire argument for doing it this way. A page that exists but is missing from the sitemap, the hub that should list it and the related links that should point at it is a page that renders perfectly and is never visited.

The route that renders a page refuses to be generous about a slug it does not know:

/**
 * The entry a route renders. Throws for an unknown slug: failing the build is
 * better than shipping a page with no title, schema or sitemap entry.
 */
export function requireEntry(kind: Kind, slug: string): Entry {
  const entry = byPath.get(`${KINDS[kind].base}/${slug}`);
  if (!entry) throw new Error(`No ${kind} page with slug "${slug}" in src/content`);
  return entry;
}
Enter fullscreen mode Exit fullscreen mode

The suite is generated, one describe block per page

for (const { path, page } of ENTRIES) {
  describe(path, () => {
    test("metadata fits the search result", () => { ... });
    test("answers first and has substance", () => { ... });
    test("ids, dates and links are valid", () => { ... });
    test("follows the voice rules", () => { ... });
  });
}
Enter fullscreen mode Exit fullscreen mode

A for loop around describe means the runner prints the failing page's own path, and a new page is in the suite the moment it is in the registry. No fixtures, no snapshot files, no test to remember to add.

Does it fit in a search result

const MAX_TITLE = 51; // " · Nakodo" makes 60
const MAX_DESCRIPTION = 158;

assert.ok(page.metaTitle.length <= MAX_TITLE, `title is ${page.metaTitle.length} chars`);
assert.ok(!page.metaTitle.includes("Nakodo"), "the layout appends the brand; don't repeat it");
assert.ok(page.metaDescription.length >= 90 && page.metaDescription.length <= MAX_DESCRIPTION, ...);
Enter fullscreen mode Exit fullscreen mode

The 51 is not a magic number, it is 60 minus the suffix the layout appends, and that arithmetic is in a comment so the next person does not "fix" it. The second assertion exists because I wrote "… · Nakodo" into a title by hand once and shipped a page whose tab said the brand name twice.

The description has a floor as well as a ceiling. A 40 character description is not a bug the typechecker can see, and it is a page that lets the search engine write its own snippet.

You can check the result from outside. The titles on the guide to finding YouTube influencers and the sponsorship cost calculator come to exactly 51 characters including the suffix, and the longest in the site is in the high fifties, which is the whole point of the cap.

Does it say anything

const n = words(page.summary);
assert.ok(n >= 40 && n <= 90, `summary is ${n} words`);
assert.ok(page.sections.length >= 3, "at least 3 sections");
const total = words([page.summary, ...bodyTexts(page)].map(plain).join(" "));
assert.ok(total >= 600, `only ${total} words: thin pages don't rank and don't help`);
assert.ok(page.faqs.length >= 3, "at least 3 FAQs");
for (const f of page.faqs) assert.ok(words(f.a) >= 15, `FAQ answer too short: ${f.q}`);
Enter fullscreen mode Exit fullscreen mode

The summary bounds are the house style written down: the first thing on the page is the answer in 40 to 80 words, not an introduction to the answer. The 600 word floor and the three FAQ minimum are the cheapest possible defence against the thing that actually happens when you write 63 pages, which is that pages 50 through 63 get thinner. plain() strips the inline markup first, so the count is words and not syntax.

Do the links go anywhere

function checkLink(href: string, from: string) {
  if (/^https?:\/\//.test(href)) {
    assert.match(href, /^https:\/\//, `${from}: outside links must be https (${href})`);
    return;
  }
  assert.ok(href.startsWith("/"), `${from}: internal links start with / (${href})`);
  const [path] = href.split("#");
  assert.ok(path === "/glossary" ? resolveLink(href) || href === "/glossary" : resolveLink(path), `${from}: link to unknown page ${href}`);
}
Enter fullscreen mode Exit fullscreen mode

Every inline link in every block of every page is resolved against the registry. An internal link to a page that does not exist fails the suite rather than the visitor, and a glossary link with an anchor has to name a term that is actually in the glossary. Same for the related list, which must have at least two entries, must not include the page itself, and must resolve.

There is also a date check that updated >= published, and one that reads better than it sounds:

if (page.howTo) {
  const s = page.sections.find((x) => x.id === page.howTo);
  assert.ok(s?.blocks.some((b) => b.type === "steps"), "howTo must name a section with a steps block");
}
Enter fullscreen mode Exit fullscreen mode

A page can mark one section as a real procedure, which renders as HowTo structured data. If that id points at a section with no steps in it, the page emits schema describing a procedure that is not on the page, which is exactly the sort of structured data that earns a manual action. The test makes the claim and the content inseparable.

Does it sound like us

This is the part I would not have predicted was testable.

// Banned in all copy (docs/brand.md section 5.6).
const BANNED = [/—/, /–/, /\bscrap(e|es|ed|ing|er|ers)\b/i, /\bhunt(s|ed|ing|er|ers)?\b/i, /\bsupercharg/i, /\bseamless/i, /\beffortless/i, /game-changing/i, /\bAI-powered\b/i, /unlock the power/i, /!/];
Enter fullscreen mode Exit fullscreen mode

Four kinds of rule in one array. The em dash and en dash, because the house style uses neither and they creep in from everywhere, including from my own editor. The marketing words, because "seamless" and "AI-powered" are what a page says when it has nothing to say. The exclamation mark, because the product's voice does not raise its own voice. And two words that are positioning rather than style: nothing we publish describes the product as scraping or hunting, since the whole discovery design is built on official APIs and public pages, and a word that implies otherwise in a guide is a promise the code does not keep.

Then the rule I like most:

// "Influencer" is the searcher's word, not ours: allowed only where the page
// meets the search (titles, description, H1, summary, keywords, FAQ questions).
const INFLUENCER = /\binfluencers?\b/i;

for (const t of bodyTexts(page).map(plain)) assert.ok(!INFLUENCER.test(t), `say "creator" in body copy: ${t.slice(0, 90)}`);
Enter fullscreen mode Exit fullscreen mode

People search for "influencer". The people being described call themselves creators. So the word is permitted exactly where the page has to meet the search, and banned in the body, where the page is just talking. One regex, split across two sets of fields, holds a distinction that would otherwise need a style guide nobody rereads.

Finally, across pages rather than within one:

test("titles, H1s and descriptions are unique across the site", () => {
  for (const key of ["metaTitle", "h1", "metaDescription", "label"] as const) { ... }
});
Enter fullscreen mode Exit fullscreen mode

With this many pages on neighbouring topics, two pages drifting into the same title is not hypothetical, and the assertion failure names both paths, which is what you need in order to fix it in one go.

The detail that keeps the gate shut

The checks read a page's text through a switch over the block union:

function blockTexts(b: Block): string[] {
  switch (b.type) {
    case "p":
    case "callout":
      return [b.text, ...(b.type === "callout" && b.title ? [b.title] : [])];
    case "list":
      return b.items;
    case "steps":
      return b.items.flatMap((s) => [s.title, s.text]);
    case "table":
      return [...(b.caption ? [b.caption] : []), ...b.head, ...b.rows.flat()];
    case "template":
      return [b.label, b.subject ?? "", b.body];
    case "terms":
      return b.items.flatMap((t) => [t.term, t.text]);
    case "tags":
      return b.items;
  }
}
Enter fullscreen mode Exit fullscreen mode

No default. The function is declared as returning string[], so if someone adds a ninth block type to the union, this switch no longer returns on every path and the build fails. Without that, the new block type would simply be invisible to the word count, the banned words and the link check, and the gate would quietly stop covering the newest content. An exhaustive switch is the difference between a test suite and a test suite that stays true.

The honest limit: the gate covers the pages in src/content. Hand-written surfaces outside it, including the app's own UI copy, do not get these assertions, and that is where my mistakes still reach production.

The publishing half

Each page carries an updated date, which feeds the sitemap, and the sitemap feeds the search engine notifications:

// No pages at all is a bug, not a site that deleted itself. Announcing every
// page as gone would ask the engines to drop the lot.
if (pages.length === 0) return { changed: [], removed: [] };
Enter fullscreen mode Exit fullscreen mode

The hourly job diffs the pages against a record in Redis of what was last announced at what date, submits the new and changed ones plus the removed ones, and records nothing when the submission is rejected so the next run retries. It reads the deployed sitemap rather than the repository, because this checkout is usually ahead of what is live and announcing a page that is not deployed yet sends a crawler to a sign-in page.

nakodo.app/sitemap.xml is the full list, 76 URLs, every one with a last-modified date that comes from the content object rather than from the deploy time. nakodo.app/llms.txt is the same registry rendered for language models, with each page's own summary line.

If you want the content rather than the machinery, the guide index is the hub, and the glossary is the 28 terms the guides are allowed to link to by anchor.

Earlier posts in this series: the public path allowlist that decides which of these pages a crawler may see at all, and the plan limits object that the pricing copy is generated from, for the same reason this content is generated from a registry.

Top comments (0)