DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our llms-full.txt is 574,626 bytes generated from the same array that renders the pages

Nakodo has 76 public pages. 63 of them are rendered by eight dynamic routes, because the page itself is not a template: it is a TypeScript object in src/content, and the route picks one and hands it to a single renderer.

That decision was made for the usual reasons: one renderer, one set of tests, no MDX build step. The payoff I did not plan for arrived when we added llms.txt and llms-full.txt, the two files llmstxt.org proposes for assistants reading a site. Writing them by hand would mean a second copy of every page, which starts accurate and is wrong within a fortnight. Generating them from the objects took an afternoon.

One page is one object

Every search intent page, which is to say every guide, niche page, audience page and tool page, has the same shape:

export type ContentPage = {
  slug: string;
  label: string;        // breadcrumbs, cards, footer links
  h1: string;
  metaTitle: string;    // at most 50 chars; the layout appends " ยท Nakodo"
  metaDescription: string;
  keywords: string[];
  cardBlurb: string;
  summary: string;      // the answer, 40 to 80 words
  published: string;
  updated: string;
  sections: Section[];
  faqs: Faq[];
  related: string[];
  howTo?: string;       // the section whose steps are a real procedure
};
Enter fullscreen mode Exit fullscreen mode

A section is an id, a heading and a list of blocks, and a block is a union of eight shapes:

export type Block =
  | { type: "p"; text: Rich }
  | { type: "list"; items: Rich[]; ordered?: boolean }
  | { type: "steps"; items: { title: string; text: Rich }[] }
  | { type: "table"; caption?: string; head: string[]; rows: Rich[][] }
  | { type: "template"; label: string; subject?: string; body: string }
  | { type: "callout"; title?: string; text: Rich }
  | { type: "terms"; items: { term: string; text: Rich }[] }
  | { type: "tags"; items: string[] };
Enter fullscreen mode Exit fullscreen mode

Across the 63 pages that is 385 sections, 853 blocks and 306 FAQs. The distribution is lopsided in a way I find reassuring: 524 paragraphs, 139 lists, 53 term lists, 50 tables, 26 tag rows, 22 step lists, 20 copyable templates, 19 callouts. If a block type had two uses, it would be a sign the type was invented for one page rather than for the writing.

Rich text is two tokens, on purpose

Rich is just string. The inline markup is one regex:

export const RICH_TOKEN = /\[([^\]]+)\]\(([^)\s]+)\)|\*\*([^*]+)\*\*/g;
Enter fullscreen mode Exit fullscreen mode

Links and bold. Nothing else. The comment above it is the whole reason:

// The inline markup content is written in: [label](href) and **bold**. Kept to
// two forms so the same string is valid markdown for llms-full.txt.
Enter fullscreen mode Exit fullscreen mode

That is the trick the markdown generator rests on. The page renderer walks the tokens and emits <a> and <strong>; the markdown generator emits the string unchanged, because the string was already markdown. There is no serialiser, no escaping pass, no HTML to markdown conversion anywhere in the codebase. The two consumers that need the text without markup, meta tags and structured data, call the same regex with a different replacer:

export function plain(text: Rich): string {
  return text.replace(RICH_TOKEN, (_, label, _href, bold) => label ?? bold ?? "");
}
Enter fullscreen mode Exit fullscreen mode

The generator is one switch

block() turns a block into markdown, and the compiler will not let it forget a case, because the union is exhaustive and the function returns string:

function block(b: Block): string {
  switch (b.type) {
    case "p":
      return b.text;
    case "list":
      return b.items.map((item, i) => `${b.ordered ? `${i + 1}.` : "-"} ${item}`).join("\n");
    case "steps":
      return b.items.map((s, i) => `${i + 1}. **${s.title}.** ${s.text}`).join("\n");
    case "table":
      return [
        ...(b.caption ? [`${b.caption}:`, ""] : []),
        `| ${b.head.map(cell).join(" | ")} |`,
        `| ${b.head.map(() => "---").join(" | ")} |`,
        ...b.rows.map((r) => `| ${r.map(cell).join(" | ")} |`),
      ].join("\n");
    ...
  }
}
Enter fullscreen mode Exit fullscreen mode

Add a ninth block type and this file fails to compile before any page using it can ship. That is the only enforcement I wanted: not a test that the output looks right, a type error if the output cannot exist.

Two details in there cost more thought than the rest of the file:

Links have to grow a host. Content links are site relative, which is right for a page and wrong for a text file somebody downloads. One regex, applied once to the finished document:

const absolute = (md: string) => md.replace(/\]\(\//g, `](${SITE_URL}/`);
Enter fullscreen mode Exit fullscreen mode

Table cells are not paragraphs. A pipe inside a cell ends the cell, and a newline inside a cell ends the row, so a cell is escaped and flattened before it is joined:

const cell = (s: string) => absolute(s).replace(/\|/g, "\\|").replace(/\n/g, " ");
Enter fullscreen mode Exit fullscreen mode

In llms.txt there is a third case. The spec's shape is a link followed by a note, and a note containing its own links reads as nonsense in a list of links, so the note is put through plain() and the summary's links are dropped:

const link = (path: string, name: string, text: string) =>
  `- [${name}](${absoluteUrl(path)}): ${plain(text)}`;
Enter fullscreen mode Exit fullscreen mode

Both files are static

The routes are eight lines each:

import { llmsFullTxt } from "@/content/markdown";

// Built once at build time from src/content.
export const dynamic = "force-static";

export function GET() {
  return new Response(llmsFullTxt(), { headers: { "Content-Type": "text/plain; charset=utf-8" } });
}
Enter fullscreen mode Exit fullscreen mode

force-static matters more than it looks. llmsFullTxt() concatenates every section of every page; it is not expensive, but it is not free either, and there is no reason for it to run twice. At build time it runs once and the result is a file on the CDN.

The sizes are the part people ask about. llms.txt is 35,064 bytes: the product pages, then every content page as a link and a one line summary, grouped by kind. llms-full.txt is 574,626 bytes, 92,301 words, 6,345 lines: the whole site as one markdown document, each page preceded by its own source URL and its updated date, separated by ---. You can diff either against the pages. Pick any page in the guides index and search the full file for its H1; the body text is identical, because it is the same strings.

The reason this is worth doing at all

The registry has one comment at the top, and it is the real argument:

// Every public page, in one place. The sitemap, llms.txt, llms-full.txt, the
// hubs, the footer and "read next" links all derive from these lists, so a new
// page can't be left out of any of them: a page missing from the sitemap still
// works, it just never gets found.
Enter fullscreen mode Exit fullscreen mode

Six consumers, one list. The failure mode this removes is not a crash. It is a page that exists, renders, reads well, and is absent from the one file that would have told anything about it. Nobody notices for a month.

The same list is what lets a test refuse a broken internal link anywhere in the content, including inside a table cell:

assert.ok(resolveLink(path), `${from}: link to unknown page ${href}`);
Enter fullscreen mode Exit fullscreen mode

A link to a page we deleted fails pnpm test, not a crawl weeks later.

The one thing I would do differently is start here. Our first pass at these pages was files, and converting them into objects was a day of tedium. The content types came out of the conversion, not before it, which is probably how it has to go: you cannot design the shape of 63 pages before you have written five.

Top comments (0)