DEV Community

Olivia
Olivia

Posted on

Build Gates for a Small Next.js 16 Site: Failing the Build on Bad Copy, Serving Markdown to AI Agents

I recently built a small service site in Turkish: villaisiksusleme.com, outdoor Christmas lighting for villas and waterfront houses in Istanbul. About 40 pages: a home page, five service pages, ten neighbourhood pages, nine guides, a pricing page and the usual legal pages. Next.js 16 App Router on Vercel, Tailwind v4, vitest.

The code was never the hard part. The hard part was keeping the content honest across 40 pages written partly with AI help: no filler phrases, no prices typed by hand that drift from the real price list, no two district pages that are the same page with the name swapped. Reviewing that by eye does not scale, even at 40 pages.

So the rules live in the build. If a page breaks one, npm run build exits with code 1 and nothing deploys. This post walks through the gates, with short excerpts from the real scripts (variable names are Turkish; I translate as I go).

TL;DR

  • Prebuild: a 124-line Node script scans app/, components/, data/ and lib/ for banned phrases, em-dashes, emoji, icon library imports and hand-typed prices. Any finding fails the build.
  • Postbuild: a second script reads the generated HTML in .next/server/app, because the copy lives in JSX and only exists as text after rendering. It checks titles, descriptions, H1 count, a "short answer" box, image count and 8-word shingle similarity between pages.
  • One source of truth: prices, company facts and the founding year live in data/. Pages import them; the scanner rejects price literals in page code.
  • AI readability: every page gets a Markdown twin generated from the build output, and the same URL returns Markdown when a client sends Accept: text/markdown.
  • IndexNow: a postbuild ping, with a fallback key so it does not depend on an environment variable I could not verify.

Why put content rules in the build at all?

Content rules belong in the build when the content is code. On this site every sentence sits in a page.tsx file, written as JSX. A style guide in a Markdown file gets read once. A script in prebuild runs on every deploy, including the deploys where I am tired and an AI assistant has just rewritten a paragraph.

The package.json wiring is plain npm lifecycle scripts:

"prebuild": "node scripts/policy-scan.mjs",
"build": "next build",
"postbuild": "node scripts/ai-mirror.mjs && node scripts/content-audit.mjs && node scripts/indexnow-ping.mjs",
Enter fullscreen mode Exit fullscreen mode

Vercel runs npm run build, so prebuild and postbuild run there too. No CI config needed.

Gate 1: scanning the source before the build

scripts/policy-scan.mjs walks four directories and checks every .ts, .tsx, .mjs and .json file line by line. The phrase list is the interesting part. It is in Turkish, so here are a few entries with translations:

const PHRASES = [
  // context.md §2.1 giriş kalıpları (stock openers)
  "günümüzde", "modern dünyada", "yılbaşı yaklaşırken", "bu yazıda", ...
  // boş vurgu (empty emphasis)
  "mükemmel", "eşsiz", "büyüleyici", "unutulmaz", "kusursuz", ...
  // kapanış kalıpları (closing clichés)
  "sonuç olarak", "özetle", "unutmayın ki",
Enter fullscreen mode Exit fullscreen mode

That is "nowadays", "in the modern world", "in this article", "perfect", "unique", "unforgettable", "in conclusion", "to sum up". Turkish marketing copy and AI-generated Turkish lean on the same phrases. The list also holds company-specific bans, for example claims that belong to a sister business and must not appear on this site.

Three details made the scanner usable instead of annoying:

Comments are not scanned. The script blanks out comments before matching, keeping line numbers intact, so I can write a comment that explains why a phrase is banned without tripping the rule:

/** Yorumları boşlukla değiştirir (satır numaraları korunur). */
// "Replaces comments with spaces (line numbers are kept)."
function stripComments(src) {
  return src
    .replace(/\/\*[\s\S]*?\*\//g, (m) => m.replace(/[^\n]/g, " "))
    .replace(/(^|[^:"'`])\/\/.*$/gm, (m, p1) => p1 + " ".repeat(m.length - p1.length));
}
Enter fullscreen mode Exit fullscreen mode

There is an escape hatch. A line containing policy-ok is skipped. The convention is that the reason goes in a comment on the same line, so an exception is visible in review.

Phrases that span lines are caught too. JSX wraps long sentences, so "sonuç olarak" can end up split across two lines. After the line pass, the script joins the whole file into one string and searches again.

The same pass rejects em-dashes and en-dashes (a style choice for this site), emoji, imports from lucide-react, react-icons and similar packages (the design uses typography instead of icons), and one more thing worth its own section.

Gate 2: no prices typed by hand

Prices live in exactly one file, data/pricing.ts:

export const PACKAGES: readonly Package[] = [
  { slug: "sade", name: "Sade", min: 22_000, max: 48_000, ... },
  { slug: "bahce-cephe", name: "Bahçe + Cephe", min: 48_000, max: 120_000, ... },
  { slug: "tam-dekor", name: "Tam Dekor", min: 120_000, max: 380_000, ... },
] as const;
Enter fullscreen mode Exit fullscreen mode

Pages import these numbers and format them. To stop anyone, me included, from typing "48.000 TL" into a page, the scanner runs a price regex over app/ and components/ only:

const PRICE_RE = /\b\d{1,3}(?:\.\d{3})+\s*(?:TL|₺)|\b\d{2,3}\s*bin\s*(?:TL|lira)/i;
Enter fullscreen mode Exit fullscreen mode

It matches Turkish number formatting ("48.000 TL") and the spoken form ("48 bin lira"). The data/ folder is excluded from this one check, because that is where prices are supposed to be.

The postbuild audit closes the loop from the other side. It reads data/pricing.ts, formats each band the Turkish way, and fails if public/llms.txt does not contain it. When the price list changes and the AI-facing summary does not, the build stops.

The same idea covers company facts. data/site.ts holds the name, address, phone numbers and founding year, and the years-in-business figure is computed rather than written:

export const YEARS_EXPERIENCE = new Date().getFullYear() - ENTITY.foundingYear;
Enter fullscreen mode Exit fullscreen mode

A hard-coded "17 years" would be wrong on 1 January.

Gate 3: measuring the rendered HTML, not the source

scripts/content-audit.mjs runs after next build. Its header comment explains the design choice: "Metin JSX içinde yaşadığı için ölçüm kaynak yerine çıktıda yapılır", which means "because the text lives inside JSX, we measure the output instead of the source". Counting words in a .tsx file counts className strings. Counting words inside the rendered <main> counts what a reader sees.

For each HTML file in .next/server/app it checks:

  • <title> between 30 and 65 characters (a warning outside 50 to 62), meta description between 120 and 165 for indexable pages, a canonical link.
  • Exactly one <h1> inside <main>.
  • Exactly one HomeAndConstructionBusiness node in the JSON-LD, and no aggregateRating, because the site has no review data to back one.
  • A "short answer" box (data-kisa-cevap) of 40 to 60 words containing at least one digit, on every page except legal and utility pages.
  • At least five <img> tags on service, district and guide pages.
  • Service and district pages under 700 words must carry noindex.
  • Every service page must contain a sentence about what the company does not do. The check is a blunt regex for Turkish verbs like "yapmıyoruz" ("we don't do") and "önermiyoruz" ("we don't recommend").
  • No broken internal links, no orphan pages, nothing deeper than three clicks from the home page.

Near-duplicate pages

Ten district pages invite templating: same layout, swap "Beykoz" for "Sarıyer". The audit compares every pair of indexable pages with 8-word shingles and Jaccard similarity:

const BENZERLIK_ESIK = 0.3; // similarity threshold

for (let i = 0; i + 8 <= ws.length; i++) s.add(ws.slice(i, i + 8).join(" "));
// ...
const jac = inter / (a.size + b.size - inter);
if (jac >= BENZERLIK_ESIK) errors.push(`benzerlik ${jac.toFixed(2)}: ${keys[i]} ~ ${keys[j]}`);
Enter fullscreen mode Exit fullscreen mode

Eight words is long enough that shared navigation phrases and normal Turkish collocations do not count, and short enough that a pasted paragraph with one name changed does. Lowercasing uses toLocaleLowerCase("tr"), because Turkish has a dotted and a dotless i, and plain toLowerCase() gets "İ" wrong.

The script prints the highest pair on every run, even when it passes, so I can see how close the site is to the line.

Making every page readable as Markdown

The site is meant to be quoted by AI assistants, so it ships three things beyond HTML.

A Markdown twin of every page. scripts/ai-mirror.mjs takes the <main> of each built page, runs it through turndown with the GFM plugin, adds front matter (title, description, canonical URL, dates, author) and writes public/md/<path>.md. It also concatenates everything into public/llms-full.txt. A few custom rules matter: image src values are pulled back out of the next/image optimizer URL, relative links become absolute so a quoted fragment still works, and lists inside table cells are flattened to a; b; c so they do not break the Markdown table. Pages with noindex are skipped.

Content negotiation on the same URL. In Next.js 16, middleware.ts is called proxy.ts. Mine compares the q values in the Accept header:

export function markdownIster(accept: string | null): boolean {
  if (!accept) return false;
  // Markdown yalnız açıkça istenirse; "*/*" tarayıcı ve botlara HTML kalır.
  // "Markdown only when asked for explicitly; */* keeps HTML for browsers and bots."
  const md = q(accept, "text/markdown", false);
  return md > 0 && md >= q(accept, "text/html");
}
Enter fullscreen mode Exit fullscreen mode

The false argument disables wildcard matching for Markdown. A crawler that sends */* gets HTML. Only a client that names text/markdown gets the rewrite to /md/....

Checked against production:

$ curl -s -D - -o /dev/null -H "Accept: text/markdown" https://www.villaisiksusleme.com/paketler
HTTP/1.1 200 OK
Content-Type: text/markdown; charset=utf-8
Vary: Accept

$ curl -s -o /dev/null -w "%{http_code} %{content_type}\n" -H "Accept: */*" https://www.villaisiksusleme.com/paketler
200 text/html; charset=utf-8
Enter fullscreen mode Exit fullscreen mode

Discovery headers. HTML responses carry a Link header pointing to the sitemap, llms.txt, llms-full.txt and the page's own Markdown twin with rel="alternate"; type="text/markdown". The /md/ files themselves are served with X-Robots-Tag: noindex, follow so the copies do not compete with the HTML pages in search.

IndexNow, with a fallback key

The last postbuild step pings IndexNow (used by Bing and Yandex) with every URL in the live sitemap. It only runs when VERCEL_ENV is production or when forced, so preview deploys do not ping.

The first version read the key only from INDEXNOW_KEY. When I went to submit the site, I could not confirm whether that variable was set in the Vercel project. A missing env var here fails silently: the script logs a line and moves on. So I generated a key, committed the key file to public/, and made the env var optional:

const DEFAULT_KEY = "5ed03e268fe03568d4c4e1d62402aba2";
const KEY = process.env.INDEXNOW_KEY || DEFAULT_KEY;
Enter fullscreen mode Exit fullscreen mode

IndexNow keys are public by design, since the key file has to be reachable at the site root, so there is nothing secret in the repo. The first manual run (INDEXNOW_FORCE=1 node scripts/indexnow-ping.mjs) returned 202 for 41 URLs, which means accepted with key validation pending. I also made the script print the response body on non-2xx, which the first version did not.

What went wrong

A layout shift on the home page. The commit log has fix: ana sayfa hero butonları mobilde alt alta (CLS 0,108 -> 0), "home hero buttons stacked on mobile". The two hero buttons sat side by side on phones; stacking them took the measured CLS from 0.108 to 0. None of my gates measure layout shift. That number came from measuring the page by hand, and it is still a manual check.

Two Link headers and a lost Vary. While writing this post I ran curl -I against an HTML page and saw two Link headers: one from headers() in next.config.ts and one set in proxy.ts. The second is a superset of the first, so nothing breaks, but it is redundant. The same response shows Vary: rsc, next-router-state-tree, ... from Next.js rather than the Vary: Accept that proxy.ts sets, so on HTML responses my header is replaced. The Markdown response keeps Vary: Accept. Both are open items; I would rather list them than pretend the setup is finished.

The phrase list keeps growing. Every new page has added one or two entries. A scanner like this is only as good as the last time someone read the output and asked "why does this sentence sound wrong?"

If you want to copy the idea

You do not need my phrase list. The pattern is small:

  1. Keep facts that appear on many pages (prices, contact details, counts) in one data file, and fail the build when a page types them by hand.
  2. Scan the source for things you never want to ship, skip comments, and allow a visible, greppable escape hatch.
  3. Measure content on the build output, not in the JSX.
  4. Compare pages pairwise with long shingles if your site has templated sections.
  5. Generate the Markdown copy from the same HTML, so it can never drift from what humans see.

The live site is at villaisiksusleme.com. Request any page with Accept: text/markdown to see the Markdown version, or read llms.txt for the site summary.

![ ](https://dev-to-uploads.s3.us-east-2.amazonaws.com/uploads/articles/knv5eylwq57okk59g7ea.jpg)

Top comments (0)