A growing share of people never see your HTML. They ask ChatGPT, Claude, Perplexity or Gemini, and an agent fetches your page on their behalf. That agent has a token budget, and your cookie banner, mega-menu and 40 KB of hydration script are not what it came for.
I run AskedTheAI, a shopping guide that asks four AI models the same buying question and has a human check every pick on Amazon. The site is about AI, so it would be odd if AI agents struggled to read it. Before launch I set up four things:
- A Markdown copy of every page, generated from the build output
- Content negotiation: the same URL returns Markdown when an agent asks for it
-
llms.txtandllms-full.txtas a map - IndexNow pings so Bing and Yandex hear about changes immediately
The stack is Next.js 16 (App Router, fully static), deployed on Vercel. Everything below is the actual code from the repo.
1. Generate the Markdown mirror from the build, not from source
My first instinct was to render Markdown from the React components. That's a trap: every page is its own composition, and I'd be maintaining two renderers.
Instead, the mirror is built from what Next already prerendered. After next build, every static page exists as HTML under .next/server/app. A small Node script walks that folder, pulls out <main>, converts it with node-html-markdown and writes public/md/<route>.md:
// scripts/md/generate.mjs (excerpt)
const APP = '.next/server/app'
const OUT = 'public/md'
const SKIP = /^\/(api|_not-found|_global-error|404|500|examples|search|templates|md)(\/|$)/
const nhm = new NodeHtmlMarkdown({
ignore: ['script', 'style', 'noscript', 'svg', 'button', 'form', 'iframe', 'video', 'template'],
maxConsecutiveNewlines: 2,
useLinkReferenceDefinitions: false,
})
Each file gets a small front matter block so an agent knows what it's holding and where it came from:
---
title: "How We Ask the AI: Our Research Method | AskedTheAI"
url: https://www.askedtheai.com/ai-research
description: "How AskedTheAI puts the same buying question to Claude, ChatGPT, Gemini and Grok..."
last_updated: 2026-10-01
source: AskedTheAI (https://www.askedtheai.com)
affiliate_disclosure: As an Amazon Associate we earn from qualifying purchases.
---
The script also writes src/data/md-routes.json, a list of every route that has a mirror. The middleware below imports it.
The pipeline is two commands, and the output is committed so Vercel serves it as static files:
npm run build && npm run ai:files
The HTML is the source of truth. The Markdown can't drift from it, because it's derived from it on every run.
2. Content negotiation in middleware
Agents that want Markdown can say so with an Accept header. The page URL stays the same; only the representation changes. Here's the whole middleware:
// src/middleware.ts
import { NextResponse, type NextRequest } from 'next/server'
import mdRoutes from '@/data/md-routes.json'
const SITE = 'https://www.askedtheai.com'
const MD_ROUTES = new Set<string>(mdRoutes as string[])
const mdPathFor = (route: string) => (route === '/' ? '/md/index.md' : `/md${route}.md`)
function prefersMarkdown(accept: string): boolean {
if (!/text\/markdown/i.test(accept)) return false
const first = accept.split(',')[0]?.trim().toLowerCase() ?? ''
return first.startsWith('text/markdown') || !/text\/html/i.test(accept)
}
export function middleware(request: NextRequest) {
const { pathname } = request.nextUrl
const route = pathname.length > 1 ? pathname.replace(/\/$/, '') : '/'
if (!MD_ROUTES.has(route)) return NextResponse.next()
const mdPath = mdPathFor(route)
if (prefersMarkdown(request.headers.get('accept') || '')) {
const url = request.nextUrl.clone()
url.pathname = mdPath
const res = NextResponse.rewrite(url)
res.headers.set('Vary', 'Accept')
res.headers.set('Content-Type', 'text/markdown; charset=utf-8')
res.headers.set('X-Robots-Tag', 'noindex, follow')
return res
}
const res = NextResponse.next()
res.headers.set('Link', `<${SITE}${mdPath}>; rel="alternate"; type="text/markdown"`)
res.headers.set('Vary', 'Accept')
return res
}
export const config = {
// pages only: skip api, Next internals, the mirror itself and any file with an extension
matcher: ['/((?!api/|_next/|md/|.*\\.[a-z0-9]+$).*)'],
}
A few decisions worth explaining:
-
prefersMarkdownis strict. Browsers never sendtext/markdown, so they always get HTML. An agent gets Markdown if it lists it first, or lists it withouttext/htmlat all. -
Vary: Accept, and what actually survives. The idea is to stop a CDN from caching the Markdown response and serving it to the next browser. When I checked the live headers, the Markdown response hadVary: Accept, but on the HTML response Next.js replaced my value with its own list (rsc, next-router-state-tree, ...). It's safe anyway: the Markdown branch is a rewrite to a different path (/md/...), so Vercel caches the two representations under different keys. Still, check your real headers withcurl -Iinstead of trusting the code. -
X-Robots-Tag: noindex, followon the Markdown. The mirror exists for agents, not for search results. I don't want/md/best-dash-cams.mdcompeting with/best-dash-camsin Google. -
A
Link: rel="alternate"header on every HTML page. Agents that fetched HTML first learn a lighter version exists. -
The matcher skips anything with a file extension. That turned out to matter later: search engine verification files like
BingSiteAuth.xmland the IndexNow key file pass straight through, untouched.
Try it yourself:
curl -s -H "Accept: text/markdown" https://www.askedtheai.com/ai-research | head -20
curl -sI https://www.askedtheai.com/ai-research | grep -i '^link'
# Link: <https://www.askedtheai.com/md/ai-research.md>; rel="alternate"; type="text/markdown"
curl -sI -H "Accept: text/markdown" https://www.askedtheai.com/ai-research | grep -i -E '^(content-type|vary|x-robots-tag)'
# Content-Type: text/markdown; charset=utf-8
# Vary: Accept
# X-Robots-Tag: noindex, follow
A note on naming: Next.js 16 renames middleware.ts to proxy.ts. The old name still works. I'm switching after launch rather than touching routing the night before.
3. llms.txt as the map
llms.txt is a plain Markdown file at the site root that tells a model what the site is and where the important pages are. Mine is generated by the same ai:files step, from the same build output.
Two things I'd recommend:
- Say what the site does NOT do, in the first lines. My file states that the site doesn't lab-test products and lists the two review formats with page counts. A model summarizing the site should get that right on the first try.
-
Ship a long version too.
llms-full.txtlists every page with a one-line summary and its Markdown URL. The short file is the table of contents; the long one is the index.
4. robots.txt that actually says yes
A lot of sites still block AI crawlers by default, then wonder why assistants never cite them. If you want to be read, say so explicitly. My robots.txt names the retrieval and search bots (OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, Applebot and others) and adds a Content-Signal line:
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=yes
Allow: /
Disallow: /api/
Whether to allow training is a business decision. For a site whose whole pitch is "we asked the AI," it would be strange to say no.
5. IndexNow: tell Bing and Yandex the moment something changes
Google ignores IndexNow, but Bing, Yandex, Seznam and Naver share it. One POST to api.indexnow.org reaches all of them. Setup is two files.
A key file in public/, named after the key and containing only the key. Here's the script that reads the live sitemap and submits every URL:
// scripts/indexnow-ping.mjs
const HOST = 'www.askedtheai.com'
const KEY = 'your-32-char-hex-key'
const ORIGIN = `https://${HOST}`
const keyRes = await fetch(`${ORIGIN}/${KEY}.txt`)
if (!keyRes.ok || (await keyRes.text()).trim() !== KEY) {
console.error(`[indexnow] key file not live (HTTP ${keyRes.status}); deploy first`)
process.exit(1)
}
let urls = process.argv.slice(2)
if (urls.length === 0) {
const xml = await (await fetch(`${ORIGIN}/sitemap.xml`)).text()
urls = [...xml.matchAll(/<loc>([^<]+)<\/loc>/g)].map(m => m[1].trim())
}
urls = urls.filter(u => u.startsWith(ORIGIN))
for (let i = 0; i < urls.length; i += 10000) {
const urlList = urls.slice(i, i + 10000)
const res = await fetch('https://api.indexnow.org/indexnow', {
method: 'POST',
headers: { 'Content-Type': 'application/json; charset=utf-8' },
body: JSON.stringify({ host: HOST, key: KEY, keyLocation: `${ORIGIN}/${KEY}.txt`, urlList }),
})
console.log(`[indexnow] ${urlList.length} URLs -> HTTP ${res.status} ${await res.text()}`)
if (!res.ok && res.status !== 202) process.exit(1)
}
With no arguments it submits the whole sitemap. With arguments it submits only those URLs, which is what I use after editing a page:
npm run indexnow -- https://www.askedtheai.com/authors/ismail-gunaydin
The script checks that the key file is live before sending anything. That check exists because of the first gotcha below.
Gotchas I actually hit
IndexNow returned 403 on the first run. The response was {"errorCode":"SiteVerificationNotCompleted"}, even though the key file was live and returned 200. IndexNow verifies a new key asynchronously. I retried five minutes later and got HTTP 200 for all 516 URLs. If your first ping fails with this error, wait, don't debug.
Regenerating the mirror produced noisy diffs. After a one-line fix to an author profile, ai:files touched 20 Markdown files. Only one had a real change. The rest were last_updated dates on a few dynamic pages and a homepage block that shows a rotating selection of deals. I committed the one file and restored the rest. If your mirror includes anything time-based or shuffled, expect this, and review the diff instead of committing the folder blindly.
next build no longer lints. In Next.js 16 the build step doesn't run ESLint. My deploy command is now npm run lint && npm run build && npm run ai:files, in that order, locally, before every push.
Was it worth it?
It's too early to show citation numbers, and I won't invent any. What I can say is that the setup costs almost nothing to maintain: the mirror, md-routes.json, and both llms files are regenerated by one command from the real build. There's no second content system to keep in sync.
If you want to poke at it:
- Site: https://www.askedtheai.com
-
llms.txt: https://www.askedtheai.com/llms.txt - How the site asks the AI: https://www.askedtheai.com/ai-research
- AI and developer access notes: https://www.askedtheai.com/developers/ai-consumption
Questions about the middleware or the mirror, ask in the comments. I'm happy to share more of the generator.

Top comments (0)