Munchable is a barcode scanner for digestive conditions, and its marketing site is 426 indexable URLs. The file that decides which of those a crawler may fetch is 29 lines of TypeScript, and the only interesting line in it is the one that is shorter than people expect:
disallow: ['/api/', '/auth/'],
Two paths. Not the checkout return pages, not the post-signup chooser, not the signed-in entry point. Those three are pages we actively do not want in anyone's index, and they are explicitly crawlable.
That looks backwards until you say the two mechanisms out loud.
Disallow stops the fetch. noindex needs the fetch.
Disallow is an instruction not to request a URL. noindex is a header or a meta tag inside the response body of a URL that was requested.
So if you Disallow a page and also give it a noindex, the crawler never fetches the page, never sees the directive, and the URL is free to sit in the index as a bare result with no title and no snippet for as long as anything links to it. You have blocked the one request that could have removed it.
Google documents this plainly: combining the two means the noindex is never seen. The pages this matters for are exactly the ones you are tempted to Disallow, because they are the ones you think of as private: checkout returns, auth entry points, anything a signed-in person is handed after a redirect. Our homepage links to one of them. Our homepage carries a href="/get-started" in its pricing section, and a homepage link is more than enough for a URL to be discovered without ever being fetched. Fetch it today and the response says noindex, nofollow, nocache and carries no canonical, which is the pair of facts the rest of this post is about.
So Munchable's app/robots.ts only disallows the two prefixes that are machine surface and have nothing in them a crawler could ever read a directive out of: the API routes and the auth callbacks. Everything else is allowed, and the pages that should not be indexed say so themselves, in the response:
// apps/web/lib/seo.ts
export const NO_INDEX: Metadata = {
robots: { index: false, follow: false, nocache: true },
};
The comment above that constant is the second thing worth copying:
Deliberately robots-only.
robotsis the one directive that is safe for a child segment to inherit, because it says the same true thing about every page underneath it; acanonicalor anopenGraph.urlis per-page by definition, so neither belongs in anything that is inherited.
That is why it is robots and nothing else. A layout in the App Router is an inheritance point, and the temptation is to park the whole metadata block there for a section. Do that with a canonical and every page under the layout starts claiming to be the parent. Do it with openGraph.url and sharing any page in the section previews as the section's front page. robots is the only part of that object that stays true as it travels down.
A noindex page gets no canonical at all
The helper that builds per-page metadata takes a noIndex flag, and the flag does something more than add the directive:
const url = absoluteUrl(opts.path);
return {
title: resolvedTitle(opts.title),
description: opts.description,
// A noindex page gets NO canonical at all. `noindex` alongside a
// rel=canonical pointing at another URL is the documented way to have the
// directive consolidated onto the canonical target instead of this page, so
// these say nothing rather than pointing anywhere, including at themselves.
...(opts.noIndex ? {} : { alternates: { canonical: url } }),
...
};
A canonical is a statement that two URLs are the same page and that the other one is the real one. A noindex is a statement about a page. Put them together and you have told a crawler that the page carrying your directive is a duplicate of something else, which is an invitation to apply the directive over there instead. The safe shape is: indexable pages carry their own canonical, noindex pages carry none.
This is also why every page on the site goes through one function instead of hand-writing metadata objects. No segment inherits a canonical, so a page's canonical is either its own URL or absent, and there is no third state where it quietly points somewhere else.
Preview deploys have to disagree with nothing
Every branch and preview deploy on Vercel gets its own hostname serving the same site. Left alone it competes with production for the same queries.
One flag decides it:
export const IS_INDEXABLE =
!process.env.VERCEL_ENV || process.env.VERCEL_ENV === 'production';
robots.ts returns a blanket disallow when that is false. The root layout flips its robots metadata to index: false, follow: false. And sitemap.ts returns an empty array:
export default function sitemap(): MetadataRoute.Sitemap {
if (!IS_INDEXABLE) return [];
...
}
The empty sitemap is the part that is easy to skip. A sitemap is a list of URLs you are asserting should be indexed. Shipping the full list from a host whose robots.txt says "disallow everything" is two files on the same deploy contradicting each other, and when two signals disagree you do not get to choose which one wins. An empty sitemap and a blanket disallow say the same thing.
The flag defaults to indexable when VERCEL_ENV is unset, so a local build or a self-hosted run behaves like production rather than silently deindexing itself. That default is a choice: the failure mode of "my staging box got indexed" is recoverable, and the failure mode of "the live site has been serving noindex for a week" is not.
What the 426 URLs are
Worth being concrete about what is on the allowed side, because the whole point of the exercise is that the list is small enough to enumerate:
- 1 homepage
- /conditions plus 7 condition pages, one per condition the app covers, for example low FODMAP and IBS
- /answers plus 373 question pages whose slug is the question, like is-onion-low-fodmap and does-coffee-cause-reflux
- /recipes plus 38 recipes, for example gentle rice congee
- /about, /privacy, /terms, /licenses
Every one of those comes from a TypeScript array at build time, which is why sitemap.ts can be the framework's file convention rather than a hand-rolled XML route. You can read the output at munchable.app/sitemap.xml and the file it pairs with at munchable.app/robots.txt. The disallow list there is still two entries long, which is the whole post.
The checklist version
If you take one thing from this, take the ordering, because it is the opposite of the instinct:
- Decide whether you want a URL fetched. That is what robots.txt answers. Machine endpoints and callbacks: Disallow. Anything that renders HTML to a human: allow.
- Decide whether you want a URL indexed. That is what
noindexanswers, and it is delivered in the response, so it requires step 1 to have allowed the fetch. - Never give the same URL both. One cancels the other, and the one that survives is the one you did not want.
- Give noindex pages no canonical.
- Make the sitemap agree with robots.txt on every deploy, including the ones nobody looks at.
None of this is clever. All of it is the kind of thing that silently costs you a page for months, which is a worse outcome than a loud mistake.
Top comments (0)