DEV Community

Daniel Pertu
Daniel Pertu

Posted on

robots.txt matching is a literal prefix, and one trailing slash would have left our dashboard crawlable

Our robots.txt is generated from one array:

export const CRAWLER_DISALLOW = [
    '/dashboard',
    '/play',
    '/api',
    '/auth',
    '/subscribe',
    '/webhook',
    '/too-many-requests',
] as const
Enter fullscreen mode Exit fullscreen mode

Not one of those has a trailing slash, and that is the first thing I would check in anybody else's robots.txt.

Disallow is a literal prefix match against the path. Disallow: /dashboard/ blocks /dashboard/settings and /dashboard/billing, and leaves the bare /dashboard perfectly crawlable, because /dashboard does not begin with /dashboard/. In our case that is not a hypothetical: /dashboard is linked from the footer of the public homepage, so it is the single most discoverable URL in the whole private area. A trailing slash would have blocked everything except the one page a crawler was guaranteed to find.

The directive is also not a glob. Disallow: /api and Disallow: /api/ are different instructions, and the shorter one is almost always what you meant.

robots.txt stops crawling, not indexing

This is the fact that most "I blocked it in robots.txt" reasoning gets wrong. Disallowing a path tells a crawler not to fetch it. It does not tell anyone not to index it. A URL that Google has seen somewhere else, in a link, in a sitemap, pasted into a public forum, can still end up in the index as a title-only result that the crawler has never actually read.

So the private segments say it again, in the page itself:

export const metadata: Metadata = {
    title: "Dashboard",
    robots: { index: false, follow: false },
}
Enter fullscreen mode Exit fullscreen mode

And now the ordering problem, which is the genuinely counterintuitive part. Those two mechanisms do not stack cleanly. Once crawling is disallowed, Google can never fetch the page, so it can never discover a noindex you added afterwards. The belt only works if it was already on before you fitted the braces.

What follows from that is a sequencing rule rather than a config rule: the noindex goes on first, you let it be crawled long enough to be seen, and the Disallow is what you add afterwards to stop wasting crawl budget on a segment nobody should be reading. Adding both in the same commit to a segment that is already indexed achieves the opposite of what you wanted, and it does so silently.

Put it at the segment boundary, not in a list

// app/dashboard/layout.tsx
robots: { index: false, follow: false },
Enter fullscreen mode Exit fullscreen mode

One declaration on the layout covers every present and future route under /dashboard. The alternative is remembering to append each new private route to an array in app/robots.ts, and forgetting to do that is invisible: no error, no failing test, just a page quietly eligible for the index.

The same reasoning applies to /play, for a reason specific to the product. Player URLs get handed out on table cards and shared phone to phone in a pub, so they escape into the wild and get crawled whether we like it or not. A segment level index: false is the only thing that scales with that.

A preview deploy is a competitor

Every preview and staging build serves the same pages as production. Letting one be indexed puts a second copy of all seventy pages in the index, competing with the first, and pointing at a sitemap full of production URLs it does not itself serve.

So robots.txt is conditional on the deployment being the real one:

export default function robots(): MetadataRoute.Robots {
  if (!IS_PRODUCTION_SITE) {
    return { rules: { userAgent: "*", disallow: "/" } };
  }
  return {
    rules: { userAgent: "*", allow: "/", disallow: [...CRAWLER_DISALLOW] },
    sitemap: absoluteUrl("/sitemap.xml"),
  };
}
Enter fullscreen mode Exit fullscreen mode

And the root metadata says the same thing a second way, for the same ordering reason as above:

robots: {
  index: IS_PRODUCTION_SITE,
  follow: IS_PRODUCTION_SITE,
  ...
}
Enter fullscreen mode Exit fullscreen mode

IS_PRODUCTION_SITE is one comparison, SITE_URL === 'https://pub-trivia.app', derived from a single module that resolves the site's own origin once. Everything that needs an absolute URL reads it from there: metadataBase, the sitemap, robots.txt, the JSON-LD graph, Stripe's redirect URLs, password reset emails, and the QR codes printed on table cards. Deriving that origin independently in each caller is how a deploy ends up serving a sitemap for one host and sending emails for another.

The sitemap is a statement of intent, not an inventory

/login, /signup and /forgot-password used to be in the sitemap. They are deliberately not any more, and they are still crawlable. The distinction: robots.txt is about what a crawler may fetch, and a sitemap is a list of the pages you want ranked. A sign-in form is not one of them. Leaving it in does not get it ranked, it just dilutes the signal in a document whose only job is to be a signal.

Two other things went the same way. changeFrequency and priority are gone, because Google documents both as ignored and they were eight invented numbers nobody could justify. And lastModified now comes from each page's own recorded content date rather than new Date() at build time, which is the subject of its own post about generating the whole content graph from one array.

The canonical that is deliberately missing

Related trap, same family. The root layout sets no alternates.canonical, on purpose:

// No `alternates` here on purpose: a canonical set at this level is inherited
// by every page that does not set its own, and each of them then declares
// itself a duplicate of the homepage.
Enter fullscreen mode Exit fullscreen mode

Metadata in the App Router inherits down the tree. A canonical URL is the one field where inheriting the parent's value is actively harmful, because the inherited value is a claim about this page, and the claim is false. Every page builds its own from its path instead.

The three lists have to agree, so a test says so

Which routes are public, which are indexable, and which crawlers should stay out of are three questions that used to be answered independently in three files. They disagreed. /about shipped in the sitemap while the auth gate bounced every crawler that followed it to /login, and robots.txt and sitemap.xml were themselves gated, which quietly made an entire SEO release inert in production. That last one has its own war story.

They are one module now, and the tests check the relationships rather than the contents:

it('every sitemap URL is reachable without a session', () => {
    for (const path of INDEXABLE_ROUTES) {
        expect(isPublicRoute(path === '' ? '/' : path), `${path} is auth-gated`).toBe(true)
    }
})

it('no sitemap URL is disallowed to crawlers', () => {
    for (const path of INDEXABLE_ROUTES) {
        const blocked = CRAWLER_DISALLOW.some((prefix) => (path || '/').startsWith(prefix))
        expect(blocked, `${path} is in the sitemap and in robots.txt disallow`).toBe(false)
    }
})
Enter fullscreen mode Exit fullscreen mode

Neither test knows what the site contains. They assert that submitting a page for ranking, gating it behind auth, and telling crawlers to stay away are mutually exclusive, which is the kind of thing that is obvious in review and invisible six months later.

Check ours

curl -s https://pub-trivia.app/robots.txt
Enter fullscreen mode Exit fullscreen mode

Seven Disallow lines, no trailing slashes, and a Sitemap: line with an absolute production URL.

curl -s https://pub-trivia.app/sitemap.xml | grep -c '<loc>'
curl -s https://pub-trivia.app/sitemap.xml | grep -c 'login'
Enter fullscreen mode Exit fullscreen mode

Seventy-two URLs, zero of which are the sign-in page. And the gate itself:

curl -sI https://pub-trivia.app/dashboard | grep -i location
Enter fullscreen mode Exit fullscreen mode

A 307 to /login?next=%2Fdashboard, which is the same answer a crawler gets, which is why the noindex on that segment has to exist independently of robots.txt.

Everything in the sitemap, by contrast, answers 200 with no session at all. Features, solutions, guides, quiz questions, tools and comparisons are all open, and that coherence between the three lists is the only property any of this was built to guarantee.

Top comments (0)