Every request to pub-trivia.app that is not excluded by one regular expression runs through a single function: a rate limit check, then a Supabase session refresh, then a decision about whether the path is public. The interesting part of that file is not the function. It is the exclusion list, because almost every entry on it was added after something went wrong.
In Next.js 16 the file is proxy.ts rather than middleware.ts, and the matcher is still the same idea: a list of path patterns, and the convention for "everything except" is one negative lookahead.
export const config = {
matcher: [
'/((?!_next/static|_next/image|favicon\\.ico|manifest\\.webmanifest|robots\\.txt|sitemap\\.xml|.*\\.(?:svg|png|jpg|jpeg|gif|webp)$|webhook|api|monitoring).*)',
],
}
Unpleasant to read, and worth reading anyway, because the ten alternatives inside that lookahead fall into three groups that have nothing to do with each other.
Group one: the ones that are only about cost
_next/static, _next/image, favicon.ico and the image extension list are the standard ones every Next.js project ships with. Nothing about a hashed JavaScript chunk needs a session, and a build emits hundreds of them. Leaving them in costs a Redis round trip and a Supabase token refresh per asset, on a page that pulls in dozens of assets at once.
These are the only entries I would call boilerplate. If you deleted them the site would still behave correctly, just slower and more expensively.
Group two: the ones where running the middleware is a bug
webhook, api and monitoring are different. For these, passing through the middleware is not slow, it is wrong.
A Stripe webhook arrives with no cookie and no session. It carries a signature header instead, and the only thing that may decide whether it is genuine is the signature check in the handler. An auth gate in front of it cannot do anything useful: it can only turn a legitimate webhook into a redirect to a login form, which Stripe records as a failed delivery and retries.
api is excluded because every route handler under it does its own auth with its own Supabase client. There is no shared assumption to enforce in one place, and a gate that redirected a JSON request to an HTML login page would be returning the wrong content type to a caller that asked for JSON.
monitoring is the one I would not have guessed. It is not ours. It is the Sentry tunnel route, which exists so that browser error reports go to our own origin instead of to Sentry's, because an ad blocker will happily eat a request to a known telemetry domain and silently take your client side error reporting with it. The Sentry config sets it in one line:
tunnelRoute: "/monitoring",
and the comment next to that option in Sentry's own setup says, in effect, check that the route you picked does not collide with your middleware, or client side error reporting will fail. That warning is doing a lot of work in a small space. If /monitoring were matched, the browser's error report would be answered with a redirect to /login, the report would never arrive, and nothing anywhere would tell you. You would simply have a dashboard that reported no client side errors, which looks exactly like an application with no client side errors.
Group three: the two that got us
robots.txt and sitemap.xml are the reason I think this list is worth a blog post.
Both of those are served by app/robots.ts and app/sitemap.ts. They are not files in public/. That means as far as the matcher is concerned they are ordinary routes, matched like any page, and so they went through the auth gate like any page. A crawler asking for the sitemap got a 307 to /login.
It is a hard bug to notice from the outside, because every tool you would reach for to check it is signed in as you. The browser you test with has a session, so you see the sitemap. The crawler does not, so it does not.
They are now excluded here, and they are also listed in a PUBLIC_ROUTES array in lib/routes.ts, which is the allowlist the gate itself consults. That is deliberate duplication. The regex saves the round trips; the array is what stops a future edit to the regex from quietly bringing the redirect back. The comment above it says as much: a change to the matcher cannot re-introduce the bug, because the gate would let them through anyway.
The one we deliberately left in
There is a fourth category with one member. The site serves an IndexNow key at a long hex filename, also from a route handler, and it is not in the regex. It is only in the PUBLIC_ROUTES array.
The reasoning is traffic. Crawlers fetch robots.txt and sitemap.xml on every pass, so saving two round trips there is worth the regex noise. The key file is read a handful of times a year, when a search engine verifies a submission. Hard coding a 32 character hex string into a regex to save a few round trips a year is a worse trade than leaving it out, and the array entry is enough to stop the gate answering the verification fetch with a redirect to a login form.
Check it from your own terminal
All of this is observable from outside, and the three responses that matter are three different status codes.
curl -sI https://pub-trivia.app/robots.txt | head -1
curl -sI https://pub-trivia.app/sitemap.xml | head -1
Both answer 200. Before the exclusion, both answered 307.
curl -sI https://pub-trivia.app/dashboard | grep -i '^location'
That one is still gated, and the header names the destination: /login?next=%2Fdashboard. The gate works, and this is what it looks like when it fires on purpose.
curl -sI https://pub-trivia.app/monitoring | head -1
That answers 404, and the 404 is the point. The tunnel route only accepts the POST that the Sentry browser SDK sends, so a GET gets nothing. What matters is which nothing. A 404 means the request reached the route and the route had no answer for it. A 307 to /login would have meant the middleware got there first, and that our client side error reporting had been quietly broken for however long.
If you want to see the public surface the exclusions exist to protect, the sitemap lists all 72 pages a crawler is invited to read, and every one of them answers without a session. The features pages and guides are the bulk of it. Sign up on the free tier if you want to see the other side of the gate: it needs no card, and /dashboard will stop answering 307 to you specifically.
The part that generalises
A middleware matcher looks like configuration and reads like a performance tweak, which is why it tends to be copied from a template once and never revisited. It is actually a security and correctness boundary written in a syntax that gives you no help: no names, no comments inside it, no way to express why any one alternative is there, and a single typo in a lookahead changes which requests are authenticated without changing anything you would notice in development.
Two habits have made it manageable. Write the reasoning in a comment above the regex, one line per alternative, so the next person knows which entries are speed and which are load bearing. And for the load bearing ones, state the same fact twice in two mechanisms, so the regex is an optimisation rather than the only thing standing between a crawler and a login form.
Top comments (0)