PubTrivia runs a fallback limiter in middleware, in front of everything, as a backstop behind the per-endpoint limits that do the real work. Its only job is to catch runaway scripts and bots.
When it fires on a browser we wanted something better than a bare JSON error, so there is a page: pub-trivia.app/too-many-requests. It says "Slow down", it explains itself in one sentence, and it prints HTTP 429 in a mono font underneath.
It is eighteen lines of JSX. Everything that made it hard is in the four decisions around it.
One client wants a page, the other wants a status code
The same middleware sees a phone loading a page and a fetch() from code already running on that page. A redirect to a nice HTML page is right for the first and useless for the second: the caller gets a 200 full of markup where it expected JSON, and whatever error handling it had does not fire.
So the branch is on content negotiation, not on path:
const isHtmlRequest = request.headers.get('accept')?.includes('text/html')
HTML requests get redirected to the page. Everything else gets a JSON body with a 429 status and a Retry-After header, which is the thing a client can actually act on.
Accept is a hint rather than a guarantee, but it is the right hint here: browsers navigating to a document send text/html in that header, and fetch() defaults to */* unless someone sets otherwise. The failure mode of guessing wrong is cosmetic in one direction and inconvenient in the other, which is a good trade for one line.
The error page cannot be subject to the error
The first version limited every path uniformly, which included /too-many-requests.
Follow what that does. A browser trips the limit, gets redirected to the page, requests the page, and that request is also over the limit, so it is redirected to the same URL. The user never sees the page. They see ERR_TOO_MANY_REDIRECTS, which is both uglier than the JSON we were trying to improve on and much harder to diagnose, because the error comes from the browser rather than from us.
Worse, with a sliding window, every hop in that loop consumed another token. The loop fed the condition that caused it, so the window never drained and waiting did not help.
The fix is a single early check that exempts the one path before the limiter is consulted at all. It is two lines and it is the entire reason the page works.
The general shape of the bug is worth naming because it is not specific to rate limits: any error page served from the same middleware that produces the error has to be exempt from it. The same applies to a maintenance page behind an auth check, or a "payment required" page behind a subscription check.
It is public, and it is also disallowed
Two lists in this codebase have to agree about that page, and they say what look like opposite things.
It is in the public-route allowlist, because middleware also runs the auth gate: without that entry, a signed-out visitor who tripped the limiter would be redirected to the error page and then redirected again to the login form.
It is also in the crawler disallow list, so it appears in robots.txt:
Disallow: /too-many-requests
Public and crawlable are different properties. "Anyone may fetch this without a session" and "we would like this in a search index" are independent questions, and almost every error, status or utility page answers yes to the first and no to the second.
An HTML 200 that says 429 is a page about nothing
Here is the subtle one. Because the limiter redirects, the page itself returns a real 200. Its body contains the characters HTTP 429, but as far as any client is concerned this is a successful response.
That matters for a crawler that trips the limiter. It asked for a page about running a pub quiz, followed a redirect, and received a successful response whose content is Slow down. Left indexable, that content can be recorded as the content of the URL it originally asked for.
So the page sets noIndex through the same metadata helper every other page uses, and you can confirm it from outside:
curl -s https://pub-trivia.app/too-many-requests | grep robots
which gives:
<meta name="robots" content="noindex, nofollow"/>
robots.txt and that meta tag are doing two different jobs, and this page needs both. robots.txt asks crawlers not to fetch the URL. The meta tag is what handles the case where a crawler arrives here anyway, by redirect, without having chosen to visit it. A disallow rule cannot help with a path you were pushed onto.
The redirect also clears the query string rather than carrying it over. There is nothing on the error page that could use it, and the alternative is a family of indexable-looking URLs that are all the same page.
When the limiter cannot answer
The limiter's store is Redis, and if Redis is unreachable the check is skipped rather than failing closed. A backstop whose outage takes the entire site down is worse than no backstop: the thing it protects against is a nuisance, and the thing it would cause is an outage.
That is a judgement about this particular limiter and not a general rule. The per-endpoint limiters guarding expensive or billable operations are the ones where failing closed deserves a second look, and the right answer there depends on what the operation costs.
The page is live and reachable directly if you want to see what a visitor sees: pub-trivia.app/too-many-requests. robots.txt shows the disallow entry beside the rest of the private segments, and the app the limiter is protecting is at pub-trivia.app, free tier, no card.
Top comments (0)