Rate limiting advice is almost always written as requests per minute per IP. It is good advice, and it rests on an assumption so ordinary nobody states it: one IP address is roughly one client.
Our app runs pub quiz nights. A hundred people in a room, all on their phones, all behind the pub's WiFi, all arriving at our servers from one exit IP. The assumption is not slightly wrong there. It is inverted.
The arithmetic that was wrong
Players answer on their phones over a WebSocket. When that connection cannot be established, which happens in venues with hostile WiFi more often than you would like, each phone falls back to polling.
So the worst legitimate case is: an entire pub behind one exit IP, with the WebSocket server unreachable, every phone polling.
A hundred players at six polls a minute is six hundred requests, before anybody has joined a session, loaded a page or submitted an answer.
Our global limit was three hundred per minute per IP.
Read that back and the failure is not "some players got a 429". The fallback path, the thing whose entire purpose is to keep the night running when the primary transport fails, was a denial of service against the venue, triggered at exactly the moment it was needed. The quiz was fine while everything worked and died the instant something degraded.
/**
* global: 1200 requests per minute per IP.
*
* Applied in proxy.ts before every page/action request. This is a bot and
* runaway-script guard, NOT the real protection: the per-endpoint limiters
* do that work, and they key on the participant or user where a shared IP
* would otherwise punish a whole room.
*
* Sized for the worst legitimate case: an entire pub behind one WiFi exit IP
* with the WebSocket server unreachable, so every phone is on the polling
* fallback. 100 players x 6 polls/min = 600, plus joins, page loads and
* answer submissions. The previous 300 was sized for the join burst alone
* and so 429'd the whole venue precisely when the fallback kicked in.
*/
global: sliding(1200, 60),
The number is not the lesson. The lesson is that the old number was derived from one scenario, the join burst, and the scenario that broke it was a different one nobody had done the sum for.
The key is the design, not the limit
Raising the ceiling is a patch. The actual fix is that most of these limiters should never have been keyed on an IP at all.
A limit exists to stop one client from doing something too often. "Client" is the thing you have to name correctly, and for a player-facing endpoint in a shared room, the client is the player, not the network they are on.
So the hot endpoints key on the participant. Answer submission is keyed on the participant id. The poll that backs the WebSocket fallback is keyed on the participant id. The host's controls are keyed on the authenticated user id. What stays IP-keyed is the stuff where an IP genuinely is the unit of abuse: sign-in attempts, password resets, and the global bot guard.
The difference in behaviour is total. An IP-keyed limit on the poll endpoint is not a limit on any client, it is a limit on the room, and it gets stricter the more popular your product is. Six hundred requests a minute from one IP is either an attack or a successful quiz night, and the request headers cannot tell you which. The participant id can.
Which also means the generous per-player limits are still tight. One player gets one accepted answer per question, enforced by a unique constraint in the database rather than by the limiter, so the limiter only has to leave room for retries and for a few questions passing inside one window.
Fail open, and say so
try {
const { success } = await ratelimit.global.limit(ip)
if (!success) {
// ...
}
} catch {
// Redis unavailable, fail open so the app stays up
}
The limiter runs on Upstash Redis. If Redis is unreachable, this skips the check rather than throwing.
That is a real decision with a real cost: during a Redis outage we have no rate limiting. The alternative is worse. A rate limiter that fails closed converts an outage of a protective dependency into an outage of the entire product, which means an attacker who can degrade your Redis can take you down without touching your app. Protection you cannot serve traffic without is not protection.
Locally there is no Redis at all, and every limit() call returns { success: true }, so the app runs in development without anybody configuring anything.
One more line in the limiter config, for a reason that only shows up on a bill:
// No analytics. It costs an extra Upstash write on every limit() call,
// doubling the command count on the hottest path in the app, for data
// nothing in this repo ever reads. proxy.ts also never awaits the
// returned `pending` promise, so on Vercel the write was liable to be
// torn down mid-flight anyway.
analytics: false,
The bug I like best: our error page rate limited itself
Browsers do not want a JSON 429. So a limited request that looks like a page load is redirected to a friendly page, and everything else gets the status code and a Retry-After header:
const isHtmlRequest = request.headers.get('accept')?.includes('text/html')
if (isHtmlRequest) {
const url = request.nextUrl.clone()
url.pathname = RATE_LIMITED_PATH
url.search = ''
return NextResponse.redirect(url)
}
return NextResponse.json(
{ error: 'Too many requests. Please slow down.' },
{ status: 429, headers: { 'Retry-After': '60' } }
)
For a while /too-many-requests was limited like every other path. Follow that through.
A browser trips the limit and is redirected to /too-many-requests. It requests /too-many-requests. That request is also over the limit, so it is redirected to /too-many-requests. Which is where it already is.
The user never saw the page. They saw ERR_TOO_MANY_REDIRECTS. And every hop burned another token, so the sliding window never drained and the loop sustained itself.
/** The page browsers are sent to when they trip the global limiter. */
const RATE_LIMITED_PATH = '/too-many-requests'
if (request.nextUrl.pathname !== RATE_LIMITED_PATH) {
// ... check the limit
}
Four words of condition. The general form is worth keeping: any page you redirect to as a consequence of a rule must be exempt from that rule. It applies to login redirects, consent walls, maintenance pages, region blocks, anything whose own URL is the destination of its own enforcement.
And it must not be auth-gated either
There is a second way to break the same page, and it is the one we were already primed to catch.
/too-many-requests appears in two lists that look like they contradict each other. It is in the auth gate's public allowlist, and it is in robots.txt under Disallow.
export const PUBLIC_ROUTES = [
// ...
'/too-many-requests',
] as const
export const CRAWLER_DISALLOW = [
'/dashboard',
'/play',
'/api',
'/auth',
'/subscribe',
'/webhook',
'/too-many-requests',
] as const
They are answers to different questions. "May a crawler spend its budget here" is no, obviously, it is an error page. "May an unauthenticated request reach it" is yes, necessarily, because a signed-out visitor hitting the limit is precisely who gets sent there. Had it not been in the allowlist, a limited visitor would be redirected to the limit page and then redirected again to /login, which is a worse version of the loop and considerably more confusing to debug.
Have a look
- pub-trivia.app/too-many-requests loads for anybody, with no session, and it is the one page on the site that is exempt from the rule it exists to explain.
- pub-trivia.app/robots.txt shows it disallowed, next to the six other prefixes crawlers have no business in.
- If you want to see the thing all of this protects, pub-trivia.app runs the quiz night. The first session is free and needs no card, which is also the point at which you discover whether your venue's WiFi is one of the hostile ones.
If you run anything that serves a room of people on one network, go and look at what your limiters are keyed on. The number is usually the thing people argue about, and the key is almost always the bug.
Top comments (1)
Sizing the fallback for the whole room is a useful correction. Does the load test also cover everyone switching to polling at the same instant, rather than six evenly spaced polls per minute? I'd test the initial synchronized burst, 429 responses with Retry-After, and recovery back to WebSockets. Per-participant limits protect fairness, while jitter in the fallback schedule could help avoid a room repeatedly hitting the global guard together.