DEV Community

137Foundry
137Foundry

Posted on

The Hidden Cost of Rate Limiting by IP Address Instead of API Key

IP-based rate limiting is often the first thing teams reach for, because it requires no authentication context and works at the edge before a request even hits application code. It's also the rate limiting decision most likely to quietly punish the wrong people, and the cost of that mistake usually doesn't show up until well after the limiter has been live for months, buried in support tickets that never obviously point back to the rate limiter as the cause.

Why IP Feels Like the Obvious Key

Every request has a source IP, so limiting on it needs no extra plumbing and no assumptions about authentication having already happened. A reverse proxy like Nginx can enforce an IP-based limit in a handful of configuration directives, entirely before your application even sees the request, which makes it attractive as a fast, low-effort first line of defense. For an unauthenticated public endpoint with no other identifier available yet, like a signup form or a public search, this is a legitimate and often genuinely necessary choice.

Where It Breaks: Shared Addresses

The problem is that an IP address doesn't map cleanly to a single user or a single client integration. Corporate networks, mobile carrier NAT, and cloud provider egress ranges all mean many distinct users, or many distinct customer integrations entirely, can share one visible IP address from your server's point of view. Limit that IP too aggressively and you throttle every one of them together because of the behavior of just one heavy account sharing the same address.

The Failure Mode Is Invisible to You

This is the part that makes IP limiting quietly dangerous over time: from your side, the metrics look completely fine. Requests from "one IP address" are being correctly capped exactly as designed, and your dashboards show the limiter doing its job. What you can't see from that vantage point is that the IP represents forty different customer accounts behind a shared office network, thirty-nine of whom are now getting throttled because of one heavy account's legitimate traffic. The support tickets, when they eventually come in, describe symptoms that rarely point back to the rate limiter without real investigation.

API Keys Tie the Limit to the Actual Responsible Party

For any authenticated endpoint, the API key or account ID is almost always the better key to limit on instead. It maps one-to-one with the entity actually generating the traffic, regardless of how many people or how many network hops sit behind it on their end. A Redis-backed counter keyed on account ID instead of IP fixes the shared-address problem entirely, at the modest cost of needing the request already authenticated before the limiter can act on it.

The Practical Migration Path

Most teams don't rip out IP-based limiting entirely when they discover this, they layer it instead. A coarse, high-ceiling IP check at the edge still catches unauthenticated floods and obvious abuse before it reaches your app servers at all. A precise, key-based limit inside the application handles everything that gets past the edge, where you actually have the authentication context to know who's genuinely responsible for a given burst of traffic. Cloudflare's edge rate limiting products are a common way to implement that first coarse layer without building the edge infrastructure yourself from scratch.

What to Check If You Suspect This Is Already Happening

If you're getting throttling complaints from customers whose actual traffic doesn't look excessive from their own side, check what key your limiter is actually keying on before assuming the limit number itself is simply too low. It's a five-minute check in your rate limiter's configuration or middleware code that's caught more than one case of "we clearly need higher limits" that turned out to actually be "we're limiting on the wrong thing entirely."

Mobile Apps Make the Shared-IP Problem Worse

Mobile carrier networks compound this problem further, since carrier-grade NAT can put thousands of unrelated mobile subscribers behind a small number of visible IP addresses at any given moment, far more aggregation than a typical office network produces. An API with any meaningful mobile client base that limits by IP is very likely already throttling legitimate users in ways that are almost impossible to diagnose from server-side metrics alone, since the affected users are scattered across carriers and geographies with nothing obviously in common except sharing an egress address none of them chose.

A Concrete Example of How This Plays Out

Picture a customer with fifty employees, all working from the same office network behind a single NAT gateway. Three of those employees happen to use an internal tool built on your API throughout the day, each making their own independent, reasonable amount of requests. From your rate limiter's perspective, keyed on IP, that's one address generating three times the traffic any single reasonable user should, and it gets throttled accordingly, degrading the experience for all three employees even though none of them individually did anything wrong.

Why This Rarely Shows Up in a Code Review

The code implementing an IP-based limiter usually looks completely correct on review, because it is correct, in the narrow sense of doing exactly what it was written to do. The problem isn't a bug in the implementation, it's a mismatch between what the code measures, requests per IP, and what the team actually intended to measure, requests per customer. That gap between intent and implementation is exactly the kind of issue that survives code review, survives testing, and only surfaces once real traffic from real shared networks starts flowing through it in production.

Testing for This Before It Becomes a Pattern

Before shipping an IP-based limit on any endpoint that authenticated users might also hit, think through whether your own team, sitting behind a shared office network, would trip it during normal internal testing. If the answer is yes, real customers behind similarly shared networks will hit the exact same wall, just without the context to know why.

Why This Deserves More Than a One-Line Fix Note

It's tempting to treat "switch from IP to API key" as a trivial one-line config change and move on, but it's worth actually documenting the decision and the reasoning behind it somewhere your team will find later. The next engineer who adds a new endpoint, or a new limiter for a different part of the API, benefits from knowing this was a deliberate choice made after real customer pain, not just an arbitrary default someone picked without thinking it through. A short comment in the code, or a line in your internal API design guidelines, is usually enough to keep the lesson from having to be relearned the hard way a second time on a different endpoint.

A Quick Way to Check Your Own API Right Now

If you maintain a public API today, it's worth checking this directly rather than assuming your setup is fine. Look at your rate limiter's configuration or middleware code and confirm, in plain terms, what value it actually keys on for authenticated requests. If the answer is "the request's IP address" for any endpoint that also requires an API key or auth token, you have an easy, low-risk win available: switch the key to the authenticated identifier and watch whether your 429 rate drops noticeably over the following week without changing the numeric limit at all.

At 137Foundry we've seen this exact pattern more than once: a client raises their rate limit twice in response to complaints, the complaints don't meaningfully stop, and the actual fix turns out to be re-keying from IP to account, not the number itself. If you want the fuller picture on algorithm choice and key selection together, we covered it in a longer guide to rate limiter design.

Top comments (0)