DEV Community

AI Dev Hub
AI Dev Hub

Posted on

Stay under API quotas with a rate limit calculator in 2026

Stay under API quotas with a rate limit calculator in 2026

Divide the quota by its window to get allowed requests per second, multiply your real demand by a retry factor, then budget only 70 to 80 percent of the allowed rate. If demand is above that budget, cut workers or batch calls before you ship. A rate limit calculator does this arithmetic in seconds and shows how much headroom you have left.

The rate limit calculator I link to below is one I built. I tried four quota calculators and spreadsheet templates before writing it, and none of them included retries in the demand figure, which is exactly what bit me. It's free and runs client-side. There's no signup, and nothing about your API gets uploaded anywhere. If you have a better one, tell me and I'll happily point people to it.

The night six workers ate the GitHub quota

In March 2026 I was running a repo sync job against the GitHub REST API. Authenticated requests get 5,000 per hour. That sounds like plenty until you actually do the division: it works out to about 1.39 requests per second, sustained, across everything that shares the token.

We had five workers. Someone (me) added a sixth because a backlog had built up. I did the math in my head, decided "six workers, a few calls each, we're fine," and went to bed.

At 2:13am the alerts started. The job had burned the whole hourly quota in 41 minutes and was getting 403s with x-ratelimit-remaining: 0. My first fix was a sleep(1) between calls. That's the classic workaround, and it failed. The sleep slowed each worker down, but the retry logic still fired on every 403. Retries don't skip the quota line. They count against it exactly like the first attempt did, so the retries kept the token pinned at zero for the rest of the window.

I looked at the logs the next morning. In the final 47 minutes before the window reset, the job made 1,247 requests, and 312 of them were retries of requests that had already failed. Honestly, that number annoyed me for a full day. A quarter of the traffic was the system punishing itself.

The real problem had nothing to do with GitHub being stingy. I never wrote down three numbers together: what the quota allows per second, what my workers demand per second (retries included), and how much gap I wanted between them. Every postmortem I've read about quota exhaustion comes down to some version of that.

How the math works under the hood

There's nothing clever here, which is why it's easy to skip. You convert everything to requests per second, because quotas come in mixed units (per minute, per hour, per day) and comparing them in your head is where mistakes creep in.

The steps:

  1. Allowed rate = quota / window in seconds.
  2. Demand rate = workers x requests per job x jobs per minute / 60, then multiply by (1 + retry rate).
  3. Budget = allowed rate x a safety target (I use 0.8).
  4. Headroom = how far demand sits below the allowed rate, as a percentage.
  5. Max safe workers = budget divided by per-worker demand, rounded down.

Here's the exact function I wrote after that incident. It runs as-is on Python 3.8 or newer:

import math

def headroom(quota, window_s, workers, req_per_job, jobs_per_min,
             retry_rate=0.05, target=0.8):
    allowed_rps = quota / window_s
    demand_rps = workers * req_per_job * jobs_per_min / 60 * (1 + retry_rate)
    budget_rps = allowed_rps * target
    per_worker = demand_rps / workers
    return {
        "allowed_rps": round(allowed_rps, 3),
        "demand_rps": round(demand_rps, 3),
        "budget_rps": round(budget_rps, 3),
        "headroom_pct": round((1 - demand_rps / allowed_rps) * 100, 1),
        "max_workers": math.floor(budget_rps / per_worker),
    }

print(headroom(quota=5000, window_s=3600, workers=6,
               req_per_job=3, jobs_per_min=4))
# {'allowed_rps': 1.389, 'demand_rps': 1.26, 'budget_rps': 1.111,
#  'headroom_pct': 9.3, 'max_workers': 5}
Enter fullscreen mode Exit fullscreen mode

Look at that output. Six workers at a modest 5% retry rate leaves 9.3% headroom. That's technically under the limit, and it's one bad minute away from tipping over. The function says five workers is the safe ceiling at an 80% target. My gut said six was fine. My gut was wrong.

Now plug in what actually happened that night. Once 403s started, the retry rate jumped past 25%. Change retry_rate to 0.25 and demand goes to 1.5 requests per second, which is above the 1.389 the quota allows. That's the whole incident in one line of arithmetic.

I got tired of opening a REPL every time someone asked "can we add another worker?" in Slack, so I put the same logic in a page: the rate limit calculator on aidevhub. You enter the quota and its window, your traffic shape and a retry rate, and it shows allowed throughput, demand and headroom side by side. Mostly I use it to paste a link into a PR description so the reviewer can check my numbers without trusting my mental math.

One thing I'm still unsure about is the right default for the safety target. I use 0.8 because it survived a few traffic spikes for us. I don't have a principled argument for 0.8 over 0.75. If your traffic is very bursty, go lower.

Calculator versus the other ways people do this

Before I built anything, I tried the usual options. Here's how they compared for the specific job of "will this config stay under the quota?"

Approach Setup time Counts retries in demand Easy to share with a reviewer Cost
Mental math 0 seconds Rarely No Free
Spreadsheet template 10 to 20 minutes Only if you add the column Sort of (permissions, versions) Free
Custom Python script 5 to 15 minutes Yes, if you remember Only if they run it Free
Load test (k6, Locust) An hour or more Yes, empirically Yes, via reports Free, but burns real quota
Rate Limit Calculator Under a minute Yes, built in Yes, it's a URL Free

A few honest notes on that table.

The load test wins on truth. It measures what your code actually does, including the retry loops you forgot existed. But load testing against a production API spends the very quota you're trying to protect, and plenty of providers don't offer a sandbox with matching limits. I'd use it to validate a design. I wouldn't use it to answer a quick question at standup.

The Python script is what I'd still recommend if you want this in CI. Drop the function above into a test that fails when headroom_pct falls below 20, and nobody can quietly bump the worker count without seeing the consequence.

The spreadsheet is fine. Really. If your team already lives in one, add a retry column and you've got 90% of what the calculator gives you. My problem with spreadsheets was drift. There were three copies of ours and they disagreed about the window size.

When not to use a rate limit calculator

A calculator assumes your limits are fixed numbers in fixed windows. Plenty of real APIs don't work that way, and that's where this kind of tool will mislead you.

Cost-based or point-based limits. GitHub's GraphQL API and Shopify's GraphQL Admin API both charge per query based on its complexity. A single request can cost 1 point or hundreds. Requests per second is the wrong unit there. You need to estimate query cost first, and a requests-based calculator can't do that for you.

Token-based LLM limits. Most LLM providers cap tokens per minute alongside requests per minute, and the token limit is usually the one you hit first. You can treat tokens as the "request" unit if your prompts are fairly uniform in size. If they vary a lot (say, 400 tokens for some calls and 12,000 for others), an average will hide exactly the spikes that get you throttled.

Concurrency limits. Some APIs cap simultaneous open connections rather than calls per window. Throughput math doesn't capture that, because ten slow requests can hit the cap while your per-second rate looks tiny.

Adaptive or undocumented limits. If the provider adjusts limits on the fly, or just doesn't publish them, any calculation is a guess with decimals attached. Read the response headers instead (x-ratelimit-remaining, retry-after) and back off based on what the server tells you.

Enforcement. A calculator tells you whether a plan fits. It doesn't stop your code from exceeding the limit. For that you still need a client-side limiter, a token bucket in process or something Redis-backed if several hosts share one key, plus retries with exponential backoff and jitter. I learned that the hard way: the math was the cheap half of the fix. The expensive half was rewriting the retry logic so a 403 meant "wait until reset" instead of "try again immediately."

FAQ

Q: What headroom percentage should I aim for?

A: I target at least 20% below the allowed rate, which is the same as budgeting 80% of the quota. If your traffic is spiky, or several services share one API key, go for 30% or more. Single-digit headroom (like the 9.3% in my example) means one retry storm will push you over.

Q: How do I convert a daily quota into requests per second?

A: Divide by 86,400. A 100,000 per day quota is about 1.157 requests per second. Be careful, though: daily quotas often allow bursts within the day, so the per-second figure is a sustained ceiling, not a hard per-second cap. Check whether the provider also has a shorter window stacked on top.

Q: Should retries really count toward demand?

A: Yes. Almost every API counts a retried request exactly like the original. That's the gap I kept hitting in other calculators, and it's the most common reason a config that "fits on paper" blows up in production. Measure your real retry rate from logs if you can. Use 5% as a starting guess if you can't.

Q: Does the calculator handle sliding windows versus fixed windows?

A: It works with average throughput over the window, which is the right first check for both. With a fixed window, though, a burst at the end of one window followed by a burst at the start of the next can briefly push you close to double the average rate. If your provider uses fixed windows and you send bursty traffic, either smooth it out with a client-side limiter or lower your target.

Written with AI assistance and human review. Try the tool at aidevhub.io/rate-limit-calculator.

Top comments (0)