DEV Community

Jeff
Jeff

Posted on Originally published at powerduck.com

Document rate limits in OpenAPI so clients actually back off: 429, Retry-After, and rate-limit headers

Every production API rate-limits. Most document it as a single sentence: "100 requests per minute." The client never learns its remaining budget, never knows when the window resets, and treats a 429 the same as a 500. Autonomous agents are worse: they retry immediately and get banned. A rate-limit contract is not the limit itself; it is the machine-readable feedback that lets a well-behaved client stay under it.

Pick the right status code

Code Meaning Typical cause
429 Too Many Requests The client is over its quota Per-token, per-user, or per-IP rate limiting
403 Forbidden Identity is known but not allowed Plan does not include the endpoint, not a rate issue
503 Service Unavailable The whole service is degraded Overload or maintenance; pair with Retry-After

Do not return 403 for throttling and do not return 429 for a plan restriction. 429 specifically means "too many requests in a window," and it is the one response that must carry retry guidance.

Retry-After is mandatory on a 429

The header takes one of two forms.

Delta seconds, the common case for a fixed window:

HTTP/1.1 429 Too Many Requests
Retry-After: 42
Enter fullscreen mode Exit fullscreen mode

An HTTP date, used when the block lasts until an absolute instant:

Retry-After: Wed, 07 Oct 2026 15:00:00 GMT
Enter fullscreen mode Exit fullscreen mode

Clients and agents implement the delta form trivially (sleep, then retry once). A date needs clock handling, so prefer seconds for ordinary throttling. Never emit Retry-After: 0 on a 429; it tells the client to slam the endpoint again immediately and defeats the limit.

Report the budget, not just the block

The RateLimit-* headers from the IETF rate-limit header draft give the client a running picture on every response, including successful ones.

RateLimit-Limit: 100; window=60
RateLimit-Remaining: 37
RateLimit-Reset: 23
Enter fullscreen mode Exit fullscreen mode
  • RateLimit-Limit is the quota and the window in seconds.
  • RateLimit-Remaining counts down to zero.
  • RateLimit-Reset is seconds until the window refreshes.

Many gateways still emit the older X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset names. Document whichever set your gateway actually sends; do not list headers your edge layer does not produce. If you control the gateway, prefer the unprefixed draft names.

Model it once, reuse it everywhere

Define the shared headers and a problem body, then reference the 429 response from every operation.

components:
  headers:
    RateLimitLimit:
      description: Requests allowed per window for this credential.
      schema: { type: integer, example: 100 }
    RateLimitRemaining:
      description: Requests left in the current window.
      schema: { type: integer, example: 0 }
    RateLimitReset:
      description: Seconds until the window resets.
      schema: { type: integer, example: 23 }
    RetryAfter:
      description: Minimum seconds to wait before retrying.
      schema: { type: integer, example: 42 }
  responses:
    TooManyRequests:
      description: Quota exceeded. Honor Retry-After with exponential backoff and jitter.
      headers:
        Retry-After: { $ref: '#/components/headers/RetryAfter' }
        RateLimit-Limit: { $ref: '#/components/headers/RateLimitLimit' }
        RateLimit-Remaining: { $ref: '#/components/headers/RateLimitRemaining' }
        RateLimit-Reset: { $ref: '#/components/headers/RateLimitReset' }
      content:
        application/problem+json:
          schema:
            allOf:
              - $ref: '#/components/schemas/Problem'
              - type: object
                properties:
                  retryAfterSeconds: { type: integer, example: 42 }
                  limit: { type: string, example: '100; window=60' }
          example:
            type: https://api.powerduck.com/probs/rate-limited
            title: Too Many Requests
            status: 429
            detail: 'Plan free allows 100 requests per 60 seconds.'
            retryAfterSeconds: 42
Enter fullscreen mode Exit fullscreen mode

Then each operation gets the response for free:

responses:
  '200': { $ref: '#/components/responses/InvoiceList' }
  '429': { $ref: '#/components/responses/TooManyRequests' }
Enter fullscreen mode Exit fullscreen mode

State the quota dimension and scope

"100 per minute" is incomplete. Say what the counter keys on, because clients share credentials and run background jobs.

Dimension Example Note
Credential 100 req/min per API key Most common; document shared-account behavior
User 60 req/min per user Relevant when one key serves many users
IP 300 req/min per IP Catches anonymous traffic; flag NAT false positives
Endpoint class 10 expensive exports/min Heavy routes often have their own bucket
Concurrency 2 concurrent exports A separate limit; return 429 or 503 consistently

Put the dimension, window, and per-plan differences in the operation description or a dedicated rate-limit section. A free and a paid tier with the same 429 shape but different limits should show both values.

Verify the backoff before production

A documented 429 is only useful if the client handles it. Drive a mock server from the spec and script the edge cases:

# First call exhausts the window, second returns the throttle contract.
curl -i https://mock.local/invoices \
  | grep -iE 'ratelimit|retry-after'

curl -i https://mock.local/invoices
# HTTP/1.1 429
# Retry-After: 42
# RateLimit-Remaining: 0
Enter fullscreen mode Exit fullscreen mode

Assert three behaviors: the client reads Retry-After instead of using a fixed sleep, it adds jitter so a fleet of throttled agents does not retry in lockstep, and it surfaces a clear error after the window still fails instead of looping forever. A spec-driven mock can return the 429 on demand, which is far easier than flooding a real gateway.

For AI agents, Retry-After is the single most important header in the contract: it turns an unstructured refusal into a scheduling instruction. Agents that parse it pause exactly as long as asked and resume without human intervention.

Common mistakes

  • 429 with no Retry-After, forcing clients to guess.
  • Headers documented in the spec but never emitted by the gateway (always capture real response headers and align the two).
  • Using 503 for per-user throttling, which conflates one noisy client with an outage.
  • Resetting a fixed window such that a client sends 100 requests at 00:59 and 100 more at 01:00. A sliding window or a documented fixed boundary fixes the burst; say which one you run.
  • Hiding plan limits. A client on a paid tier should see its higher quota in RateLimit-Limit, not discover it by being blocked.

Checklist

  1. Throttling returns 429, never 403; overload returns 503.
  2. Every 429 carries a non-zero Retry-After in seconds.
  3. Successful responses carry the RateLimit-* budget headers your gateway actually sends.
  4. A reusable 429 response with a problem body is referenced by every operation.
  5. The spec states the quota dimension, window, and per-plan values.
  6. A mock reproduces the 429 and the client is tested for backoff, jitter, and a final give-up.

When the limit is visible before the block, well-behaved clients never hit the wall.

Describe the throttled response, run it through a spec-driven mock, and confirm your client backs off correctly, all in the browser app. The error body above follows the standard problem-details format; see the RFC 9457 error response guide for the full envelope.

Top comments (1)

Collapse
 
officialmailkr profile image
오피셜메일 •

429를 단순 오류가 아니라 다음 시도 시각을 알려 주는 계약으로 설명한 부분이 좋습니다. 특히 공유 API 키를 쓰는 백그라운드 작업이라면 Retry-After뿐 아니라 제한 기준이 키·사용자·엔드포인트 중 무엇인지 문서와 실제 게이트웨이 응답이 일치하는지 확인해야겠네요. 모의 서버에서 지터와 최종 중단까지 시험하라는 점도 운영에 바로 도움이 됩니다.