Treat a feature flags API 429 Too Many Requests rate limit as a shared-capacity signal, not as permission for every Node.js process to start its own retry loop. Use one poller per deployment boundary, honor Retry-After, add bounded jitter when the server does not supply a delay, and let application processes read a last-known-good snapshot locally. For a B2B SaaS experiment, attach every refresh attempt and evaluated snapshot to a tenant cohort and a stable configuration version; otherwise the retry traffic is impossible to attribute and the experiment comparison is suspect.
Short answer: centralize polling, preserve the last valid configuration through transient 429s, and charge network work to the cohort set served by that poller. The operational limit is an SLO decision: the maximum acceptable configuration age must be shorter than the product's rollout tolerance, while the request budget must remain below the API's demonstrated capacity. If those two conditions cannot both hold, faster retries are not a fix.
This guide uses a small Go polling controller beside a Node.js service because the controller is an operational boundary rather than application business logic. The same state machine can live inside a Node.js worker, but it should still have a single owner, a shared snapshot, and measurable retry behavior.
How should a feature flags API client handle a 429 rate limit?
Polling load scales as instances x cohorts x requests per interval. Forty application instances polling 12 cohort-specific documents every 30 seconds produce 16 requests per second before retries, deployments, or health-check synchronization are counted. That arithmetic is a capacity-planning input, not a benchmark or a claim about what any flag API supports.
A 429 response means the origin is rate limiting the client, and it may include Retry-After; RFC 6585 defines the status, while HTTP semantics defines Retry-After as either a delay in seconds or an HTTP date. A client that ignores that field substitutes its own opinion for the server's explicit recovery signal. A client that retries every cohort immediately makes the overload worse.
Synchronization is the quieter failure. Identical intervals wake replicas together, and identical exponential backoff keeps them together. Then an autoscaling event multiplies the crest. Random jitter spreads attempts over time, but jitter does not create capacity, so cap retries and retain the old snapshot instead of waiting indefinitely in a request path.
Stale flags have a cost too. A five-minute-old snapshot may be acceptable for a gradual UI experiment and unacceptable for a kill switch. Define separate age objectives before choosing the polling interval:
| Signal | Decision it supports | Example policy, not a universal default |
|---|---|---|
| Snapshot age | Can evaluations continue locally? | Alert before the experiment's agreed staleness limit |
| 429 ratio | Is the poller exceeding available capacity? | Burn-rate alert against a refresh-success SLO |
| Attempts by cohort set | Who caused network cost? | Allocate requests by tenants represented in the snapshot |
| Version divergence | Are cohorts being compared on the same input? | Exclude windows with mismatched versions |
Do the multiplication first. It often changes the design.
Build one bounded polling state machine
The controller below polls one configuration document, accepts only successful responses, parses both legal Retry-After forms, and continues serving the last valid body. It uses full jitter when the response supplies no delay. All code paths put a deadline on network work.
package main
import (
"context"
"errors"
"fmt"
"io"
"math/rand"
"net/http"
"strconv"
"strings"
"sync"
"time"
)
type Snapshot struct {
Body []byte
Version string
FetchedAt time.Time
}
type Poller struct {
client *http.Client
url string
interval time.Duration
maxDelay time.Duration
mu sync.RWMutex
snapshot Snapshot
attempt int
}
func retryAfter(value string, now time.Time) (time.Duration, bool) {
value = strings.TrimSpace(value)
if seconds, err := strconv.ParseInt(value, 10, 64); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second, true
}
when, err := http.ParseTime(value)
if err != nil {
return 0, false
}
if delay := when.Sub(now); delay > 0 {
return delay, true
}
return 0, true
}
func (p *Poller) fallbackDelay() time.Duration {
shift := p.attempt
if shift > 6 {
shift = 6
}
ceiling := time.Second * time.Duration(1<<shift)
if ceiling > p.maxDelay {
ceiling = p.maxDelay
}
if ceiling <= 0 {
return time.Second
}
return time.Duration(rand.Int63n(int64(ceiling) + 1))
}
func (p *Poller) refresh(ctx context.Context) (time.Duration, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, p.url, nil)
if err != nil {
return p.fallbackDelay(), err
}
resp, err := p.client.Do(req)
if err != nil {
p.attempt++
return p.fallbackDelay(), err
}
defer resp.Body.Close()
if resp.StatusCode == http.StatusTooManyRequests {
p.attempt++
if delay, ok := retryAfter(resp.Header.Get("Retry-After"), time.Now()); ok {
if delay > p.maxDelay {
delay = p.maxDelay
}
return delay, errors.New("configuration endpoint rate limited the poller")
}
return p.fallbackDelay(), errors.New("configuration endpoint rate limited the poller")
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
p.attempt++
return p.fallbackDelay(), fmt.Errorf("configuration endpoint returned status %d", resp.StatusCode)
}
body, err := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
if err != nil {
p.attempt++
return p.fallbackDelay(), err
}
p.mu.Lock()
p.snapshot = Snapshot{
Body: append([]byte(nil), body...),
Version: resp.Header.Get("ETag"),
FetchedAt: time.Now(),
}
p.mu.Unlock()
p.attempt = 0
return p.interval, nil
}
func (p *Poller) Current() Snapshot {
p.mu.RLock()
defer p.mu.RUnlock()
copy := p.snapshot
copy.Body = append([]byte(nil), p.snapshot.Body...)
return copy
}
Retry-After wins because it conveys the origin's requested delay. The local maximum prevents an extreme or malformed policy from silencing refresh forever; reaching that cap should page only when snapshot age threatens the SLO, not merely because a retry occurred. A production implementation should also validate the response schema before replacing the snapshot. A 200 response containing unusable configuration is not success.
Keep polling out of request handlers. The run loop below has one timer, a per-attempt timeout, and cancellation for clean deployment shutdown.
func (p *Poller) Run(ctx context.Context) error {
delay := time.Duration(0)
for {
timer := time.NewTimer(delay)
select {
case <-ctx.Done():
timer.Stop()
return ctx.Err()
case <-timer.C:
}
attemptCtx, cancel := context.WithTimeout(ctx, 5*time.Second)
next, _ := p.refresh(attemptCtx)
cancel()
delay = next
}
}
Ignoring the returned refresh error is deliberate only in this loop: evaluation uses Current, while metrics and structured logs must record the outcome. Returning the error to a Node.js request would couple user latency to control-plane availability and encourage parallel retries.
Attribute cost without corrupting the cohort comparison
Tenant identity should not determine polling cardinality unless the upstream document is genuinely tenant-specific. If 800 tenants share the same experiment definition, fetch it once, then evaluate locally with tenant attributes. Per-tenant polling creates 800 units of network work where one could suffice, and it also lets large cohorts dominate retry traffic.
For each attempt, record a low-cardinality cohort-set identifier, outcome, delay source, response status, and snapshot version. Do not put raw tenant IDs into metric labels; that creates a cardinality problem and may expose customer identifiers to systems with broader readership. Keep tenant-level allocation in a controlled event stream or warehouse table, keyed by an opaque tenant identifier.
A defensible allocation rule is proportional to evaluations served from a snapshot during its validity window. Suppose cohort A accounts for 7,000 evaluations and cohort B for 3,000 while the poller makes 20 refresh attempts. Allocate 14 request-equivalents to A and 6 to B. These numbers illustrate the rule; they are not measured results. If a cohort requires a dedicated configuration document, charge its actual attempts directly instead.
The comparison dataset needs four fields at minimum: opaque tenant key, cohort assignment, configuration version, and evaluation timestamp. Join business outcomes only after rejecting records produced with an unknown or over-age snapshot. Otherwise a rate-limited refresh can leave one deployment on version v17 while another observes v18, and the analysis labels infrastructure drift as experiment effect.
Cost attribution follows the unit of shared work. Charging every tenant equally is easy, but it hides heavy evaluators; charging only successful refreshes hides retry amplification. Count all attempts, retain their outcomes, and allocate shared attempts using one declared rule throughout the experiment.
Choose the ownership boundary
There are three reasonable shapes, and none wins without the deployment numbers.
| Shape | On-call effect | Cost attribution | Lock-in and maintenance |
|---|---|---|---|
| Poller inside every Node.js instance | More synchronized clients and more retry state | Simple per instance, weak for shared cohorts | Little extra infrastructure; duplicated control logic |
| One poller per cluster or region | One observable retry budget and shared snapshot | Natural allocation across the served cohort set | Requires leader election or a separately deployed worker |
| Managed relay or proxy | Smaller application surface | Depends on exported metrics and tenant metadata | Transfers operations but adds a service boundary and exit work |
For a platform team, the buy-versus-build question is narrower than a feature checklist. Can the chosen boundary export attempt counts, snapshot age, version, and rate-limit delay without tenant labels exploding metric cardinality? Can it preserve a last-known-good document during control-plane errors? Can the team test and replace it without changing flag semantics? A managed component can reduce on-call work, while a small self-operated controller gives direct control over allocation records; either choice is poor if its evidence cannot support the cohort comparison.
Set a request budget before rollout. With one regional poller at a 30-second steady interval, the baseline is two requests per minute per region. Add an explicit retry allowance and deployment overlap, then test that envelope against a controlled endpoint. Do not infer available capacity from the absence of 429s during a quiet hour.
Verify, deploy, and roll back
Test the state machine with a local server that returns a successful snapshot, then a 429 with Retry-After, then another success. Use a fake clock in unit tests so a one-hour delay does not create a one-hour test. Property tests are useful for three invariants: fallback delay never exceeds the cap, a failed response never replaces the last valid snapshot, and cancellation stops the loop.
During a staged deployment, watch distributions rather than averages: snapshot age by region, refresh attempts by outcome, selected retry delay, and configuration versions serving evaluations. The acceptance criterion should name both sides of the trade-off, for example: the configured percentage of evaluations uses a snapshot younger than the agreed limit, and refresh attempts remain within the planned request budget. Pick the percentage and age from the experiment's risk tolerance; inventing universal SLO numbers would be theater.
Rollback has two separate levers. First, stop the new poller while leaving the last-known-good snapshot readable. Second, restore the previous polling owner only after confirming the new lease or leader has released ownership; overlapping old and new loops can recreate the 429 spike during the rollback itself. Preserve version and cohort allocation records through the change so the experiment analysis can exclude the transition window.
Test the ugly path.
Simulate absent and malformed Retry-After values, an HTTP-date in the past, timeouts, oversized bodies, invalid configuration, process restarts, and two pollers briefly believing they are leader. Confirm that none of these cases clears the snapshot. Then run a capacity test at the planned region and cohort count, because a retry algorithm proved with one process says little about a synchronized fleet.
The final operational rule is plain: serve locally, refresh through one accountable owner, and stop the experiment comparison when snapshot provenance is missing. A lower 429 count is useful, but trustworthy cohort evidence is the outcome that matters.
References
- https://www.rfc-editor.org/rfc/rfc6585.html#section-4
- https://www.rfc-editor.org/rfc/rfc9110.html#name-retry-after
- https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/
- https://martinfowler.com/articles/feature-toggles.html
- https://nodejs.org/api/globals.html#class-abortcontroller
Top comments (0)