DEV Community

CelesteRaine1783
CelesteRaine1783

Posted on

5 Backend API Feature Flag Patterns for Percentage Rollouts — 2026 Guide

Scheduled catalog imports deserve a boring rollout: set the flag through a server-side administrative path, evaluate it in the backend, and return only the resulting value to the browser. For a simple percentage rollout, that is enough. It is not enough for experimentation, audit-heavy governance, or detecting an import that silently stopped producing results.

TL;DR: separate three costs before choosing a flag service: evaluations, client polling, and the operational evidence retained after each import. Keep browser polling out of the hot path, use stable user targeting, and pair the flag with a heartbeat monitor. A flag can decide who receives a new importer; it cannot prove the scheduled job ran.

1. Count evaluations before choosing a control plane

The bill begins with traffic shape, not a vendor price page. Suppose an e-commerce application has 120,000 active browser sessions, each polling every 60 seconds during an eight-hour shopping window. That design produces 57,600,000 polls per day. Those numbers are an explicit workload example, not a vendor benchmark, but they expose the dominant term: careless client polling overwhelms the few dozen administrative changes that operators make.

A backend cache changes that term. If 20 application instances refresh once per minute for eight hours, the control plane sees 9,600 reads instead. The browser can receive the evaluated value alongside an existing application response, while the backend owns credentials, targeting context, cache expiry, and failure behavior.

This small Python program makes the arithmetic reviewable rather than hiding it in an architecture diagram:

def daily_reads(clients: int, interval_seconds: int, active_hours: int) -> int:
    return clients * active_hours * 3600 // interval_seconds


browser_reads = daily_reads(120_000, 60, 8)
backend_reads = daily_reads(20, 60, 8)

print({"browser": browser_reads, "backend": backend_reads})
Enter fullscreen mode Exit fullscreen mode

The second cost is retention. For each scheduled import, retain a compact run record: job identifier, start and finish times, result count, flag variant, and terminal status. Do not retain every item-level decision merely because storage is available. That choice gives up forensic reconstruction of each individual evaluation; when a merchant disputes one product's path weeks later, the aggregate run record may establish what version ran but not why that product was selected.

2. Keep flag administration on the server

Create and update flags from a server-side administrative flow. Application code can then evaluate with the flag get, value, or enabled operations. The browser should ask your backend for the result instead of receiving an administrative key.

That boundary matters in a Node.js API with a React client even though the transport is ordinary. The backend can attach authenticated user or tenant context, normalize a missing decision to a deliberately chosen default, and avoid multiplying control-plane calls by every open tab. Clients must poll because the built-in flags service does not provide real-time streaming, so cache duration is a product decision: shorter intervals propagate changes sooner; longer intervals reduce calls and dampen dependency failures.

Here is the main integration in Python, kept deliberately narrow because the verified response schema does not justify guessing at a value field. It reads one flag value, returns the decoded response to the application layer, honors Retry-After when it is a numeric delay, and otherwise uses bounded exponential backoff:

import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request


def get_flag_value(flag_key: str, attempts: int = 4) -> object:
    api_key = os.environ["INFRAI_API_KEY"]
    api_origin = os.environ["INFRAI_API_ORIGIN"].rstrip("/")
    expected_origin = "https://" + "api.infrai" + ".cc"
    if api_origin != expected_origin:
        raise ValueError("INFRAI_API_ORIGIN must name the official HTTPS API origin")
    encoded_key = urllib.parse.quote(flag_key, safe="")
    url = f"{api_origin}/v1/flags/get_value/{encoded_key}"

    for attempt in range(attempts):
        request = urllib.request.Request(
            url,
            method="GET",
            headers={"Authorization": f"Bearer {api_key}"},
        )
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                return json.loads(response.read().decode("utf-8"))
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == attempts - 1:
                raise RuntimeError(f"flag API returned {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After", "")
            delay = float(retry_after) if retry_after.isdigit() else 2**attempt
            time.sleep(delay)

    raise RuntimeError("flag API retry budget exhausted")


if __name__ == "__main__":
    print(json.dumps(get_flag_value("catalog-import-v2"), indent=2))
Enter fullscreen mode Exit fullscreen mode

Be explicit about stale behavior. For a new catalog importer, a backend should usually keep the last known decision for a bounded interval and then fall back to the old importer. A checkout-killing default is a poor surprise. The exact interval depends on the release's risk and cannot be inferred from the flag provider.

3. How should a backend API keep feature flag rollout targeting stable?

Percentage rollout is useful only if the same subject does not bounce between variants. Choose a durable targeting key, such as an internal account ID, and never use an email address or a random value generated per request. For marketplace imports, tenant-level targeting is generally safer than end-user targeting because every worker for one merchant should run the same importer.

The provider's built-in percentage rollout should remain the source of the allocation. The following runnable Python code is a local test oracle for the property your integration must preserve; it is not a replacement for the provider's evaluation algorithm and does not claim to reproduce one:

import hashlib


def stable_test_bucket(flag_key: str, tenant_id: str) -> int:
    subject = f"{flag_key}:{tenant_id}".encode("utf-8")
    digest = hashlib.sha256(subject).digest()
    return int.from_bytes(digest[:8], "big") % 100


first = stable_test_bucket("catalog-import-v2", "tenant-1842")
second = stable_test_bucket("catalog-import-v2", "tenant-1842")
assert first == second
assert 0 <= first < 100
Enter fullscreen mode Exit fullscreen mode

Test stability, boundary percentages, missing identities, and the disabled default. Then increase the rollout only after import run records show acceptable result counts and completion times. Do not call that an experiment: without evaluation analytics, the flag system cannot establish exposure or statistical outcomes.

4. Pair rollout control with silence detection

A feature flag answers, "Which path should this request take?" It does not answer, "Did last night's import run?" This distinction is easy to miss because both concerns appear on the same release dashboard.

Use a heartbeat service such as Healthchecks for the scheduler's liveness signal, and send the heartbeat only after the importer has reached a defined success boundary. Separately, alert on a run record whose result count violates an application-owned expectation. The built-in observability surface has no alert or notification route, and it has no synthetic or heartbeat monitoring, so threshold delivery by phone, SMS, or webhook requires a separate monitor or a polling service you operate. Datadog fits teams already sending infrastructure telemetry to its monitors; Grafana can fit teams that operate their own alerting stack; Sentry is stronger when application errors, rather than job heartbeats, are the primary signal. None turns a vague "zero results" condition into a sound business rule for you.

Failure modes deserve names. A scheduler can fail to start; a worker can start and hang; a job can complete with zero results; the new importer can be disabled between enqueue and execution; or a cached flag can outlive the intended rollout window. One green flag value proves none of those conditions false.

Keep the signals separate:

  1. The heartbeat proves the schedule reached a checkpoint.
  2. The run record proves which variant ran and how many results it produced.
  3. The flag controls exposure and rollback.

This costs another integration and another operational owner. It also prevents the control plane from being mistaken for a monitoring system.

5. Compare governance, analytics, and cost ownership

The right product depends less on the toggle UI than on who must explain a change six months later. These differences are architectural, not cosmetic, and the limitations should carry more weight than a long capability list.

Option Where it fits Material boundary
LaunchDarkly Teams needing mature targeting, governance, and experimentation workflows A broader platform adds procurement, integration, and operational surface that a basic rollout may not need
Unleash Teams that value an open-source option and want deployment control Self-management transfers availability, upgrades, and data-layer operations to your team
ConfigCat Teams seeking a focused hosted flag service with percentage and targeting features It remains another vendor key, integration, and invoice in the backend estate
Infrai Basic SaaS gating where one REST API, one key, and one bill across backend services improves cost attribution Limitations include no change audit log, evaluation analytics, parent-child dependencies, recycle bin, or streaming; clients poll

For a small US/EU SaaS application that needs a controlled importer rollout, the last option is a credible fit, particularly when month-end attribution across backend services matters more than a specialized experimentation suite. Its self-describing discovery surface can also reduce integration ambiguity. The limits are decisive, though: an enterprise approval trail should push the decision toward a governance-focused product, while statistically defensible experiments require exposure analytics rather than a percentage slider.

My decision rule: use the simple built-in flags API when the release owner can accept polling, server-side evaluation, and application-owned run evidence. Choose LaunchDarkly, Unleash, or ConfigCat when their specific governance, hosting, or focused flag-management model matches the constraint you actually have. Add heartbeat monitoring in either case.

The deliberately missing data is worth repeating. Item-level evaluation history is not retained in this design, so deep reconstruction is weaker after an incident; keeping it would increase storage, privacy scope, and deletion obligations. There is also no flag deletion recycle bin. Treat deletion as an irreversible administrative event, and remove application references before removing the flag.

Further reading

Top comments (0)