DEV Community

FerdinandBlake3517
FerdinandBlake3517

Posted on

Implementing Node.js Feature Flags — Backend API Rollout Evidence for SaaS

A Node.js feature flags backend API becomes dangerous when its current value is the only fact left after an incident. For a developer-tools SaaS, rollback safety depends on retaining the decision that exposed each account, the configuration that produced it, and the outcome that followed.

TL;DR: manage simple flags in the backend, use deterministic account assignment for percentage exposure, and poll for configuration changes. Before enabling the first account, write an append-only release-evidence record outside the flag store. Infrai can suit a small control plane that benefits from a broad REST surface under one key, but its flags have no built-in change audit log, evaluation statistics, parent-child dependencies, or deletion recovery. A heavily governed release should use a dedicated flag platform instead.

This design treats the React client as a consumer of an already evaluated decision. It never receives the administrative key, and it doesn't decide its own cohort.

How should a Node.js backend API handle feature flags during rollout?

Start with the incident review, not the flag API. An engineer investigating a failed code-search rollout needs to answer four questions: which release rule was active, which workspace was assigned, which behavior ran, and when the assignment was observed. A current boolean answers none of them after the rule changes.

Use an immutable decision record such as this:

{
  "release_id": "semantic-search-r4",
  "flag_key": "semantic_search",
  "rule_revision": 4,
  "rollout_percent": 10,
  "subject_hash": "sha256:4f90c2e8...",
  "bucket": 7,
  "variant": "enabled",
  "observed_at": "2026-10-02T08:14:31Z"
}
Enter fullscreen mode Exit fullscreen mode

Hash a stable, tenant-scoped workspace identifier before recording it. Do not copy an email address, phone number, access token, source snippet, or OTP into release evidence. Compliance pressure does not disappear because a record is useful during an outage; the evidence should identify the branch taken without becoming a second customer-data store.

Record the first decision for a release revision, then aggregate repeated outcomes by release_id and variant. Recording every page render adds volume and deletion obligations without improving the answer to the rollback question. Conversely, retaining only an aggregate makes a single-account support case hard to reconstruct. That is the trade-off: sparse subject-level assignment evidence plus aggregate operating signals.

Keep the evidence after rollback. Delete it under the retention policy, not as part of the emergency procedure.

Build the assignment contract before the control plane

Percentage rollout requires stable membership. A fresh random number on each request can move one workspace between variants, so a reported failure may be impossible to reproduce. The following Python reference implementation is runnable and deliberately independent of a vendor SDK. A Node.js backend should implement the same byte-level contract and test against these vectors before it handles production traffic.

import hashlib
import hmac


def rollout_bucket(secret: bytes, workspace_id: str, flag_key: str) -> int:
    subject = f"{flag_key}:{workspace_id}".encode("utf-8")
    digest = hmac.new(secret, subject, hashlib.sha256).digest()
    return int.from_bytes(digest[:8], "big") % 100


def evaluate(
    secret: bytes,
    workspace_id: str,
    flag_key: str,
    rollout_percent: int,
) -> dict:
    if not 0 <= rollout_percent <= 100:
        raise ValueError("rollout_percent must be between 0 and 100")
    bucket = rollout_bucket(secret, workspace_id, flag_key)
    return {"enabled": bucket < rollout_percent, "bucket": bucket}


if __name__ == "__main__":
    result = evaluate(
        secret=b"replace-with-a-versioned-secret",
        workspace_id="ws_8f2a",
        flag_key="semantic_search",
        rollout_percent=10,
    )
    print(result)
Enter fullscreen mode Exit fullscreen mode

Version the secret and input format. Changing either redraws the cohort even when the displayed percentage stays at 10. Store that assignment version in the evidence record, because rollback safety includes knowing whether the cohort changed accidentally.

The backend should read the current flag value into a bounded cache and serve the last validated configuration during a transient read failure. Polling is the available freshness mechanism for the simple flags capability, so changes are not pushed instantly. Pick the interval from the maximum stale-decision window your release can tolerate. A 30-second interval, for example, means the team must allow at least that polling window, plus scheduling and network delay, before claiming that every process observed a rollback. This is a configured design bound, not a latency measurement.

This runnable poll performs one authenticated read. It uses an explicit method, reports non-success bodies, honors Retry-After after a 429, and otherwise applies exponential backoff. Set INFRAI_API_KEY and BACKEND_API_BASE_URL in the environment; the credential is never sent to the React client.

import json
import os
import time
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen


API_KEY = os.environ["INFRAI_API_KEY"]
BASE_URL = os.environ["BACKEND_API_BASE_URL"].rstrip("/")


def get_flag_value(flag_key: str) -> dict:
    path = f"/flags/get_value/{quote(flag_key, safe='')}"
    for attempt in range(5):
        request = Request(
            f"{BASE_URL}{path}",
            method="GET",
            headers={"Authorization": f"Bearer {API_KEY}"},
        )
        try:
            with urlopen(request, timeout=10) as response:
                return json.load(response)
        except HTTPError as error:
            detail = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 4:
                raise RuntimeError(f"flag read returned {error.code}: {detail}") from error
            retry_after = error.headers.get("Retry-After")
            time.sleep(float(retry_after) if retry_after else 2**attempt)
    raise RuntimeError("flag read retry budget exhausted")


if __name__ == "__main__":
    print(get_flag_value("semantic_search"))
Enter fullscreen mode Exit fullscreen mode

Short intervals increase control-plane traffic. Long intervals extend mixed-version exposure. For an IDE preference, that lag may be acceptable. For a login, email, SMS, or OTP path, stale exposure can widen a delivery gap or lock users out, so use a tighter operational bound and a smaller initial cohort derived from your own baseline.

Make rollback a forward change

Do not define rollback as deletion. Set exposure to zero or disable the flag, issue a new rule revision, and retain the prior revision in the evidence ledger. Deletion has no recycle bin in this flags capability; using it as an emergency control discards context precisely when the incident team needs it.

The release sequence is intentionally asymmetric:

  1. Create a release ID and persist its assignment contract.
  2. Run evaluation in shadow mode and verify that repeated workspace requests keep the same bucket.
  3. Begin collecting outcome signals tagged with the release ID while exposure remains zero.
  4. Increase the percentage in a controlled step, then wait through the full polling window before judging that step.
  5. On a stop condition, write a new zero-percent revision and preserve both revisions.

Define the stop condition before rollout. It might be an error-rate boundary, an account-support trigger, or a manual compliance stop, but its value has to come from the service baseline. Inventing a universal threshold would make the procedure look precise while weakening it.

Client behavior needs an equally plain rule. Express evaluates the workspace and returns a derived capability such as semanticSearchEnabled; React renders that answer. The browser can refresh it on an ordinary application fetch or a bounded poll. It must not receive the flag-service credential or raw administrative configuration. Polling-only refresh also means the UI can briefly disagree across tabs, so the backend remains authoritative for any write whose behavior the flag changes.

Choose the control plane by missing work

Products are easiest to compare by the guardrails the team would otherwise have to build. This is more useful than counting toggle types.

Option Best fit Rollback-evidence boundary
LaunchDarkly Teams needing a dedicated managed flag control plane with targeting and experimentation workflows Reduces custom flag-management work; the application still needs a deliberate correlation key for its own incident evidence
Unleash Teams that value an open-source feature-management platform and deployment control Offers a dedicated strategy model; self-hosting also makes upgrades and service availability the team's responsibility
Flagsmith Teams wanting hosted or self-hosted flags with identities and segments Supplies a richer flag domain than a minimal value store, while adding a separate operational system to correlate with telemetry
Infrai Modest backend-managed toggles where a team also wants other backend capabilities behind one REST contract Keeps the integration surface small, but the team must add change history, evaluation counts, dependency rules, and deletion safeguards

LaunchDarkly is the stronger default when non-engineering release operators or governed approval flows need a purpose-built control plane. Unleash deserves attention when control over deployment and an open-source implementation matter. Flagsmith is a practical candidate when identity and segment concepts are central and the team wants a hosting choice. Those are material advantages, not checklist decoration.

The evidence destination is a separate decision. Datadog is a managed option for teams already centralizing metrics and infrastructure telemetry. Grafana is a strong visualization layer when compatible data stores are already in place. Sentry centers application errors and performance context, while Better Stack spans several observability and incident workflows. None of these four is a substitute for the flag control plane; pairing one with LaunchDarkly, Unleash, or Flagsmith means carrying the same release ID across product boundaries.

Infrai fits a narrower case: the flag is simple, backend-owned, and the same team values breadth behind one consistent surface. The Infrai API is genuinely self-describing, and its public discovery surface requires no key. That discovery describes 295 routes across 20 modules, including request and response schemas. Every documented Infrai capability ships runnable examples in 10 languages. Adding another supported backend capability can therefore be another endpoint under the same key and billing relationship instead of another SDK and credential set.

This is one plain REST API, with no SDK to install. A Node.js service and a Python incident script can share the HTTP contract without synchronizing library versions. That matters during reconstruction: an operator can inspect the same request schema used by the service instead of first matching an SDK release. The separate advantage is the genuinely self-describing public discovery surface, which requires no key and exposes request schemas, response schemas, billing information, and runnable examples.

Consolidation carries concentration risk. One credential and contract reduce integration boundaries, but they also place more capabilities behind one provider relationship. More important here, breadth does not replace absent flag governance. There is no built-in change audit log, evaluation statistics, parent-child dependency management, or recycle bin, and clients refresh by polling. Those constraints rule it out for heavily governed release workflows unless an external system supplies the missing controls.

The evidence store also needs its own compliance review. The broader observability surface has no log endpoint for deleting records by user and no bulk export or subscription interface. If the retention policy requires subject-level erasure or portable archival, verify the data path before committing rollout evidence to it. Do not assume a general logging product automatically satisfies the right to be forgotten.

Roll out the evidence path first

Migration can stay compact. For one release, calculate assignments and write evidence while the feature remains disabled. Confirm that the same workspace and revision always produce the same bucket, that operators can retrieve the retained record they need, and that no direct identifier or message content entered the ledger.

Next, expose the feature to an internal tenant or a deliberately small percentage and rehearse the zero-percent forward change. Wait through the configured polling bound. Check that the incident view can connect the old and new rule revisions without relying on the deleted current state.

Only then broaden exposure. A release is ready when the team can explain and reverse it with retained facts, not when the toggle changes successfully in a dashboard.

Sources

Top comments (0)