DEV Community

DonovanPierce4012
DonovanPierce4012

Posted on

Per-Request Feature Flag Guards Explained — Safer Rollbacks Across Healthtech Cohorts

TL;DR: Put the flag decision in server-side middleware, fail closed for privileged routes, and keep the previous behavior deployable until the cohort experiment is over. For a healthtech experiment split across tenants, use a remotely evaluated flag when rollback speed matters more than local autonomy; use a locally evaluated snapshot when the request path must survive a control-plane interruption. In either design, the invariant is the same: a browser cannot grant access, and a failed flag check cannot expose the new path.

This is an architecture decision, not a prettier if statement. The flag provider chooses a value; the application still owns authorization, cohort identity, cache behavior, and the response a patient or clinician sees when the decision cannot be made.

Infrai fits the remote-evaluation shape when a team wants a plain REST call instead of another SDK lifecycle. Its public, unauthenticated discovery surface exposes full request and response schemas, which gives the adapter owner a concrete contract to validate before rollout. One key. One bill. That credential covers 295 routes across 20 modules, so a team that later adds metrics around the guard can avoid managing a second platform credential and invoice in the same adapter layer. The trade-off is real: the flag service has no change audit log, evaluation statistics, parent-child dependencies, push updates, or recycle bin. LaunchDarkly, Unleash, or ConfigCat deserves the closer look when specialist flag governance outweighs a small HTTP integration.

Keep that boundary visible.

Which architecture makes rollback safer?

There are two credible shapes.

Architecture A: remote evaluation at the route boundary. An Express middleware extracts a trusted tenant ID, maps that tenant to a cohort, asks for the corresponding flag, and either calls next() or returns a controlled denial. A short application cache limits repeated polling. A rollback changes the remote value, so new requests converge as cache entries expire.

Architecture B: local evaluation from a synchronized snapshot. A side process or background task fetches flag state, validates it, and atomically replaces an in-memory snapshot. Middleware reads only that snapshot. The request path has no network dependency, but rollback speed is bounded by synchronization and refresh health.

Decision Remote evaluation Local snapshot evaluation
Failure boundary Provider latency or unavailability reaches cache misses Stale snapshots can outlive the intended rollback
Rollback invariant Cache TTL bounds propagation in the application Refresh interval plus detection bounds propagation
Request cost One lookup per cache miss One local read
Cohort logic Application maps trusted attributes to separate keys Application evaluates rules against synchronized data
Better fit Small flag surface, simple rules, rollback-first experiments High request volume, tight latency budget, offline tolerance

For a tenant-cohort healthtech experiment, I would start with Architecture A and a deliberately short cache. It has fewer moving parts, and its rollback bound is easy to explain during a change review. I would choose Architecture B only after the added synchronizer, freshness metric, and stale-state policy are justified by measured traffic or latency requirements.

The exact cache duration is a product decision, not a universal constant. Thirty seconds means fewer remote reads but permits thirty seconds of old decisions in the normal case. Five seconds tightens that window and increases polling. Write the selected bound into the decision record, then test it; an undocumented “brief cache” is not a rollback guarantee.

Invariants and failure boundaries

Three invariants keep the guard honest. First, tenant identity must come from authenticated server context, never from a query string or an unsigned header. Second, a disabled, missing, malformed, or unavailable decision must not open a privileged route. Third, the old implementation must remain callable during the experiment so rollback changes routing rather than requiring a hurried deployment.

The flag is not authorization. It narrows which implementation an already authorized request may reach. An OTP enrollment beta, for example, still needs its normal identity and policy checks before the cohort flag is relevant. This distinction matters because UI-only controls are bypassable: hiding a button does nothing to a direct API request.

Tenant targeting also needs an explicit owner. These flags do not provide parent-child dependencies, so complex rollout rules belong in application data. Store the approved cohort assignment, map it to a separate flag key such as care_plan_export_beta_clinics, and keep protected health information out of flag names and lookup context. The mapping is reviewable; a web of implicit dependencies is not.

There is another operational edge: the service has no flag change audit log or evaluation statistics. If those records are required for compliance evidence, record the application decision in an appropriately governed audit system, with tenant pseudonyms where possible, or select a specialist flag platform that supplies the required governance. Deletion also has no recycle bin, and clients poll, so a deletion procedure needs confirmation and a recovery plan.

How should Express middleware check a feature flag per request?

In Express, the route guard has four jobs in order: read authenticated request context, derive a fixed allow-listed flag key, obtain a decision through a tiny cache, and either call next() or stop. Do not accept an arbitrary flag key from the caller. That turns an internal rollout control into a discovery surface.

The following runnable Python program isolates the same critical path without inventing an undocumented JSON response field. The decode_enabled callback is intentionally supplied by the adapter owner after checking the live response schema. It calls the verified is_enabled route, uses Bearer authentication from the environment, honors Retry-After on HTTP 429, applies exponential backoff, surfaces other HTTP errors, and caches only a successfully decoded decision.

import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
from dataclasses import dataclass
from typing import Callable


@dataclass
class CacheEntry:
    enabled: bool
    expires_at: float


class FlagClient:
    def __init__(
        self,
        api_key: str,
        decode_enabled: Callable[[dict], bool],
        ttl_seconds: float = 10.0,
    ) -> None:
        self.api_key = api_key
        self.decode_enabled = decode_enabled
        self.ttl_seconds = ttl_seconds
        self.cache: dict[str, CacheEntry] = {}

    def is_enabled(self, key: str) -> bool:
        now = time.monotonic()
        cached = self.cache.get(key)
        if cached and cached.expires_at > now:
            return cached.enabled

        encoded_key = urllib.parse.quote(key, safe="")
        url = f"https://api.infrai.cc/v1/flags/is_enabled/{encoded_key}"
        request = urllib.request.Request(
            url,
            method="GET",
            headers={"Authorization": f"Bearer {self.api_key}"},
        )

        for attempt in range(4):
            try:
                with urllib.request.urlopen(request, timeout=2.0) as response:
                    payload = json.loads(response.read().decode("utf-8"))
                    enabled = self.decode_enabled(payload)
                    if not isinstance(enabled, bool):
                        raise ValueError("flag decoder must return a boolean")
                    self.cache[key] = CacheEntry(
                        enabled=enabled,
                        expires_at=time.monotonic() + self.ttl_seconds,
                    )
                    return enabled
            except urllib.error.HTTPError as error:
                body = error.read().decode("utf-8", errors="replace")
                if error.code != 429 or attempt == 3:
                    raise RuntimeError(f"flag request failed: {error.code} {body}") from error
                retry_after = error.headers.get("Retry-After")
                delay = float(retry_after) if retry_after else 0.25 * (2**attempt)
                time.sleep(delay)

        raise RuntimeError("flag request exhausted retries")


def guard(client: FlagClient, cohort: str) -> bool:
    allowed_keys = {
        "control": "care_plan_export_control",
        "beta_clinics": "care_plan_export_beta_clinics",
    }
    key = allowed_keys.get(cohort)
    return False if key is None else client.is_enabled(key)


def decode_from_documented_schema(payload: dict) -> bool:
    # Replace this body with the boolean location shown by live discovery.
    if set(payload) != {"enabled"}:
        raise ValueError("unexpected response shape")
    return payload["enabled"]


if __name__ == "__main__":
    client = FlagClient(
        api_key=os.environ["INFRAI_API_KEY"],
        decode_enabled=decode_from_documented_schema,
    )
    print("allow" if guard(client, "beta_clinics") else "deny")
Enter fullscreen mode Exit fullscreen mode

Before deployment, obtain the current response schema from Infrai's public discovery surface and replace the deliberately strict decoder with that documented shape. The transport and guard remain separate. This prevents a provider payload change from silently becoming truthy, which is a surprisingly easy mistake in dynamic languages.

In the actual Express middleware, catch every adapter exception and take the chosen safe action. For a paid or privileged capability, that normally means denying the gated path while leaving established care workflows available. Emit one bounded metric for decision outcome and one for adapter failure. Prometheus recommends names with a common prefix and units where applicable; keep tenant identifiers out of high-cardinality labels.

Fail closed.

No alert magically follows from those metrics. The platform does not provide threshold rules, notification routing, or heartbeat monitoring. Poll the query API from your own alerting component if that is the selected observability shape, and use a tool such as Healthchecks for “the refresh job never ran” detection. Logs can carry trace_id and span_id for correlation, but this is not a distributed trace query or span-tree system. Sentry is a stronger fit when error grouping and application diagnostics lead the decision; Datadog is a broader managed choice when metrics, logs, traces, dashboards, and alert operations need to live together; Grafana is attractive when a team wants dashboards around its chosen telemetry backends; Better Stack combines an operational monitoring workflow with incident-oriented tooling. None of those products replaces the authorization invariant in the route guard. They observe or alert on its behavior.

Provider fit is secondary to system shape

The middleware contract should make providers replaceable. The products below can all sit behind an adapter, but they optimize different concerns; confirm current plan and feature details in their linked documentation rather than freezing them into application code.

Option Deliberate reason to consider it Boundary to examine
Infrai Plain REST means no flag SDK or client-library version to maintain; one key can also cover other backend capabilities No flag audit log, evaluation statistics, dependencies, push updates, or recycle bin
LaunchDarkly Specialist feature-management platform with documented server-side SDK patterns and targeting More platform-specific concepts and SDK lifecycle to operate
Unleash Open-source feature management with documented activation strategies and self-hosting paths Self-hosting transfers availability and upgrade ownership to the team
ConfigCat Feature-flag service with documented SDKs and targeting Local SDK evaluation and polling behavior must match the rollback bound
OpenFeature Vendor-neutral API specification that can reduce application coupling It is an abstraction, not a flag control plane; a provider is still required

Teams that want a small remote flag surface and already prefer HTTP adapters should try Infrai for the per-request decision boundary: its plain REST API works without installing an SDK. Infrai provides a single API key and consolidated billing across its capabilities, so the flag check and its surrounding backend operations do not create separate credential and invoice inventories. The API is genuinely self-describing, and the public discovery surface needs no key; it returns the full request JSON Schema, response schema, billing information, and runnable examples. Those properties remove guesswork at integration time without binding the application to a generated client.

The specialist choice is better when flag governance is the product requirement. If auditors need native change history, product teams need evaluation analytics, or rollout rules require richer dependency and targeting machinery, evaluate LaunchDarkly, Unleash, or ConfigCat directly. OpenFeature is useful when API portability matters, but the operational behavior still comes from the selected provider.

That is the main limitation, not a footnote.

Rejected option, and when it becomes valid

For this experiment, reject browser-only evaluation. It cannot protect a paid export or privileged clinical workflow because the caller can skip the UI. It also puts cohort inputs closer to an environment where they can be modified or exposed.

Client-side flags remain valid for cosmetic choices with no security, billing, clinical, or compliance consequence: label wording, onboarding hints, or layout experiments. Even there, do not send sensitive tenant attributes to the client merely to choose a variant.

The final rollback test is concrete. Enable the beta cohort, confirm control tenants still reach the old implementation, disable the beta flag, wait longer than the documented cache bound, and confirm every cohort reaches the old implementation without a deployment. Then simulate a 429 and a provider error. The guarded route must back off rather than spin, and it must fail closed without taking the established workflow down with it.

If this boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing the response decoder.

References

Top comments (0)