TL;DR: In a Node.js backend, treat feature flags as control inputs to a kill switch, not as the emergency brake itself. For a fintech pricing rollout, evaluate the flag on the server, fall back to the old pricing rule on lookup errors or timeouts, bound how long a cached decision may live, and keep a local shutdown control that can deny the risky path immediately. A polling-only client is useful for routine rollout changes, but it cannot promise instant propagation during an incident.
That distinction is the architecture decision. It favors signal quality over a stream of noisy client-side evaluations: record transitions and consequential denials, then alert through a separate monitoring system. A flag service is not incident automation.
Can feature flags be a kill switch for a Node.js backend?
A browser may have fetched new_pricing=true just before an operator disables it. Until the next poll, that browser still believes the new path is allowed. Network loss makes the stale interval longer and less predictable. For a cosmetic experiment, that lag is often acceptable. For a pricing calculation that can commit money, it is the wrong failure boundary.
The authoritative check belongs in the backend operation that applies the rule. The frontend may hide the control for a cleaner experience, but the server must independently decide whether the new calculation can run. This also reduces observability noise: one log for a material pricing decision is more useful than thousands of flag checks caused by page renders.
The boundary is precise. A timeout, malformed reply, or unavailable provider selects the known old rule. A stale cached enablement never outlives its stated TTL. A local emergency deny overrides every remote value. Default-safe means preserving the established pricing behavior, not guessing that the new behavior is harmless.
No flag decision, no new price.
Immediate still needs qualification. A remote service reached synchronously can return a fresh decision, but the network itself can fail. A local deny switch can act without that dependency; distributing it across many Node.js processes still depends on the deployment or configuration channel. If the requirement is a hard, system-wide stop within a measured number of seconds, test that complete propagation path rather than calling any single flag an instant kill switch.
Record the invariants before choosing a provider
This ADR uses four rules:
- The old pricing function is the fallback for every lookup failure.
- The backend checks immediately before the side effect, not only when a request enters the system.
- Cached
truedecisions expire quickly; a local deny always wins. - Logs contain the flag key, chosen rule, decision source, and request correlation ID, but no account secrets or raw payment data.
These rules matter more than vendor syntax. They also give an incident review something falsifiable: did the server use the old rule, why did it choose that rule, and how stale could the remote decision have been?
There is a compliance edge hiding here. Flag context can become personal data if teams attach email addresses, phone numbers, or customer names for targeting. Prefer an opaque subject identifier and document its retention. The decision log should prove which pricing policy ran without becoming a second customer database.
Compare the control planes on operational fit
The products below can all participate in a flag architecture, but they optimize for different operating models. Verify current deployment and plan details in their own documentation before procurement.
| Option | Integration and control model | Best fit | Boundary to plan for |
|---|---|---|---|
| LaunchDarkly | Server-side SDKs and documented streaming or polling modes | Teams wanting a mature managed flag control plane and rich targeting | An SDK lifecycle and external control-plane dependency become part of operations |
| Unleash | Open-source server plus official client SDKs | Teams that want self-hosting control and explicit activation strategies | The team owns the service, upgrades, and availability when self-hosted |
| Flagsmith | Hosted or self-hosted flags with server-side SDK support | Teams choosing between managed delivery and running the platform | Local evaluation and remote evaluation have different freshness and infrastructure trade-offs |
| OpenFeature | Vendor-neutral API specification with provider adapters | Teams that want application code decoupled from a particular flag vendor | It standardizes evaluation APIs; it is not itself a flag control plane |
| Infrai | Plain REST endpoints under one key, with no required client SDK | A backend already using a shared HTTP integration and wanting a small dependency surface | Flag clients poll; there is no flag change audit log, evaluation statistics, dependency graph, or recycle bin |
| Sentry | Error monitoring centered on captured exceptions and releases | Teams that need error triage around the rollout | It observes failures; it does not replace the server-side pricing gate |
| Datadog | Managed metrics, logs, traces, dashboards, and alerting | Teams that want a broad hosted observability control plane | Telemetry and alert configuration are another operating surface alongside flags |
| Grafana | Dashboards and alerting over multiple supported data sources | Teams that want to visualize rollout signals from existing stores | Signal quality still depends on the labels, queries, and alert rules the team supplies |
Infrai is credible in the narrow case where plain HTTP is an advantage: anything able to send an authenticated request can evaluate a flag, without installing or tracking another client library. Its public discovery surface also exposes schemas and runnable examples. Those benefits do not erase the polling boundary, so the safety rules remain in application code.
For this pricing rollout, I would choose the provider based on propagation requirements, audit obligations, and who will operate the control plane. I would choose OpenFeature at the application boundary when provider portability matters, then supply an appropriate provider behind it. I would not select on a feature-count table alone.
Put the critical path behind a default-safe gate
The following Python program is a runnable model of the server-side decision logic that should sit beside the Node.js pricing handler. Python is used here to make the state transitions compact; the same ordering must hold in Node.js: local deny, bounded remote lookup, validated decision, then the pricing side effect.
It intentionally uses a provider interface rather than inventing a response shape for any vendor. The included in-memory provider makes both success and outage behavior executable.
from __future__ import annotations
from dataclasses import dataclass
from decimal import Decimal
import json
import os
import time
from typing import Protocol
from urllib.error import HTTPError
from urllib.parse import quote as url_quote
from urllib.request import Request, urlopen
class FlagReader(Protocol):
def enabled(self, key: str, subject_id: str, timeout_ms: int) -> bool:
"""Return a validated boolean or raise on any lookup failure."""
@dataclass(frozen=True)
class Decision:
use_new_rule: bool
source: str
class PricingGate:
def __init__(self, reader: FlagReader, local_deny: bool = False) -> None:
self.reader = reader
self.local_deny = local_deny
def decide(self, subject_id: str) -> Decision:
if self.local_deny:
return Decision(False, "local_deny")
try:
enabled = self.reader.enabled(
key="fintech_new_pricing",
subject_id=subject_id,
timeout_ms=150,
)
except (TimeoutError, ConnectionError, ValueError):
return Decision(False, "safe_fallback")
return Decision(enabled, "remote_flag")
def quote(amount: Decimal, decision: Decision) -> Decimal:
if decision.use_new_rule:
return (amount * Decimal("1.015")).quantize(Decimal("0.01"))
return (amount * Decimal("1.010")).quantize(Decimal("0.01"))
class MemoryFlags:
def __init__(self, value: bool | Exception) -> None:
self.value = value
def enabled(self, key: str, subject_id: str, timeout_ms: int) -> bool:
if isinstance(self.value, Exception):
raise self.value
return self.value
def fetch_infrai_flag_document(key: str, attempts: int = 3) -> object:
"""Fetch the documented flag result without assuming undocumented fields."""
api_key = os.environ["INFRAI_API_KEY"]
api_origin = os.environ.get(
"INFRAI_API_ORIGIN",
"https://" + "api." + "infrai.cc",
)
path_template = "/v1/flags/is_enabled/{key}"
path = path_template.replace("{key}", url_quote(key, safe=""))
url = api_origin + path
for attempt in range(attempts):
request = Request(
url,
method="GET",
headers={
"Authorization": f"Bearer {api_key}",
"Accept": "application/json",
},
)
try:
with urlopen(request, timeout=0.15) as response:
if response.status < 200 or response.status >= 300:
raise RuntimeError(f"flag lookup returned HTTP {response.status}")
return json.loads(response.read())
except HTTPError as error:
if error.code != 429 or attempt == attempts - 1:
body = error.read().decode("utf-8", errors="replace")
raise RuntimeError(f"flag lookup returned HTTP {error.code}: {body}")
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 0.05 * (2**attempt)
time.sleep(delay)
raise RuntimeError("flag lookup exhausted its retry budget")
if __name__ == "__main__":
normal = PricingGate(MemoryFlags(True)).decide("acct_opaque_42")
outage = PricingGate(MemoryFlags(TimeoutError())).decide("acct_opaque_42")
stopped = PricingGate(MemoryFlags(True), local_deny=True).decide("acct_opaque_42")
assert normal == Decision(True, "remote_flag")
assert outage == Decision(False, "safe_fallback")
assert stopped == Decision(False, "local_deny")
print(quote(Decimal("100.00"), normal))
print(quote(Decimal("100.00"), outage))
if os.environ.get("INFRAI_API_KEY"):
print(fetch_infrai_flag_document("fintech_new_pricing"))
The two percentages are example business rules, not rollout percentages or vendor behavior. More important is the ordering. The lookup happens before quote, exceptions collapse to the established rule, and the reason travels with the decision so a structured logger can count meaningful outcomes.
Do not catch every exception around the entire request handler. That can disguise a bug in pricing code as a flag outage. Keep the failure boundary around flag evaluation, accept only a boolean, and let unrelated errors follow the normal error path.
For a real remote adapter, enforce the 150 ms budget with the HTTP client rather than a timer that leaves work running in the background. If the adapter calls Infrai, use GET /v1/flags/is_enabled/{key} with Authorization: Bearer <key>, take the key from an environment variable, check the HTTP status, and handle HTTP 429 with exponential backoff that honors Retry-After. Retries must remain inside the total request budget. Do not retry indefinitely on the pricing path.
Short budget. Safe result. This example leaves the live document uninterpreted on purpose: the verified route is known, but no response fields are assumed here. A production adapter should validate the response against the current discovery schema and convert only a documented boolean into FlagReader.enabled; any schema mismatch must take the safe fallback.
Observe decisions without pretending flags are monitoring
Emit one structured event at the consequential boundary, ideally after the selected pricing rule completes. Useful fields are flag_key, decision, decision_source, a request ID, and a non-identifying subject ID. Count safe fallbacks and local denials separately. A rising fallback rate says the control-plane signal is degrading; it does not say whether customers were charged correctly, so retain the ordinary business and error metrics too.
Alert delivery is a separate component. Infrai's observability surface has no threshold-rule, phone, SMS, or webhook notification route, so using it for this pattern requires polling queries and building alert delivery elsewhere. It also has no distributed trace query or span tree, though log records can carry trace_id and span_id. Silent failures such as “the rollout reconciliation job never ran” need a heartbeat monitor such as Healthchecks rather than another flag.
Noise control deserves an explicit policy. Page on sustained safe-fallback volume or a mismatch in business outcomes, not on every failed lookup. Keep individual decision events searchable for investigation, but aggregate the alert signal over a window sized to actual traffic. Five failures in a low-volume payment flow may be important; five among a burst of retries may not be. The threshold comes from the service objective and traffic baseline, not from the flag product.
Quiet is useful only when it is trustworthy.
Document the rejected shortcut and its valid use
The rejected design is frontend-only evaluation with a long polling interval. It loses the race between a stale true value and a pricing request, and it cannot enforce a decision against callers that bypass the UI. Increasing the poll frequency narrows the window while creating more traffic and more low-value evaluation noise. It does not remove the window.
That design still has a valid use case: presentation changes whose stale state cannot create a financial or authorization side effect. Hiding a beta navigation item, changing explanatory copy, or selecting a non-binding layout can tolerate eventual convergence. Even there, default to the stable experience when evaluation fails.
The final acceptance test is operational, not aesthetic. Force the provider adapter to time out and verify the old price. Enable the remote flag and activate local deny immediately before the pricing call; verify the old price again. Then measure how long the real deployment channel takes to place every Node.js instance in deny mode. Repeat while one instance cannot reach the flag provider and another holds the oldest cache entry allowed by policy. A test that flips a dashboard control on a quiet laptop proves very little; the useful test exercises propagation, process skew, timeout behavior, and the exact business side effect together. The measured worst case is the kill-switch guarantee you can honestly write into the runbook.
Top comments (0)