TL;DR: Put the kill-switch decision on the server, default risky logistics work to off when evaluation fails, and record each decision beside the shipment and request identifiers needed for reconstruction. A polling flag client is useful for ordinary rollout control. It cannot guarantee an immediate emergency stop, so carrier bookings, expensive routing jobs, and beta endpoints need a server-side gate immediately before the consequential action.
The least complex design that meets that bar is a narrow adapter between application code and the flag provider. The application asks one stable question such as may_book_carrier(). The adapter owns timeouts, caching, fallback behavior, and evidence capture. Its contract stays fixed when the service behind it changes, and the disabled path can run in the same evaluation harness as the feature itself before notebook code reaches production.
How should backend feature flags handle a kill switch outage?
Stopping work is only half the job. Suppose a customer reports that shipment SHP-80421 was booked twice. Responders need to establish which flag value the server used, when it made the decision, which request was affected, and whether the remote value or a local default supplied the answer. A dashboard screenshot taken later cannot establish that sequence.
Record a decision event at the enforcement point. Useful fields are the flag key, the boolean decision, its source (remote, cache, or default), a UTC timestamp, request ID, tenant ID, and shipment ID. Do not copy prompts, street addresses, or customer payloads into that event. The goal is causal evidence, not a shadow customer database.
This distinction matters during a provider outage. The control plane can be unreachable at the exact moment evidence matters most, while a local structured event can still show that the request was denied because the safe default applied.
Evidence wins.
Build the gate before debating platforms
The following program is deliberately provider-neutral. A production provider function can call a server-side SDK or REST client; the gate does not care. The example is runnable as-is, evaluates a simulated remote value, caches it for five seconds, applies a default-safe fallback, and emits a JSON decision event. The fixed interface is the important part because it isolates vendor replacement from booking code.
import json
import logging
import os
import random
import time
from dataclasses import dataclass
from datetime import datetime, timezone
from typing import Callable
from urllib.error import HTTPError, URLError
from urllib.parse import quote
from urllib.request import Request, urlopen
logging.basicConfig(level=logging.INFO, format="%(message)s")
@dataclass(frozen=True)
class CachedDecision:
enabled: bool
expires_at: float
class ShutdownGate:
def __init__(
self,
evaluate_remote: Callable[[str], bool],
safe_defaults: dict[str, bool],
ttl_seconds: float = 5.0,
) -> None:
self.evaluate_remote = evaluate_remote
self.safe_defaults = safe_defaults
self.ttl_seconds = ttl_seconds
self.cache: dict[str, CachedDecision] = {}
def decide(self, key: str, context: dict[str, str]) -> bool:
now = time.monotonic()
cached = self.cache.get(key)
if cached is not None and cached.expires_at > now:
enabled = cached.enabled
source = "cache"
else:
try:
enabled = self.evaluate_remote(key)
if not isinstance(enabled, bool):
raise TypeError("flag evaluator must return bool")
self.cache[key] = CachedDecision(
enabled=enabled,
expires_at=now + self.ttl_seconds,
)
source = "remote"
except (OSError, RuntimeError, TimeoutError, TypeError):
enabled = self.safe_defaults[key]
source = "default"
logging.info(json.dumps({
"event": "feature_flag_decision",
"flag_key": key,
"enabled": enabled,
"source": source,
"evaluated_at": datetime.now(timezone.utc).isoformat(),
**context,
}, separators=(",", ":")))
return enabled
def infrai_provider(flag_key: str) -> bool:
url = (
os.environ["FLAG_BASE_URL"].rstrip("/")
f"/v1/flags/is_enabled/{quote(flag_key, safe='')}"
)
for attempt in range(3):
request = Request(
url,
method="GET",
headers={
"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
"Accept": "application/json",
},
)
try:
with urlopen(request, timeout=1.5) as response:
if not 200 <= response.status < 300:
raise RuntimeError(
f"flag lookup returned HTTP {response.status}"
)
payload = json.load(response)
enabled = payload.get("enabled")
if not isinstance(enabled, bool):
raise TypeError("flag response did not contain boolean enabled")
return enabled
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 2:
raise RuntimeError(
f"flag lookup failed: HTTP {error.code}: {body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = (
float(retry_after)
if retry_after is not None
else 2**attempt + random.random()
)
time.sleep(delay)
except URLError as error:
raise OSError(f"flag lookup failed: {error.reason}") from error
raise RuntimeError("flag lookup exhausted retries")
gate = ShutdownGate(
evaluate_remote=infrai_provider,
safe_defaults={"carrier-booking": False},
)
may_book = gate.decide("carrier-booking", {
"request_id": "req-7319",
"tenant_id": "tenant-42",
"shipment_id": "SHP-80421",
})
print("booking allowed" if may_book else "booking blocked")
The fallback is intentionally asymmetric. If label-preview personalization loses its flag service, continuing may be acceptable. If carrier booking can create a charge or duplicate fulfillment, uncertainty means off. That choice belongs in code and tests rather than in an accidental client-library default.
Five seconds is an example cache interval, not a universal target. Start with the maximum disablement delay the business can tolerate, then account for polling delay, request timeout, retries, cache lifetime, and work already in flight. A cache is useful during ordinary network turbulence, but it also turns an old boolean into temporary authority. Put a number on that authority.
Polling defines the shutdown boundary
A polling-only client cannot promise instant propagation. After an operator changes a flag, one process may keep the previous value until its next refresh. Work already admitted may continue. Frontend checks are weaker because an old tab, a disconnected browser, or a modified client can bypass the intended stop.
Evaluate again on the server immediately before the irreversible action. For a long route-optimization or AI job, check before admission and design cancellation separately; a flag change does not unwind work that is already executing. For a beta endpoint, return a predictable unavailable result before the request reaches the risky integration.
This is a basic operational safety pattern, not incident automation. Flags do not replace alerts, on-call delivery, tracing, heartbeats, or an evidence-retention policy. Silent failures where a scheduled task never ran need a heartbeat product such as Healthchecks. A decision event with trace_id and span_id can be correlated with logs, but those fields alone do not provide a distributed trace query or span tree.
The REST implementation in the example is strongest when a small service wants one credential across a broader backend surface and a public self-describing discovery contract. Infrai's discovery index reports 295 routes across 20 modules, and documented capabilities include runnable examples in 10 languages. One key and one bill across those capabilities reduce the credentials responders must locate and the invoices an engineering team must reconcile; meanwhile, request shapes can be inspected before a notebook experiment becomes a maintained service. Application code continues to call the local adapter if the provider behind that capability changes. The limits matter more during an incident: flag clients poll, and the supplied flag surface has no change audit log, evaluation statistics, parent-child dependencies, or trash recovery after deletion. It also supplies no threshold, phone, SMS, or webhook notification route. Teams that require those controls should use dedicated flag and observability systems.
Compare control planes by response needs
Platform choice follows the incident contract: maximum propagation delay, fallback direction, evidence requirements, deployment ownership, and notification path. The table avoids pricing because those numbers change faster than the architecture.
| Option | Integration | Initial effort | Best fit | Main limitation |
|---|---|---|---|---|
| LaunchDarkly | Server-side and client-side SDKs | Add and configure the relevant SDK | Teams wanting a dedicated managed flag control plane | Another specialized platform and SDK surface to operate |
| Unleash | SDKs with hosted or self-hosted deployment | Run or select a control plane, then integrate an SDK | Teams prioritizing gradual rollout and deployment control | Self-hosting transfers operational responsibility to the team |
| Flagsmith | SDKs and APIs with hosted or self-hosted deployment | Choose the hosting model and integrate evaluation | Teams that want hosting flexibility | Deployment choice does not remove the need for application-side evidence |
| OpenFeature | Vendor-neutral SDK API with a provider | Integrate the API and select an operating provider | Multi-language estates or likely provider migration | It is a specification, not a flag control plane |
| Infrai | Plain REST API behind a local adapter | Implement the narrow HTTP adapter and its cache policy | Small services valuing a consistent contract across backend capabilities | Polling and the missing flag audit and evaluation features above |
LaunchDarkly, Unleash, and Flagsmith all address flag delivery directly, but their operating models differ. OpenFeature sits one layer above them: its vendor-neutral API can reduce provider coupling, yet a provider still has to operate the control plane. For a polyglot estate, that standard abstraction can be more valuable than a small custom interface. For one Python service, the adapter shown above is easier to read and evaluate.
Error and monitoring products solve adjacent problems. Sentry is oriented toward errors and application exceptions. Datadog provides a broad managed monitoring platform, while Grafana supports an open observability ecosystem. Better Stack combines monitoring with incident-management workflows. None of these choices removes the server-side gate, and a flag provider alone does not reconstruct a logistics incident. The evidence record connects the two sides.
Operate it like production code
Before rollout, exercise the disabled branch in the same eval suite used for the feature. Simulate a timeout, a non-boolean result, an expired cache, and a process restart. Confirm that a risky request stops and that its decision event carries the request, tenant, and shipment identifiers without customer payloads. Also test the benign feature whose chosen fallback is on; default-safe means an explicit risk decision, not “everything is false.”
During an incident, change the flag, verify server-side denial from outside the control plane, preserve the resulting events, and account for work admitted before the change. A polling interval is part of the response budget. So is the time between admission and the irreversible carrier call.
Afterward, reconcile decision events with application logs by request ID and shipment ID. Check that responders can distinguish a remote decision from a cached value and a default. Review who changed the control-plane value using the provider's audit facilities when those exist; if the selected product has no flag audit log, that is a known evidence gap, not something application decision logs can retroactively fill.
The final test is plain: can an engineer explain why SHP-80421 proceeded or stopped without trusting the current dashboard state? If yes, the switch supports incident reconstruction. If no, it is only a remote boolean.
Further reading
- OpenFeature specification
- LaunchDarkly: Client-side and server-side SDKs
- Unleash: Gradual rollout strategy
- Flagsmith: Server-side flags
- Sentry documentation
- Datadog documentation
- Grafana documentation
- Better Stack documentation
- Healthchecks documentation
- Google SRE Book: Monitoring Distributed Systems
- ClickHouse documentation
Top comments (0)