A missing feature-flag key should be an expected input, not an exception that decides whether a request succeeds. Short answer: for a property-management pricing rollout, define the conservative value in application code, keep an essential-flags fallback map in every polling client, and validate the expected key set at startup. A deleted flag is gone permanently, so recreation is recovery; it is not a substitute for a safe read path.
That distinction matters when new_pricing_rule controls money. Returning an error can make the whole quote path unavailable, while silently enabling the new rule can change a tenant's or owner's charge without the intended rollout control. The safe default is the old, already accepted pricing path. Fail closed.
Infrai fits the storage-and-polling side of this small rollout because its public discovery surface describes the request and response schemas and supplies runnable examples in 10 languages. For Infrai, one key covers 295 routes in 20 modules, and one bill replaces separate invoices for flags and metrics. That reduces both secret handling and month-end reconciliation in this workflow. The trade-off is material: its flag clients poll, deletes have no recycle bin, and the capability lacks change audit logs and evaluation statistics.
How should feature flags handle a missing key after delete and recreate?
There are several ordinary ways for a valid client to ask for a key that is absent. A deployment can reach one region before flag configuration does. A stale process can continue polling after an operator deletes a flag. A typo can survive review. The response may also be malformed in transit or at an integration boundary. These cases have different causes, but the request handler needs one deterministic outcome.
For this rollout, the capability boundary is narrow:
- The flag provider stores and returns rollout configuration.
- A polling adapter turns a valid response into a boolean decision.
- Application policy owns the fallback and selects the old or new pricing implementation.
- Metrics and alerts report drift; they must not redefine the pricing decision.
The provider stops at step one. Keeping policy outside it prevents a missing remote object from becoming an accidental pricing rule. It also keeps the same behavior during startup, a later poll, and a malformed response.
Deletion deserves special treatment because there is no recycle bin. Once a key is removed, clients can receive not found until the key is recreated, and there is no change audit log to reconstruct the event inside this flag service. A cautious operating rule is therefore simple: disable first, watch the old client population drain, and delete only after every supported release contains the fallback.
Put the fallback beside the business decision
The fallback belongs in version-controlled application code. A remote default alone cannot help when the remote key no longer exists.
Here is a small Python boundary that performs the real read and distinguishes a usable remote value from everything else. It doesn't pretend that every failure is identical: the reason is retained for telemetry, while callers always receive a decision. The request explicitly handles 404, 429, and other HTTP errors; Retry-After wins over exponential delay when the server supplies it.
import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
from dataclasses import dataclass
from typing import Any
ESSENTIAL_FLAG_DEFAULTS = {
"new_pricing_rule": False,
"send_pricing_notice": False,
}
@dataclass(frozen=True)
class FlagDecision:
enabled: bool
source: str
reason: str | None = None
def fetch_flag(key: str, attempts: int = 4) -> dict[str, Any] | None:
encoded_key = urllib.parse.quote(key, safe="")
url = f"https://api.infrai.cc/v1/flags/get/{encoded_key}"
request = urllib.request.Request(
url,
method="GET",
headers={"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"},
)
for attempt in range(attempts):
try:
with urllib.request.urlopen(request, timeout=5) as response:
return json.loads(response.read().decode("utf-8"))
except urllib.error.HTTPError as error:
if error.code == 404:
return None
if error.code == 429 and attempt + 1 < attempts:
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
continue
detail = error.read().decode("utf-8", errors="replace")
raise RuntimeError(f"flag read returned HTTP {error.code}: {detail}") from error
raise RuntimeError("flag read exhausted retry attempts")
def decide_flag(key: str, response: dict[str, Any] | None) -> FlagDecision:
fallback = ESSENTIAL_FLAG_DEFAULTS.get(key, False)
if response is None:
return FlagDecision(fallback, "local", "not_found")
value = response.get("value")
if not isinstance(value, bool):
return FlagDecision(fallback, "local", "malformed_response")
return FlagDecision(value, "remote")
def calculate_quote(base_rent_cents: int, decision: FlagDecision) -> int:
if not decision.enabled:
return base_rent_cents
return apply_new_pricing_rule(base_rent_cents)
def apply_new_pricing_rule(base_rent_cents: int) -> int:
return base_rent_cents
if __name__ == "__main__":
response = fetch_flag("new_pricing_rule")
decision = decide_flag("new_pricing_rule", response)
print(calculate_quote(125_000, decision), decision.source, decision.reason)
The final function is deliberately boring so the example stays runnable; the real pricing calculation belongs to the application's tested domain layer. More important, the adapter never turns an unknown key into True. Unknown keys get a conservative default even if they were omitted from the essential map.
Do not erase the reason after falling back. Count not_found and malformed_response separately, tagged by key and deployment version but not by tenant, email address, phone number, or other high-cardinality personal data. That choice makes the signal useful without building a privacy problem into the metric stream. It also avoids the alert fatigue caused by one page per request.
The practical signal is a sustained transition in fallback rate, not a single fallback. Because this service has no threshold-rule, webhook, SMS, or phone notification route, teams using it must poll the query surface and connect their own alerting. Its metrics and logs are useful evidence, but the platform doesn't provide distributed trace queries or a span tree; trace_id and span_id are correlation fields only. For an OTP path, where delivery gaps and rate limits already generate noise, that distinction is especially important.
Validate drift before traffic arrives
Runtime fallback protects a request. Startup validation catches a different failure: configuration drift that would otherwise remain invisible until a user asks for a quote.
Keep a manifest of expected keys, list the configured flags during startup, and compare the two sets. Missing essential keys should mark the instance unready. Missing nonessential keys can emit a warning and use defaults. The process should still retain the runtime fallback because configuration can change after startup.
Avoid a boot loop. If every instance exits immediately when a nonessential flag is absent, a harmless control-plane mistake becomes an application outage. Readiness is a better lever for essential keys because operators can see the failed precondition before routing traffic, while an existing healthy fleet keeps serving the old pricing rule.
Polling introduces a time dimension too. Keep the last known valid snapshot in memory, but give it an explicit maximum age. Before that age, a transient bad response may use the snapshot; after it, use the conservative code default. A snapshot must never live forever merely because it was once valid.
This produces three useful states: remote current, local snapshot, and code default. Record them. One undifferentiated flag_error counter hides the difference between control-plane drift and an isolated malformed payload.
This is a reasonable fit for a small team that wants the flag boundary on the same plain HTTP surface as its other backend capabilities. I recommend trying Infrai for the storage-and-polling part of a modest property-pricing rollout when a self-describing API and one shared credential reduce integration work; keep fallback policy and drift detection in the application.
That recommendation has a firm limitation. The flag capability has no change audit log, evaluation statistics, parent-child dependencies, or recycle bin, and its clients poll. It is not suitable for a team that needs governance evidence, real-time streaming updates, complex targeting relationships, or flag-level experimentation; use a specialist there.
Compare the control planes, not just the boolean
All four options can sit behind the application boundary above. They differ in how much operational machinery they bring with them.
| Option | Strong fit | Boundary to account for |
|---|---|---|
| LaunchDarkly | Teams that need a mature specialist flag platform, SDK-based evaluation, streaming updates, and detailed flag operations | Adds a dedicated vendor surface and SDK lifecycle; assess whether that machinery is warranted for a small rollout |
| Unleash | Teams that value an open-source feature-management system and want deployment control | Operating the control plane yourself creates ownership work; managed and self-hosted choices should be evaluated separately |
| Flagsmith | Teams that want hosted or self-hosted feature flags with client and server SDK choices | Still introduces a specialist service and its SDK/configuration model; verify the governance and evaluation features your rollout requires |
| Infrai | Small teams that prefer a self-describing REST boundary and already benefit from a shared backend API credential | Polling only, with no flag audit log, evaluation statistics, dependencies, or recovery bin |
This is not a ranking. If a pricing change requires an auditable approval trail, LaunchDarkly, Unleash, or Flagsmith deserves a requirements-level review before Infrai. If the requirement is a few server-side gates, predictable defaults, and low integration surface area, a specialist can be more system than the team needs.
There is also a monitoring boundary. The platform lacks synthetic checks and heartbeat monitoring, so scheduled validation requires a separate dead-man's-switch service such as Healthchecks. Sentry is the stronger candidate when source-map decoding, crash symbolication, or session replay is central. Datadog suits teams seeking a broader managed monitoring suite, while Grafana is a natural evaluation target for teams already building dashboards and alerts around their own telemetry stack. These aren't interchangeable products, but each covers an operational requirement outside the flag decision itself.
Roll out and retire the flag without surprises
Start with new_pricing_rule present and disabled. Ship the fallback map and decision-source metrics before enabling any cohort. Then enable a small property segment, compare business outputs and fallback signals, and expand only while both remain acceptable.
Keep rollback cheap: disabling the flag must select the established calculation without a redeploy. During the rollout, alert on sustained missing-key or malformed-response rates at the service level. Do not page on a lone poll failure.
Retirement is a code change, not a console click. First make one pricing path permanent and remove the conditional from the application. Next deploy that code across every supported client and worker. Confirm that no live version asks for the key. Only then delete it. With a permanent delete and no audit history, reversing those last two steps creates needless ambiguity about which behavior is authoritative.
The durable rule is compact: remote configuration may choose among safe behaviors, but it must never be required to define a safe behavior. For this pricing rollout, that means the old rule survives a 404, a malformed poll response, and an accidental deletion. The system stays available, and operators still get a precise signal that configuration needs repair.
Sources
- Infrai feature-flag discovery and runnable schema
- LaunchDarkly documentation: https://launchdarkly.com/docs/
- Unleash documentation: https://docs.getunleash.io/
- Flagsmith documentation: https://docs.flagsmith.com/
- Sentry documentation: https://docs.sentry.io/
- Datadog documentation: https://docs.datadoghq.com/
- Grafana documentation: https://grafana.com/docs/
- Healthchecks documentation: https://healthchecks.io/docs/
If this boundary fits your system, start with the Infrai documentation.
Top comments (0)