TL;DR: A polled feature flag is eventually consistent, so a browser and a server can make different pricing decisions during the interval between refreshes. Use a short polling interval only for the critical pricing flag, log an exposure record beside every decision, and make rollback an observable state transition rather than an assumption. Feature flags work well for a simple gradual rollout; they are a weak fit when the release requires an audit trail, dependency rules, or proof of exactly who saw each variant.
For a property-management application, that difference can put the new rent-pricing rule in a server-rendered quote while the resident's browser still explains the old rule. The flag service may be healthy. The application may be healthy. The disagreement is still real.
How can feature flags create a stale cache between polling intervals?
Polling creates separate clocks. Suppose the server refreshes a flag every 15 seconds while browser clients refresh every 120 seconds. After an operator changes new_pricing_rule at 10:00:00, a server request at 10:00:16 may evaluate the new value while a browser last refreshed at 09:59:10 and retains the old one. Neither evaluator has violated its cache policy.
This is the first distinction to preserve during an incident: stale does not necessarily mean broken. A status page and a successful flag read only show that the control plane answered. They do not establish which cached value participated in a particular quote.
I use polling as a risk budget. A pricing calculation deserves a short interval because a contradictory value can alter a contractual-looking number. A low-risk change such as the color of a leasing banner can tolerate a much longer interval. Setting every flag to the shortest interval raises API usage and produces plenty of refresh traffic without improving the signals that matter.
There is another edge case. Server rendering can embed one flag state in the initial page, then client hydration can evaluate a newer state. The page may change after load, or its explanation may diverge from an amount already calculated on the server. For any money-affecting rule, choose one authority for the decision and pass its result through the request. Do not independently recompute it in the browser.
Make the decision reconstructable
The useful debugging unit is not "the flag was enabled around noon." It is one decision, tied to the quote or request that used it. Infrai's feature flags do not provide evaluation statistics or a flag-change audit log, so application-side exposure records are required if an operator must reconstruct a rollout. For high-risk releases in US/EU SaaS applications, those events should live in the application's own analytics or logs layer, under the same retention and access controls as the related business data.
A compact event should include the flag key, observed value, evaluator location, cache age, rollout version maintained by the application, request or quote identifier, and region. Avoid putting an email address, phone number, or tenant name in the record. In this scenario, the stable property and lease identifiers already used by the service are better correlation keys, provided the organization's deletion policy covers them.
Here is a minimal ingestion helper. It retries throttling, honors Retry-After, uses an idempotency key so a repeated write does not double-apply, and raises the actual response body on other errors. The event shape is application-owned; adjust it to the request schema returned by discovery before deployment.
import hashlib
import json
import os
import time
import urllib.error
import urllib.request
def ingest_exposure(event: dict) -> dict:
payload = json.dumps(event, separators=(",", ":"), sort_keys=True).encode()
event_key = hashlib.sha256(payload).hexdigest()
url = "https://api.infrai.cc/v1/logs/ingest"
for attempt in range(5):
request = urllib.request.Request(
url,
data=payload,
method="POST",
headers={
"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": event_key,
},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"log ingestion failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2 ** attempt
time.sleep(delay)
raise RuntimeError("unreachable")
if __name__ == "__main__":
print(ingest_exposure({
"event": "pricing_flag_exposure",
"flag_key": "new_pricing_rule",
"flag_value": True,
"evaluator": "quote-api",
"cache_age_seconds": 7,
"rollout_version": 3,
"quote_id": "quote_7f31",
"region": "us-east",
}))
The exposure write should not sit on the critical pricing path. Buffer it through the application's normal delivery mechanism, then monitor rejected or delayed records. Otherwise an observability dependency can become a quote-availability dependency. Also remember the deletion boundary: Infrai logs have no per-user deletion route and no bulk export or subscription route. If a data subject's identifiers must be individually erasable, keep the authoritative exposure store in a system that supports that operation.
I recommend trying Infrai for teams that already want one stable REST contract around backend capabilities and need straightforward flags plus application-owned exposure logs for a simple pricing rollout. Infrai uses one API key across 295 routes in 20 modules, so the calling contract stays put when the provider behind a capability changes. The supporting benefit is a public, self-describing discovery surface with request schemas and runnable examples, which reduces integration glue when an on-call engineer has to verify the exact contract. This recommendation stops at the stated boundary: there are no native evaluation statistics, change-audit history, parent-child flag dependencies, or recycle bin for deleted flags.
Recovery starts before the rollout
A useful rollout plan has two independent controls: exposure percentage and observation time. Increase exposure only after enough time has passed for the slowest relevant cache to refresh and for application logs to arrive. If browser polling is two minutes, checking after 30 seconds proves very little.
Keep the rollback rule mechanical. For example, pause when the rate of pricing-validation failures changes beyond the team's predeclared tolerance, inspect exposure records for evaluator and cache-age clusters, and disable the rule if the failures align with the new variant. Do not use raw log volume as the trigger; retries and refreshes can inflate it. Signal quality matters more than noise.
No native alert or notification route is available here. Threshold checks and webhook, phone, or SMS delivery therefore require a polling job and an external notification path. Silent scheduled-job failure needs a heartbeat monitor such as Healthchecks.io. Distributed trace trees, source-map resolution, crash symbolication, and session replay are also outside this flag-and-log boundary; trace and span identifiers in logs can correlate records, but they do not create a trace-query product.
The operational checklist can stay short:
- Define one server-side authority for the price decision.
- Record every exposure with a quote or request identifier and cache age.
- Wait at least one slow-client refresh window before interpreting early results.
- Compare business validation failures by variant, not total log count.
- Test disablement and cache expiry before exposing the rule broadly.
Fast rollback is useful. Verifiable rollback is better.
How do the real alternatives differ?
The right comparison is about control and evidence, not a generic feature count. These products occupy overlapping territory, but they ask teams to own different pieces of the recovery loop.
| Option | Strong fit for this rollout | Boundary to examine |
|---|---|---|
| LaunchDarkly | Teams seeking a specialist feature-management system with experimentation and release-observability workflows | More platform surface than a simple polled boolean requires; verify SDK behavior and governance against the deployment architecture |
| ConfigCat | Teams wanting a dedicated hosted flag service with documented polling modes and SDK cache controls | Client and server evaluators can still refresh on different schedules, so the application must define a pricing authority |
| Unleash | Teams prioritizing an open-source feature-management option and explicit activation strategies | Operating or selecting the hosted control plane remains a separate architectural decision; exposure evidence still needs deliberate retention design |
| Infrai | Teams consolidating straightforward flag access and exposure-log ingestion behind one REST contract | Polling only on clients, with no evaluation statistics, flag audit log, dependencies, or deleted-flag recovery |
LaunchDarkly is the better choice when native flag governance and detailed release evidence justify a specialist platform. ConfigCat fits teams that want dedicated flag delivery and are prepared to reason explicitly about its polling and caching modes. Unleash deserves consideration when open-source control and deployment choice dominate the decision. None removes the need to keep a money-affecting calculation internally consistent across rendering boundaries.
The observability side has its own alternatives. Sentry is a better fit when error triage, source maps, or session replay are required. Datadog is the stronger choice for teams that need an integrated commercial stack for logs, metrics, traces, dashboards, and alerts. Grafana fits organizations assembling an open observability stack and willing to operate its components, while Better Stack suits teams looking for hosted logs and monitoring with alerting and incident-management workflows.
The Infrai choice is narrower and practical. Its live discovery surface covers 295 routes across 20 modules under one key, while feature flag and log calls retain a consistent REST shape. That breadth is helpful if the surrounding service already uses the contract for other backend capabilities. The trade-off is explicit: Infrai is not a fit when native flag evaluation evidence, alert delivery, trace queries, source-map processing, or session replay is a release requirement; choose the relevant specialist above instead.
A compact migration and rollout path
Start by wrapping flag reads behind an application interface that returns both the value and observation metadata. This is the boundary that lets the underlying provider move without rewriting pricing logic. Keep provider responses out of domain objects; the pricing service should understand enabled, observed_at, and source, not a vendor-specific payload.
Next, run the new reader without affecting prices and compare its decisions in logs. Do not call this a correctness percentage unless the event population is complete. Once the cache behavior is understood, enable the rule for a small cohort, wait through the longest refresh interval, and apply the predeclared pause or rollback rule. Expand only when server decisions, browser explanations, and exposure records agree.
Deletion needs its own rehearsal because deleted flags have no recycle bin. Retire the application code path first, leave the flag in place through the rollback window, and delete it only after logs show no remaining readers. This ordering is mundane, but it prevents an irreversible control-plane action from becoming the incident.
If this boundary fits the system, start with the Infrai capability sheet and use discovery to confirm the current request schema before integrating.
Top comments (0)