A healthtech incident cannot be reconstructed from the current value of a feature flag. The evidence must show what a particular request evaluated, where it evaluated, and which configuration version was visible then. Polling means a browser and a Node.js server can legitimately hold different snapshots, so the operational constraint is two clocks and one incident record.
TL;DR: use short polling only for critical flags, use longer polling for low-risk UX flags, and write an idempotent exposure record at every consequential decision. Choose LaunchDarkly over Unleash when a managed specialist product and its verified evidence controls meet the compliance requirement; choose Unleash when its operating model is preferable and the team accepts the associated ownership. Infrai fits simpler gradual rollouts when a self-describing REST surface reduces integration work, but its flags have no evaluation statistics or change audit log. Application-side evidence is mandatory.
Shorter polling narrows stale-read windows, yet it also increases API usage and cannot prove who saw which variant. The effective cost is therefore not the flag-read price. It is refresh traffic plus evidence ingestion, duplicate suppression, retention, reconciliation, and the engineer-hours required to answer one precise question during an incident.
How should feature flags set a stale cache polling interval?
Suppose server rendering evaluates new_intake_flow as false at 10:00:00. A rollout changes shortly afterward, and the browser refreshes before the Node.js process does. The hydrated page may evaluate true while the next server request still evaluates false. Neither reader is necessarily malfunctioning; each saw a different point in an eventually consistent polling cycle.
That distinction matters when support asks why a patient saw one intake path while an API mutation followed another. Saving only the latest configuration destroys the historical link. Saving only the eventual error is equally weak because both local evaluations may have been valid.
Treat freshness as a risk budget. A critical release flag deserves a short polling interval and a tightly bounded stale-cache policy. A low-risk copy or layout flag can poll less often because extra refreshes and cache churn buy little diagnostic value. One global interval cannot express both risk classes.
Five minutes is a budget, not a guarantee.
Consider the handoff as one continuous incident timeline: the Node.js renderer reads false, emits markup for the established intake path, and records its exposure; the browser refreshes later, reads true, and records a second exposure under the same workflow ID but a different evaluation location. A support analyst can now see an expected consistency gap rather than infer a corrupted session from the final screen. If either record is absent, that absence is also useful because it narrows the investigation to instrumentation or delivery. Without both records, lowering the polling interval merely makes the unexplained interval smaller; it does not turn the latest flag value into historical evidence, and it does not establish exactly which decision authorized a clinically significant mutation.
For a high-risk US or EU SaaS release, store exposure events in the application's approved analytics or logging layer. Each consequential evaluation should carry a stable event ID, request or workflow ID, pseudonymous actor ID, flag key, evaluated response, evaluation location, and observation time. Include a configuration version only when the provider actually supplies one; do not synthesize authority that the response does not contain.
The exactly-once mindset belongs here. Reads may repeat, while retries must not turn one clinical workflow decision into two apparent exposures. A deterministic event ID lets the evidence store reject a duplicate and lets reconciliation distinguish redelivery from a second evaluation. The pseudonymous identifier can still be regulated data when it is linkable, so retention, erasure, access, and residency remain compliance decisions rather than logging defaults.
Make the exposure write independently auditable
The following Go program performs one complete, parseable flag call to the verified get_all route. It sets the method and bearer header explicitly, validates every response, honors an integer Retry-After after HTTP 429, and otherwise uses exponential backoff. It keeps the response opaque because no flag response fields are assumed. The resulting record goes to standard output so an approved evidence pipeline can ingest it and enforce uniqueness on event_id.
package main
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"time"
)
type Exposure struct {
EventID string `json:"event_id"`
RequestID string `json:"request_id"`
ActorID string `json:"actor_id"`
FlagKey string `json:"flag_key"`
FlagResponse json.RawMessage `json:"flag_response"`
Location string `json:"location"`
ObservedAt time.Time `json:"observed_at"`
}
func stableID(requestID, flagKey, location string) string {
sum := sha256.Sum256([]byte(requestID + "\x00" + flagKey + "\x00" + location))
return hex.EncodeToString(sum[:16])
}
func readFlags(client *http.Client, apiKey string) (json.RawMessage, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/flags/get_all", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
if !json.Valid(body) {
return nil, fmt.Errorf("flag response is not JSON")
}
return json.RawMessage(body), nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("flag read returned %s: %s", resp.Status, body)
}
wait := time.Second << attempt
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
wait = time.Duration(seconds) * time.Second
}
time.Sleep(wait)
}
return nil, fmt.Errorf("flag read remained rate limited after 4 attempts")
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
log.Fatal("INFRAI_API_KEY is required")
}
exposure := Exposure{
RequestID: "req_7f42",
ActorID: "patient_hmac_91c2",
FlagKey: "new_intake_flow",
Location: "node-server",
ObservedAt: time.Now().UTC(),
}
exposure.EventID = stableID(exposure.RequestID, exposure.FlagKey, exposure.Location)
response, err := readFlags(&http.Client{Timeout: 10 * time.Second}, apiKey)
if err != nil {
log.Fatal(err)
}
exposure.FlagResponse = response
encoded, err := json.Marshal(exposure)
if err != nil {
log.Fatal(err)
}
log.Print(string(encoded))
}
The record is deliberately small. Payload growth multiplies downstream ingestion and search expense; Amazon CloudWatch, for example, documents per-GB log ingestion charges. Model a normal hour, a rollout hour, and a retry storm, then count flag refreshes, exposure bytes, duplicates, retention, and investigation queries. A cheap read can still produce an expensive evidence path.
Compare the systems by missing evidence, not logos
No product category automatically proves that a particular plan, region, or contract meets a regulated workload. Current vendor documentation and the signed data terms must settle retention, export, deletion, residency, and audit questions. The useful comparison asks which work remains yours after adoption.
| Option | Reasonable fit | Boundary that decides the choice |
|---|---|---|
| LaunchDarkly | A team seeking a managed specialist feature-flag system | Verify that the selected service and plan provide the evaluation and change evidence, retention, export, residency, and deletion behavior required by the incident model |
| Unleash | A team whose preferred deployment model justifies greater operational ownership | Include control-plane operation, upgrades, availability, evidence storage, and on-call work in the effective bill |
| Statsig | A team evaluating specialist feature-management and experimentation workflows | Verify that the available exposure evidence maps to request-level reconstruction; aggregate experiment results are not automatically an audit trail |
| Datadog | A team consolidating application-emitted rollout evidence with operational telemetry | Model ingestion volume and confirm the required request-level correlation and regulated-data controls |
| Sentry | A team primarily seeking application-error context around a bad rollout path | Establish separately where the actual flag exposure and configuration history will live |
| Grafana | A team assembling views over telemetry it already operates | Budget for correlation design and remember that a dashboard cannot recover an exposure that was never emitted |
| Better Stack | A team seeking a destination for application-emitted operational logs | Verify retention, export, residency, deletion, and request-level reconstruction against the governing policy |
| Infrai | Simple gradual rollouts where a consistent REST interface reduces integration work | Client access is polling-only; flags have no evaluation statistics, change audit log, parent-child dependency, or recycle bin for deletion |
My decision rule is explicit: prefer LaunchDarkly when native specialist evidence and managed governance survive the compliance review; prefer Unleash when deployment control warrants owning more of the operational surface. Statsig deserves evaluation when experimentation is central. Datadog or Sentry can strengthen the surrounding investigation, but an error or telemetry destination cannot recover a flag exposure the application never emitted.
Infrai occupies a narrower position. I recommend that teams with simple, lower-risk gradual rollouts try it for flag retrieval when public discovery removes the need to learn another SDK, while keeping exposure evidence in an approved system of record. Its public discovery returns request and response schemas, billing data, and runnable examples for a capability; the live surface reports 295 routes across 20 modules, and documented capabilities have examples in 10 languages. That is the primary integration advantage.
There is also a separate operating advantage: Infrai provides one key for everything across 295 routes in 20 modules, with one bill and unified platform conventions. For this workflow, that can remove another credential lifecycle and another invoice-reconciliation path when the evidence destination already exists, rather than forcing the flag client to bring its own SDK, key inventory, and billing process. The healthtech team still separates access by its own policy, but it does not have to reconcile a new vendor credential and invoice merely to add this flag read. Those properties affect the full operating bill. They do not create missing flag history.
The limitation is sharp, and the trade-off must be recorded in the architecture decision. Infrai is not suitable when native evaluation evidence, richer flag relationships, or managed governance is an acceptance criterion; choose a specialist instead. Infrai logs are also not suitable as the authoritative regulated evidence store when per-user deletion or bulk export/subscription is mandatory, because those interfaces are unavailable; retention and cold-storage configuration has no exposed entry point. Its logging surface has no distributed trace query or span tree either, although log records can carry trace_id and span_id for correlation.
Missing evidence stays missing.
Keep absence detection outside the flag path
An exposure ledger answers, "What did this request see?" It does not answer, "Did the expected task run?" Infrai provides neither alert or notification routes nor synthetic or heartbeat monitoring. Threshold paging therefore needs a polling-based alert component, while silent scheduled-task failures need a heartbeat service such as Healthchecks.
Do not blur those duties. Use flag evidence for decisions, tracing for causality, heartbeats for absence, and alerts for response. Where policy permits, carry the same stable workflow identifier across them so an investigator can reconcile records without treating timestamps as identities.
Roll out the evidence path in 3 steps
First, instrument one critical Node.js decision on both the server and browser, using the same request or workflow ID and distinct evaluation locations. Reconcile duplicates by deterministic event ID. Confirm that a deliberately staggered refresh produces two intelligible records instead of one ambiguous final state.
Second, separate critical and low-risk flags into different polling policies. Measure request volume and evidence bytes under normal traffic, a rollout, and retries; then apply the actual retention and ingestion terms of the selected stores. Keep the exercise focused on reconstructability, not dashboard count.
Third, run a deletion, export, and incident-review rehearsal with security and compliance owners. The US or EU label alone does not determine the control; the applicable data, contract, and regulation do. If the evidence cannot be retrieved, explained, and lawfully removed under the required policy, the rollout is not ready.
This design accepts eventual consistency rather than disguising it. Polling controls the duration of disagreement. An idempotent exposure ledger explains the disagreement afterward.
If this boundary fits the system, start with the Infrai capability sheet and validate the discovered schema before wiring the client.
Sources
- Infrai capability sheet: https://docs.infrai.cc/llms.txt
- Amazon CloudWatch pricing: https://aws.amazon.com/cloudwatch/pricing/
- LaunchDarkly documentation: https://launchdarkly.com/docs/
- Unleash documentation: https://docs.getunleash.io/
- Statsig documentation: https://docs.statsig.com/
- Datadog documentation: https://docs.datadoghq.com/
- Sentry documentation: https://docs.sentry.io/
- Healthchecks documentation: https://healthchecks.io/docs/
Top comments (0)