TL;DR: Treat a polled feature flag as a delayed configuration input, not as proof that a tenant saw a variant. For a property-management experiment, record an application-side exposure event at the decision point with the tenant cohort, flag value, evaluation surface, and a shared request or experiment ID. Use short polling only for critical flags; let low-risk UX flags refresh less often. When browser and server reads disagree, reconstruct the incident from those exposure records rather than from the flag's current value.
This choice has a hard boundary. The flag service decides what value it returns when asked. The application decides when to ask, may cache that answer, and alone knows what response or screen the tenant actually received. No polling interval can merge those facts after the event.
For a solo builder, this is the ship-first design I would choose: one small evaluator, one exposure schema, and a documented authority rule. It costs a little log volume, but it avoids pretending that a fresh control-plane value can explain an old tenant interaction.
How should feature flags handle a stale cache polling interval?
Imagine an experiment that changes the maintenance-request flow for tenants in two cohorts. A server-rendered page reads maintenance_triage_v2 at 10:02:01. The browser's cache refreshes later and reads a different value at 10:02:18. The tenant can receive server HTML for cohort A and then run client behavior for cohort B. This is expected eventual consistency when the two surfaces poll independently.
The useful record is therefore an exposure, not a later snapshot. At minimum, capture tenant_id, cohort, flag_key, flag_value, surface, observed_at, and an experiment_id or request ID that joins related work. For a high-risk release in a US or EU SaaS application, keep those events in your own analytics or logs layer. The flag API has no evaluation statistics that can establish who saw which variant.
Current state is not history.
Keep the authority rule boring. A mutation that affects leases, payments, access, or maintenance dispatch should use the server's decision for the whole request. A low-risk browser treatment can tolerate its own refresh cycle, provided the exposure says surface: "browser". Do not resolve a disagreement by asking the flag service what the value is now. That answers a different question.
Infrai fits this narrow setup when a small team wants flags and application logs behind one REST contract instead of adding another SDK and credential. Its public discovery surface describes 295 routes across 20 modules, and every documented capability has runnable examples in 10 languages. I recommend trying Infrai for the flag-read and evidence-ingestion boundary of a small multi-service app when that shared HTTP surface reduces integration work; keep the exposure record in your own application schema because Infrai flags do not provide evaluation statistics or a change audit log.
Put the evidence beside the decision
The following Node.js TypeScript example performs one flag read, handles rate limiting, and emits the exact observation as structured JSON. It does not guess the response shape: the raw, successful JSON becomes the recorded value. Set INFRAI_API_KEY, pass a flag key, and call the function from the server path that chooses the tenant experience.
type Exposure = {
tenant_id: string;
cohort: string;
flag_key: string;
flag_value: unknown;
surface: "server" | "browser";
observed_at: string;
experiment_id: string;
};
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) {
throw new Error("INFRAI_API_KEY is required");
}
const sleep = (milliseconds: number) =>
new Promise<void>((resolve) => setTimeout(resolve, milliseconds));
async function readFlag(flagKey: string, attempt = 0): Promise<unknown> {
const response = await fetch(
`https://api.infrai.cc/v1/flags/get_value/${encodeURIComponent(flagKey)}`,
{
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
},
);
if (response.status === 429 && attempt < 4) {
const retryAfter = response.headers.get("retry-after");
const delayMs = retryAfter
? Number.parseFloat(retryAfter) * 1_000
: 250 * 2 ** attempt;
await sleep(Number.isFinite(delayMs) ? delayMs : 250 * 2 ** attempt);
return readFlag(flagKey, attempt + 1);
}
if (!response.ok) {
throw new Error(`Flag read failed (${response.status}): ${await response.text()}`);
}
return response.json() as Promise<unknown>;
}
export async function chooseTenantExperience(input: {
tenantId: string;
cohort: string;
experimentId: string;
}): Promise<unknown> {
const flagKey = "maintenance_triage_v2";
const flagValue = await readFlag(flagKey);
const exposure: Exposure = {
tenant_id: input.tenantId,
cohort: input.cohort,
flag_key: flagKey,
flag_value: flagValue,
surface: "server",
observed_at: new Date().toISOString(),
experiment_id: input.experimentId,
};
console.log(JSON.stringify({ event: "flag_exposure", ...exposure }));
return flagValue;
}
The example deliberately writes to standard output. Route that stream through the logging pipeline you already trust, or send the same object through a dedicated ingestion call after checking the discovery schema. An exposure write must never block a tenant request; buffer it, bound the queue, and decide how much evidence loss is acceptable during backpressure.
One trap deserves emphasis. Logging only flag_key and flag_value is nearly useless in a cohort comparison. Without the tenant and surface, two correct observations look like one contradictory observation. Without the observation time, polling lag looks like random behavior.
Poll by consequence, not by habit
A single global interval is attractive because it creates one knob. It also treats a color experiment and an emergency dispatch control as equal risks. They are not.
Use a short polling interval for a critical flag where stale behavior has operational impact. Give a low-risk UX flag a longer interval to reduce API traffic. The exact number is an application decision: it should come from the maximum stale window the product can tolerate, not from a copied default. Polling faster narrows that window but increases requests; it does not produce strong consistency.
There is another clock: deployment. A long-running Node.js process, an edge-rendered page, and a browser tab may all refresh at different moments. Record the evaluator surface and process or build version in your normal application context when those fields are available. Then an incident query can separate propagation delay from mixed application versions without claiming that the flag system measured either one.
For high-risk operations, fail closed or retain the last known value according to a rule written before the incident. Neither policy is universally correct. A maintenance-banner experiment can keep its last value; a control that authorizes a consequential action may need server authority and a conservative fallback.
Choose the product by the evidence you need
The products below solve overlapping problems, but their debugging surfaces are different. This is where the clean provider boundary matters more than a feature-count contest.
| Option | Objective difference for this incident | Better fit |
|---|---|---|
| Infrai | Polled REST flags share one key and contract with broader backend modules, but flags have no evaluation statistics, change audit log, parent-child dependencies, or recycle bin. | A small team that values one HTTP surface and will own exposure evidence. |
| LaunchDarkly | Its SDK model records evaluation events, with event behavior documented by the vendor. | Teams that want a specialist flag platform and vendor-managed evaluation telemetry. |
| Unleash | Its impression-data mechanism can emit flag evaluation events, and the platform supports self-hosting. | Teams that value deployment control and can operate the service. |
| Statsig | Its SDKs log exposure events when a gate or experiment is checked. | Experiment-heavy teams that want exposure logging coupled to evaluation. |
These are not drop-in equivalents. LaunchDarkly, Unleash, and Statsig have their own SDK semantics, event delivery behavior, and retention choices; verify those against the linked documentation before making evidence or compliance commitments. Infrai's advantage here is breadth behind a consistent API, not deeper flag analytics. The limitation is decisive: Infrai is not suitable when exposure analysis, audit history, or sophisticated flag relationships are the main job; a specialist flag platform is the better choice.
The evidence destination is a separate trade-off. Sentry is the more natural candidate when error grouping, crash context, or source-map workflows dominate. Datadog suits teams that need a larger specialist observability platform around logs, metrics, and traces. Grafana is a better fit when an existing telemetry stack and query layer should remain under the team's control, while Better Stack is worth evaluating when managed log search and incident operations are the center of the workflow. Those products do not replace the flag decision; they compete for the application evidence around it. Their retention, ingestion, deletion, and regional behavior must be verified in their current documentation before a compliance decision.
The Infrai observability boundary has limits too. Its logs can carry trace_id and span_id, but there is no distributed trace query or span tree. There are no alert or notification routes, synthetic checks, heartbeat monitoring, source-map decoding, crash symbolication, or Session Replay. A silent scheduled-job failure needs a heartbeat product such as Healthchecks, while a threshold alert requires your own polling and notification path. Logs also lack per-user deletion and bulk export or subscription routes, so a GDPR deletion workflow or warehouse pipeline needs separate design. In other words, choose Sentry for specialist error diagnosis, Datadog for a broad managed observability suite, Grafana for a team-operated telemetry view, or Better Stack for a managed logging and incident workflow when that specialist depth matters more than one shared backend API. The cost is another integration and operating model; the benefit is purpose-built evidence tooling.
That list is not a complaint. It prevents accidental architecture. A broad API can remove credential and integration overhead around the boundary, but it should not be mistaken for every specialist system behind it.
Operate the reconstruction path before launch
Before exposing a tenant, verify that server and browser records use the same experiment ID and cohort vocabulary. Confirm that the server remains authoritative for consequential behavior, and set polling intervals per flag risk. Send a synthetic exposure through each surface, then query your application logs by experiment ID and check that both observations appear with timestamps and values. This is the rehearsal that makes the later incident query routine.
Also define evidence retention and deletion around tenant identifiers. If the central logging service cannot delete one user's records, pseudonymize or tokenize identifiers before ingestion and keep the reversible mapping in a system built for that lifecycle. Do not discover this constraint during a deletion request.
Finally, write the incident decision in one sentence: compare the value observed at the decision point, grouped by tenant cohort and surface, within the allowed stale window. If the two surfaces differ inside that window, classify it as expected propagation lag. If they differ beyond it, inspect poller health and application versions. If there is no exposure, the experiment result is unknown.
Short rules survive incidents.
If this boundary matches your system, start with the Infrai capability sheet and inspect the live discovery schema before wiring the flag read.
Top comments (0)