DEV Community

SladeBarrett9642
SladeBarrett9642

Posted on

How to Debug Stale Feature Flag Cache: Polling Interval Mismatches

A longer polling interval reduces control-plane traffic, but it also extends the period in which two processes can make different decisions from the same feature flag. For a property-management pricing rollout, the practical answer is to record decision context at evaluation time, compare cohorts rather than isolated log lines, and treat cache age as a bounded operating condition. Short answer: trace the flag key, evaluated value, configuration version, cache age, and evaluation location through one pricing request; then separate propagation lag from targeting errors and code-version mismatch before changing the polling interval.

Do not begin by flushing every cache. That destroys the evidence.

How can stale feature flag cache polling cause client-server mismatches?

A flag evaluation is a decision made from inputs. The server may evaluate new_pricing_rule using a locally cached configuration, while a browser or worker evaluates with a different configuration version, subject attributes, default value, or point in time. Eventual consistency makes brief disagreement possible; an unbounded or unexplained disagreement is the incident.

In the property-management example, suppose the rule applies only to a pilot portfolio. A request can reach the new pricing calculation on the API while an asynchronous lease-document worker retains the old decision. The symptom looks like bad arithmetic, yet the arithmetic may be correct on both sides. The inputs differ. A pricing decision should also be reproducible without logging tenant names, phone numbers, email addresses, or full lease payloads.

The useful mental model has four clocks: configuration publication time, local fetch time, request evaluation time, and business-record commit time. If those timestamps collapse into one generic timestamp, the investigation becomes guesswork. Preserve them separately and use UTC.

Step 1: Instrument the decision

Start with a compact, structured event where the pricing branch is selected. High-cardinality personal data adds exposure and search noise; it does not explain why a flag evaluated differently. A pseudonymous correlation key and a stable property-group identifier are enough to join the path.

from dataclasses import asdict, dataclass
from datetime import datetime, timezone
import hashlib
import json


@dataclass(frozen=True)
class FlagDecision:
    request_id: str
    flag_key: str
    value: bool
    config_version: str
    config_published_at: str
    cache_fetched_at: str
    evaluated_at: str
    evaluation_location: str
    subject_key_hash: str
    used_default: bool


def pseudonymize(value: str, audit_salt: str) -> str:
    material = f"{audit_salt}:{value}".encode("utf-8")
    return hashlib.sha256(material).hexdigest()


def emit_flag_decision(decision: FlagDecision) -> None:
    print(json.dumps({"event": "flag_decision", **asdict(decision)}, sort_keys=True))


now = datetime.now(timezone.utc).isoformat()
emit_flag_decision(FlagDecision(
    request_id="req_7f31",
    flag_key="new_pricing_rule",
    value=True,
    config_version="cfg_1842",
    config_published_at="2026-10-06T08:00:00+00:00",
    cache_fetched_at="2026-10-06T08:00:12+00:00",
    evaluated_at=now,
    evaluation_location="pricing_api",
    subject_key_hash=pseudonymize("portfolio_42", "rotated-audit-salt"),
    used_default=False,
))
Enter fullscreen mode Exit fullscreen mode

The values are illustrative identifiers, not a promised schema. The contract is semantic: config_version identifies what was evaluated, while cache_fetched_at says when this process received it. Record the service deployment version in surrounding telemetry too. Without it, a changed evaluator and changed flag configuration can masquerade as one failure.

Logging only the boolean is nearly useless. true from version A and true from version B may have arrived through different rules. Keep the provenance.

Step 2: Classify mismatch before tuning polling

Take one affected request and reconstruct both decisions. Use this order:

Observation Likely class Next check
Same subject, different configuration versions Propagation lag or stalled poller Compare publish, fetch, and evaluation times
Same version, different subject context Context construction mismatch Trace attribute source and normalization
Same version and context, different result Evaluator or build mismatch Compare service build and rule semantics
One side used a default Fetch, parse, initialization, or missing-key path Inspect error reason and last successful fetch
Decisions match, stored prices differ Downstream pricing or write-path fault Trace calculation inputs

This ordering protects signal quality. Widening logging too early can produce many repetitive events, making the rare mismatched pair harder to find. Capture every mismatch during rollout, but sample routine matching decisions only after confirming the sampler retains configuration versions and default use.

This method has limits. It is unsuitable when a decision must become globally atomic at publication time: polling plus local caches cannot provide that guarantee, regardless of telemetry quality. In that case, put the authoritative decision in the synchronous transaction path and accept its added dependency and latency. Conversely, central evaluation is a poor trade-off for disconnected workers that must continue operating; those need a defined stale-data policy and reconciliation after connectivity returns.

A small analyzer can expose the lag distribution:

from datetime import datetime
from statistics import median


def parse_time(value: str) -> datetime:
    return datetime.fromisoformat(value.replace("Z", "+00:00"))


def summarize(events: list[dict]) -> dict:
    lags = []
    for event in events:
        if event.get("event") != "flag_decision":
            continue
        published = parse_time(event["config_published_at"])
        fetched = parse_time(event["cache_fetched_at"])
        lags.append((fetched - published).total_seconds())
    if not lags:
        return {"samples": 0, "median_seconds": None, "max_seconds": None}
    return {
        "samples": len(lags),
        "median_seconds": median(lags),
        "max_seconds": max(lags),
    }
Enter fullscreen mode Exit fullscreen mode

Do not interpret the maximum alone. A single sleeping worker and a fleet-wide shift demand different responses. Break the distribution down by evaluation location, process build, and configuration version, while keeping dimensions controlled.

Step 3: Set a staleness budget from business risk

The polling interval should follow the tolerated inconsistency window, not habit. A pricing rule that changes quoted rent needs a stricter bound than a cosmetic flag because a tenant can see one amount and receive a document containing another. The appropriate number is a business decision involving pricing owners, compliance, and operations; no universal interval is defensible here.

Define the maximum configuration age allowed in a new price decision. Reserve time within it for publication, polling jitter, retries, and request processing. The poll interval is only one component of observed staleness. Shortening it cannot repair a stopped poller, a worker that never refreshes, or a client evaluating another subject.

Use alerts with distinct meanings. Warn when cache age approaches the agreed budget. Escalate when it exceeds the budget, and route pricing work according to a tested failure policy. That policy might preserve the last known configuration or decline to apply the new rule; choose explicitly. Silent defaults convert an observable dependency problem into apparently valid pricing.

Telemetry cost belongs in the design discussion, but it is not the main decision. Log ingestion is metered in common observability systems, so full evaluation context on every request can create material volume. Remove sensitive and redundant fields first, aggregate healthy-path metrics, sample matching events, and retain complete mismatch events. Dropping configuration provenance first saves bytes while deleting the answer.

Step 4: Prove the diagnosis and roll out

Build deterministic tests around a generic flag interface. Inject configuration versions and time instead of depending on a live control plane. Exercise a fresh cache, one version behind, an expired cache, a missing flag, malformed context, and a worker restarting with no cached state. For the property pilot, include a subject just inside the cohort and one just outside it.

Then stage the rollout. Shadow-evaluate the new rule without changing stored prices. Enable a small named property cohort, requiring paired evidence across the API and document worker. Expand only while mismatch and cache-age distributions remain inside agreed bounds. Stop on unexplained defaults, version splits beyond the budget, or divergent persisted prices.

Add jitter so a fleet does not poll in lockstep, use bounded backoff after failures, and expose last successful fetch separately from last attempt. Test by pausing configuration delivery and observing whether cache age crosses the warning threshold. A rollback is another configuration publication and travels through the same consistency window, so measure it with the same care.

The resulting migration is compact: add configuration version and evaluation location, add publication and fetch times, build cache-age and mismatch views, agree on the staleness budget, then begin shadow decisions. Optimize for explainable disagreement. Once the mismatch is classified, polling becomes a measured capacity-versus-latency choice instead of a ritual response to stale-cache reports.

References

Top comments (0)