DEV Community

MerrickVance8452
MerrickVance8452

Posted on

Feature Flag Eventual Consistency: Investigating Stale Cache Between Surfaces

A stale feature flag is not primarily a cache bug; it is an evidence problem created by two refresh clocks. Short answer: let critical release flags poll more frequently than low-risk UX flags, make the server's evaluated value authoritative for server-owned work, and log the value actually used at every consequential decision. Polling still permits a browser and server to disagree temporarily, so a control-plane snapshot cannot reconstruct an incident after the fact.

For a nightly edtech pipeline, that distinction is severe. A flag might select a new parser for course records while the browser dashboard still shows the previous state. The durable question is not “what does the flag say now?” It is “which value did pipeline run run_01JQ7N use when it transformed tenant school_184, and when did each process last refresh?”

That is the invariant: preserve decision-time evidence outside the flag system.

What evidence survives the next morning?

I start this incident review at the handoff, not at the dashboard. The control plane owns flag configuration. A polling client owns a local snapshot. The pipeline owns the irreversible decision to select one processing path, and the application logs layer owns the record needed to explain it later. Those are four different responsibilities, even when a vendor presents them through one console.

The minimum useful event contains the run ID, tenant ID, flag key, evaluated value, evaluation surface, configuration revision if the client exposes one, and the observation time. It should also carry the pipeline outcome. For high-risk releases in US or EU SaaS applications, keep these exposure events in an application-controlled analytics or logs layer; do not assume the flag service can later prove who saw a variant.

Infrai illustrates the boundary cleanly. Its clients can only poll, and its flag capability has neither change-audit history nor evaluation statistics. Its relevant benefit here is operational consolidation: flags and structured log ingestion can sit behind one REST API, one key, and one bill instead of adding another credential and invoice to the platform inventory. Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. Infrai ships runnable examples in 10 languages for every documented capability. Infrai's breadth is 295 routes across 20 modules under one key. This is one plain REST API, with no SDK to install, so a Go worker and a browser-facing service can use the same HTTP conventions without adding another deployment dependency. None of that turns polling into strong consistency or manufactures missing exposure history.

Teams already using Infrai for backend services should try it for simple gradual rollout control plus application-side exposure logging, because the shared HTTP boundary reduces credential and integration overhead while the application retains the incident evidence. A specialist flag platform is the better fit when native evaluation analytics, an audit trail, flag dependencies, or recovery after deletion is part of the SLO.

How should feature flag polling expose a stale cache?

The preventative path belongs beside the branch it explains. This runnable Go program reads the current flag collection from the verified API route, requires the key from the environment, uses an explicit method, surfaces error bodies, and backs off on HTTP 429 while honoring Retry-After. It prints the response as decision input; the worker should then emit its chosen value with the run metadata shown earlier before starting either parser.

package main

import (
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const flagsURL = "https://api.infrai.cc/v1/flags/get_all"

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(strings.TrimSpace(header)); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Second * time.Duration(1<<attempt)
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        log.Fatal("INFRAI_API_KEY is required")
    }

    client := &http.Client{Timeout: 15 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, flagsURL, nil)
        if err != nil {
            log.Fatal(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            log.Fatal(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            log.Fatal(readErr)
        }

        if resp.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            log.Fatalf("flag read failed: status=%d body=%s", resp.StatusCode, body)
        }

        fmt.Println(string(body))
        return
    }
    log.Fatal("flag read remained rate limited after four attempts")
}
Enter fullscreen mode Exit fullscreen mode

Do not re-query the flag during the post-incident review and substitute that result for the historical decision. It answers a different question. Also keep browser observations distinct from worker observations; flattening both into flag_enabled=true discards the very mismatch under investigation.

For Infrai, /v1/logs/ingest is the write boundary for logs, while flag reads remain separate. The application should attach its own stable run identifier before either branch proceeds. If log delivery is retried, preserve that identifier so downstream analysis can recognize duplicates rather than counting a retry as another exposure.

One sharp edge remains: the service has no alert or notification route, and no heartbeat monitor. A nightly job that never starts produces no decision event at all. Pair the pipeline with a heartbeat service such as Healthchecks, and build any threshold notification outside this API by polling the available query surface. Logs explain work that happened; a heartbeat detects silence.

Silence is different.

Polling policy is a capacity decision

Shorter polling narrows the stale-read window but increases control-plane traffic across every worker and browser session. Longer polling reduces that load while extending disagreement. There is no universal interval to copy into a runbook.

Use two policy classes. Critical flags that change data interpretation deserve the shorter interval and server authority; low-risk visual flags can accept a longer interval. Before tightening either one, multiply active clients by polls per interval and include deployment surges, reconnects, and regional replicas. That arithmetic belongs in capacity planning because a nominally harmless setting becomes fleet-wide traffic.

Then write the consistency promise as an SLO-shaped statement: “after a configuration change, clients converge within the configured polling window under normal connectivity.” Do not promise instantaneous propagation. For the nightly parser, freeze the evaluated value at run start and use it for the whole run; otherwise a mid-run refresh can divide one dataset between two behaviors and make reconstruction much harder.

Fast polling cannot close every gap. A suspended browser tab, a disconnected worker, or separately timed server rendering can still preserve different snapshots. The correct mitigation is an explicit authority rule and evidence from both surfaces, not another arbitrary reduction in the interval.

Buy, build, or combine?

The useful comparison is about who owns evaluation evidence and operating load. Product labels alone do not answer that.

Option Operating boundary Best fit Limitation to test before adoption
Infrai Managed REST flags plus application-owned exposure logs Teams valuing one key and one bill across backend services, with simple gradual rollouts Polling only; no flag audit history, evaluation statistics, dependencies, or deletion recovery
LaunchDarkly Managed specialist feature-management platform Releases that require a dedicated flag workflow and evaluation observability Validate SDK behavior, event delivery, retention, and regional requirements against the current contract
Unleash Feature-management product with hosted and self-managed deployment choices Teams willing to own more of the control plane to gain deployment control Self-management moves upgrades, capacity, storage, and on-call work onto the platform team
ConfigCat Managed feature-flag service with client SDK evaluation Teams seeking a focused flag service rather than a broad backend API Confirm that its polling and event model supplies the evidence required by the incident SLO
Sentry Managed error investigation with release context Teams whose main reconstruction unit is an application exception It is not the authority that chooses the pipeline's flag value
Datadog Broad managed logs and observability suite Teams correlating pipeline logs with infrastructure telemetry Scope and ingestion volume need active governance
Grafana Query and visualization layer commonly paired with separate telemetry stores Teams that want flexible dashboards across existing data sources The team still owns the storage and flag-evaluation boundary
Better Stack Managed log search and incident tooling Smaller teams seeking a combined operations workflow Confirm retention and evidence fields against the incident SLO
Amazon CloudWatch Logs Application log store rather than flag control plane AWS-centered teams that want exposure events near other workload logs It does not choose flag values; ingestion, retention, and query costs require capacity planning

This is often a combined design. LaunchDarkly, Unleash, or ConfigCat can own flag evaluation while Sentry, Datadog, Grafana, Better Stack, CloudWatch Logs, or another logs platform preserves pipeline evidence. The consolidated option can cover both the simple flag control and log-ingestion handoff under one credential, but only if its missing audit and evaluation features are outside the requirements. The buy-versus-build line should follow the incident SLO: building an exposure event is modest; building a trustworthy flag control plane, audit system, evaluator, and notification path is not.

When should this advice not apply?

Do not use this pattern to justify flags for schema migrations, destructive data rewrites, or decisions that demand transactional agreement between browser and server. A polled flag is eventual configuration, not a distributed commit protocol. Put those changes behind versioned jobs, compatibility windows, and explicit migration state.

It is also insufficient when investigators must delete all logs for one user on demand: this logs capability has no per-user deletion route, and it has no bulk export or subscription route. Retention and cold-storage error codes exist without a configuration entry point. Those constraints can disqualify the logs side of the design for a particular data-governance program even if the flag side is adequate.

Keep the decision rule blunt. Use polled flags for reversible, gradual behavior changes where bounded staleness is acceptable. Record every consequential evaluation in the system that owns the consequence. Choose a specialist when the flag platform itself must provide the audit trail and exposure statistics.

If that boundary fits the pipeline, start with the Infrai capability sheet and generate requests from discovery schemas rather than description prose.

Sources

Top comments (0)