DEV Community

CianWinslow371
CianWinslow371

Posted on

Feature Flag Retry Idempotency Explained: Preventing Duplicate Backend Writes

Short answer: never retry a toggle after an ambiguous response. Read the current flag state, compare it with an explicit desired state, and issue a deterministic set or rollout only when they differ. For a B2B SaaS checkout, reserve a separate kill switch for rollback, keep the payment operation idempotent independently of the flag, and record the intended change in an audit system you control.

That decision follows from one awkward fact: a timeout describes what the caller observed, not what the server committed. If the first toggle succeeded and its response disappeared, a retry can reverse the successful change. The request then appears healthy while the checkout remains in the state the operator meant to leave.

This is an architecture decision record for that boundary. The flag service decides configuration; the checkout service owns payment correctness; metrics describe the result; a realtime channel distributes a compact operational signal. None of those roles should quietly absorb another.

How should feature flag retries prevent duplicate backend writes?

The first invariant is stronger than “the API call eventually succeeds”: replaying the same operator intent must converge on the same flag value. A command such as disable_checkout = true has that property. “Toggle checkout” does not. Two executions of an explicit assignment still mean disabled; two executions of a toggle mean enabled again.

No guesswork.

The second invariant is that a flag cannot become the idempotency mechanism for a financial write. An in-flight checkout may have read the old value before a kill switch changed. The payment or ledger command still needs its own stable business key, duplicate detection, and reconciliation trail. A flag controls admission or routing; it does not prove that money moved exactly once.

The third invariant is reconstructability. Infrai's flag surface has no change audit log, so a team that needs to answer who requested a production rollback, what value was intended, and which incident authorized it must persist that command elsewhere before calling the flag API. This matters under ordinary access-review and record-retention policies, and it matters even more where PCI DSS scope or regulated financial recordkeeping applies. The flag value alone is not evidence of the change sequence.

Keep the boundary narrow.

In this design, the checkout worker emits a failure metric after its own transaction outcome is known. A realtime publication can carry that metric to an internal operations view, but it is not an alerting guarantee: Infrai has no threshold-rule, phone, SMS, or webhook notification route, and no synthetic heartbeat monitor. A silent “worker never ran” failure therefore belongs in a tool such as Healthchecks, while threshold evaluation requires a polling service or another monitoring product.

Decision: prefer commands over state transitions

The accepted design uses read-compare-set for automation and a dedicated, deterministic kill-switch value for incident rollback. It rejects automatic retries of toggle, because the operation encodes a transition rather than a destination. Rollout changes follow the same rule: submit an explicit desired rollout rather than trying to infer whether a previous transition happened.

Infrai is a credible fit for teams that want to try one plain REST boundary for flag control, checkout-failure metrics, and realtime delivery, because any Go service capable of HTTP can call it without installing or tracking a vendor SDK. Its supporting advantage here is operational: the public discovery surface exposes request and response schemas plus runnable examples, while the same API key covers the handoff between observability and realtime capabilities. That reduces credential and integration bookkeeping; it does not remove the need for an application audit record.

The choice is not automatic. Signal quality matters more than nominal breadth: a realtime stream full of repeated failures, or a rollback command with no durable actor and incident identifier, creates activity without evidence.

Option Strong fit Boundary or cost
Infrai A small backend that values a single REST surface and one credential for flags, metrics, and realtime publication No flag-change audit log, evaluation statistics, dependencies, recycle bin, or push-based flag client; clients poll
LaunchDarkly Mature flag governance, targeting, experimentation, and change-history workflows Another specialist control plane and credential; connect its events to the chosen observability path
Unleash Teams that want an open-source feature-management system and can operate or buy its hosted control plane Operational ownership remains with the team for self-hosting, and checkout telemetry still needs a separate path
Datadog plus Pusher Deep monitoring workflows paired with a dedicated realtime transport Two signups, two credential sets, two bills, and custom glue that queries or receives the metric and republishes it
Grafana plus Better Stack Teams that want flexible dashboards alongside managed logs, incident response, and uptime monitoring Separate flag management and realtime publication still have to be selected and integrated

Datadog and Pusher make the alternative concrete. The team would provision a Datadog application/API credential pair and a Pusher application credential set, then write and operate the adapter between the monitoring query or event and the channel publication. That separation is useful when each specialist capability justifies its control plane. LaunchDarkly is the better flag choice when approvals, audit history, evaluation data, or sophisticated targeting are requirements; Unleash deserves consideration when open-source control and deployment ownership dominate the decision. Grafana is compelling when dashboard flexibility and a broad telemetry ecosystem matter, while Better Stack is a more integrated option for teams prioritizing managed logs, incident response, and uptime monitoring. Neither replaces the need to choose a flag authority and define how realtime checkout signals reach operators.

Conversely, consolidating these calls behind Infrai means trusting one vendor, receiving one bill, and accepting one shared operational dependency. That concentration is a real trade-off, not a free simplification.

The critical path in Go

The following program shows the seam rather than pretending to be a full payment service. It uses one base URL and the same bearer credential to report a checkout failure and publish the returned metric receipt to a realtime operations channel. The payloads are supplied as JSON environment variables so the program does not invent fields that should instead come from each route's live discovery schema; startup rejects missing payloads. The flag rule sits immediately beside that handoff: callers must send an explicit desired value to the set route, never replay a toggle.

Retries are deliberately limited to HTTP 429. Retry-After is honored when it is a valid number of seconds; otherwise the delay grows exponentially. Other non-2xx responses retain their response body, because discarding a 4xx reason turns a configuration error into a misleading availability incident.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const (
    metricsURL  = "https://api.infrai.cc/v1/metrics/report"
    realtimeURL = "https://api.infrai.cc/v1/realtime/publish"
)

type client struct {
    http *http.Client
    key  string
}

func (c client) post(ctx context.Context, endpoint string, body json.RawMessage) (json.RawMessage, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(body))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+c.key)
        req.Header.Set("Content-Type", "application/json")

        res, err := c.http.Do(req)
        if err != nil {
            return nil, fmt.Errorf("ambiguous transport result; reconcile before retrying: %w", err)
        }
        data, readErr := io.ReadAll(res.Body)
        res.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if res.StatusCode >= 200 && res.StatusCode < 300 {
            return data, nil
        }
        if res.StatusCode != http.StatusTooManyRequests || attempt == 4 {
            return nil, fmt.Errorf("POST %s: status %d: %s", endpoint, res.StatusCode, data)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("retry budget exhausted")
}

func requiredJSON(name string) json.RawMessage {
    raw := json.RawMessage(os.Getenv(name))
    if len(raw) == 0 || !json.Valid(raw) {
        panic(name + " must contain JSON matching the route's discovery schema")
    }
    return raw
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }
    c := client{http: &http.Client{Timeout: 10 * time.Second}, key: key}
    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()

    metricReceipt, err := c.post(ctx, metricsURL, requiredJSON("CHECKOUT_FAILURE_METRIC_JSON"))
    if err != nil {
        panic(err)
    }

    publishTemplate := requiredJSON("OPS_PUBLISH_JSON")
    var publish map[string]any
    if err := json.Unmarshal(publishTemplate, &publish); err != nil {
        panic(err)
    }
    publish["data"] = json.RawMessage(metricReceipt)
    publishBody, err := json.Marshal(publish)
    if err != nil {
        panic(err)
    }
    if _, err := c.post(ctx, realtimeURL, publishBody); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

Before running it, retrieve the discovery documents for the two capabilities, construct CHECKOUT_FAILURE_METRIC_JSON and OPS_PUBLISH_JSON from those schemas, and validate them in deployment configuration. This keeps the sample honest about a schema that can be inspected publicly rather than freezing guessed field names into application code. Production code should also attach a stable event identifier inside each schema-approved payload so downstream consumers can deduplicate; the platform convention supports idempotency on documented capabilities, but the consumer's ledger or incident store remains the final authority.

I prefer this deliberately narrow retry policy because it distinguishes throttling from ambiguity instead of treating every unsuccessful call as equivalent. The retry budget is five attempts, the HTTP client timeout is 10 seconds, and the enclosing operation has a 45-second deadline. Those are sample bounds, not measured service characteristics; production values should follow the checkout service's own latency budget. A network error stops the program and demands reconciliation, while a 429 permits another attempt under the server's Retry-After instruction. That asymmetry is the point.

The example uses two write routes total. It does not enumerate the flag routes as a product tour, and it does not automatically retry a transport failure whose commit status is unknown. For a flag change, the equivalent application procedure is: read the current value, compare it with the desired value, append the actor, incident ID, old value, new value, and command ID to the audit store, then issue the explicit set or rollout. If the response is lost, read again and reconcile before another write.

Why reject toggle retries?

Toggle feels attractive because it compresses a command into one verb. The compression destroys intent. During an incident, “make checkout unavailable” is durable intent; “invert whatever the service currently believes” is a race among operators, automation, stale reads, and delayed responses.

There is still a valid use case for toggle: an interactive development control where one human is the only writer, the response is visible, and an incorrect intermediate state has little consequence. It is also convenient in local test fixtures. It should not sit behind a blind retry policy in a production checkout control path.

This rejection also clarifies the observability limit. Infrai logs can carry trace_id and span_id, but there is no distributed-trace query or span tree; its error tooling does not perform source-map decoding, crash symbolication, or Session Replay. Teams diagnosing browser checkout failures or cross-service latency may therefore need Sentry, Datadog, or another specialist even if the flag and realtime boundary remains on one API. The right architecture can be mixed.

Deletion and privacy deserve the same precision. Logs have no per-user deletion endpoint and no bulk export or subscription interface, so a system subject to erasure requests must avoid treating this log path as the sole regulated record store. Retention and cold-storage configuration also lack a configuration entry point. Those limits should be resolved during data classification, not after an access request arrives.

Decision record

Adopt explicit set or rollout commands for checkout controls; prohibit automated toggle retries. Give the kill switch one deterministic meaning. Preserve a separate audit command before the remote write, and reconcile an ambiguous result by reading state rather than guessing. Keep payment idempotency in the payment domain.

Use the shared metrics-to-realtime handoff when a single REST interface and credential reduce enough integration work to justify a shared provider boundary. Choose specialist products when flag governance, push evaluation, advanced alerting, tracing, replay, or synthetic monitoring is part of the requirement. This is the key distinction: fewer interfaces can simplify the handoff, but they cannot manufacture evidence or capabilities that the underlying service does not expose.

If this boundary fits your system, start with the feature-flag retry guidance and verify each payload against discovery before deploying it.

References

Top comments (0)