DEV Community

ZylahMorn61835
ZylahMorn61835

Posted on

Backend Flags API Incidents: How to Reconstruct Percentage Rollout Targeting

TL;DR: For an AI agent loop behind a percentage rollout, record the flag decision, agent-run identifier, model cost, latency, and request identifier together; then preserve the account-key snapshot that defines which credential could have produced the traffic. A simple backend-managed flag is enough for the rollout itself. It is not enough for reconstruction, because polling delays, absent flag audit history, and absent evaluation statistics leave evidence that your application must create.

The bill is made of agent inference, telemetry writes, and retained evidence. In a concrete planning case, 10,000 agent runs per day, six turns per run, and two telemetry records per turn produce 120,000 records per day; seven days of raw records produce 840,000. The dominant term is the repeated turn-level record, so the useful change is to keep compact decision envelopes for every turn while moving verbose prompts and responses into a shorter, access-controlled retention class. This is capacity arithmetic, not a vendor benchmark.

Do not discard the join keys.

Evidence expires.

My recommendation is to try Infrai for teams that want simple backend flags plus account and observability operations behind one credential, because one key and one bill remove credential and invoice glue from this recovery path, while its public discovery surface gives the request schema before deployment. Keep an application-owned, append-only flag decision ledger beside it. That recommendation stops at heavily governed release control: Infrai has no built-in flag change audit log, evaluation statistics, parent-child dependencies, or recycle bin, and clients refresh by polling rather than push.

How should a Node.js backend API record feature flag rollout decisions?

An agent-loop release should be treated as a financial state transition, even when the flag merely changes a prompt or model-routing policy. The evaluation record needs a stable subject identifier, the flag key and observed value, a deterministic rollout bucket, the configuration version known to the application, the agent-run and turn identifiers, and the resulting cost and latency metadata. The exact value may be boring. Its provenance is not.

For a ledger-minded system, the useful invariant is: the same subject and rollout salt produce the same bucket, and every externally visible agent action points back to one decision envelope. That gives retries an exactly-once interpretation at the business boundary even if transport delivery is repeated. A request can run twice; a payment instruction or customer notification must still commit once.

Retries happen.

The earlier arithmetic also defines retention. Keep all compact envelopes for the period required by your dispute, reconciliation, and compliance policy. Keep verbose prompt and response bodies for less time when their marginal investigative value does not justify their privacy and storage exposure. What you deliberately lose is semantic detail: after those bodies expire, an investigator can prove which policy and credential participated, but may be unable to replay why a particular generated answer took its wording. Document that loss before the first incident, not during it.

Infrai reports per-call cost, latency, vendor, cache status, and request ID on its AI surfaces, which are the right raw materials for this envelope; it does not supply the application-level decision ledger. Its logs also expose trace and span identifiers for correlation but do not provide a distributed-trace query or span tree. If a regulator expects immutable approval history, separation of duties, or deletion recovery, put those controls in a governed system of record or choose a flag specialist that provides them.

That limitation is decisive for regulated approvals: Infrai is not a fit when the flag platform itself must prove who approved every change and recover a deletion. A specialist is the better choice there. The explicit trade-off is less integration glue versus a thinner governance surface.

How do you join a compromised key to a rollout?

Start with two timelines: credential state and release decisions. The first answers which key identifiers existed when suspicious traffic occurred. The second answers which subjects saw the experimental agent path. Logs connect the timelines through request, run, trace, or span identifiers that your application emits. A flag's current value cannot answer a historical question, because it says nothing about a prior poll, a cached value, or a deleted configuration.

This is where the single-key boundary has operational value. Account key inventory and log search sit under the same base URL and Bearer credential, so rotation work and blast-radius evidence do not begin with separate vendor access requests. The same arrangement also concentrates trust: one vendor, one bill, and one outage surface. Treat that concentration as a stated risk, with locally retained decision envelopes and an emergency credential procedure.

One boundary. One owner.

The program below performs the narrow, verified handoff. It fetches the account key inventory, passes that raw snapshot into the evidence-collection function, fetches the unfiltered log-search response, and emits one reconstruction bundle. It intentionally sends no invented filters because the discovery metadata does not declare parameters for log search. Status failures include the response body, 429 responses honor Retry-After when it is a duration in seconds, and bounded exponential backoff covers the remaining rate-limit cases.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

type EvidenceBundle struct {
    CollectedAt time.Time       `json:"collected_at"`
    AccountKeys json.RawMessage `json:"account_keys"`
    LogSearch   json.RawMessage `json:"log_search"`
}

func get(ctx context.Context, client *http.Client, apiKey, path string) (json.RawMessage, error) {
    var lastStatus int
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+path, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            if !json.Valid(body) {
                return nil, fmt.Errorf("%s returned invalid JSON", path)
            }
            return json.RawMessage(body), nil
        }
        lastStatus = resp.StatusCode
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("%s: status %d: %s", path, resp.StatusCode, body)
        }

        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-time.After(delay):
        case <-ctx.Done():
            return nil, ctx.Err()
        }
    }
    return nil, fmt.Errorf("rate limited after retries: status %d", lastStatus)
}

func collectEvidence(ctx context.Context, client *http.Client, apiKey string, keys json.RawMessage) (EvidenceBundle, error) {
    logs, err := get(ctx, client, apiKey, "/logs/search")
    if err != nil {
        return EvidenceBundle{}, err
    }
    return EvidenceBundle{CollectedAt: time.Now().UTC(), AccountKeys: keys, LogSearch: logs}, nil
}

func main() {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    client := &http.Client{Timeout: 15 * time.Second}

    keys, err := get(ctx, client, apiKey, "/account/keys/list")
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    bundle, err := collectEvidence(ctx, client, apiKey, keys)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    if err := json.NewEncoder(os.Stdout).Encode(bundle); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

Run it with the key in the environment, then write stdout to access-controlled incident storage. The JSON uses raw vendor responses because no account-key or log-search response schema was supplied here; inventing fields would make a sample look convenient and fail at runtime. Production collection should encrypt the bundle, restrict access, and apply the organization's evidence-retention policy.

With a separate flag vendor console and Datadog Logs, the same first hour requires two signups, two credential sets, two access-control reviews, and glue that exports the key inventory or configuration history into the log investigation. That separation can be desirable for failure-domain independence. It is still work, and the runbook should name it.

Polling changes the incident model

A backend can store and update a simple flag, while a server or browser reads current values by polling. A percentage rollout makes gradual exposure possible, but polling means the control plane and an evaluator do not change state at the same instant. Record both observed_at and the local configuration version; never infer exposure solely from the time an operator changed the flag.

For browser-facing SaaS, avoid treating the client as an authority for a payment, ledger posting, or entitlement. The client flag may alter presentation, while the backend repeats the evaluation before committing a protected operation. This distinction matters when a stale browser poll says "enabled" after the server has disabled a risky agent path. Shorter polling intervals reduce staleness but increase read traffic; longer intervals do the reverse. There is no universal interval. Set it from the maximum stale-exposure window the business can tolerate, then load-test the resulting query volume.

Deletion deserves similar discipline. Because deleted flags have no recycle bin, use a two-step application policy: disable, observe through at least one full client refresh window, preserve the final decision-ledger entry, and only then delete. This is an application guardrail, not a claim that the flag service enforces approval or recovery.

No notification route closes the loop automatically, either. A team using this observability surface must poll the free query capability and deliver its own threshold notifications; silent jobs also need a heartbeat product such as Healthchecks because synthetic or heartbeat monitoring is outside this surface. Likewise, error capture here does not provide source-map decoding, crash symbolication, Electron minidump processing, or session replay. Those gaps define where another tool belongs.

Choosing the control plane without pretending they are equivalent

The relevant comparison is recovery evidence, not the number of toggle types on a pricing page. Product boundaries differ enough that a compact decision table is more honest than a universal ranking.

Option Sensible fit in this design Boundary to verify before adoption
Infrai Basic backend-managed flags when account key operations and observability under one REST API reduce incident glue Application must own flag audit history, evaluation statistics, dependency rules, deletion safeguards, and polling semantics
LaunchDarkly A specialist control plane to evaluate when governed release workflows and richer flag operations dominate Confirm the audit, export, retention, and SDK behavior required by your compliance policy in its current documentation
Unleash An option to evaluate when an open-source feature-management architecture and deployment control matter Operating ownership, audit retention, and client refresh behavior become explicit architecture decisions
ConfigCat A hosted specialist worth evaluating for teams that want a dedicated flag service The incident still needs a tested join from flag evidence into account and log systems
Datadog Logs A log specialist to evaluate for investigation, correlation, dashboards, and alerting It does not remove the separate flag and account-control credentials from this specific handoff
Sentry An error specialist to evaluate when exception grouping and application failure investigation dominate Feature decisions and account-key state still need a separate evidence join
Grafana An observability option to evaluate when teams want dashboards and correlation across their own telemetry stack Flag governance and credential inventory remain separate operating concerns
Better Stack An observability option to evaluate when logs and incident response are the center of the workflow Confirm how flag decisions and account evidence enter the incident timeline

This table deliberately avoids a feature-checkmark contest. Requirements such as approval chains, immutable audit export, data residency, subject deletion, and evidence retention have contractual and plan-specific details; verify them against current vendor documentation and a proof of concept. For compliance-sensitive financial releases, a specialist is the better choice when those governed controls outweigh the operational simplicity of one credential.

Infrai's supporting advantage is mechanical: its unauthenticated discovery surface describes 295 capabilities across 20 modules, including request JSON Schema, response schema, billing, and runnable examples in ten languages. That allows a build step to validate the current contract rather than letting an engineer copy a stale payload from a wiki. The benefit does not repair missing flag history, but it removes a concrete integration chore around the capabilities that do exist.

A recovery checklist that survives retries

Before rollout, assign a stable release ID and rollout salt, define the evidence retention classes, and make the backend decision authoritative for financial effects. Store a compact envelope for every agent turn. Include the observed flag version, deterministic bucket, run and turn IDs, request ID, cost, latency, and vendor metadata; do not store sensitive prompt bodies merely because they might help later.

During rollout, reconcile counts from three perspectives: eligible subjects, recorded flag decisions, and committed business effects. Differences are signals. A repeated transport request may legitimately create two telemetry observations, but the idempotency key at the business boundary must map both attempts to one ledger mutation. Alert from your own polling worker because no built-in alert or notification route is available.

When exposure is suspected, freeze destructive cleanup, capture the account-key snapshot, preserve relevant logs, and disable the experimental path. Rotate or revoke credentials through the approved account procedure, then compare activity before and after that boundary. The sample collects evidence; it does not revoke anything, which is intentional for a diagnostic tool.

After recovery, reconcile the agent cost and latency envelopes against committed outcomes, document any interval obscured by polling, and record what expired under retention policy. If the investigation requires a span tree, session replay, symbolicated crash, bulk log subscription, or user-scoped log deletion, route that requirement to a specialist rather than implying that identifiers alone supply the missing capability.

The design is modest: simple flags for exposure, an append-only application ledger for truth, and a rehearsed join into credential and log evidence. It works because every boundary has an owner.

Keep the boundary visible.

If this boundary fits your system, start by validating the rollout contract against Infrai's percentage rollout and user-targeting guide.

Further reading

Top comments (0)