The page says the AI agent control plane is breaching its write-latency SLO. The on-call opens the trace and sees two successful feature-flag mutations carrying the same rollout intent, separated by a client timeout and a retry. The second write is fast, so a latency-only dashboard looks healthy; the resulting duplicate audit event and repeated rollout evaluation are the actual failure.
TL;DR: instrument a flag change as one logical operation, give that operation a stable idempotency key, and correlate four signals: end-to-end agent-loop latency, mutation attempts, idempotency outcomes, and durable-write cost. Alert on exhausted error-budget burn or an abnormal ratio of duplicate outcomes, not raw retry count. Retries are expected. Unidentifiable side effects are not.
This is a signal-quality problem before it is a storage problem. A useful trace must let the on-call prove whether a timeout happened before the server accepted the mutation, after it committed, or while the response was returning. Without that boundary, every retry is a guess, and every page invites a manual reconstruction from unrelated logs.
How can feature flag retries avoid duplicate writes with idempotency?
The earlier warning should have been a change in outcomes at the mutation boundary: the rate of requests returning replayed or conflict for an idempotency key rises relative to accepted mutations. A retry counter alone is noisy because network interruptions, explicit backoff, and transient dependency failures can all increase attempts without changing flag state twice. Conversely, a duplicate write can occur while every individual request stays below the latency objective.
Four signals are enough to begin, provided their meanings are narrow.
| Signal | Unit and useful dimensions | Operational question |
|---|---|---|
| Agent-loop latency | Histogram in milliseconds; operation and result | Is the user-visible loop consuming its latency budget? |
| Mutation attempts | Counter; logical operation and attempt result | How much retry pressure reaches the endpoint? |
| Idempotency outcomes | Counter; accepted, replayed, conflict
|
Did repeated delivery preserve one logical effect? |
| Durable-write cost | Counter or histogram in workload units; operation class | Did one loop create unexpected storage or model work? |
Do not attach the raw idempotency key, flag identifier, user identifier, prompt, or trace ID to metric labels. Those values have high or unbounded cardinality. Keep them on traces and structured logs, where an investigator can query a single operation; keep metrics aggregated over bounded attributes such as route class, deployment environment, and outcome.
The page policy then follows the SLO rather than the instrumentation. A duplicate-outcome ratio is a diagnostic signal, while fast error-budget burn on agent-loop latency or failed mutations is page-worthy. A slow rise in replayed requests may deserve a ticket because idempotency is doing its job. A rise in conflicts deserves investigation because the same key is arriving with different intent.
Short alerts are better.
Ambiguity is the enemy.
Trace one intent across retries and side effects
The root span represents the whole AI agent loop, not one HTTP attempt. Beneath it, each attempt gets its own client span, and the backend creates a server span plus spans for the idempotency lookup, flag-state transaction, audit append, and any downstream evaluation. This hierarchy preserves the distinction that matters during an incident: one user intent can produce several transports but must produce one committed mutation.
The instrumentation below uses Go throughout and relies on OpenTelemetry's API shape. It deliberately records bounded outcomes on spans and metrics while putting the operation key only on the trace. The handler is abbreviated around storage because transaction syntax depends on the chosen database; the invariants do not.
package flags
import (
"context"
"errors"
"time"
"go.opentelemetry.io/otel/attribute"
"go.opentelemetry.io/otel/metric"
"go.opentelemetry.io/otel/trace"
)
type Mutation struct {
Flag string
Percentage int
}
type Result struct {
Outcome string // accepted, replayed, or conflict
}
type Store interface {
ApplyOnce(ctx context.Context, key string, mutation Mutation) (Result, error)
}
type Handler struct {
store Store
tracer trace.Tracer
attempts metric.Int64Counter
outcomes metric.Int64Counter
latency metric.Float64Histogram
}
func (h *Handler) Toggle(ctx context.Context, key string, m Mutation) (Result, error) {
start := time.Now()
ctx, span := h.tracer.Start(ctx, "feature_flag.mutate")
defer span.End()
// Keep the unique key searchable on the trace, never as a metric label.
span.SetAttributes(attribute.String("operation.idempotency_key", key))
h.attempts.Add(ctx, 1, metric.WithAttributes(
attribute.String("operation", "flag_mutation"),
))
result, err := h.store.ApplyOnce(ctx, key, m)
if err != nil {
span.RecordError(err)
return Result{}, err
}
if result.Outcome != "accepted" && result.Outcome != "replayed" && result.Outcome != "conflict" {
return Result{}, errors.New("invalid idempotency outcome")
}
h.outcomes.Add(ctx, 1, metric.WithAttributes(
attribute.String("operation", "flag_mutation"),
attribute.String("outcome", result.Outcome),
))
h.latency.Record(ctx, float64(time.Since(start).Milliseconds()),
metric.WithAttributes(attribute.String("operation", "flag_mutation")),
)
return result, nil
}
ApplyOnce needs an atomic claim on the key and a binding between that key and the normalized request. The first request stores the request fingerprint and result in the same transactional boundary as the state change. A later request with the same key and fingerprint returns the stored result; the same key with a different fingerprint returns a conflict. Expiry is a capacity decision, not a correctness shortcut: the retention window must exceed the maximum period during which clients can legitimately retry, including queued agent work and recovery after an outage.
This also changes how errors should be read. A client timeout is ambiguous evidence. It proves that the client did not receive a response within its deadline; it does not prove that the backend rejected or failed to commit the operation. The follow-up attempt, carrying the same key and intent, turns ambiguity into an observable replayed result rather than another state transition.
Make the storage model answer the incident question
The operational record needs enough information to reconstruct one logical mutation without treating the analytics store as the source of truth. Persist an operation key, request fingerprint, final result, commit timestamp, and trace correlation alongside the transactional state. Export a separate append-only event for analysis only after the mutation result is known. If delivery to analytics is repeated, its event identity must also be stable.
ClickHouse is documented as an analytical database, which makes it a reasonable example of the read-optimized side of this split, not proof that it should arbitrate a toggle request. An event table can answer questions such as “how many agent loops produced more than one transport attempt?” and “which deployment preceded a rise in conflicts?” The correctness decision remains at the transactional boundary that owns the flag state.
For an AI agent loop, include model-call duration and workload units as child-span measurements, but do not mistake correlation for causation. A loop can become expensive because a replay restarted model work before reaching the mutation endpoint; it can also retry only the final mutation while reusing earlier results. The trace graph distinguishes those paths. A single cost total cannot.
Capacity planning starts with logical operations and expands by retry amplification. If peak traffic is L loops per second, a fraction M reaches a flag mutation, and the mean number of transport attempts is A, the endpoint sees L × M × A attempts per second. For a deliberately hypothetical planning case, 120 loops per second, a mutation fraction of 0.25, and 1.4 attempts per mutation produce 42 endpoint attempts per second; these are input assumptions, not measured performance. The idempotency index grows with the logical mutation rate and retention window, while the telemetry pipeline grows closer to the attempt rate. Size both with headroom for retry storms, then test the SLO when an exporter is slow or unavailable; application correctness must not depend on synchronous telemetry delivery.
Buy or build the idempotency and telemetry path?
The decision is less about feature count than ownership during the next page. Managed and self-hosted paths can both emit standards-based telemetry and enforce idempotency. Their failure domains, staffing costs, and exit costs differ.
This approach has limits. Idempotency cannot repair a non-transactional side effect that occurs before the key is claimed, and trace correlation cannot prove exactly-once execution across dependencies that ignore the operation identity. The trade-off is extra write-path state, retention capacity, and operational complexity in exchange for deterministic replay behavior; a low-impact, naturally commutative update may not justify that machinery.
| Concern | Managed path | Self-hosted path | Decision evidence |
|---|---|---|---|
| On-call load | Provider operates more of ingestion and storage | Platform team owns scaling, upgrades, and recovery | After-hours staffing and tested runbooks |
| Signal control | Defaults may accelerate adoption but constrain retention or processing | Full control, plus responsibility for every queue and schema | Required sampling, residency, and deletion rules |
| Lock-in | Proprietary queries or alert semantics can raise migration effort | Infrastructure coupling replaces service coupling | Export test using standard trace and metric formats |
| Capacity | Contracted limits and service behavior define the ceiling | Hardware and operator skill define the ceiling | Peak attempt rate, retention, and retry-storm load test |
I would require a replay test before accepting either path: send the same key and request twice, then send the same key with a changed percentage; verify one state change, a replayed result, a conflict, and a trace that joins all three attempts without high-cardinality metric labels. That is a concrete acceptance test, not a product preference.
Tune the threshold without training the team to ignore it
Start by recording outcomes without paging. Establish the normal replay ratio by environment and deployment phase, because rollout automation may create retry patterns that interactive traffic does not. Then inject three controlled conditions in staging: a response lost after commit, a timeout before commit, and a repeated key with changed intent. The traces and stored results should make each condition distinguishable.
For production, alert the on-call on fast burn against the service's published latency or mutation-success SLO. Route a sustained increase in conflicts to investigation, and keep ordinary replays on a dashboard unless they coincide with cost amplification or SLO burn. Fixed universal percentages would be invented precision; the threshold needs the service's baseline, traffic volume, and acceptable false-positive budget.
The false-positive cost is real. If every retry pages, the team learns that the alert describes transport noise and stops treating it as evidence of user harm. Set the threshold too high, however, and duplicate side effects can inflate audit volume and agent-loop work before latency moves. The defensible middle is an alert tied to an SLO, enriched with the bounded idempotency outcomes that tell the responder where to look.
The closing rule is plain: one intent gets one identity, one committed effect, and as many observable attempts as delivery requires. Measure the attempts. Page on harm.
Top comments (0)