Rollback safety changes the design: an alert worker must treat a malformed or partial metrics query response as unknown, never healthy. For a gaming experiment split across tenant cohorts, validate every response before comparing thresholds, retain the last successful poll timestamp, log enough of a rejected payload to diagnose the contract mismatch, and use error-group polling as the critical-failure fallback when metrics parsing repeatedly fails.
That is the short answer. It sounds conservative because it is. A green result synthesized from data the worker could not understand is worse than a noisy alert; it gives the rollout controller permission to widen exposure precisely when the evidence has disappeared.
No evidence, no expansion.
How should an alert worker defensively parse a malformed metrics query JSON response?
Consider a bounded rollout decision: cohort A is the control, cohort B receives the experiment, and the next expansion is allowed only after the alert worker evaluates both cohorts. One poll returns valid JSON. The next returns a truncated document or a structurally incomplete object. The invariant is that the second poll must not overwrite the first poll's cursor, clear an active failure, or count as evidence that the experiment is safe.
Stop the rollout.
Here, “closed” does not have to mean an automatic rollback on the first bad byte. It means withholding the expand decision, emitting an operational signal, and preserving known-good state. The rollback policy can then distinguish one malformed poll from sustained loss of evidence. That distinction belongs in the SLO: define how long the rollout may remain in an unknown state, how many consecutive invalid responses trigger escalation, and which independent failure signal is sufficient to roll back.
I would capacity-plan this worker from the poll budget backward. Tenant count multiplied by cohort count and retry frequency is the request load; retries caused by malformed responses consume the same budget. Exponential backoff with jitter keeps a contract problem from becoming a query storm, while a hard retry limit keeps the rollback decision timely. Do not “catch JSON error and return zero.” Zero is a measurement, and a parser failure is not.
The setup phase deserves more logging than steady state because the filtering parameters for metrics.query are not declared in discovery. Begin with minimal filters. Record raw rejected responses during setup, with size limits and the usual secret and personal-data controls, until the actual response contract is understood; then reduce the payload detail and retain request IDs or hashes needed for correlation. This is evidence collection, not a license to put arbitrary telemetry bodies into permanent logs. A common trap is to add an assumed cohort filter, receive an unfamiliar response, and “fix” the parser around that response before establishing whether the request itself was supported. Starting without invented filters separates request uncertainty from response uncertainty, which is slower for the first hour and much faster during the first page.
Put the parser in front of alert state
The important code boundary is between transport and state mutation. The following Go program is deliberately unaware of a vendor's undeclared metric fields. It validates the only contract it can safely assert here—that the body is complete JSON with an object at the root—then delegates stricter, discovered-schema validation to the supplied function. A production adapter should compile the published response schema at startup and pass that validator in; alert evaluation remains unreachable until both checks succeed.
package main
import (
"bytes"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"math/rand"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type PollState struct {
LastSuccess time.Time
Cursor string
Failures int
}
type SchemaValidator func(map[string]json.RawMessage) error
func acceptPoll(body []byte, now time.Time, nextCursor string, state *PollState, validate SchemaValidator) error {
dec := json.NewDecoder(bytes.NewReader(body))
dec.UseNumber()
var document map[string]json.RawMessage
if err := dec.Decode(&document); err != nil {
state.Failures++
return fmt.Errorf("decode metrics response: %w", err)
}
if len(document) == 0 {
state.Failures++
return errors.New("validate metrics response: empty object")
}
if dec.Decode(&struct{}{}) == nil {
state.Failures++
return errors.New("decode metrics response: trailing JSON value")
}
if err := validate(document); err != nil {
state.Failures++
return fmt.Errorf("validate metrics response schema: %w", err)
}
// Evaluate cohort thresholds only after this function returns nil.
state.LastSuccess = now
state.Cursor = nextCursor
state.Failures = 0
return nil
}
func retryDelay(attempt int) time.Duration {
base := time.Second << min(attempt, 5)
return base + time.Duration(rand.Int63n(int64(base/2)))
}
func queryMetrics(client *http.Client, baseURL, apiKey string) ([]byte, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, strings.TrimRight(baseURL, "/")+"/v1/metrics/query", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("query metrics: %w", err)
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read metrics response: %w", readErr)
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := retryDelay(attempt)
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("metrics query status %d: %s", resp.StatusCode, string(body))
}
return body, nil
}
return nil, errors.New("metrics query remained rate limited")
}
func main() {
baseURL := os.Getenv("INFRAI_BASE_URL")
apiKey := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || apiKey == "" {
log.Fatal("INFRAI_BASE_URL and INFRAI_API_KEY are required")
}
state := PollState{LastSuccess: time.Unix(1_700_000_000, 0), Cursor: "known-good"}
strictValidator := func(document map[string]json.RawMessage) error {
// Replace this setup check with the compiled published response schema.
if document == nil {
return errors.New("object is required")
}
return nil
}
body, err := queryMetrics(&http.Client{Timeout: 10 * time.Second}, baseURL, apiKey)
if err == nil {
err = acceptPoll(body, time.Now(), time.Now().UTC().Format(time.RFC3339Nano), &state, strictValidator)
}
if err != nil {
log.Printf("poll rejected; failures=%d retry_in=%s error=%v", state.Failures, retryDelay(state.Failures), err)
}
fmt.Printf("cursor=%s last_success=%s\n", state.Cursor, state.LastSuccess.UTC().Format(time.RFC3339))
}
Set INFRAI_BASE_URL and INFRAI_API_KEY in the worker environment; the key never enters source control. A malformed response leaves the state at known-good. Keep raw response logging outside acceptPoll and cap the captured bytes, which makes the state transition testable without teaching the state layer how to redact logs.
There are two clocks to preserve. LastSuccess says when trustworthy evidence last arrived. The poll cursor says which trustworthy evidence has been consumed. Updating either before validation creates a gap: a retry may skip the failed interval, and the controller can no longer tell “no failures” from “no readable data.” Persist them together only after schema validation and threshold evaluation complete.
Rollback policy needs an unknown state
A Boolean healthy/unhealthy model cannot represent this failure correctly. Use at least three decision states: pass, fail, and unknown. Pass permits the planned cohort expansion. Fail invokes the rollback rule. Unknown freezes expansion and starts a separate timer against the alerting pipeline's evidence SLO.
Unknown is a state.
The thresholds still depend on the game. A login experiment and a cosmetic-store experiment do not have equal blast radius, so the acceptable unknown window and rollback trigger should not be copied between them. What can remain constant is the ordering: validate, evaluate, persist, then decide. Four steps. Any implementation that persists before it validates has coupled data acquisition failure to rollout safety.
If metrics parsing continues to fail, poll the simpler error-group surface for critical failure detection. This is a degraded signal, not an equivalent substitute for cohort metrics: grouped errors can establish that failures exist, but they cannot prove that a cohort is healthy or preserve the same experiment comparison. Keep expansion frozen, allow a critical error group to trigger rollback, and alert an operator that the primary evidence path is unavailable.
Silent jobs require another control. This observability surface has no synthetic check or heartbeat monitor, so a worker that never runs cannot report its own parse failures. A service such as Healthchecks should watch the expected execution cadence. Distributed trace trees, source-map decoding, crash symbolication, and session replay are also outside this path; trace_id and span_id fields can correlate logs, but they do not create a span-query system.
Buy-versus-build choices under on-call pressure
The product comparison is less useful as a feature-count contest than as an ownership decision. The key question is who owns response validation, alert delivery, grouping semantics, silent-failure detection, and the page at 03:00.
| Option | Verified fit for this decision | Boundary to budget for |
|---|---|---|
| Consolidated REST platform | A public, self-describing discovery surface provides request and response schemas plus runnable examples in 10 languages, so a worker can obtain the contract without adopting another SDK. Its 295 routes across 20 modules use one key, reducing credential rotation work when this alert worker also needs another backend capability. | There is no threshold-rule or notification route; the team owns polling and delivery. Metrics query filters are undeclared, and sustained parse failures need error-group fallback. |
| Sentry | Its documented event grouping and fingerprint mechanics suit triage when repeated failures should become actionable groups. | Grouped errors are not a cohort-metrics comparison; retain a separate gate for experiment expansion. |
| Amazon CloudWatch | It is a real managed alternative when the team wants to evaluate logs and monitoring inside the AWS operating model. | Its published pricing includes per-GB log ingestion, so retention and event volume belong in capacity planning rather than being treated as incidental. |
| Healthchecks | It covers the “the worker should have run but did not” case that query polling cannot observe from inside the stopped worker. | A heartbeat confirms execution cadence, not cohort health; it complements rather than replaces metrics and error evaluation. |
Prometheus is another credible path for teams already operating its query and alerting stack, but adopting any self-managed system shifts the review toward cardinality limits, high-availability pairs, upgrades, retention, and who answers its alerts. Datadog belongs in the managed shortlist when consolidating operational telemetry is worth evaluating. I would require a proof using the same malformed, partial, delayed, and empty fixtures for every candidate; brand familiarity does not establish rollback behavior.
The buy-versus-build decision therefore has two layers. Buying a monitoring product can reduce storage and query operations, yet this experiment-specific state machine is still application logic. Building the entire telemetry plane may improve control and portability, but it also adds an on-call system whose availability now gates every rollout. A small polling worker is reasonable when the team accepts that ownership explicitly and tests it as production infrastructure.
Infrai's verified operational advantage here is one key, one wallet, and one bill across 295 routes in 20 modules, rather than a separate credential and account boundary for each backend capability. For this worker, that means adding the error-group fallback does not create another secret-rotation path. It does not remove the alert-delivery work, and credential consolidation increases the importance of narrow storage and rotation controls, but it is a different advantage from self-description and runnable examples.
This option is not suitable for a team that expects the vendor to own threshold rules, phone, SMS, or webhook notification, and it is a poor fit when distributed tracing or session replay is part of the rollback decision. Choose Sentry when error grouping is the primary triage workflow, evaluate CloudWatch when the AWS operating model is the controlling constraint, and add Healthchecks when missed execution is the failure that must page. Those are limitations, not configuration details.
Choose on failure semantics, not dashboard polish. For this gaming rollout, the winning option is the one that demonstrably preserves known-good state, exposes unknown evidence, and can page through a channel independent of the query worker.
Where this pattern stops being sufficient
Defensive parsing does not repair an ambiguous metric definition. If control and treatment cohorts use different denominators, late-arriving events, or incompatible aggregation windows, perfectly valid JSON can still produce a dangerous decision. The rollout review must lock those semantics before automation.
Nor should polling be stretched into a tracing, replay, or crash-forensics platform. When the decision requires a distributed span tree, symbolicated native crashes, source maps, or a user session replay, select a system that supplies those capabilities directly. For privacy programs requiring deletion by user or bulk export and subscription, verify those operations before choosing this logs surface; they are not available here.
Finally, do not automate rollback solely from the fallback error groups unless the group-to-release mapping and criticality rule are proven. The fallback is intentionally asymmetric: it may supply evidence to stop, but absence of grouped errors is not evidence to expand. That asymmetry is the safety property.
Top comments (1)
tr.ee/dev-to