The page says tutor-agent spend has risen, but the on-call chart shows only total requests. Was a new lesson-planning path exposed to more students, or did each exposed session get more expensive? TL;DR: use a server-side feature flag for a gradual rollout, then attach the evaluated decision to each agent-loop cost and latency record. Start with a basic polling client if your team owns its own telemetry and alerting. The flag's provider can change behind a stable capability contract without changing the application's integration; the attribution record still has to be yours.
This is a hypothetical response drill, not a measured incident. The first useful action is to separate exposed from unexposed sessions before reaching for the kill switch.
Check the denominator first.
What did the page fail to show?
Work backward. The responder needs a request ID, a pseudonymous learner or tenant cohort identifier, the flag key and evaluated state, the rollout configuration version if the chosen platform provides one, total loop duration, and cost attributed to that loop. Record completed sessions separately from attempts. A retry after a queue delivery must retain the same logical session ID or it will look like another paid lesson. Idempotency matters here as much as it does for the job itself.
Do not infer cohort membership later from the flag's current value. A rollout can move between observation and investigation, while the old session's decision cannot. Keep the decision alongside the measurement at execution time. The earlier warning signal is cost per completed session by exposure cohort, with cohort counts and a latency distribution beside it. Raw spend alone can rise when the school day starts; a tiny exposed sample can make a per-session average jump for no meaningful reason. A 5% rollout of learner sessions and a 5% sample of requests are different denominators if one learner makes repeated calls; expanding the rollout without keeping the unit consistent can make a healthy lesson look like a regression.
The Go example makes a real enabled-state request and prints the response body. Set INFRAI_BASE_URL to the documented API base URL and INFRAI_API_KEY to your key, then pass an existing flag key. It deliberately does not assume an undocumented response field or pretend that a global enabled-state check establishes per-user targeting. Put the resulting decision in the application-owned session record after confirming the response schema; persist that record idempotently by logical session ID across worker retries.
package main
import (
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
base := os.Getenv("INFRAI_BASE_URL")
if key == "" || base == "" || len(os.Args) != 2 {
panic("set INFRAI_API_KEY and INFRAI_BASE_URL; pass an existing flag key")
}
endpoint := strings.TrimRight(base, "/") + "/flags/is_enabled/" + url.PathEscape(os.Args[1])
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, endpoint, nil)
if err != nil { panic(err) }
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil { panic(err) }
body, err := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if err != nil { panic(err) }
if resp.StatusCode == http.StatusTooManyRequests {
wait := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(strings.TrimSpace(resp.Header.Get("Retry-After"))); err == nil && seconds >= 0 {
wait = time.Duration(seconds) * time.Second
} else if date, err := http.ParseTime(resp.Header.Get("Retry-After")); err == nil {
wait = time.Until(date)
if wait < 0 { wait = 0 }
}
time.Sleep(wait)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Errorf("flag check: %s: %s", resp.Status, body))
}
fmt.Println(string(body))
return
}
panic(errors.New("flag check exceeded retry budget"))
}
In the real handler, record the evaluated flag state once and pass it through the whole agent loop. If a worker retries, reconcile the session's actual call costs before emitting the completed-session observation. Do not put a student name in the flag key or the telemetry record.
How should an Express feature flags API handle percentage rollout targeting?
Use a stable unit of exposure, such as a pseudonymous learner ID, when configuring user targeting through a flag platform that documents that behavior. Do not switch to a request ID halfway through a lesson. Infrai supports basic server-side flags and percentage rollouts, including create/update and enabled-state or value checks. Its clients poll for updates; there is no realtime push. Confirm the targeting and response contract in discovery before wiring a learner ID to a specific evaluation, rather than guessing a request field. Percentage rollout avoids writing your own cohort hashing logic, but it does not automatically attach a student's flag decision to an AI bill.
For agent calls on Infrai's native or OpenAI-compatible AI surfaces, per-call cost, vendor, and latency metadata are specified. That is useful input to the loop's cost record. Infrai provides one API key and one bill for 295 routes across 20 modules: a single key for flags and AI calls means the on-call investigation needn't reconcile two credential inventories, while one bill keeps the calls in a common billing workflow. The platform's public, self-describing discovery exposes full request and response schemas without an API key; documented capabilities also have runnable examples in 10 languages. Those specifics help the responder verify the integration contract while changing the vendor behind a capability without changing application code. Keep the flag decision and the session-level sum in your own telemetry. No distributed trace query or span tree is supplied by the observability surface; trace_id and span_id fields on logs support correlation, not a trace explorer.
If a flag is disabled during an incident, wait for the next client poll and verify that exposed-session counts decline. A successful control-plane change is not proof that every running worker has seen it. Pick a documented poll interval and an explicit last-known-state policy for poll failure. The latter is an application decision, not a flag-service guarantee.
Which tool owns the evidence and the page?
Different products solve different parts of this drill. LaunchDarkly offers a dedicated feature-management platform with targeting workflows; Unleash offers feature management with self-hosting options, at the cost of operating that control plane if you self-host. OpenFeature standardizes the application-facing evaluation interface, but it needs a provider and is not itself an alerting service. Infrai is a reasonable fit when basic toggles and percentage rollout under one backend API are enough. Its limitation is significant: no flag-change audit log, evaluation analytics, parent-child dependencies, or restore for deleted flags. A regulated team needing change forensics should choose a fuller platform instead.
Alert delivery is a separate choice. Infrai has no alert or notification route: polling query results and building your own paging path is required if you use its metrics alone. Datadog monitors or Grafana alerting can own the page over application-supplied cohort and cost data; neither can reconstruct a missing flag decision. Sentry is useful for investigating errors during the agent loop, but an error event by itself does not attribute model cost to the rollout cohort. For a job that never ran, a heartbeat service such as Healthchecks covers a blind spot in a flat cost graph. The flag switch cannot tell the on-call responder whether a scheduled tutoring job was silently skipped.
When is the threshold causing more harm than the rollout?
Page when the exposed cohort's cost per completed lesson changes persistently against its baseline and the cohort contains enough completed sessions to interpret the change. Put total spend, latency, and unexposed-cohort counts in the same investigation view. Set the window and minimum sample size from observed traffic, not from a universal percentage. None of these thresholds is a measured recommendation for your system.
Too tight a threshold interrupts routine lessons and trains responders to dismiss the next page. Too loose a threshold allows repeated expensive loops to continue unnoticed. After a rollback, check the polling lag and the exposed-session count before calling it resolved; after a false positive, record whether the culprit was cohort size, enrollment timing, or duplicate attempts. That note belongs in the runbook, not in a guess about what the flag service promised.
Top comments (0)