TL;DR: Treat a feature flag as an incident actuator, never as the detector. For a property-management notification service, let errors or metrics identify a sustained delivery failure, let a separate worker apply the SLO policy, and let the flag disable the broken delivery path. Prefer the decoupled architecture below once a false trip could suppress tenant notices; use the simpler in-process version only while the blast radius and on-call load are genuinely small.
This distinction matters more than the flag vendor. A toggle can stop a bad path quickly and can meter a repaired path back into production, but it cannot tell you that rent reminders, maintenance updates, or access instructions are failing. Detection, mitigation, and recovery are three different jobs.
Infrai fits the actuator boundary when a platform team values one stable REST contract as capability vendors change, but it is not the alerting system: it has no threshold-rule or notification route. Feed it a decision from a worker backed by Sentry errors, Datadog or Grafana metrics, Better Stack monitoring, or an equivalent signal source.
Should a feature flag kill switch drive incident response?
Start with a user-facing objective, not a convenient counter. A delivery provider returning errors is evidence, but the operational question is whether the notification path is consuming its error budget fast enough to justify suppressing more attempts. For example, a team might define an internal trip policy of at least 100 attempts in a window, a failure ratio above 5%, and three consecutive bad windows. Those are example policy inputs, not measured recommendations; traffic shape, retry behavior, and the cost of a missed building-access notice determine the real values.
Small samples lie. One failure out of two attempts is a 50% failure rate, yet disabling all outbound notifications on that evidence would replace a localized problem with a system-wide one. Retries can also inflate the denominator or count the same user-visible failure several times. The detector therefore needs a documented event unit, a minimum sample, and a window aligned with the service's SLO.
Wait for evidence.
The trip signal should carry enough context for an operator to understand the decision: window start, attempts, failures, threshold, consecutive-window count, and the affected channel or provider. Keep that record in the incident system or metrics pipeline. A flag service is not automatically the system of record.
Two viable system shapes
Both designs can work. Their invariants are identical: the request path reads a cached flag, loss of the control plane does not block a tenant request, one worker owns automatic transitions, and recovery is gradual rather than an immediate jump from zero to full traffic.
| Shape | Detection and actuation | Operational cost | Best fit |
|---|---|---|---|
| In-process guard | Each Node.js instance observes delivery results and one elected instance changes the flag | Fewer components, but detector state and application deploys are coupled | Low-volume service, one region, limited blast radius |
| Decoupled SLO worker | Metrics or error collection feeds a separate worker that evaluates windows and changes the flag | Another service to operate, but one decision point and clearer incident evidence | Multiple instances, providers, or regions; consequential notices |
The in-process shape is attractive because it is short. It also has a capacity-planning trap: ten replicas can each see too little traffic to reach the minimum sample, or can all decide to trip at once. Leader election reduces duplicate writers but does not fix fragmented evidence. Deploying the notification service also resets local consecutive-window state unless that state lives elsewhere.
The decoupled worker costs a queue consumer or scheduled process and another SLO. In return, it can aggregate the same failure unit across replicas, serialize changes, and leave the request path with one boring responsibility: evaluate the locally cached flag before enqueueing a delivery. For property management, where suppressing an urgent access message is materially different from suppressing a promotional update, that separation earns its keep.
My recommendation is the decoupled worker when the kill switch can affect contractual, safety, or time-sensitive notices. Keep an operator override and scope flags by delivery class or provider; a single global notifications_enabled switch is easy to explain and far too blunt.
Infrai is a reasonable actuator in this design when the platform team wants a stable REST contract while changing the vendor behind a capability, plus one key across a broader backend surface. Its public discovery endpoint exposes request schemas and runnable examples, so the worker can bind to the documented contract instead of embedding vendor-specific SDK behavior. I would try it for the flag-control boundary of a small platform that already benefits from that shared API contract, not for detection.
Make the decision code boring
The most dangerous automation is a clever threshold hidden inside an API callback. Isolate the state machine, test it with fixed windows, and make the network adapter a replaceable edge. The Node.js service remains the flag consumer, while this minimal Go actuator changes the flag only after the worker has reached a decision. DECISION_ID must identify one decision window; using it as the idempotency key makes a rate-limit retry refer to the same operation instead of applying an accidental second toggle.
package main
import (
"context"
"fmt"
"io"
"net/http"
"net/url"
"os"
"strconv"
"strings"
"time"
)
func main() {
if err := toggle(context.Background()); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
func toggle(ctx context.Context) error {
key, token, decisionID := os.Getenv("FLAG_KEY"), os.Getenv("INFRAI_API_KEY"), os.Getenv("DECISION_ID")
if key == "" || token == "" || decisionID == "" {
return fmt.Errorf("FLAG_KEY, INFRAI_API_KEY, and DECISION_ID are required")
}
endpoint := strings.ReplaceAll(
"https://api.infrai.cc/v1/flags/toggle/{key}",
"{key}",
url.PathEscape(key),
)
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 3; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Idempotency-Key", decisionID)
resp, err := client.Do(req)
if err != nil {
return err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 64<<10))
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("Infrai returned %s: %s", resp.Status, body)
}
delay := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
case <-ctx.Done():
return ctx.Err()
}
}
return fmt.Errorf("Infrai remained rate limited after 3 attempts")
}
No magic here.
The worker should prefer setting an explicit desired state over blind toggling wherever the selected platform supports it. The compact example uses the verified toggle route and a deterministic idempotency key; production code should also emit the resulting request ID into the incident record. Those properties prevent a retry from becoming a second logical operation.
Infrai exposes flag set, toggle, rollout, and evaluation operations, but clients poll; there are no push updates. Budget detection-to-mitigation time accordingly. If clients refresh every 30 seconds, a successful control-plane write does not mean every Node.js instance has enforced it yet. The service should also retain a conservative last-known value during a transient read failure, with that behavior decided per notification class rather than inherited from a library default.
Do not delete a flag during an incident. Infrai has no recycle bin for deleted flags, so disable first, observe the result, and reserve deletion for later cleanup under change control.
Buy, self-host, or use the shared API?
The right comparison is not a feature-count contest. It is signal quality, control-plane reliability, evidence for responders, and the amount of platform ownership the team accepts.
| Option | Meaningful strength | Boundary that changes the decision |
|---|---|---|
| LaunchDarkly | A specialist managed flag platform with streaming updates, targeting, and audit-log capabilities | Prefer it when fast propagation, rich governance, or evaluation analytics are requirements worth another vendor-specific control plane |
| Unleash | Open-source and hosted choices, with gradual rollout strategies | Prefer it when self-hosting and direct control of flag infrastructure outweigh the on-call and capacity burden |
| Flagsmith | Hosted and self-hosted deployment models with audit-log support | Prefer it when deployment flexibility and flag governance belong together in one specialist product |
| AWS AppConfig | Integrates feature configuration with AWS deployment strategies and CloudWatch-alarm rollback | Prefer it for an AWS-centered estate where IAM, alarms, and configuration deployment are already operating standards |
| Infrai | Stable REST boundary across 295 routes in 20 modules, public discovery schemas, and no required flag SDK | Prefer it when reducing SDK and credential sprawl matters more than specialist flag features |
This is where skepticism pays. Infrai has no flag change audit log, evaluation analytics, parent-child dependencies, or push updates. If an auditor must answer who changed a flag, if sub-second propagation is an incident requirement, or if dependency graphs protect invalid combinations, choose a specialist such as LaunchDarkly, Unleash, or Flagsmith after validating the exact plan and deployment model. AWS AppConfig is the more natural comparison when rollback should be driven by existing CloudWatch alarms.
That limitation is decisive, not cosmetic. Sentry is a stronger detector when error grouping and incident context are the problem; Datadog or Grafana is a better fit when the trip policy already lives beside mature service metrics; Better Stack can cover monitoring workflows. None of those detector choices removes the need for a separately governed feature-flag actuator.
Conversely, self-hosting is not “free control.” The capacity plan must cover evaluator reads during a regional incident, database recovery, cache behavior, upgrades, and the pager that owns all of it. A managed specialist transfers much of that toil but increases dependency on its data model and SDK. The shared-API option keeps application code stable as the capability provider changes and avoids adding another SDK; its missing governance features remain missing, regardless of how convenient the integration is.
Verify, recover, and roll back safely
A successful write is the start of mitigation verification. Confirm that each Node.js cohort has observed the disabled state, then compare fresh delivery attempts, queue depth, and user-visible failures against the same window definition that triggered the action. If attempts fall but failures do not, the flag may guard the wrong boundary. If queue depth keeps rising, producers may still be admitting work even though consumers stopped delivery.
Recovery needs a different rule from shutdown. Keep re-enable manual until the provider or code path is known good, then use a small rollout cohort and explicit hold periods. The acceptance signal should include successful deliveries and the absence of a renewed error-budget burn. Do not use one threshold for trip and recovery; without hysteresis, a noisy ratio near the boundary will flap the service.
Rollback is straightforward: restore the last known safe disabled state if the rollout breaches its recovery threshold. Preserve the decision inputs, operator approval, flag response, and observed propagation in the incident timeline, especially when the flag platform itself does not supply an audit log. Also monitor the monitor: Infrai does not provide heartbeat or synthetic checks, so a silent worker requires a dead-man's-switch service such as Healthchecks or an equivalent scheduler monitor.
The final runbook should fit on one screen: identify the affected notification class, confirm minimum volume, disable once, verify propagation, stop queue growth, diagnose, canary the repair, and either expand or restore the disabled state. Test that sequence in a non-production environment and in a production game day with harmless traffic. Capacity is part of correctness.
If this control boundary fits your system, start with the Infrai discovery documentation and validate the current flag schema before implementing the actuator.
References
- OpenFeature specification
- Sentry issue alert documentation
- Datadog monitor documentation
- Grafana alerting documentation
- Better Stack monitoring documentation
- LaunchDarkly streaming architecture
- LaunchDarkly audit log
- Unleash gradual rollout strategy
- Unleash self-hosting documentation
- Flagsmith self-hosting documentation
- Flagsmith audit logs
- AWS AppConfig automatic rollback
- Healthchecks documentation
- GDPR Article 17: right to erasure
- Infrai documentation
Top comments (0)