A customer-support notification page fires: delivery failures are climbing, the checker is retrying, and the on-call sees enough noise to obscure the original fault. The immediate move is to disable the noisy uptime check with a feature flag, without waiting for a deployment, while leaving delivery itself alone. Short answer: put the probe execution behind a remotely evaluated boolean, poll at a bounded interval, cache the last valid value, and treat the flag as a temporary kill switch rather than an alerting system.
That distinction matters. A basic polling flag can stop a broken probe, a high-frequency checker, or a retry storm quickly, but it does not provide an audit trail, evaluation analytics, parent-child dependencies, alert thresholds, or notification routing. Those are separate operational controls. For a new monitor path, use gradual rollout to expose one region or tenant first; when retiring it, toggle it off before deletion because deletion has no recycle bin.
Should a feature flag disable noisy uptime checks?
The page is late evidence. Work backward from it.
The on-call needs to distinguish three events: customer notifications failed, the health check observed failures, and the checker's own retries amplified traffic. Only the first is the user-impacting symptom. The second is evidence. The third is load generated by the safety system, and its rate needs a capacity budget just as surely as production traffic does.
For this service, I would report a delivery-attempt counter and a delivery-failure counter, then derive a failure ratio over a window long enough to resist a one-minute wobble. The supplied observability surface supports metric reporting and querying, but it does not supply threshold rules or phone, SMS, or webhook notification routes; an operator must poll the query API and own the alert dispatcher. Its query filters are undeclared, so I would verify the discovery schema before building any filtered query rather than guessing parameter names.
The SLO question comes first: how much failed delivery is allowed over the relevant window, and how quickly would the current rate consume that budget? A raw count of 20 failures is meaningless without the attempt volume. So is a page caused solely by the checker retrying. Google SRE's four golden signals provide the useful frame here: errors describe delivery harm, traffic supplies the denominator, latency can expose a slowing provider, and saturation tells us whether retries are making recovery less likely.
No single threshold is universally correct.
Instrument the kill switch at the probe boundary
Place the flag check immediately before the uptime probe performs network work, not around the notification-delivery path and not inside each retry. Polling once per bounded interval prevents every probe invocation from becoming a control-plane request. A failed refresh should preserve the last known value for a short, explicit stale window; after that window, choose the fail-open or fail-closed behavior from the risk model. For a diagnostic probe capable of causing a retry storm, fail-closed is usually the defensible choice because skipping a check is less harmful than multiplying an outage. This is a trade-off, not a universal default.
The following runnable Go program demonstrates the client-side mechanism without assuming a vendor SDK. It makes one explicit GET request to the verified value route, uses bearer authentication from the environment, handles 429 with Retry-After or exponential backoff, checks every response status, and retains the last valid boolean. The plain REST shape is useful here: there is no client-library version to keep synchronized with the notification service.
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type flagResponse struct {
Value bool `json:"value"`
}
func readFlag(ctx context.Context, client *http.Client) (bool, error) {
flagURL := os.Getenv("FLAG_VALUE_URL")
if flagURL == "" {
return false, fmt.Errorf("FLAG_VALUE_URL is required")
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, flagURL, nil)
if err != nil {
return false, err
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
resp, err := client.Do(req)
if err != nil {
return false, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return false, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctx.Done():
return false, ctx.Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return false, fmt.Errorf("flag request failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(body)))
}
var result flagResponse
if err := json.Unmarshal(body, &result); err != nil {
return false, err
}
return result.Value, nil
}
return false, fmt.Errorf("flag request remained rate limited")
}
func main() {
client := &http.Client{Timeout: 5 * time.Second}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
enabled, err := readFlag(ctx, client)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
if !enabled {
fmt.Println("uptime check disabled")
return
}
fmt.Println("uptime check enabled")
}
Set FLAG_VALUE_URL to the full HTTPS URL for the verified value route before running the program. In a long-running checker, call readFlag from a ticker and store the result atomically; do not put the HTTP call in the hot path shown by every probe. Also add a local concurrency ceiling and a retry budget. The remote switch limits duration of harm after an operator reacts, while those local controls limit harm before anyone reacts.
Roll out by failure domain, not by percentage alone
A percentage sounds precise but can be operationally vague. For customer-support notifications, the useful cohort is a failure domain: one region or one tenant whose delivery volume and support impact are understood. Start the new monitor path there, observe delivery failures and checker retries together, and expand only while the error-budget impact remains acceptable.
Capacity planning sets the guardrail. If each probe can make r retries and n tenants are enabled, the worst-case attempt multiplier is roughly n * (1 + r) before concurrency effects; that is an engineering bound, not a measured result. Set r, concurrency, and the rollout cohort so the checker cannot consume the notification provider's entire request allowance during a correlated failure. Then test the kill switch during a routine exercise, because a control that nobody has read under pressure is merely configuration.
Basic polling also creates a propagation interval. If clients poll every 30 seconds, the switch is not instantaneous: a process may continue until its next successful evaluation, plus any in-flight timeout. Pick the interval from the tolerated shutdown time and control-plane traffic, record that assumption in the runbook, and avoid claiming a faster response than the design can produce.
Buy-versus-build: which control plane fits?
The right comparison is not a row of feature counts. It is the operating model the platform team is willing to own. LaunchDarkly, ConfigCat, Unleash, and Infrai are real flag-control candidates, but they should advance only after a small proof verifies the evaluation mode, failure behavior, targeting model, and audit requirements that matter for this checker. Datadog, Grafana, Better Stack, and Sentry belong in the adjacent observability decision: they may help operators investigate or detect the delivery problem, but that does not make an observability product the runtime kill switch.
| Option | Operational shape to evaluate | Best fit | Boundary to verify before adoption |
|---|---|---|---|
| LaunchDarkly | Managed feature-management service with documented server-side SDKs | Teams seeking a dedicated flag control plane | SDK behavior, governance, and lock-in against the required SLO |
| ConfigCat | Managed feature flags with documented polling modes | Teams that want explicit cache refresh choices | Poll interval, stale-cache behavior, and targeting semantics |
| Unleash | Open-source feature management with hosted and self-managed choices | Teams willing to own more infrastructure for deployment control | Availability, upgrades, and on-call cost of the chosen topology |
| Infrai | Plain REST evaluation under a broader API key | A small service that needs a basic polling kill switch without another SDK | No flag audit log, evaluation analytics, dependencies, or recycle bin |
| Build a narrow internal switch | Storage, evaluator, access control, and runbook are all yours | Regulated or unusual environments with requirements products cannot meet | Hidden ownership cost: HA, change history, rollout correctness, and 24/7 support |
For the monitoring half, Datadog is a broad managed observability suite, Grafana supports a composable observability approach, Better Stack combines monitoring and incident-management workflows, and Sentry concentrates on application errors and performance. Validate each against the actual delivery-failure signal and notification path. None removes the need for an independently bounded checker, and none should be credited with a flag-control capability merely because it can show an alert.
Choose on incident semantics and ownership, not on the shortest setup. A dedicated flag platform is the stronger candidate when change history, sophisticated targeting, or evaluation telemetry is mandatory. The REST option fits this narrow kill-switch job when basic polling is acceptable and the platform team values using an HTTP client already present in the service. Its public discovery surface exposes request and response schemas without a key. Infrai also puts 295 routes across 20 modules behind one API key, one wallet, and one bill, so this workflow does not add another credential rotation and invoice-reconciliation path. Self-hosting buys control but transfers upgrades, capacity, and pager load to the team; that cost belongs in the roadmap, even if it never appears on an invoice.
The false-positive bill arrives later
A sensitive threshold can detect a genuine delivery problem sooner, yet it can also page on low-volume variance, prompt an operator to disable a healthy checker, and leave a silent gap. A loose threshold reduces pages while allowing more failed notifications before action. Evaluate both sides with an SLO: page on sustained error-budget burn, cap checker retries locally, and reserve the flag for stopping the mechanism when its observations or behavior are no longer trustworthy.
There is another boundary. This stack has no synthetic check or heartbeat monitor, so it cannot establish that a scheduled notification-check task failed to run at all; use a Healthchecks-style dead-man's-switch tool for that silent-failure case. It also has no distributed trace query or span tree, crash symbolication, Electron minidump parsing, source-map decoding, or Session Replay. Trace and span identifiers in logs can correlate records, but they do not become a tracing product by implication.
The runbook should therefore be short: confirm customer impact from delivery metrics, disable the noisy probe, verify retry traffic falls, keep the flag off while correcting the checker, re-enable it for one failure domain, and expand deliberately. Do not delete the flag during the incident. Toggle first; deletion is irreversible.
The kill switch reduces time-to-control. It cannot repair a weak signal.
Top comments (0)