Short answer: page on a sustained rise in checkout failures, not on every missing flag lookup. Treat a missing key after deletion or recreation as an expected evaluation state, return a conservative local default, and emit a low-cardinality metric that joins the evaluation result to the checkout outcome.
The page says checkout_failure_ratio > 0.02 for 10m. On-call sees the affected region, checkout stage, current ratio, and a graph overlaying failed checkouts with flag_evaluation_total{reason="not_found"}. That answers the first operational question: did a stale toggle coincide with customer-visible failures, or is the lookup noise incidental?
A raw 404 page cannot answer that. Deleting and later recreating a key can leave application instances, caches, or rollout configuration referring to different lifecycle states. The evaluator should degrade predictably while telemetry preserves the reason.
No panic.
What should alert when feature flags delete or recreate a missing key?
Page on the symptom that consumes the logistics checkout SLO: unsuccessful checkout attempts over a sustained window. A missing-key counter is supporting evidence, because a deleted experiment that cleanly falls back may produce evaluations without dropping a single order. Paging on that counter alone converts harmless cleanup into on-call work and trains responders to distrust the alert.
Use two severities. A ticket or dashboard annotation fits a rising not-found rate while checkout remains healthy. A page fits checkout failures burning error budget fast enough to require action, with the missing-key series as one correlated signal. Correlation is not proof, so the first response still checks payment, inventory reservation, address validation, and carrier quoting.
| Signal | Action | Reason |
|---|---|---|
| Missing evaluations rise; checkout is healthy | Investigate in business hours | Fallback contains the fault |
| Checkout failures rise; misses are flat | Inspect other checkout stages | The flag signal does not explain the symptom |
| Both rise in one region and interval | Inspect lifecycle and deployment state | Shared timing makes the flag path a useful lead |
Traffic matters.
With low request volume, one failure can produce a frightening ratio; require a minimum event count before evaluating it. A ten-minute window can damp single-request noise, but that number is only a starting policy. Derive the production threshold from the checkout SLO and its error budget. This method has a limitation: a ratio-and-count rule reacts more slowly in a quiet region, while a single-event rule reacts quickly by accepting more false positives. Choose that trade-off explicitly in the SLO review, and use a synthetic checkout when low traffic would otherwise make detection unacceptably slow.
Instrument the fallback at the decision point
Instrumentation belongs where the application converts an evaluation failure into behavior. Logging only the remote response loses the chosen default; logging only the final checkout result loses the reason. This Go example uses a generic evaluator, recognizes absence without coupling business code to an HTTP status, and records bounded attributes.
package checkout
import (
"context"
"errors"
)
var ErrFlagNotFound = errors.New("flag not found")
type Evaluator interface {
Bool(context.Context, string, bool) (bool, error)
}
type Metrics interface {
AddFlagEvaluation(reason, outcome string)
}
type ToggleReader struct {
eval Evaluator
metrics Metrics
}
func (r ToggleReader) ExpressCheckout(ctx context.Context) bool {
const safeDefault = false
value, err := r.eval.Bool(ctx, "express-checkout", safeDefault)
switch {
case err == nil:
r.metrics.AddFlagEvaluation("resolved", "value_returned")
return value
case errors.Is(err, ErrFlagNotFound):
r.metrics.AddFlagEvaluation("not_found", "fallback_returned")
return safeDefault
default:
r.metrics.AddFlagEvaluation("evaluation_error", "fallback_returned")
return safeDefault
}
}
The default disables the optional express path and retains ordinary checkout. That choice is reviewable: availability of the core order path outranks rollout of the accelerated path. A flag controlling fraud rejection, tax calculation, or legal consent could require a fail-closed default. There is no globally safe Boolean.
Do not attach customer IDs, order IDs, or arbitrary missing keys to metric labels. Put identifiers in sampled logs or traces under the controls already applied to checkout data. Keep a fixed metric vocabulary such as resolved, not_found, and evaluation_error; otherwise a cleanup mistake can also create a cardinality problem.
Work backward from the page
The page should link to four aligned series: checkout attempts, failed checkouts, fallback evaluations, and release markers. Region and checkout stage narrow ownership. A request identifier belongs in traces or logs, where a responder can follow one attempt without multiplying time-series count.
Test the causal path. Can a responder find a failed checkout trace, see the evaluated key, confirm not_found, observe fallback_returned, and identify the next operation that failed? If the trace ends at the evaluator, the instrumentation is incomplete. The useful boundary is the business decision, not the network request.
Contract-test that behavior before deployment:
package checkout
import (
"context"
"testing"
)
type missingEvaluator struct{}
func (missingEvaluator) Bool(context.Context, string, bool) (bool, error) {
return false, ErrFlagNotFound
}
type captureMetrics struct{ reason, outcome string }
func (m *captureMetrics) AddFlagEvaluation(reason, outcome string) {
m.reason, m.outcome = reason, outcome
}
func TestMissingExpressCheckoutFlagUsesFallback(t *testing.T) {
metrics := &captureMetrics{}
reader := ToggleReader{eval: missingEvaluator{}, metrics: metrics}
if reader.ExpressCheckout(context.Background()) {
t.Fatal("expected ordinary checkout fallback")
}
if metrics.reason != "not_found" || metrics.outcome != "fallback_returned" {
t.Fatalf("unexpected telemetry: %#v", metrics)
}
}
Run the contract test in CI, then exercise deletion and recreation outside production with the same cache policy used in production. Acceptance is behavioral: deletion returns the documented fallback and records not_found; recreation eventually returns the configured value and records resolved; neither transition crashes checkout. The exact convergence time depends on evaluation and caching design, so measure it rather than promising an invented bound.
Separate control-plane lag from application failure
A 404 is an HTTP response, not a diagnosis. RFC 9110 defines it as the origin server not finding a current representation, or not being willing to disclose one. Responders still need to distinguish a genuinely absent key from stale configuration, a wrong environment, a malformed key, or a cache that has not refreshed.
Compare deployment version, environment, region, and evaluation reason. If one release reports absence while another resolves the key in the same environment, inspect configuration distribution and cache refresh. If every instance reports absence immediately after an intentional deletion, fallback is doing its job; the cleanup process should remove callers before the key.
Recreation deserves suspicion. Reusing a readable key can make an old caller appear valid while its intended semantics, targeting population, or default have changed. Prefer a new key for a new decision, and treat retirement as a migration: inventory readers, ship fallback-capable code, observe usage reach zero, then delete.
Decide what to operate and what to buy
Signal quality depends more on the application contract and alert design than on the brand behind evaluation. The operating model still changes who owns storage, delivery, upgrades, access control, and incident response. Capacity planning should include evaluation request rate, cache churn during rollout, telemetry volume, and behavior when the control plane is unreachable.
| Decision | Managed service | Self-hosted service |
|---|---|---|
| On-call load | Provider operates its boundary; the team owns defaults and alerts | Platform owns health, upgrades, backup, and capacity |
| Lock-in | Proprietary targeting may enter application code | Internal APIs can become private lock-in |
| Cost model | Service consumption plus contract review | Engineering time, compute, storage, and paging |
| Evidence | Export behavior, failure semantics, auditability, exit path | Load tests, recovery drills, staffing, upgrade history |
Do not select from a feature checklist. Run deletion, stale-cache, unavailable-evaluator, and rollback tests against the candidate boundary, then account for the resulting on-call obligation as honestly as infrastructure spend.
The final step is to tune the alert without hiding the failure.
The dangerous threshold looks precise without an error-budget argument. Alert too aggressively and intentional retirement pages people. Alert only on a large checkout ratio and a low-volume region can fail quietly for too long. Use a minimum request count, sustained window, and checkout SLO to set the page; keep missing evaluations available for diagnosis and a lower-severity cleanup workflow. This design is not suitable when every individual failed decision carries catastrophic consequences; in that case, fail closed where the domain requires it and alert on the single event, accepting the added noise. It is also limited by telemetry delivery: metrics can show association, but only traces, logs, and a reproducible lifecycle test can establish the path through a particular request.
Review the rule after deployments and planned deletions. Count pages that required action, pages where fallback contained impact, and checkout incidents the rule missed. False positives consume attention; false negatives consume error budget.
The operational goal is modest.
Missing keys should be visible, bounded, and boring. Checkout failures should be loud enough to act on, with evidence pointing toward the failing decision rather than merely the nearest 404.
Top comments (0)