A pricing-rule rollout has a harder requirement than detecting a busy error graph: the team must be able to prove which rule was active, which requests failed, and why an alert fired. I would keep the request path in Node.js, record low-cardinality counters at that boundary, and let a small scheduled worker evaluate a durable Postgres window. The flag decision belongs in the evidence, not in a label that can disappear when a process restarts.
TL;DR: Count failed pricing evaluations by stable dimensions, persist each polling window and alert transition, and page only when a minimum event count and a failure ratio are both breached. A count catches volume; a ratio supplies context. Make retries idempotent, delay evaluation beyond ingestion lag, and reconstruct an incident from stored windows plus request logs.
How should a Node.js worker poll metrics for a failure count alert?
Consider a multiplayer game rolling out a new regional pricing rule behind a flag. During the first cohort, the pricing API can fail because the rule rejects an input, a downstream dependency times out, or the application returns an unexpected server error. Those outcomes do not mean the same thing. A single aggregate named failures erases the distinction that the incident review needs.
The invariant is narrower: for every evaluation window, retain the population, failed population, rule revision, cohort, region, and final alert state. Keep player IDs out of metric labels; they create unbounded cardinality and belong in access-controlled logs or traces. The metric dimensions should be finite enough that an operator can estimate their series count before rollout. If there are 3 rule revisions, 4 cohorts, 6 regions, and 5 outcome classes, that is already 360 possible combinations before replicas and routes enter the picture. Capacity planning starts there.
The alert should also describe a service-level symptom. A pricing endpoint that processed 20,000 requests with 120 failures has a 0.6% failure ratio; a quiet cohort with 2 requests and 1 failure has a 50% ratio. Paging on either number alone produces a bad policy. Use a minimum count to reject tiny samples, a ratio to normalize traffic, and a sustained-window condition to avoid treating one scrape gap as an incident. The exact threshold is an SLO decision, derived from the error budget and the risk of showing an invalid price, rather than a universal constant.
Traffic lies.
This is the point many dashboards miss.
A durable polling path
The request service should emit monotonically increasing counters for total pricing evaluations and failures, partitioned only by the bounded dimensions needed during response. The scheduled worker then queries increases over a closed interval, waits long enough for late samples, and writes one evaluation row with a unique key such as (policy_id, window_end, cohort, region). On retry, the worker updates or ignores that same row. It must not create a second notification.
A database-backed state machine is useful because the alert itself becomes reconstructable: healthy can move to firing after the configured number of breached windows, and firing can move to resolved after the recovery condition. Notification delivery is a separate, idempotent step keyed by the transition ID. That separation handles the awkward case where evaluation succeeds but delivery times out; the next run retries delivery without inventing a new incident.
Persist the verdict.
The following Go example shows the preventative boundary. It assumes QueryWindow returns counter increases for a closed interval and that the repository enforces a unique evaluation key. The threshold values are examples, not operational defaults.
package alerting
import (
"context"
"fmt"
"time"
)
type Window struct {
PolicyID string
Cohort string
Region string
End time.Time
}
type Counts struct {
Total uint64
Failed uint64
}
type Evaluation struct {
Window Window
Counts Counts
Ratio float64
Breach bool
}
type Metrics interface {
QueryWindow(context.Context, Window, time.Duration) (Counts, error)
}
type Repository interface {
UpsertEvaluation(context.Context, Evaluation) error
}
type Worker struct {
Metrics Metrics
Repo Repository
Lookback time.Duration
MinTotal uint64
MinFails uint64
MaxRatio float64
}
func (w Worker) Evaluate(ctx context.Context, key Window) error {
counts, err := w.Metrics.QueryWindow(ctx, key, w.Lookback)
if err != nil {
return fmt.Errorf("query closed metrics window: %w", err)
}
ratio := 0.0
if counts.Total > 0 {
ratio = float64(counts.Failed) / float64(counts.Total)
}
result := Evaluation{
Window: key,
Counts: counts,
Ratio: ratio,
Breach: counts.Total >= w.MinTotal &&
counts.Failed >= w.MinFails &&
ratio >= w.MaxRatio,
}
if err := w.Repo.UpsertEvaluation(ctx, result); err != nil {
return fmt.Errorf("persist alert evaluation: %w", err)
}
return nil
}
A real implementation also validates that the queried interval is complete, records query failures as worker-health signals, and limits catch-up work after downtime. If a worker was paused for six hours, blindly evaluating every missed minute may overload the metrics backend and deliver obsolete pages. Set a catch-up budget, store a visible gap, and run incident analysis on that gap deliberately. Silent interpolation is worse than missing data because it creates false confidence.
Reconstruct the rollout, not merely the outage
During an incident, start from the persisted transition: its policy version, evaluation interval, count, ratio, and query timestamp. Join that evidence to the deployment record and flag audit log, then locate representative request traces or structured logs using a correlation ID. Severity fields in syslog have defined numerical meanings, but severity alone does not establish business impact; a pricing-rule rejection can be a handled warning while repeated inability to calculate a price can consume the availability budget.
Feature flags change the code path without requiring a deployment, which is precisely why the flag revision and cohort must be recorded alongside operational evidence. The rollback decision becomes testable: disable or reduce exposure when breached windows correlate with the new rule and the control cohort remains within its objective. If both cohorts degrade, investigate shared dependencies before blaming the rule.
I use this incident record as a compact proof, not as a warehouse for raw telemetry:
- window start and end, including timezone;
- rule revision and bounded cohort key;
- total, failed, and unknown evaluations;
- threshold policy version and resulting state transition;
- links or identifiers for the deployment, flag change, and notification attempt.
Unknown deserves its own counter. Treating timeout, missing telemetry, or an incomplete interval as success makes the ratio look healthy exactly when the observation path is least trustworthy. The alert evaluator should fail closed for its own state update: if the query is incomplete, persist unknown and raise a separate worker-health signal instead of manufacturing a healthy result.
Choosing the operating model
The buy-versus-build question is mostly about who owns state correctness and the 03:00 failure modes, not the chart renderer. A managed alerting service may reduce maintenance, while a self-hosted scheduler can provide tighter control over retention and transition logic. Neither choice removes the need to test duplicate execution, delayed samples, database unavailability, and notification retries.
| Approach | What the team owns | Incident-reconstruction strength | Main trade-off |
|---|---|---|---|
| Metrics-native rule | Labels, query, routing, and runbooks | Good when rule history and evaluations are retained | Simple path, but business rollout context may live elsewhere |
| Scheduled worker plus Postgres | Query semantics, schema, locking, retries, and recovery | Strong when every window and transition is durable | More code and on-call surface |
| Managed evaluator | Integration, policy, identity, and export strategy | Depends on audit retention and evidence export | Lower routine maintenance, with portability constraints |
Before choosing, estimate active series, windows evaluated per hour, rows retained per day, query latency at peak traffic, and the maximum acceptable detection delay. Then test the failure path. Two worker replicas should race safely; a repeated window should produce one transition; a failed notification should retry without reevaluating history; and a flag rollback should be visible without waiting for a fresh deployment.
The scheduled-worker pattern does not fit every system. If the SLO requires detection within seconds, a minute-scale poller is the wrong control loop. If traffic is extremely sparse, a fixed-window ratio remains statistically noisy and a direct event workflow may be clearer. If existing rule evaluation already retains query results, policy revisions, and notification state for the required incident window, another Postgres state machine adds on-call load without adding evidence.
The release gate
A pricing rollout is ready when the team can rehearse the whole chain: send known successes and failures to a test cohort, observe the closed metrics window, verify one durable transition, retry the worker, retry notification delivery, and reconstruct the decision from retained records. The gate should also assert bounded label values and reject an unknown rule revision.
No page can compensate for missing provenance. The useful alert is the one that tells the responder what changed, establishes that the threshold was evaluated over complete data, and survives long enough to support the review after the rollout has been stopped or expanded.
Sources
- Martin Fowler, "Feature Toggles": https://martinfowler.com/articles/feature-toggles.html
- IETF RFC 5424, "The Syslog Protocol": https://datatracker.ietf.org/doc/html/rfc5424
Top comments (0)