The page says tenant_experiment stalled, but the on-call needs a more useful answer: which cohort stopped, which scheduled evaluation observed it, and whether the notification represents a new failure or a replay. The least complex design that preserves those answers is a scheduled poller with a small durable state record. It queries one failure-oriented metric, evaluates freshness and consecutive failures, then sends an idempotent webhook containing the evidence used for the decision.
TL;DR: For a small property-management application, polling a metrics query endpoint is a reasonable failure-alerting mechanism when the poll schedule, query window, cohort labels, evaluation result, and notification attempt share one incident key. Do not let a stateless function translate every nonzero sample directly into a page. Require a short run of failed evaluations, detect missing or late data separately, and persist enough state to suppress duplicate deliveries. This makes the alert reconstructable after the tenant-cohort experiment has moved on.
The distinction matters in operations. A webhook is transport, not incident state. Metrics are observations, not proof that a page was delivered. Joining those two with a durable evaluation record is the small piece of architecture that turns a script into an alerting path.
What should a failure alert from a metrics API query prove?
Start at the receiving end. A useful page for this experiment might say that cohort renewal-reminder-b has failed two evaluations, that the latest query covered an explicitly named time interval, and that the newest underlying sample is older than expected. It also carries an incident key, such as a stable hash of the rule and cohort, plus an evaluation ID unique to this scheduled run.
Those fields answer different questions. The incident key groups retries and repeated evaluations into one operational problem. The evaluation ID lets an engineer locate the exact decision. The query interval explains what data was considered. Sample freshness distinguishes a real bad result from a pipeline that has stopped reporting. Without all four, a later timeline is guesswork.
This is where cohort experiments make a generic failures > 0 rule dangerous. A building may have tenants in control and treatment cohorts, while move-ins and lease renewals change the eligible population over time. Aggregate failure counts can rise because the cohort got larger. A ratio can also look healthy while both numerator and denominator are stale. The alert payload should therefore name the cohort and retain the numerator, denominator, and newest sample timestamp used by the evaluator. Use a narrow label set: property ID, unit ID, tenant ID, email address, and lease ID do not belong in metric dimensions; besides privacy concerns, they create an operationally awkward number of series. A cohort label with a bounded vocabulary and a rule identifier is usually enough for the page. Per-tenant evidence belongs in logs or an audit store with appropriate access controls, joined later through the evaluation ID. After being paged for missed jobs and duplicate deliveries, I treat notification delivery as at-least-once and make the receiver interaction idempotent. Those failures create different evidence, yet both can produce the same human-visible symptom if the state model is vague.
Should a scheduled app poll a metrics API for failure alerting?
The scheduled evaluator has four jobs: query, validate, decide, and deliver. Each deserves its own outcome. Collapsing them into a single success counter makes a query timeout indistinguishable from a healthy zero, and makes a rejected webhook look like a rule that never fired.
Yes, within limits. The runtime can be a Node.js app, a Lambda function, or the Go binary illustrated below; that choice does not repair a weak state model. A scheduled poll is a practical alternative when the rule count is small, the detection objective tolerates the polling interval, and the team can own durable state plus delivery retries. It is not appropriate when second-level detection is required, when hundreds of independently changing rules need lifecycle management, or when the team cannot staff the notification path. In those cases, use an alert manager or managed on-call system that owns evaluation, grouping, escalation, and delivery state.
That boundary is deliberate.
For each run, record a compact evaluation envelope before attempting delivery:
| Evidence | Why it survives the incident review |
|---|---|
| Scheduled time and actual start time | Reveals a missed or delayed invocation |
| Query interval and newest sample time | Exposes gaps, lag, and boundary mistakes |
| Cohort, numerator, and denominator | Replays the rule decision without guessing |
| Prior and next alert state | Explains opening, suppression, and recovery |
| Incident key and evaluation ID | Joins retries without merging separate runs |
| Webhook attempt and response class | Separates detection from notification delivery |
Persist the envelope with a conditional update on the incident key. Only the transition from pending to firing should create an opening notification. Later failed evaluations update the evidence and may send a deliberately configured reminder, but they must not create a fresh incident every time the scheduler runs. A transition to resolved sends one recovery notification.
Short outages complicate the boundary. Suppose a poll runs every five minutes and queries exactly the previous five minutes. Scheduler delay, ingestion lag, and interval endpoint semantics can leave a sliver of data unseen or counted twice. Query with a deliberate overlap, then deduplicate by sample timestamp or evaluation identity. The overlap is an engineering parameter, not a magic constant; set it from observed ingestion delay and retain it in rule configuration.
No data is not zero. It needs its own branch.
A missing series may mean the experiment has no eligible tenants, the application stopped emitting, the metrics path is delayed, or the query is wrong. The evaluator should first test freshness and expected activity, then evaluate the failure ratio only when the denominator is meaningful. If “no eligible tenants” is normal outside business processing windows, encode that schedule explicitly rather than teaching the on-call to ignore night pages.
Make the evaluator boring and replayable
The core does not need a commercial paging SDK. It needs deterministic inputs and a state store with compare-and-set behavior. The following Go sketch leaves the metrics API query endpoint, persistence, and webhook delivery behind interfaces so the decision can be replayed in a test or incident review. The numbers are example policy values for this tenant experiment, not universal thresholds.
package alerting
import (
"context"
"fmt"
"time"
)
type Observation struct {
Cohort string
Failures int
Eligible int
NewestAt time.Time
WindowFrom time.Time
WindowTo time.Time
}
type State struct {
Status string // healthy, pending, firing
ConsecutiveFailures int
Version int64
}
type Decision struct {
IncidentKey string
EvaluationID string
Next State
Notify bool
Reason string
}
func Evaluate(now time.Time, evaluationID string, obs Observation, prior State) (Decision, error) {
if obs.Eligible < 0 || obs.Failures < 0 || obs.Failures > obs.Eligible {
return Decision{}, fmt.Errorf("invalid cohort counts")
}
d := Decision{
IncidentKey: "tenant-experiment:" + obs.Cohort,
EvaluationID: evaluationID,
Next: prior,
}
if obs.NewestAt.IsZero() || now.Sub(obs.NewestAt) > 10*time.Minute {
d.Next.ConsecutiveFailures++
d.Reason = "telemetry_stale"
} else if obs.Eligible == 0 {
d.Next = State{Status: "healthy", Version: prior.Version}
d.Reason = "no_eligible_tenants"
return d, nil
} else if float64(obs.Failures)/float64(obs.Eligible) >= 0.05 {
d.Next.ConsecutiveFailures++
d.Reason = "failure_ratio"
} else {
d.Next = State{Status: "healthy", Version: prior.Version}
d.Reason = "within_policy"
d.Notify = prior.Status == "firing"
return d, nil
}
if d.Next.ConsecutiveFailures >= 2 {
d.Next.Status = "firing"
d.Notify = prior.Status != "firing"
} else {
d.Next.Status = "pending"
}
return d, nil
}
type Store interface {
Load(ctx context.Context, incidentKey string) (State, error)
CompareAndSet(ctx context.Context, incidentKey string, oldVersion int64, next State) error
}
The production wrapper should generate an evaluation ID before the query, record query errors as outcomes rather than fabricated zeros, and write the decision before calling the webhook. If the process stops after the write but before delivery, a retry can see an undelivered notification attached to the same evaluation. If it stops after delivery but before acknowledging success, the webhook receiver uses the incident key plus transition as its idempotency key. Exactly-once delivery is not required for exactly-once incident opening.
Test the state machine with table-driven cases: first failure, second failure, continued failure, recovery, stale telemetry, empty cohort, invalid counts, and a concurrent update. Then test the adapter with recorded query responses and a fake receiver that deliberately times out after accepting a request. That last case catches the duplicate most happy-path suites miss.
Deployment deserves the same restraint. Roll out a changed rule in shadow mode, where evaluations and would-notify transitions are retained but no page is sent. Compare the new decisions with the current rule over representative cohort activity. Promote the configuration separately from the evaluator binary so a threshold rollback does not require a code rollback.
The signal that should fire before the tenant-impact page
Working backward from the first page usually reveals an earlier signal: the alerting path itself is unhealthy. The Google SRE monitoring guidance separates symptoms from causes and frames monitoring around latency, traffic, errors, and saturation. For this design, the tenant experiment failure ratio is a symptom-facing signal. Scheduler delay, query errors, stale newest samples, state-write conflicts, and webhook delivery failures describe the alerting machinery.
Do not page on every internal wobble. Track them, then choose the response according to user impact and urgency. A single query timeout can be retried with bounded backoff inside the schedule period. Consecutive scheduler misses or telemetry age beyond the experiment's detection objective deserve escalation because the primary rule has become blind. Webhook rejection should enter a durable retry queue and raise a delivery-path signal; reevaluating the metric will not repair delivery.
This produces two timelines that can be joined. The data timeline says when tenant-facing outcomes occurred and arrived. The control timeline says when evaluations were scheduled, executed, persisted, and delivered. Incident reconstruction needs both.
Logging follows the same structure. Emit one event per evaluation with stable field names, not a paragraph assembled from conditionals. A logging appender or handler can route events to another destination, but alert correctness should not depend on parsing prose logs. The Logback appender documentation is a useful illustration of this separation: appenders deliver logging events, while filtering and lifecycle behavior remain explicit concerns. The language used by the application does not change that boundary.
Choosing the threshold means choosing interruption cost
The tempting rule is immediate: one failed tenant action, one page. It optimizes detection latency on paper and spends human attention freely. In a cohort experiment, isolated application errors, late samples, small denominators, and retried work can all produce a brief nonzero ratio. I choose two consecutive failed evaluations here, trading one polling interval of detection time for resistance to a transient observation. A longer window smooths noise but delays evidence and can blur a sharp regression.
Make that trade explicit in the runbook. Define the target detection time, polling cadence, freshness limit, minimum meaningful denominator, failure ratio, consecutive count, and recovery condition. Also define where low-urgency results go. A ticket or daytime notification can be correct for a slowly degrading experiment; a page is reserved for an actionable condition that cannot wait.
The configuration must account for cohort size. Five failures among 100 eligible tenants and one failure among two tenants both meet a five-percent ratio, but they do not carry the same evidence. A minimum denominator or an absolute-count guard can prevent tiny cohorts from dominating. Avoid inventing statistical certainty from a monitoring threshold, though. The operational rule detects a condition worth investigating; experiment analysis still needs its own statistical method. This is a limitation of threshold alerting, not an implementation bug, and adding more polling cannot remove it.
Watch the alert after launch. Count evaluations by outcome, transitions into firing, suppressed repeats, recoveries, notification retries, and pages that operators mark non-actionable. Review those results by rule and cohort with bounded labels. If most pages close without action, the threshold is consuming attention instead of protecting the tenant workflow.
That cost is the closing constraint. A poller plus webhook can be entirely adequate for a small application, but only when its state and evidence make duplicates harmless and missed runs visible. Set the threshold too low and the system trains people to distrust it. Set it too high and incident reconstruction becomes an explanation of why the signal arrived late. The right setting is the one that meets the response objective with a false-positive rate the on-call team can actually sustain.
Top comments (0)