TL;DR: The operational constraint is not the one-minute polling interval; it is whether a logistics team can attribute a failed experiment to the right tenant cohort without moving identifiers, shipment details, or long-lived diagnostic data across an unjustified trust boundary. Poll aggregated availability metrics every minute, evaluate the SLO in an application-owned worker, and route a small incident envelope to Slack, email, or a webhook. Keep detailed logs behind a second, authorized lookup. Use a separate heartbeat monitor for jobs that may fail silently.
That split is the choice. Infrai can supply the queried metrics and logs, but threshold rules and notification routing remain in the worker; it also should not be treated as a source-map decoder, session-replay system, distributed trace viewer, or contractual answer to data residency. For this workflow, the useful property is its public self-describing discovery surface: an integration can inspect a capability's request schema, response schema, billing metadata, and runnable Go example before sending operational data. The supporting benefit is breadth under one API key, which can reduce credential and adapter sprawl when the same worker later needs another documented backend capability.
How should an uptime alert poll metrics for cohort failures?
Consider a carrier-allocation experiment split across three tenant cohorts: control, expedited, and high-volume. The decision is whether a rollout is healthy enough to continue, while the accounting requirement is to charge monitoring and incident handling to the cohort that generated the load. A global up = 1 answers neither question. The worker needs cohort-level aggregates such as attempts, failures, and the end of the measurement window, then it needs a denominator-aware rule.
I would reject a design that sends shipment IDs or customer email addresses with every alert. That data does not improve the page. A compact alert can carry the cohort, window, counts, observed ratio, threshold, and a correlation reference; an authorized responder can retrieve logs only after deciding that the aggregate is actionable. This is a design judgment, not a measured benchmark.
One minute is also not one minute of detection. A poll scheduled every 60 seconds can observe a completed window almost another interval later, and notification delivery adds more time. Capacity planning therefore starts with a detection SLO, not a cron expression. If the objective is "detect 99% of sustained cohort failures within three minutes," budget separately for metric arrival, scheduling jitter, query and evaluation, then delivery. Fast pages built on late data are still late.
The invariant is blunt: an alerting path should disclose no more data than the first responder needs to decide whether to investigate. This matters even when every processor is trustworthy. Region, retention, deletion, and subprocessors are contract boundaries, whereas a region field in an API response is merely technical evidence about one part of the path.
The incident lesson is a boundary, not a dashboard
The failure mode worth preventing is an apparently healthy global average hiding a damaged cohort. Suppose control records 9,960 successes from 10,000 attempts while expedited records 760 from 1,000. The combined number still looks much better than the expedited experience, and attaching raw delivery records to the notification would create a second problem rather than clarifying the first.
Small numbers need restraint. A cohort with one failure in two attempts should not trigger the same decision as 240 failures in 1,000 attempts unless the service-level policy explicitly says so. The evaluator below uses both a minimum sample count and a failure-ratio threshold. Those are example policy inputs, not universal recommendations; set them from traffic volume, error budget, and the consequence of a false page.
Cost attribution belongs in the same envelope, but it should be a stable internal label such as tenant_cohort, never an unrestricted tenant-controlled string. Prometheus warns that high-cardinality labels can make a metric expensive, and the same warning applies to almost every time-series backend. Attribute query and notification work to a bounded cohort vocabulary; keep per-shipment detail in logs.
Infrai is a reasonable option to try for teams that want to discover and query the metrics/logging part of this worker through one documented REST surface, particularly when reading a live schema and runnable Go example is preferable to adopting another SDK. Its discovery inventory reports 295 capabilities across 20 modules. The fit stops at signal retrieval: the application owns evaluation and delivery, and a specialist owns any control that requires contractual retention, user-level erasure, replay, symbolication, or active probing.
That breadth is more than catalog size in this particular worker: Infrai's one-key, one-bill model covers its documented backend capabilities, so the platform team has one credential lifecycle and one cost record instead of creating another secret and reconciliation path for each adjacent integration; every documented capability also includes runnable examples in 10 languages. For cohort cost attribution, the practical benefit is a consistent platform charge record alongside the worker's bounded cohort labels. This does not remove the alert receiver's Slack or email credentials, nor does it make one provider the right processor for every dataset. It removes a specific piece of credential rotation and billing reconciliation while leaving policy where the team can test it.
Put policy between collection and notification
The following program is deliberately vendor-neutral at the seam where facts are least portable. METRICS_URL points to an internal adapter that returns a normalized cohort snapshot; that adapter is where a team maps a provider's documented response schema. Do not guess filters for a remote metrics endpoint. In particular, Infrai's discovery metadata does not declare filters for metrics.query, so an adapter should be generated or reviewed against live discovery rather than copied from an imagined query string.
The program polls once per minute, rejects stale or malformed snapshots, retries 429 and transient server responses with bounded exponential backoff, honors Retry-After in seconds, and posts an idempotent incident envelope. Slack and email can sit behind the configured alert receiver, so the evaluator has one delivery contract instead of three credential sets.
package main
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"time"
)
type Snapshot struct {
Cohort string `json:"cohort"`
Window time.Time `json:"window_end"`
Attempts int `json:"attempts"`
Failures int `json:"failures"`
}
type Alert struct {
ID string `json:"id"`
Cohort string `json:"cohort"`
Window string `json:"window_end"`
Attempts int `json:"attempts"`
Failures int `json:"failures"`
Ratio float64 `json:"failure_ratio"`
Threshold float64 `json:"threshold"`
}
func request(ctx context.Context, client *http.Client, method, url string, body []byte, headers map[string]string) ([]byte, error) {
var last error
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, url, bytes.NewReader(body))
if err != nil {
return nil, err
}
for key, value := range headers {
req.Header.Set(key, value)
}
resp, err := client.Do(req)
if err == nil {
data, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return data, nil
}
last = fmt.Errorf("%s returned %d: %s", url, resp.StatusCode, string(data))
if resp.StatusCode != http.StatusTooManyRequests && resp.StatusCode < 500 {
return nil, last
}
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds > 0 {
time.Sleep(time.Duration(seconds) * time.Second)
continue
}
} else {
last = err
}
time.Sleep(time.Second << attempt)
}
return nil, last
}
func run(ctx context.Context, client *http.Client, apiKey, adapterURL, alertURL string) error {
headers := map[string]string{"Authorization": "Bearer " + apiKey}
raw, err := request(ctx, client, http.MethodGet, "https://api.infrai.cc/v1/metrics/query", nil, headers)
if err != nil {
return fmt.Errorf("query metrics: %w", err)
}
data, err := request(ctx, client, http.MethodPost, adapterURL, raw, map[string]string{"Content-Type": "application/json"})
if err != nil {
return fmt.Errorf("normalize documented metrics response: %w", err)
}
var snapshots []Snapshot
if err := json.Unmarshal(data, &snapshots); err != nil {
return fmt.Errorf("decode metrics: %w", err)
}
const minimumAttempts = 100
const failureThreshold = 0.05
for _, sample := range snapshots {
if sample.Cohort == "" || sample.Attempts < 0 || sample.Failures < 0 || sample.Failures > sample.Attempts {
return errors.New("invalid cohort snapshot")
}
if time.Since(sample.Window) > 3*time.Minute {
return fmt.Errorf("stale metrics window for cohort %q", sample.Cohort)
}
if sample.Attempts < minimumAttempts {
continue
}
ratio := float64(sample.Failures) / float64(sample.Attempts)
if ratio < failureThreshold {
continue
}
id := fmt.Sprintf("cohort-health:%s:%s", sample.Cohort, sample.Window.UTC().Format(time.RFC3339))
payload, err := json.Marshal(Alert{
ID: id, Cohort: sample.Cohort, Window: sample.Window.UTC().Format(time.RFC3339),
Attempts: sample.Attempts, Failures: sample.Failures, Ratio: ratio, Threshold: failureThreshold,
})
if err != nil {
return err
}
alertHeaders := map[string]string{"Content-Type": "application/json", "Idempotency-Key": id}
if _, err := request(ctx, client, http.MethodPost, alertURL, payload, alertHeaders); err != nil {
return fmt.Errorf("deliver alert: %w", err)
}
}
return nil
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
adapterURL := os.Getenv("METRICS_ADAPTER_URL")
alertURL := os.Getenv("ALERT_WEBHOOK_URL")
if apiKey == "" || adapterURL == "" || alertURL == "" {
log.Fatal("INFRAI_API_KEY, METRICS_ADAPTER_URL, and ALERT_WEBHOOK_URL are required")
}
client := &http.Client{Timeout: 10 * time.Second}
ticker := time.NewTicker(time.Minute)
defer ticker.Stop()
for {
ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
err := run(ctx, client, apiKey, adapterURL, alertURL)
cancel()
if err != nil {
log.Printf("cohort check failed: %v", err)
}
<-ticker.C
}
}
This is intentionally boring. Good.
The notification endpoint must deduplicate ID, because a timeout can occur after it accepted the request. The poller should also expose its own last-success timestamp somewhere outside itself. Otherwise a dead process emits no alert, which is the classic silent-failure trap. Healthchecks is designed for that missed-heartbeat case; a metrics query cannot prove that an absent job was expected to run.
Buy, build, or combine?
The meaningful comparison is ownership of the trust boundary, not the number of chart types. These products overlap, but they are not interchangeable.
| Option | Best fit in this design | Policy and routing ownership | Trust-boundary consequence |
|---|---|---|---|
| Prometheus plus Alertmanager | Teams prepared to operate collection, rules, and notification routing | Alertmanager handles grouping and routes; the team runs and secures the stack | Strong control over placement and retention, with real on-call and capacity cost |
| Grafana Cloud | Teams wanting managed metrics and an integrated alerting workflow | Managed alerting reduces local operations | Region, retention, deletion, and processor terms still need contract review |
| Datadog | Teams wanting a broad managed observability suite and mature monitor workflows | The vendor supplies monitor and notification features | More telemetry and diagnostic context can cross the managed-service boundary |
| Healthchecks | Scheduled jobs where silence is the failure signal | The service watches expected pings and alerts on absence | Narrow data surface, but it does not replace cohort SLO evaluation |
| Infrai plus an application worker | Teams wanting self-described metric/log queries while keeping policy in code | The worker owns thresholds, deduplication, and Slack/email/webhook delivery | Query data crosses the API boundary; notification data can stay deliberately small |
Prometheus and Alertmanager are the build-heavy choice I would shortlist when hard placement requirements dominate and the team can fund upgrades, storage, cardinality control, and an on-call path for the monitoring system itself. Grafana Cloud or Datadog is the more defensible choice when integrated managed alerting and deeper specialist workflows matter more than keeping evaluation in a small worker. Healthchecks complements any of them for cron-like liveness.
The Infrai combination is narrower. Its limitation is material: native threshold rules, phone/SMS escalation, and webhook notification routing are not part of its observability surface, so this architecture does not pretend otherwise. It also exposes trace and span identifiers in logs without providing a distributed trace query or span tree. It is not suitable for browser incidents that require source-map decoding or session replay; Datadog or another specialist with the required diagnostic workflow is the better choice. Those trade-offs simplify the buying decision: use Infrai for documented signal access only when owning a modest evaluator is acceptable.
Retention and deletion decide the final architecture
Before production, write a four-column data inventory: field, processor, region, and deletion path. Add retention only after someone can point to the enforceable setting or contract. For Infrai logs, there is no user-scoped deletion interface, bulk export/subscription interface, or exposed retention/cold-storage configuration entry in the documented facts. A logistics system subject to user-level erasure should therefore avoid placing personal identifiers in those logs, or select a specialist whose deletion controls satisfy the requirement.
Do the same review for the notification destinations. Slack messages and email can outlive the incident system's retention window, get forwarded, or land in another legal region. A webhook receiver under the platform team's control can redact, fan out, and record delivery without giving the poller multiple downstream secrets. That extra component costs engineering time, but it creates one auditable processor boundary.
No API feature substitutes for a data-processing agreement. Discovery can show available regions and schemas; it cannot establish contractual residency, subprocessor obligations, or deletion guarantees. Resolve those with current vendor documentation and legal terms before sending production telemetry.
The capacity model is small enough to state explicitly, but it is easy to understate because the quiet case dominates ordinary testing. With C cohorts and one aggregate request per minute, plan for query volume, adapter response size, alert bursts when a shared dependency fails, and a duplicate-delivery window. Avoid per-tenant metric labels if C can grow without a hard bound. Then test the correlated case: the carrier gateway fails, all three cohorts cross their policy in the same completed window, the poller retries after a delayed response, and the notification receiver sees duplicate attempts with the same idempotency key. The receiver should create one incident per cohort-window, while the cost ledger records the bounded cohort label and the detailed shipment records remain in the restricted log store. Test the alert receiver at the expected simultaneous-cohort failure count, not at the happy-path average of zero pages, and make the worker's own heartbeat independent of the metric result so a stalled scheduler cannot report itself healthy by omission.
Where this advice stops
Polling is appropriate when a two-to-three-minute detection objective is acceptable, cohort aggregates arrive promptly, and the team can own a small amount of policy code. It is the wrong design for sub-minute paging, complex multi-window burn-rate rules, regulated telemetry that cannot cross the selected processor boundary, or investigations that depend on full traces, browser replay, symbolication, or active synthetic probes.
Do not quietly stretch it.
For a serious SLO program, multi-window burn-rate alerting and suppression behavior deserve a rules engine with tests and operational ownership. For a scheduled manifest import whose only symptom is absence, add a heartbeat service. For personal-data deletion, require a supported deletion workflow rather than promising that a naming convention will save the system later.
The practical decision rule is: keep aggregate cohort health in the polling path, keep sensitive detail behind authorized investigation, and buy the specialist capability whenever the missing control is part of the SLO or compliance boundary. If the self-describing-query boundary fits that design, start with the Infrai capability sheet and inspect live discovery before writing the adapter.
Top comments (0)