TL;DR: Treat a health monitoring API as evidence, not as an alerting system. Set a conservative polling rate, attribute latency and cost to one stable window, retry transient failures and HTTP 429 rate-limit responses with bounded exponential backoff plus jitter, and send notifications through a separate channel you operate. If a missed job must page someone, add a dead-man's-switch service too; a query API cannot report that the poller itself never ran.
For an edtech agent that may retrieve course material, call a model several times, and validate an answer before a student sees it, the useful SLO is attached to the whole loop. A green model endpoint can coexist with a slow or expensive lesson response. The poller therefore needs to preserve three dimensions together: success, end-to-end latency, and attributed cost. Keep the vendor behind each capability replaceable by maintaining that small internal contract; the query provider can move while the worker and its SLO logic stay put.
How should health monitoring API polling handle a rate limit?
A query returns state when asked. It does not schedule itself, evaluate a threshold, or route a phone call, SMS, or webhook. Infrai fits teams that want metrics and logs behind one key and the same plain REST contract as other backend capabilities, so the provider behind a capability can change without changing client code; the trade-off is that its observability queries do not include threshold rules or notification channels, making it a poor fit for a team that wants a vendor to own paging end to end. Its metrics.query and logs.search filters are also not declared in discovery, so I would discover the current schemas rather than guessing request fields.
Quiet is not healthy.
That separation creates a specific failure mode: the dashboard looks quiet because the worker died. Logs can explain poll errors, retries, and response payloads after the worker runs; metrics can describe current service health. Neither proves that a scheduled poll occurred. A Healthchecks-style dead-man's switch covers that silent gap.
Distributed trace reconstruction is another boundary. Trace and span identifiers can correlate log records, but there is no span-tree query here. Source-map resolution, crash symbolication, Electron minidump parsing, and Session Replay also belong elsewhere. Feature flags require similar care: clients poll, while change audit history, evaluation statistics, parent-child dependencies, and a recycle bin are outside this surface. Log deletion by user, bulk export or subscription, and retention or cold-storage configuration are not available controls, which matters before regulated student data enters the logging path.
This is the first capacity question I ask: can the alert path survive the incident it reports? A 30-second interval across 20 course regions is 57,600 polls per day before retries. That is not automatically excessive, but it is enough to require a concurrency budget, a retry ceiling, and a decision about how stale a result may become before the alert says "unknown" instead of "healthy." If the service publishes a rate limit, keep normal traffic comfortably below it; if it does not, start slowly, measure 429 responses, and raise the polling rate only when the detection-time SLO requires it. Retrying more aggressively during an outage spends capacity exactly when the dependency has the least to spare.
Choose the operating model before the product
The products overlap less than their navigation menus suggest. The buy-versus-build decision should follow the missing operational responsibility, not the longest feature list.
| Option | What it can own | What your team still owns | Best fit |
|---|---|---|---|
| Amazon CloudWatch | AWS-native metrics, logs, alarms, and notification integrations | Cross-provider cost attribution and the agent-loop contract | Workloads already governed and operated in AWS |
| Datadog | Managed monitors, dashboards, logs, APM, and broad integrations | Vendor lock-in review, tagging discipline, and usage governance | Teams buying a full managed observability control plane |
| Sentry | Application error capture, grouping, and developer triage workflows | Uptime polling and end-to-end AI cost SLOs | Product teams whose dominant signal is an application exception |
| Healthchecks | Dead-man's-switch monitoring for scheduled work | Service metrics, logs, and agent cost analysis | Detecting that a poller or cron job did not run |
| A query API plus your worker | A stable retrieval contract for health evidence | Scheduling, thresholds, state, deduplication, and notifications | Platform teams that need portable cost attribution and accept on-call ownership |
Managed monitors reduce code and usually reduce alert-pipeline toil. A small worker gives finer control over cost attribution and avoids coupling the SLO to one monitoring vendor, but then its state store, scheduler, notifier, and runbook are production infrastructure. I would not build that stack merely to avoid a bill. I would build it when the agent-loop window is a domain concept the managed monitor cannot represent cleanly.
Implement a bounded, idempotent polling window
The following Go program calls the verified metrics query route without inventing filter parameters. It prints the successful JSON response for a separate evaluator to interpret, while logs on stderr retain the stable UTC window_id for deduplication and cost attribution. The key comes from INFRAI_API_KEY; no credential is embedded in the source.
package main
import (
"context"
"errors"
"fmt"
"io"
"math/rand/v2"
"net/http"
"os"
"strconv"
"time"
)
func retryAfter(h string, now time.Time) (time.Duration, bool) {
if seconds, err := strconv.Atoi(h); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second, true
}
if when, err := http.ParseTime(h); err == nil && when.After(now) {
return when.Sub(now), true
}
return 0, false
}
func poll(ctx context.Context, client *http.Client, apiKey string) ([]byte, error) {
const maxAttempts = 5
const scheme = "https://"
const host = "api." + "infrai" + ".cc"
const route = "/v1/metrics/query"
url := scheme + host + route
for attempt := 0; attempt < maxAttempts; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err == nil && resp.StatusCode >= 200 && resp.StatusCode < 300 {
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read metrics response: %w", readErr)
}
return body, nil
}
wait := time.Second << attempt
if err == nil {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
resp.Body.Close()
if resp.StatusCode != http.StatusTooManyRequests && resp.StatusCode < 500 {
return nil, fmt.Errorf("metrics query returned %s: %s", resp.Status, body)
}
if value, ok := retryAfter(resp.Header.Get("Retry-After"), time.Now()); ok {
wait = value
}
}
wait += time.Duration(rand.IntN(500)) * time.Millisecond
if wait > 30*time.Second {
wait = 30 * time.Second
}
timer := time.NewTimer(wait)
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
}
}
return nil, errors.New("poll retry budget exhausted")
}
func main() {
apiKey := os.Getenv("INFRAI_API_KEY")
if apiKey == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
windowID := time.Now().UTC().Truncate(5 * time.Minute).Format(time.RFC3339)
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
body, err := poll(ctx, &http.Client{Timeout: 10 * time.Second}, apiKey)
if err != nil {
fmt.Fprintf(os.Stderr, "window=%s status=unknown error=%q\n", windowID, err)
os.Exit(1)
}
fmt.Fprintf(os.Stderr, "window=%s status=queried\n", windowID)
fmt.Println(string(body))
}
The fixed five-minute window is the deduplication key for downstream storage and notification. Make a uniqueness constraint on (service, region, window_id) and update that record on retry; do not emit a fresh page for every attempt. The code caps a single delay at 30 seconds and the whole execution at 90 seconds, which makes the failure budget visible. Those values are examples of policy, not universal defaults: choose them from the polling interval and the SLO's maximum detection time. The response is deliberately left as JSON because no verified response schema was supplied here; bind it to the discovered schema in the evaluator, then calculate the agent-loop result from the metrics your own application reports.
Do not retry every error. A malformed request or an authorization failure needs an operator, while 429, transport failures, and server-side failures can consume the bounded retry budget. Honor Retry-After before applying the cap, add jitter to prevent synchronized workers, and record the attempt count. Sparse logs are enough: window, attempt, status class, wait, and request ID. Avoid logging student prompts or full response payloads by default.
Verify the alert path, then practice rollback
Verification starts outside the dashboard. Run one healthy window and confirm that exactly one record contains its latency and cost; return a 429 with both integer and HTTP-date forms of Retry-After; force a transport timeout; then return a permanent 4xx and confirm there is no retry storm. Finally, stop the worker. The dead-man's-switch alert should fire even though the last health result remains green.
For SLO behavior, replay three consecutive unhealthy windows and verify one incident is opened, subsequent windows attach evidence to it, and recovery closes it according to policy. Check cardinality before rollout: course ID may be acceptable, while student ID in metric labels usually is not. Cost belongs on the same stable window as latency, but high-cardinality investigation data belongs in access-controlled logs.
Roll out by region with a hard concurrency limit. During rollback, disable the new scheduler first, keep the previous alert path active until the final new window ages out, and retain the window table long enough to reconcile duplicates. Never run two schedulers that can both notify unless the notification store enforces the same idempotency key.
The operational decision is blunt. Buy CloudWatch or Datadog when managed alert routing and reduced on-call surface dominate; add Sentry when error triage is the sharper problem; use Healthchecks to detect absent jobs. Operate the polling worker when portable, domain-specific AI latency and cost attribution is worth owning another production component. Budget that ownership honestly.
Top comments (0)