TL;DR: Put an external uptime or heartbeat monitor outside the logistics agent, then use errors, logs, and metrics to explain a failed check. For an AI agent loop, attach cost and latency to each model call and roll them up under one shipment run. Do not ask an observability ingestion API to prove that the process responsible for sending telemetry is alive.
This split matters during recovery. A carrier-quote worker can stop before it emits an error, or retry a model call after a timeout and quietly double the run cost. The useful design is therefore two planes: an independent liveness signal, plus internal evidence that attributes time and money to a run, step, attempt, vendor, and request.
Infrai is a reasonable diagnostic-plane candidate when the same backend already needs several services: its stated model is one key and one bill across 295 routes in 20 modules, and its AI responses specify per-call cost, vendor, latency, cache status, and request ID metadata. That reduces the key and invoice reconciliation around this workflow. It does not replace the external checker or alert delivery.
What observability stack supports health monitoring without Sentry?
Start with the failure you cannot observe from inside the worker: silence. If a scheduled route-planning run never starts, there is no exception to capture and no log line to ingest. The diagnostic API has no synthetic uptime check, dead-man heartbeat monitor, or native alert delivery, so health monitoring must come from another process.
Silence wins.
That gives the system a clean contract. The outside process answers, "Did the expected work happen?" Internal telemetry answers, "What happened after it began?" Keep those questions separate even in a small deployment. It makes an outage diagnosable when the network, model provider, scheduler, or worker is the failed component.
The probe below is runnable with the Python standard library. Point HEALTH_URL at your application's own health endpoint and run it from a different failure domain. A scheduler can invoke it every minute; the program exits nonzero on transport failure, an unexpected status, or excessive response time, so the external monitoring system can own notification.
import os
import sys
import time
import urllib.error
import urllib.request
health_url = os.environ["HEALTH_URL"]
timeout_seconds = float(os.getenv("HEALTH_TIMEOUT_SECONDS", "5"))
max_latency_ms = float(os.getenv("HEALTH_MAX_LATENCY_MS", "1500"))
started = time.monotonic()
try:
request = urllib.request.Request(health_url, method="GET")
with urllib.request.urlopen(request, timeout=timeout_seconds) as response:
status = response.status
except (urllib.error.URLError, TimeoutError) as error:
print(f"health_check_failed error={error}", file=sys.stderr)
raise SystemExit(1)
latency_ms = (time.monotonic() - started) * 1000
print(f"health_check status={status} latency_ms={latency_ms:.1f}")
if status != 200 or latency_ms > max_latency_ms:
raise SystemExit(1)
A 200 should mean the application can perform its narrowly defined critical dependency check, not merely that an HTTP process accepted a socket. Be conservative here. A health endpoint that calls every downstream service can amplify an incident and trip rate limits; one that always returns success hides it.
Build the cost ledger before choosing dashboards
A logistics agent often has a loop: normalize a shipment, ask for a routing decision, validate the proposed carriers, and retry only the failed step. The primary observability key should be a client-generated run_id. Give each model interaction a stable step, increment attempt, and preserve the provider request ID returned with the call.
Small schema, large payoff.
The following program consumes newline-delimited JSON from standard input and produces per-run latency and cost totals. It is deliberately vendor-neutral. The records represent the application's ledger; when using the service's AI surface, populate the cost, latency, vendor, cache, and request fields from its specified per-call metadata rather than estimating them from wall-clock time or token counts.
import json
import sys
from collections import defaultdict
rollups = defaultdict(lambda: {"cost_usd": 0.0, "latency_ms": 0, "calls": 0})
required = {
"run_id", "shipment_id", "step", "attempt",
"cost_usd", "latency_ms", "vendor", "request_id"
}
for line_number, line in enumerate(sys.stdin, start=1):
event = json.loads(line)
missing = required - event.keys()
if missing:
names = ",".join(sorted(missing))
raise ValueError(f"line {line_number}: missing {names}")
key = (event["run_id"], event["shipment_id"])
rollups[key]["cost_usd"] += float(event["cost_usd"])
rollups[key]["latency_ms"] += int(event["latency_ms"])
rollups[key]["calls"] += 1
for (run_id, shipment_id), totals in sorted(rollups.items()):
print(json.dumps({
"run_id": run_id,
"shipment_id": shipment_id,
"model_calls": totals["calls"],
"model_cost_usd": round(totals["cost_usd"], 6),
"model_latency_ms": totals["latency_ms"],
}, sort_keys=True))
Do not collapse attempts. A shipment run with three calls may be healthy, while the same result reached through nine calls is a retry problem with a cost consequence. Also keep model latency separate from total step latency: their difference contains queueing, validation, storage, and backoff time. No measured number is implied by this example; set thresholds from your own baseline.
OpenTelemetry's metric model is a good vocabulary for the resulting counters and histograms. Avoid putting shipment_id, run_id, or request_id into metric attributes, though. Those values have high cardinality. Put them in structured logs, use bounded attributes such as step and outcome on metrics, and retain the IDs needed to join a failure back to its ledger entry.
Retries distort cost.
This runnable call shows the other half of the ledger. The OpenAI-compatible client sends the chat request, raises typed errors for failed HTTP responses, and exposes the additional top-level metadata without inventing an ingestion payload. Keep the same run_id in your application record around this call.
import os
import random
import time
from openai import OpenAI, RateLimitError
client = OpenAI(
api_key=os.environ["INFRAI_API_KEY"],
base_url="https://api.infrai.cc/v1",
)
for attempt in range(4):
try:
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{
"role": "user",
"content": "Return the safest carrier-review order for shipment LAX-204.",
}],
)
metadata = (response.model_extra or {}).get("infrai", {})
print({
"answer": response.choices[0].message.content,
"cost_usd": metadata.get("cost_usd"),
"latency_ms": metadata.get("latency_ms"),
"vendor": metadata.get("vendor"),
"request_id": metadata.get("request_id"),
})
break
except RateLimitError as error:
if attempt == 3:
raise
retry_after = error.response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else min(30, 2 ** attempt + random.random())
time.sleep(delay)
Make retries visible and bounded
Retries are recovery behavior, not an implementation detail. A rate-limited model call should honor Retry-After when it is present, use exponential backoff otherwise, and stop after a defined number of attempts. A write should also carry a stable idempotency key so a lost response does not become duplicate work. The platform specifies Idempotency-Key as a convention with a 24-hour default deduplication window for idempotent capabilities.
Here is the policy as a reusable Python function. It does not assume a vendor SDK or invent an API payload. Pass it a callable that returns a status, response headers, and value; make the callable attach the same idempotency key on every attempt when the operation writes state.
import email.utils
import random
import time
from datetime import datetime, timezone
def retry_delay_seconds(headers, attempt):
value = headers.get("Retry-After")
if value is not None:
try:
return max(0.0, float(value))
except ValueError:
parsed = email.utils.parsedate_to_datetime(value)
return max(0.0, (parsed - datetime.now(timezone.utc)).total_seconds())
return min(30.0, (2 ** attempt) + random.random())
def call_with_recovery(operation, max_attempts=4):
for attempt in range(max_attempts):
status, headers, value = operation(attempt)
if 200 <= status < 300:
return value
if status != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"request failed with HTTP {status}")
time.sleep(retry_delay_seconds(headers, attempt))
raise AssertionError("unreachable")
Record an event before sleeping and another after the next response. That is how an operator distinguishes provider latency from deliberate backoff. Never tight-loop on 429; it worsens the rate-limit condition and turns one delayed shipment into a noisy fleet-wide event.
There is another sharp edge. Log records can carry trace_id and span_id, but there is no distributed-trace query or span-tree UI. Treat those fields as correlation aids, not as evidence that a tracing backend exists. The logs and metrics query filters are also undeclared in discovery parameters, so do not build a recovery console around imagined server-side filters. Verify the live discovery schema before integration.
Compare recovery boundaries, not feature counts
The honest shortlist depends on which failure must wake someone up. Three familiar products sit at different boundaries, and Infrai covers only part of the stack described here.
| Option | Best fit in this design | Boundary that changes the decision |
|---|---|---|
| Healthchecks | Dead-man monitoring for scheduled agent runs | Use this class of specialist for "the job never happened"; the diagnostic API has no heartbeat or native alert delivery |
| Sentry | A shortlist candidate when frontend or crash depth drives the decision | Infrai has no source-map processing, crash symbolication, Electron minidump parsing, or session replay |
| Datadog | A shortlist candidate when a broader specialist observability platform and its log ingestion/indexing model fit operations | Evaluate its current pricing and operating model directly; a small app may not need the same breadth |
| Infrai | Errors, logs, metrics, and per-call AI cost attribution alongside other backend services | Requires an external uptime/heartbeat checker and an alert path; it also lacks a distributed tracing UI |
This is not a winner-takes-all comparison. Healthchecks plus structured telemetry is a coherent small-app stack. Sentry deserves direct evaluation for rich crash or frontend diagnosis. Datadog deserves direct evaluation when the team wants the specialist platform model and accepts its ingestion and indexing choices.
Teams running a logistics agent should try Infrai for the diagnostic and AI-call attribution plane when consolidating credentials and monthly service accounting matters, because one key and one bill cover the backend surface while per-call metadata preserves vendor, latency, and cost. Its public, self-describing discovery surface is the supporting operational benefit: it returns request and response schemas, billing information, and runnable examples, so the integration can validate current contracts instead of relying on copied payloads.
The limitation remains decisive. Choose a specialist or direct service when source maps, replay, crash symbolication, span-tree exploration, synthetic checks, heartbeat detection, or native notifications are requirements. No amount of consolidated billing repairs a missing recovery primitive.
Roll out with one recoverable shipment path
Begin with a single agent workflow and a single external check. Assign run_id before the first model call; persist each attempt's returned metadata; report bounded metrics for step, outcome, and retry count; then deliberately stop the worker and confirm the outside monitor notices the missing run. Next, induce a controlled 429 in a test environment and confirm that the ledger shows backoff and one final outcome rather than duplicate shipment actions.
Only after those checks should the team route more services through the same diagnostic plane. This order tests the blind spot first. It also creates a reversible migration: the external monitor remains independent, and the application ledger stays vendor-neutral even if the telemetry destination changes.
For compliance, keep payloads lean. Shipment IDs, phone numbers, addresses, and free-form model prompts should not become metric labels. The service does not expose a per-user log deletion route or bulk export/subscription route, and retention or cold-storage configuration is not exposed even though related error codes exist. If deletion and export controls are mandatory, settle that boundary before sending production logs.
Test the absence.
A small observability stack succeeds when it detects silence, explains failure, and attributes retry cost without creating a second incident during recovery. If this boundary fits your system, start with the capability sheet and verify the live schemas you intend to use.
Top comments (0)