A media app has a stricter constraint than merely detecting downtime: after a rollback, the team still needs enough evidence to explain what a customer saw. I would choose a small, independent monitoring path with US and EU probes, a public incident page, and a separate cron heartbeat, then preserve deployment and application evidence outside the release being rolled back.
TL;DR: test one shallow availability endpoint from both regions, test one narrow user journey, send scheduled-job heartbeats independently, and make incident publication survive an application rollback. Treat the lowest monthly price as a filter, not the decision rule. The winning setup is the one that can prove which release, region, dependency class, and customer-visible symptom overlapped.
This is an experiment note, so the evaluation constraint comes first: can an operator reconstruct a customer incident after the suspected release has been removed? A simple ping against the home page fails that test. It says that an HTTP response happened; it does not connect playback or article delivery to a release, a region, or a delayed background job.
1. What should cheap uptime monitoring plus a status page preserve?
The monitoring control plane should not share the fate of the FastAPI deployment it watches. Keep probe execution, heartbeat receipt, incident publication, and evidence retention outside the application release boundary. If a bad deploy is rolled back, its observations must remain queryable.
That boundary matters more than a long feature list. For a media service, I would retain a compact evidence record for every failed check: check kind, observation time, probe region, release identifier, HTTP status, latency, and a correlation identifier. Do not put access tokens, article bodies, user prompts, or model output in this record. Evidence should answer the incident question without becoming a second content store.
Green is not evidence.
Here is a focused Python shape for the evidence emitted by a probe. The two-second timeout and the field values are example evaluation settings, not universal targets.
from dataclasses import asdict, dataclass
from datetime import datetime, timezone
import json
import urllib.request
@dataclass(frozen=True)
class CheckEvidence:
check: str
observed_at: str
region: str
release: str
status: int
latency_ms: int
correlation_id: str
def encode_evidence(item: CheckEvidence) -> bytes:
return (json.dumps(asdict(item), separators=(",", ":")) + "\n").encode()
def check_health(url: str, region: str, release: str) -> CheckEvidence:
started = datetime.now(timezone.utc)
request = urllib.request.Request(url, headers={"User-Agent": "availability-probe/1"})
with urllib.request.urlopen(request, timeout=2) as response:
finished = datetime.now(timezone.utc)
return CheckEvidence(
check="shallow-health",
observed_at=finished.isoformat(),
region=region,
release=release,
status=response.status,
latency_ms=int((finished - started).total_seconds() * 1000),
correlation_id=response.headers.get("X-Correlation-ID", "missing"),
)
The example deliberately records missing rather than inventing an identifier. It also keeps the shallow check shallow. A health endpoint that synchronously walks every database, queue, feed, and model dependency can convert one dependency slowdown into noisy global alerts.
2. Separate reachability from the customer journey
Use two layers. The frequent layer asks whether the edge can reach a cheap FastAPI endpoint and receive the expected response. A less frequent synthetic journey requests a known public media item, validates a stable property, and stops before it creates customer data or expensive AI work.
This is where notebook-to-prod habits help. Start with a tiny fixture and an explicit assertion, then promote the same assertion into the probe. Do not copy an exploratory notebook that downloads a full video, invokes a model, or depends on a changing headline. The production check needs a bounded request and a stable oracle.
For example, a synthetic check can verify that a designated test article returns the expected content type and a release header. It should not assert the exact recommendation order; ranking changes are application behavior, not necessarily availability failures. For AI-assisted summaries, keep quality in an eval harness and availability in the uptime path. Combining them makes every prompt or model change look like an outage and consumes tokens during an incident.
Fast is useful. Diagnosable is better.
Core Web Vitals are also a different signal. LCP, CLS, and INP describe user experience, and web.dev explains that the recommended assessment uses the 75th percentile, segmented across mobile and desktop. Those field-oriented measures can reveal a customer-visible regression that a health check misses, but they should not replace direct availability checks.
3. Make US and EU probes disagree usefully
Two regions earn their keep only when the alert preserves their separate observations. If the US probe succeeds while the EU probe fails, collapsing both into a single red status discards the first useful clue. Store the region on every sample and open an incident according to a written confirmation policy.
A practical experiment might run the same shallow check from each region and require repeated failure before paging, while still retaining the first failure as evidence. The exact interval and failure count belong to the team's recovery objective and traffic pattern; there is no defensible universal number in the available evidence. Test the policy with injected DNS failure, an origin timeout, and an edge response that is syntactically successful but serves the wrong fixture.
The public incident page can stay simpler than the internal evidence. Publish affected capability, affected geography when known, investigation state, and timestamps. Avoid pasting internal traces or speculative root causes. Customers need an accurate service narrative; responders need the richer private record.
The trade-off is duplication. Independent probes, evidence storage, and incident publication create three small operational surfaces instead of one convenient application plugin. This pattern is a poor fit for a prototype with no on-call owner, no meaningful rollback process, and no customer promise to publish incidents; a single external check may be the more honest choice there. It is also insufficient for a media workflow whose main risk is wrong or unsafe content rather than availability. That case needs an eval suite and editorial controls beside monitoring, because a perfectly reachable endpoint can still return a bad summary. Independence buys rollback safety, but the team must test credentials, retention, and ownership for each surface.
Write that cost down.
4. Treat cron silence as its own failure mode
A scheduled media job can stop producing thumbnails, transcripts, feeds, or index updates while every HTTP health check remains green. Give each critical schedule an independent heartbeat identity, and alert on absence after a defined grace period.
Silence is the signal.
Send the heartbeat only after the useful work and its durable commit complete. A ping at job start proves that a scheduler fired, not that the feed was published. For a multi-stage job, record the run identifier and release in internal evidence, but keep the heartbeat payload free of customer content.
There is a prompt-cost angle here too. If the job calls an AI model, separate three outcomes: scheduler liveness, pipeline completion, and eval quality. Retrying an expensive generation just to satisfy a liveness signal hides the original failure and can amplify spend. The heartbeat reports completion; the eval harness decides whether the output was acceptable.
5. Measure the rollback drill before choosing
The simplest candidate often looks best during setup because one green check is quick to obtain. I would reject that comparison. Run a controlled release drill for every candidate architecture and score the evidence after rollback, not the polish before it.
| Drill question | Evidence that should remain | Failure signal |
|---|---|---|
| Which release served the request? | Release and correlation identifiers | Missing or conflicting release data |
| Was impact regional? | Separate US and EU observations | Aggregated result hides disagreement |
| Did scheduled work finish? | Completion heartbeat and run identifier | Start-only heartbeat reports false health |
| What did customers see? | Stable journey assertion and timestamps | Shallow endpoint is the only evidence |
| Can updates continue during rollback? | Incident page outside the app release | Publication fails with the application |
Feature flags can reduce rollback pressure, but they add operational choices of their own. Martin Fowler's treatment of feature toggles distinguishes categories with different longevity and dynamics; that is a useful warning against one undifferentiated flag system. Record the relevant flag state or configuration version with the release evidence, then test both old and new paths before deployment. Do not assume that disabling a flag erases the need to preserve observations from the disabled path.
Before copying this design, measure alert detection delay, false pages, evidence completeness, incident-page independence, and the time required to associate a failed customer journey with a release. Also count probe-generated AI calls and tokens; the target for a basic availability path should usually be zero. Price can break a tie after those checks. It cannot restore evidence that disappeared during rollback.
Top comments (0)