How we collect health metadata from Airflow, Glue, dbt and Tableau into S3, serve it through a thin API, and let an AI agent write the daily report.
Every data team knows the feeling. A stakeholder pings you mid-morning: "The dashboard looks off. Is the data updated?" You open Airflow and everything looks green. Then you check Glue and find a failed job. Next is dbt, where a model errored silently. Finally, in Tableau, an extract failed because its upstream table was empty.
The information was all there, just spread across four tools. Project Sentinel is how we gathered it into one view, with a single message waiting for us every morning.
The stack we're monitoring
Our pipeline is a fairly standard AWS setup:
┌──────────────── Airflow (orchestrates everything) ────────────────┐
▼ ▼
APIs / DBs ──► AWS Glue ──► Redshift ──► dbt (transform) ──► Redshift ──► Tableau
- Airflow is the orchestrator. It triggers every step.
- AWS Glue pulls data from external APIs and source databases.
- Redshift is the data warehouse.
- dbt transforms raw data into reporting models.
- Tableau visualizes the results for the business.
Every layer can fail on its own, and each layer's failure looks different:
| Layer | What "broken" looks like |
|---|---|
| Airflow | A pipeline failed, or it's still running past its SLA |
| Glue | A job failed or timed out, with the error buried in logs |
| dbt | A model errored, or a source is stale |
| Data volume | The table updated, but with half the usual rows |
| KPIs | The numbers loaded, but a KPI dropped sharply vs. last week |
| Reports | Numbers have drifted away from the north-star benchmark |
| Tableau | An extract refresh failed |
The last four are the dangerous ones. Everything shows green, but the data is wrong.
The core idea: metadata as data
Sentinel is built on one principle: every tool writes its health metadata as JSON to S3, and every consumer reads from S3.
Airflow ──┐
Glue ─────┤
dbt ──────┼──► collectors ──► S3 (JSON) ──► Gateway API ──► AI agent ──► Slack digest
Tableau ──┤
KPI QA ───┘
This gives us:
- Decoupling. Collectors don't know who consumes their output, and consumers don't care how the data was collected.
- Simplicity. Everything is plain files, so it's easy to inspect, debug and replay.
- One contract. Anything that can make an HTTP call can check pipeline health.
Layer 1: Collectors
The collectors are Airflow jobs that run a few times each morning, spaced a few minutes apart so they don't compete for resources. Each one covers a single tool or check.
Airflow
The collector reads the Airflow API and keeps one summary per pipeline, based on its latest run:
- state
- duration compared with its average
- whether any tasks retried
- owner
- SLA status (met, missed, pending, or no SLA)
SLAs are simple time-of-day deadlines. If a pipeline hasn't finished by its deadline, it's flagged.
Capturing the real error was the hardest part. A red pipeline with "see log" doesn't help anyone. We solved it with the failure callback every pipeline already uses. When a task fails, the callback writes the first meaningful error line to S3, and the collector joins it onto the pipeline's record. We never have to search logs by hand. The callback is best-effort, so monitoring can never break the pipeline it's monitoring.
AWS Glue
The collector calls the Glue API for recent job runs and records status, duration and error message. Some jobs run many times a day, once per partition or file. For those, it keeps every run with its parameters, so the digest can say which run failed rather than just "the job failed".
dbt
A dedicated monitoring job runs dbt and exports two kinds of results:
- Model executions: what ran, what errored, and why.
-
Source tests, two checks per source:
- Freshness: did the table update when it should have?
- Volume: is the latest row count close to the same weekday last week?
The volume check catches the classic silent failure, where a load "succeeds" with a fraction of the usual rows.
KPI monitoring
A table can be fresh and full-sized and still be wrong. So we also monitor the business numbers themselves. For each brand, every core KPI (costs, leads, sessions, orders and so on) is compared with the same day last week and flagged as healthy or partially updated. Very small values are skipped, because a drop from 4 to 2 isn't a signal.
North Star KPI
This check compares reports against the north-star KPI benchmark. It catches cases where a dashboard's numbers have drifted from the source of truth, or where historical months were quietly restated.
Tableau
The collector uses the Tableau API to find the day's failed extract refreshes, along with each datasource's owner, so the right person gets tagged.
Layer 2: A thin gateway
All that JSON sits in S3, served by a small internal HTTP API. The API does one thing: given a file, it reads it from S3 and returns it as JSON.
Because it's minimal, anything can consume it: a dashboard, a script, a CLI, or an AI agent.
Layer 3: The AI digest
This is where Sentinel became useful and not just available.
Every morning, right after the collectors finish, Airflow starts a short-lived AI agent session. The agent gets a version-controlled prompt and a shell tool. It:
- Calls every gateway endpoint.
- Filters to what matters: failures, SLA misses, stale or low-volume sources, KPI drops, north-star alerts and failed extracts.
- Maps each issue to its owner and tags them.
- Trims noisy errors down to the root cause.
- Writes one Slack message: either "all clear" or a grouped list of what's broken.
Airflow posts the message to the team channel and then tears the agent session down.
A typical morning looks like this:
Good morning team!
Airflow Failures
• Cost ingestion pipeline — missing upstream relation @owner
dbt Sources
• Web sessions source — volume check failed (~50% of last week)
KPI QA
• Brand A / Paid Search — Leads partially updated
Overall: 3 issues across 3 sources. Everything else is green.
Lessons from putting an LLM in the loop
1. Never let the agent truncate its input. Some payloads are large. Early on, the agent "helpfully" cut responses down to save context, dropped most of the records, and reported a clean morning that wasn't clean. Now it must filter the full payload deterministically (with jq), which reads every record and returns only the ones that qualify. The context stays small and nothing is missed.
2. Fail loudly on empty replies. An agent session can look healthy and then return nothing. Treat an empty or broken response as a failure, not a quiet success.
3. Don't blindly retry billable, non-idempotent steps. If the agent run fails, the next scheduled run is the recovery. A retry could double-post.
4. Always clean up. Session teardown runs whether the job succeeded or not. Otherwise, leaked sessions eat into your concurrency budget.
5. Keep the prompt in git. The agent is created fresh from a version-controlled prompt on every run, so prompt changes get reviewed like any other code change.
6. Degrade gracefully. If one source is unavailable, the agent marks it as such and still sends the rest. A partial report beats silence.
What changed
Before Sentinel, we found problems when stakeholders did, or later. Now:
- Morning triage takes one Slack message, not four browser tabs.
- Silent failures are caught, including low volume, KPI drops and report drift, not just red pipelines.
- Owners are tagged automatically, so issues go straight to the right person.
- The metadata is reusable. The same files feed other reports and agents.
Build your own
You don't need our exact stack. The pattern is portable:
- Pick one sink (S3, GCS, a table) and have every tool write health JSON to it.
- Capture the root-cause error when the failure happens, in the callback, not afterwards from logs.
- Monitor data, not just jobs: freshness, volume compared with last week, and the business KPIs themselves.
- Put a thin API in front so any consumer can read the health data.
- Let the LLM summarize, not filter. Filter deterministically first, then let the model group, trim, tag and phrase the result.
Green pipelines don't mean correct data. Sentinel gives us one view of both.
How does your team monitor data quality across tools? I'd love to hear what you use in the comments.
Top comments (0)