<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: gentjan likaj</title>
    <description>The latest articles on DEV Community by gentjan likaj (@gentjan_likaj).</description>
    <link>https://dev.to/gentjan_likaj</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148745%2Fa273056b-ac1b-4bf6-8158-0b195c5f29ca.jpg</url>
      <title>DEV Community: gentjan likaj</title>
      <link>https://dev.to/gentjan_likaj</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gentjan_likaj"/>
    <language>en</language>
    <item>
      <title>Project Sentinel: One Morning Digest for Every Failure in Our Data Stack</title>
      <dc:creator>gentjan likaj</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:21:32 +0000</pubDate>
      <link>https://dev.to/gentjan_likaj/project-sentinel-one-morning-digest-for-every-failure-in-our-data-stack-2i73</link>
      <guid>https://dev.to/gentjan_likaj/project-sentinel-one-morning-digest-for-every-failure-in-our-data-stack-2i73</guid>
      <description>&lt;p&gt;How we collect health metadata from Airflow, Glue, dbt and Tableau into S3, serve it through a thin API, and let an AI agent write the daily report.&lt;/p&gt;

&lt;p&gt;Every data team knows the feeling. A stakeholder pings you mid-morning: &lt;em&gt;"The dashboard looks off. Is the data updated?"&lt;/em&gt; You open Airflow and everything looks green. Then you check Glue and find a failed job. Next is dbt, where a model errored silently. Finally, in Tableau, an extract failed because its upstream table was empty.&lt;/p&gt;

&lt;p&gt;The information was all there, just spread across four tools. &lt;strong&gt;Project Sentinel&lt;/strong&gt; is how we gathered it into one view, with a single message waiting for us every morning.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack we're monitoring
&lt;/h2&gt;

&lt;p&gt;Our pipeline is a fairly standard AWS setup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;          ┌──────────────── Airflow (orchestrates everything) ────────────────┐
          ▼                                                                   ▼
 APIs / DBs ──► AWS Glue ──► Redshift ──► dbt (transform) ──► Redshift ──► Tableau
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Airflow&lt;/strong&gt; is the orchestrator. It triggers every step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AWS Glue&lt;/strong&gt; pulls data from external APIs and source databases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redshift&lt;/strong&gt; is the data warehouse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;dbt&lt;/strong&gt; transforms raw data into reporting models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tableau&lt;/strong&gt; visualizes the results for the business.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every layer can fail on its own, and each layer's failure looks different:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What "broken" looks like&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Airflow&lt;/td&gt;
&lt;td&gt;A pipeline failed, or it's still running past its SLA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Glue&lt;/td&gt;
&lt;td&gt;A job failed or timed out, with the error buried in logs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dbt&lt;/td&gt;
&lt;td&gt;A model errored, or a source is stale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data volume&lt;/td&gt;
&lt;td&gt;The table updated, but with half the usual rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KPIs&lt;/td&gt;
&lt;td&gt;The numbers loaded, but a KPI dropped sharply vs. last week&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reports&lt;/td&gt;
&lt;td&gt;Numbers have drifted away from the north-star benchmark&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tableau&lt;/td&gt;
&lt;td&gt;An extract refresh failed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last four are the dangerous ones. &lt;strong&gt;Everything shows green, but the data is wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The core idea: metadata as data
&lt;/h2&gt;

&lt;p&gt;Sentinel is built on one principle: &lt;strong&gt;every tool writes its health metadata as JSON to S3, and every consumer reads from S3.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; Airflow ──┐
 Glue ─────┤
 dbt ──────┼──► collectors ──► S3 (JSON) ──► Gateway API ──► AI agent ──► Slack digest
 Tableau ──┤
 KPI QA ───┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives us:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Decoupling.&lt;/strong&gt; Collectors don't know who consumes their output, and consumers don't care how the data was collected.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Simplicity.&lt;/strong&gt; Everything is plain files, so it's easy to inspect, debug and replay.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One contract.&lt;/strong&gt; Anything that can make an HTTP call can check pipeline health.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Layer 1: Collectors
&lt;/h2&gt;

&lt;p&gt;The collectors are Airflow jobs that run a few times each morning, spaced a few minutes apart so they don't compete for resources. Each one covers a single tool or check.&lt;/p&gt;

&lt;h3&gt;
  
  
  Airflow
&lt;/h3&gt;

&lt;p&gt;The collector reads the Airflow API and keeps one summary per pipeline, based on its latest run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;state&lt;/li&gt;
&lt;li&gt;duration compared with its average&lt;/li&gt;
&lt;li&gt;whether any tasks retried&lt;/li&gt;
&lt;li&gt;owner&lt;/li&gt;
&lt;li&gt;SLA status (met, missed, pending, or no SLA)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;SLAs are simple time-of-day deadlines. If a pipeline hasn't finished by its deadline, it's flagged.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Capturing the real error&lt;/strong&gt; was the hardest part. A red pipeline with "see log" doesn't help anyone. We solved it with the failure callback every pipeline already uses. When a task fails, the callback writes the first meaningful error line to S3, and the collector joins it onto the pipeline's record. We never have to search logs by hand. The callback is best-effort, so monitoring can never break the pipeline it's monitoring.&lt;/p&gt;

&lt;h3&gt;
  
  
  AWS Glue
&lt;/h3&gt;

&lt;p&gt;The collector calls the Glue API for recent job runs and records status, duration and error message. Some jobs run many times a day, once per partition or file. For those, it keeps every run with its parameters, so the digest can say &lt;em&gt;which&lt;/em&gt; run failed rather than just "the job failed".&lt;/p&gt;

&lt;h3&gt;
  
  
  dbt
&lt;/h3&gt;

&lt;p&gt;A dedicated monitoring job runs dbt and exports two kinds of results:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Model executions&lt;/strong&gt;: what ran, what errored, and why.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Source tests&lt;/strong&gt;, two checks per source:

&lt;ul&gt;
&lt;li&gt;
&lt;em&gt;Freshness&lt;/em&gt;: did the table update when it should have?&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Volume&lt;/em&gt;: is the latest row count close to the same weekday last week?&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The volume check catches the classic silent failure, where a load "succeeds" with a fraction of the usual rows.&lt;/p&gt;

&lt;h3&gt;
  
  
  KPI monitoring
&lt;/h3&gt;

&lt;p&gt;A table can be fresh and full-sized and still be wrong. So we also monitor the &lt;strong&gt;business numbers&lt;/strong&gt; themselves. For each brand, every core KPI (costs, leads, sessions, orders and so on) is compared with the same day last week and flagged as healthy or partially updated. Very small values are skipped, because a drop from 4 to 2 isn't a signal.&lt;/p&gt;

&lt;h3&gt;
  
  
  North Star KPI
&lt;/h3&gt;

&lt;p&gt;This check compares &lt;strong&gt;reports against the north-star KPI benchmark&lt;/strong&gt;. It catches cases where a dashboard's numbers have drifted from the source of truth, or where historical months were quietly restated.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tableau
&lt;/h3&gt;

&lt;p&gt;The collector uses the Tableau API to find the day's failed extract refreshes, along with each datasource's owner, so the right person gets tagged.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 2: A thin gateway
&lt;/h2&gt;

&lt;p&gt;All that JSON sits in S3, served by a small internal HTTP API. The API does one thing: given a file, it reads it from S3 and returns it as JSON.&lt;/p&gt;

&lt;p&gt;Because it's minimal, anything can consume it: a dashboard, a script, a CLI, or an AI agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 3: The AI digest
&lt;/h2&gt;

&lt;p&gt;This is where Sentinel became &lt;em&gt;useful&lt;/em&gt; and not just &lt;em&gt;available&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Every morning, right after the collectors finish, Airflow starts a short-lived AI agent session. The agent gets a version-controlled prompt and a shell tool. It:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Calls every gateway endpoint.&lt;/li&gt;
&lt;li&gt;Filters to what matters: failures, SLA misses, stale or low-volume sources, KPI drops, north-star alerts and failed extracts.&lt;/li&gt;
&lt;li&gt;Maps each issue to its owner and tags them.&lt;/li&gt;
&lt;li&gt;Trims noisy errors down to the root cause.&lt;/li&gt;
&lt;li&gt;Writes one Slack message: either "all clear" or a grouped list of what's broken.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Airflow posts the message to the team channel and then tears the agent session down.&lt;/p&gt;

&lt;p&gt;A typical morning looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Good morning team!

Airflow Failures
  • Cost ingestion pipeline — missing upstream relation   @owner
dbt Sources
  • Web sessions source — volume check failed (~50% of last week)
KPI QA
  • Brand A / Paid Search — Leads partially updated

Overall: 3 issues across 3 sources. Everything else is green.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Lessons from putting an LLM in the loop
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Never let the agent truncate its input.&lt;/strong&gt; Some payloads are large. Early on, the agent "helpfully" cut responses down to save context, dropped most of the records, and reported a clean morning that wasn't clean. Now it must &lt;strong&gt;filter&lt;/strong&gt; the full payload deterministically (with &lt;code&gt;jq&lt;/code&gt;), which reads every record and returns only the ones that qualify. The context stays small and nothing is missed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Fail loudly on empty replies.&lt;/strong&gt; An agent session can look healthy and then return nothing. Treat an empty or broken response as a failure, not a quiet success.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Don't blindly retry billable, non-idempotent steps.&lt;/strong&gt; If the agent run fails, the next scheduled run is the recovery. A retry could double-post.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Always clean up.&lt;/strong&gt; Session teardown runs whether the job succeeded or not. Otherwise, leaked sessions eat into your concurrency budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Keep the prompt in git.&lt;/strong&gt; The agent is created fresh from a version-controlled prompt on every run, so prompt changes get reviewed like any other code change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Degrade gracefully.&lt;/strong&gt; If one source is unavailable, the agent marks it as such and still sends the rest. A partial report beats silence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Before Sentinel, we found problems when stakeholders did, or later. Now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Morning triage takes one Slack message&lt;/strong&gt;, not four browser tabs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent failures are caught&lt;/strong&gt;, including low volume, KPI drops and report drift, not just red pipelines.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Owners are tagged automatically&lt;/strong&gt;, so issues go straight to the right person.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The metadata is reusable.&lt;/strong&gt; The same files feed other reports and agents.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Build your own
&lt;/h2&gt;

&lt;p&gt;You don't need our exact stack. The pattern is portable:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick one sink&lt;/strong&gt; (S3, GCS, a table) and have every tool write health JSON to it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capture the root-cause error when the failure happens&lt;/strong&gt;, in the callback, not afterwards from logs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitor data, not just jobs&lt;/strong&gt;: freshness, volume compared with last week, and the business KPIs themselves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put a thin API in front&lt;/strong&gt; so any consumer can read the health data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let the LLM summarize, not filter.&lt;/strong&gt; Filter deterministically first, then let the model group, trim, tag and phrase the result.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Green pipelines don't mean correct data. Sentinel gives us one view of both.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;How does your team monitor data quality across tools? I'd love to hear what you use in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>dataengineering</category>
      <category>airflow</category>
      <category>aws</category>
      <category>ai</category>
    </item>
    <item>
      <title>North Star KPI Tests</title>
      <dc:creator>gentjan likaj</dc:creator>
      <pubDate>Tue, 29 Sep 2026 06:36:30 +0000</pubDate>
      <link>https://dev.to/gentjan_likaj/north-star-kpi-tests-4h4c</link>
      <guid>https://dev.to/gentjan_likaj/north-star-kpi-tests-4h4c</guid>
      <description>&lt;p&gt;How saved KPI totals and a simpler reference calculation expose reporting errors that a full historical rebuild can hide.&lt;/p&gt;

&lt;p&gt;You change a report, update some transformation logic, and rebuild its history.&lt;/p&gt;

&lt;p&gt;The pipeline succeeds. Every month refreshes. The dashboard still shows a familiar trend.&lt;/p&gt;

&lt;p&gt;It is tempting to treat that consistency as evidence that the change worked.&lt;/p&gt;

&lt;p&gt;But what if the new logic counts every lead twice?&lt;/p&gt;

&lt;p&gt;If the same mistake affects the entire historical period, the chart can still look reasonable. January, February, and March all move together. The growth rates stay the same. The totals are wrong.&lt;/p&gt;

&lt;p&gt;This is the problem behind a check we call &lt;strong&gt;North Star&lt;/strong&gt; &lt;br&gt;
: keeping an anchor to reality when the reports themselves can change.&lt;/p&gt;
&lt;h2&gt;
  
  
  How rebuilding history can hide an error
&lt;/h2&gt;

&lt;p&gt;Consider a report that counts leads by month. A change introduces a join that returns two rows for each lead. The final aggregation sums those rows, and a historical rebuild applies that logic to every month.&lt;/p&gt;

&lt;p&gt;Here is a simplified example:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Month&lt;/th&gt;
&lt;th&gt;Previously captured total&lt;/th&gt;
&lt;th&gt;Report after rebuilding&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;January&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;February&lt;/td&gt;
&lt;td&gt;110&lt;/td&gt;
&lt;td&gt;220&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;March&lt;/td&gt;
&lt;td&gt;120&lt;/td&gt;
&lt;td&gt;240&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These numbers are illustrative.&lt;/p&gt;

&lt;p&gt;Both versions show 10% growth from January to February and about 9.1% from February to March. A review focused on the growth pattern could miss the problem completely.&lt;/p&gt;

&lt;p&gt;The rebuilt report agrees with its own rebuilt history because both contain the same error.&lt;/p&gt;

&lt;p&gt;Row-level uniqueness tests and checks on join relationships can catch this kind of duplication. They remain valuable. A comparison of headline totals adds another way to detect a problem, especially when the final aggregated report has one row per month and therefore still passes a uniqueness test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A plausible trend does not establish that the underlying total is correct.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Two paths to the same KPI
&lt;/h2&gt;

&lt;p&gt;The design takes inspiration from the general fat-versus-lean DAG idea associated with Netflix's Checksum approach. A DAG is the chain of processing steps that turns source data into a result.&lt;/p&gt;

&lt;p&gt;The full reporting path does the work needed for analysis: joins, enrichment, attribution, business rules, and detailed breakdowns. That complexity serves a purpose, but it also creates places where a total can change unexpectedly.&lt;/p&gt;

&lt;p&gt;The lean path calculates a small set of headline KPIs with fewer transformations.&lt;/p&gt;

&lt;p&gt;For a lead-count KPI, that might mean counting eligible lead IDs from a source or core dataset, without joining every dimension needed by the dashboard.&lt;/p&gt;

&lt;p&gt;Conceptually, the two paths look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Business data
  |
  +-- Full reporting path ------&amp;gt; Report KPI total
  |
  +-- Simpler reference path ---&amp;gt; Reference KPI total
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The totals should agree within an explicit tolerance when they describe the same period and population.&lt;/p&gt;

&lt;p&gt;That last condition matters. Both calculations need compatible definitions for eligibility, dates, and geography. A simpler calculation that measures a different population will produce noise instead of a useful check.&lt;/p&gt;

&lt;p&gt;The reference should also avoid the report transformation it is supposed to validate. Copying the same problematic join into both paths would allow both calculations to make the same mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  A saved comparison point survives the rebuild
&lt;/h2&gt;

&lt;p&gt;A second calculation addresses one part of the problem. We also need a record of what the report said before its history changed.&lt;/p&gt;

&lt;p&gt;North Star preserves captured KPI totals after a defined settling period. Each capture belongs to a metric, reporting date, and country, and records when the capture happened.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A routine report rebuild must not overwrite those saved values.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Otherwise, the rebuild would replace both the number under investigation and the evidence we need to assess it.&lt;/p&gt;

&lt;p&gt;This gives us two distinct comparisons:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;Comparison&lt;/th&gt;
&lt;th&gt;What it helps reveal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reconciliation&lt;/td&gt;
&lt;td&gt;Report total versus reference total at capture&lt;/td&gt;
&lt;td&gt;The two calculation paths disagree&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Historical change&lt;/td&gt;
&lt;td&gt;Today's report value versus its saved value&lt;/td&gt;
&lt;td&gt;A previously reported number has changed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The arithmetic is straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;reconciliation_difference = report_at_capture - reference_at_capture
historical_difference     = report_now - report_at_capture
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the double-counting example, the saved January value remains 100 while the rebuilt report returns 200. The historical comparison exposes a change that the current dashboard's trend does not reveal.&lt;/p&gt;

&lt;p&gt;A saved value is evidence of what we observed at a particular time. It can still contain an error. Its value comes from preserving the comparison point rather than letting it move automatically with every report change.&lt;/p&gt;

&lt;p&gt;There is also a practical limit when introducing this system: capturing old periods for the first time gives you today's view of those periods. It cannot recover the values people saw before a previous rebuild unless another historical record exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alerts need context
&lt;/h2&gt;

&lt;p&gt;A difference is a reason to investigate. It does not identify the cause by itself.&lt;/p&gt;

&lt;p&gt;Late-arriving records, legitimate corrections, and deliberate definition changes can all move historical totals. A useful alert therefore needs more than a red status.&lt;/p&gt;

&lt;p&gt;It should identify the KPI and period, show the compared values, and make the absolute and relative differences visible. It should also explain whether the comparison is ready to evaluate.&lt;/p&gt;

&lt;p&gt;Several rules help keep the result useful:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Wait for the agreed settling period.&lt;/strong&gt; Recently reported data may still be incomplete.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use metric-specific tolerances.&lt;/strong&gt; Percentage changes and absolute differences matter differently at different volumes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check completeness.&lt;/strong&gt; A missing capture must not quietly become a successful comparison.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep an expiry and a reason for accepted differences.&lt;/strong&gt; Acknowledging an issue should preserve the evidence and its explanation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A deliberate metric-definition change needs an explicit decision about comparability. Version the definition, keep the earlier evidence, and record when the new baseline begins. Silently replacing the old baseline would remove the history needed to explain the change.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the anchor can and cannot tell us
&lt;/h2&gt;

&lt;p&gt;The strength of the check depends on where the two paths separate.&lt;/p&gt;

&lt;p&gt;If both paths read the same incorrect upstream total, they can agree and still be wrong. Such a check can validate downstream aggregation without proving that the upstream source is correct.&lt;/p&gt;

&lt;p&gt;Headline totals also cannot expose every problem. An overcount in one segment could cancel an undercount in another. Comparing by country or another meaningful dimension helps, while more detailed tests remain necessary.&lt;/p&gt;

&lt;p&gt;North Star gives us a durable point of comparison: what we captured, what the simpler calculation produced, and what the report says now.&lt;/p&gt;

&lt;p&gt;That matters whenever a team can rebuild history. A familiar chart can hide a changed total. Preserved evidence makes that change visible and gives the investigation somewhere concrete to start.&lt;/p&gt;

&lt;p&gt;Before the next historical rebuild, ask: &lt;strong&gt;which number will remain unchanged so we can tell what the rebuild changed?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>data</category>
      <category>dbt</category>
      <category>analytics</category>
      <category>dataengineering</category>
    </item>
  </channel>
</rss>
