DEV Community

gentjan likaj
gentjan likaj

Posted on

North Star KPI Tests

How saved KPI totals and a simpler reference calculation expose reporting errors that a full historical rebuild can hide.

You change a report, update some transformation logic, and rebuild its history.

The pipeline succeeds. Every month refreshes. The dashboard still shows a familiar trend.

It is tempting to treat that consistency as evidence that the change worked.

But what if the new logic counts every lead twice?

If the same mistake affects the entire historical period, the chart can still look reasonable. January, February, and March all move together. The growth rates stay the same. The totals are wrong.

This is the problem behind a check we call North Star
: keeping an anchor to reality when the reports themselves can change.

How rebuilding history can hide an error

Consider a report that counts leads by month. A change introduces a join that returns two rows for each lead. The final aggregation sums those rows, and a historical rebuild applies that logic to every month.

Here is a simplified example:

Month Previously captured total Report after rebuilding
January 100 200
February 110 220
March 120 240

These numbers are illustrative.

Both versions show 10% growth from January to February and about 9.1% from February to March. A review focused on the growth pattern could miss the problem completely.

The rebuilt report agrees with its own rebuilt history because both contain the same error.

Row-level uniqueness tests and checks on join relationships can catch this kind of duplication. They remain valuable. A comparison of headline totals adds another way to detect a problem, especially when the final aggregated report has one row per month and therefore still passes a uniqueness test.

A plausible trend does not establish that the underlying total is correct.

Two paths to the same KPI

The design takes inspiration from the general fat-versus-lean DAG idea associated with Netflix's Checksum approach. A DAG is the chain of processing steps that turns source data into a result.

The full reporting path does the work needed for analysis: joins, enrichment, attribution, business rules, and detailed breakdowns. That complexity serves a purpose, but it also creates places where a total can change unexpectedly.

The lean path calculates a small set of headline KPIs with fewer transformations.

For a lead-count KPI, that might mean counting eligible lead IDs from a source or core dataset, without joining every dimension needed by the dashboard.

Conceptually, the two paths look like this:

Business data
  |
  +-- Full reporting path ------> Report KPI total
  |
  +-- Simpler reference path ---> Reference KPI total
Enter fullscreen mode Exit fullscreen mode

The totals should agree within an explicit tolerance when they describe the same period and population.

That last condition matters. Both calculations need compatible definitions for eligibility, dates, and geography. A simpler calculation that measures a different population will produce noise instead of a useful check.

The reference should also avoid the report transformation it is supposed to validate. Copying the same problematic join into both paths would allow both calculations to make the same mistake.

A saved comparison point survives the rebuild

A second calculation addresses one part of the problem. We also need a record of what the report said before its history changed.

North Star preserves captured KPI totals after a defined settling period. Each capture belongs to a metric, reporting date, and country, and records when the capture happened.

A routine report rebuild must not overwrite those saved values.

Otherwise, the rebuild would replace both the number under investigation and the evidence we need to assess it.

This gives us two distinct comparisons:

Check Comparison What it helps reveal
Reconciliation Report total versus reference total at capture The two calculation paths disagree
Historical change Today's report value versus its saved value A previously reported number has changed

The arithmetic is straightforward:

reconciliation_difference = report_at_capture - reference_at_capture
historical_difference     = report_now - report_at_capture
Enter fullscreen mode Exit fullscreen mode

In the double-counting example, the saved January value remains 100 while the rebuilt report returns 200. The historical comparison exposes a change that the current dashboard's trend does not reveal.

A saved value is evidence of what we observed at a particular time. It can still contain an error. Its value comes from preserving the comparison point rather than letting it move automatically with every report change.

There is also a practical limit when introducing this system: capturing old periods for the first time gives you today's view of those periods. It cannot recover the values people saw before a previous rebuild unless another historical record exists.

Alerts need context

A difference is a reason to investigate. It does not identify the cause by itself.

Late-arriving records, legitimate corrections, and deliberate definition changes can all move historical totals. A useful alert therefore needs more than a red status.

It should identify the KPI and period, show the compared values, and make the absolute and relative differences visible. It should also explain whether the comparison is ready to evaluate.

Several rules help keep the result useful:

  • Wait for the agreed settling period. Recently reported data may still be incomplete.
  • Use metric-specific tolerances. Percentage changes and absolute differences matter differently at different volumes.
  • Check completeness. A missing capture must not quietly become a successful comparison.
  • Keep an expiry and a reason for accepted differences. Acknowledging an issue should preserve the evidence and its explanation.

A deliberate metric-definition change needs an explicit decision about comparability. Version the definition, keep the earlier evidence, and record when the new baseline begins. Silently replacing the old baseline would remove the history needed to explain the change.

What the anchor can and cannot tell us

The strength of the check depends on where the two paths separate.

If both paths read the same incorrect upstream total, they can agree and still be wrong. Such a check can validate downstream aggregation without proving that the upstream source is correct.

Headline totals also cannot expose every problem. An overcount in one segment could cancel an undercount in another. Comparing by country or another meaningful dimension helps, while more detailed tests remain necessary.

North Star gives us a durable point of comparison: what we captured, what the simpler calculation produced, and what the report says now.

That matters whenever a team can rebuild history. A familiar chart can hide a changed total. Preserved evidence makes that change visible and gives the investigation somewhere concrete to start.

Before the next historical rebuild, ask: which number will remain unchanged so we can tell what the rebuild changed?

Top comments (0)