DEV Community

Cover image for Why Your Orchestrator Task "Succeeded" But Your Data Is Wrong
Vaishnav Prabhu
Vaishnav Prabhu

Posted on

Why Your Orchestrator Task "Succeeded" But Your Data Is Wrong

A green checkmark means your task ran without throwing an exception. It does not
mean the data it produced is correct. These are two very different guarantees, and
confusing them is one of the most common ways bad data reaches production
dashboards undetected.

The gap between "ran" and "right"

Most orchestrators judge success on exit code: the process finished, nothing threw,
mark it green. But a task can exit cleanly while:

  • loading half the expected rows (an upstream feed was incomplete),
  • duplicating rows because a dimension join fanned out,
  • writing values that are internally consistent but wrong (a silent unit or timezone bug),
  • or succeeding on a retry that silently reprocessed a partial payload.

None of these trip an exception. All of them produce a green run and wrong numbers.

Make correctness a first-class, separate signal

The fix is to stop treating "the job ran" as a proxy for "the data is good," and add
an explicit quality gate as its own step that can fail the pipeline:

  1. Row-count / volume checks — did we land roughly the expected number of rows for this period, versus a recent baseline? A 30%+ swing is worth blocking on.
  2. Uniqueness checks — assert the natural key is unique at the table's grain. This is your early-warning system for fan-out and duplicate loads.
  3. Freshness checks — is the newest record actually from the period you expected?
  4. Reconciliation checks — does this table agree (within tolerance) with an independent source of truth?

Run these after the load, as a gate the downstream steps depend on. A failing gate
should turn the run red and stop propagation — a loud failure now is far cheaper than
a silent one that someone finds in a dashboard next week.

Watch the "succeeded after N retries" trap

A task that fails twice and passes on its third attempt still shows as success. If
the underlying action isn't idempotent, those earlier partial attempts can leave
duplicated or inconsistent data behind. Two defenses:

  • Make the operation idempotent (truncate-and-reload a partition, or MERGE on a key) so a retry is always safe.
  • Alert on "succeeded only after retries" as a distinct signal, so fragility surfaces before it becomes an outage.

Takeaways

  • Exit code answers "did it run," not "is it right." Instrument both.
  • Add uniqueness, volume, freshness, and reconciliation gates as blocking steps.
  • Treat retries and partial reloads as correctness risks, not just reliability ones.

Next in this series: a repeatable RCA framework for when one of these gates finally fires.

Top comments (0)