The launch went well. Dashboards render, numbers reconcile, the stakeholder who sponsored the project says something appreciative in a meeting. The team moves on to the next thing.
Four months later somebody notices that a chart has been flat since March. Not zero, just flat, in a way that looked plausible enough that nobody questioned it. The pipeline feeding it stopped working eleven weeks ago and reported success every single day.
This is the part of the work that doesn't get discussed much, because it isn't the interesting part. But it's where most of the total lifetime effort goes, and knowing what's coming changes how you build.
Here's the inventory.
Upstream systems change without telling you
Your pipelines read from systems owned by other teams and vendors. Those teams ship changes on their own schedule and have no idea you exist.
A column gets renamed. A field that was always populated becomes optional. Someone adds a new status value to an enum that your logic branches on. A vendor updates their API and deprecates the version you're calling, with a notice that went to an email address belonging to someone who left.
None of these are unreasonable acts. They're normal software maintenance from the perspective of the team doing them. From your side they're breakage, and the ones that hurt most are the ones that don't cause an error. A renamed column throws an exception you'll notice. A field that quietly starts arriving null produces numbers that are wrong and look fine.
The mitigation is unglamorous: contract-style checks at ingestion that validate shape and expected ranges, and a relationship with the teams upstream so you hear about changes before they land. The second one is organizational and does more good than any amount of defensive engineering.
Failures are silent by default
The nightmare scenario is not a pipeline that crashes. A crash gets attention. The nightmare is one that succeeds while doing nothing useful.
A source returns an empty result set because of a permissions change, and your load runs successfully with zero rows. A partial extract completes and looks like a normal day with lower volume. A transformation drops records that don't match an expected pattern, and the count in the log looks reasonable because you never established what reasonable is.
Monitoring that only asks "did the job finish" catches almost none of this. What catches it is monitoring the data: row counts against historical ranges, freshness checks that assert the latest record is recent enough, null-rate tracking on fields that shouldn't have nulls, and totals that should reconcile against a known source.
If you build one operational thing after launch, build this.
Data arrives late, and sometimes arrives twice
The clean mental model is that yesterday's data lands overnight and never changes. Reality is messier.
Transactions get backdated. A correction posted this week applies to last quarter. A source system had an outage and delivered three days at once. A record you already processed shows up again with different values because someone edited it.
Every one of these means the numbers for a period you already reported can change after the fact. Which is fine, as long as your architecture expects it. It's not fine if you build assuming append-only immutability, because then reprocessing a closed period means either a manual intervention or an incorrect number nobody will notice until an auditor does.
The related question is one somebody will eventually ask in a heated meeting: why does the report I ran on Tuesday show a different figure than the one I ran today? You need a good answer, and ideally a way to reproduce what a report said on a given date.
Backfills take longer than the original load
At some point you'll need to reprocess history. New field, corrected logic, a source that finally provided the archive you asked for eight months ago.
Backfills are their own category of pain. They compete with production loads for resources, they're slow enough that they span multiple days, they fail partway through and need to resume rather than restart, and while they're running the affected tables are in an inconsistent state that somebody will inevitably query.
Teams that plan for this build reprocessing as a first-class capability rather than something improvised under pressure. Teams that don't spend a weekend writing one-off scripts and then delete them, guaranteeing the next backfill is equally painful.
Definitions drift and nobody updates the documentation
The warehouse launched with agreed definitions. Then the business changed.
A new product line doesn't fit the existing category logic. A reorg means the old territory mapping is wrong. Somebody in finance revises how a metric is calculated, implements it in their own report, and doesn't mention it. Now the official number and the finance number diverge, and the credibility you spent a year building starts eroding.
The technical fix is centralizing metric logic so there's one implementation. The organizational fix is harder and more important: a named owner for each significant definition, and a process where a change to any of them is a decision rather than an edit.
Documentation, incidentally, is always out of date. The realistic goal isn't perfect documentation, it's making the actual logic readable enough that someone can determine the truth from the system itself.
Costs grow in ways nobody modeled
Cloud warehouse pricing rewards efficient queries and punishes casual ones, and after launch you have a population of users writing casual ones.
The classics: a dashboard set to auto-refresh every five minutes that nobody looks at, scanning the full history each time. An analyst's exploratory query that accidentally cross-joins. A scheduled job someone built for a one-week analysis and never turned off. Storage accumulates because no retention policy was ever defined and deleting things feels risky.
None of these are individually large. Collectively they produce a bill that gets escalated, and the escalation usually arrives without any per-team visibility into what caused it. Cost attribution and a monthly review of the most expensive recurring queries takes an hour and prevents that conversation.
Ownership evaporates
The most common root cause behind everything above.
The project had a team. After launch that team gets reassigned, and the warehouse enters a state where everyone uses it and nobody is responsible for it. Requests go to whoever answered last time. Breakages get patched by whoever notices. Nobody has allocated hours, so nothing preventive happens.
This is why organizations that treat data warehouse development as a delivery with an end date tend to be rebuilding within three years, while the ones that budget for a steady operational function keep the thing useful. The recurring work isn't large — it's usually a fraction of a role — but it has to belong to someone specific.
What to put in place before you need it
None of this requires exotic tooling. Most of it is decisions made early.
Validate incoming data, not just job completion. Alert on freshness and volume anomalies rather than only on exceptions. Design for reprocessing from the start, including partial reruns. Assume records will arrive late and out of order. Centralize metric definitions so there's one place to change them. Attribute compute cost to teams so the feedback loop exists. Name an owner, with hours, before the project team disperses.
And set the expectation with whoever funded it: this isn't a system that gets built and then works. It's a system that gets built and then gets maintained, and the maintenance is what determines whether anyone still trusts the numbers in year three.
The flat chart nobody questioned for eleven weeks is not a story about a bug. It's a story about what happens when a thing is finished but not owned.
Top comments (0)