DEV Community

Cover image for Analytics Engineer: The Job That Didn't Exist 5 Years Ago (and Why It Does Now)
James Anderson
James Anderson

Posted on

Analytics Engineer: The Job That Didn't Exist 5 Years Ago (and Why It Does Now)

Here's a scene that plays out at basically every company that has ever had more than one dashboard.

Someone in a meeting pulls up the sales dashboard. It says last month's revenue was $2.1M. Someone else pulls up the finance dashboard. Their number says $1.9M. Same company, same month, same word — "revenue" — and two different answers. The room spends twenty minutes trying to figure out which one is right. Nobody can say for sure. The meeting moves on, quietly trusting neither.

If you've worked anywhere near company data, you've lived this. And here's the thing: it's not that someone made a mistake. It's that the way most companies handle data makes this outcome almost inevitable. There's a reason it happens — and over the last few years, a whole discipline (and one very popular tool) grew up specifically to fix it.

Let me walk you through why the dashboards disagree, and what "analytics engineering" actually is — because it's one of the more useful things to understand in tech right now, whether or not you work in data.

Why the two numbers are different

The dashboards disagree because somewhere, two different people wrote two different definitions of "revenue" — and neither of them knew about the other.

Think about how a number gets onto a dashboard. Raw data lands in a database — orders, refunds, discounts, taxes, failed payments, test transactions, canceled subscriptions. To turn that mess into "revenue," someone writes a SQL query that decides a hundred small things: Do we count refunds? Do we subtract discounts? Is tax included? What about an order that was placed but not yet paid? Does a canceled subscription count for the month it was canceled?

Every one of those choices changes the final number. And here's the problem: those choices usually live inside individual queries, scattered everywhere, written by different people at different times. The sales dashboard's query counts refunds one way. The finance dashboard's query counts them another. Both authors were reasonable. Neither wrote down their assumptions. Nobody reviewed either one. And now you have two "revenues."

Multiply that across every metric a company tracks — "active user," "churn," "conversion," "monthly recurring revenue" — and you get the normal state of most companies' data: a pile of undocumented, untested, uncoordinated SQL, where the definition of any given number lives in six different places and nobody's quite sure which is authoritative.

Now here's the part that should make a developer's eye twitch.

Data transformation was the last corner of software with no engineering discipline

Think about how you write application code. You use version control, so every change is tracked and reviewed. You write tests, so you find out when something breaks before your users do. You document things. You reuse code instead of copy-pasting it. You have a review process. These habits are so standard we don't even think about them — they're just how software is built.

Now look at how that revenue query got written. No version control — it lived in a BI tool or someone's saved file. No tests — nobody asserted "revenue should never be negative" or "every order should have exactly one customer." No documentation — the assumptions were in the author's head. No reuse — the next person wrote their own slightly different version from scratch. No review — it went straight to a dashboard the whole company trusts.

For years, this was just how data transformation worked. The entire last step — turning raw data into the numbers people make decisions on — had none of the engineering discipline we take for granted everywhere else in software. It's genuinely strange when you say it out loud: the most decision-critical code in the company was the least engineered.

That gap is exactly what analytics engineering fills.

So what is analytics engineering?

It's a role — and a discipline — that emerged in the last few years to sit in a gap nobody was covering. The easiest way to see it is against the two roles it sits between:

  • Data engineers build the pipelines that move raw data into the warehouse. (Plumbing.)
  • Data analysts take clean data and answer business questions, build dashboards, find insights. (Analysis.)
  • Analytics engineers are the missing middle: they take the raw, messy data that landed in the warehouse and transform it into clean, trustworthy, well-defined datasets that the analysts can actually rely on.

The one-line version: a data analyst spends their time analyzing data; an analytics engineer spends their time transforming, testing, and documenting it — so that when an analyst asks for "revenue," there's exactly one definition, it's correct, and everyone's using it.

And what makes analytics engineering a discipline rather than just a job title is the core idea behind it: bring software engineering practices to data. Version control. Tests. Documentation. Modularity. Code review. The exact habits that data transformation never had.

Which brings us to the tool that basically created the whole movement.

dbt: the tool that brought engineering discipline to data

dbt ("data build tool") is an open-source framework, and it's become the backbone of analytics engineering. What it does sounds almost mundane until you realize nobody was doing it: it lets you write your data transformations as modular, version-controlled, tested, documented SQL.

Here's what that actually looks like in practice — and why each piece fixes the two-dashboards problem:

Transformations become code in Git. Instead of a query hidden in a BI tool, your "revenue" definition is a SQL file in a version-controlled repository. Every change is tracked, reviewed, and attributable. There's one file that defines revenue, and you can see its entire history. The scattered-definitions problem disappears because there's now one place the definition lives.

You write tests on your data. This is the part developers love once they see it. dbt lets you write assertions that run every time your data builds — unique (this ID never repeats), not_null (this field is never empty), accepted_values (status is only ever 'active'/'canceled'/'paused'), relationships (every order points to a real customer). If the data violates an assumption, the build fails — and you find out, instead of a dashboard quietly showing a wrong number to an executive for three weeks.

Models are modular and reusable. Instead of copy-pasting the same logic, you build small models that reference each other. Define "clean orders" once; every downstream metric that needs orders references that one model (in dbt you literally write {{ ref('stg_orders') }} instead of hardcoding a table name). One definition, reused everywhere — so sales and finance are now mathematically guaranteed to be counting the same orders the same way.

Documentation is generated automatically. dbt produces docs — including a visual map of how every dataset derives from every other (data lineage). When someone asks "what does this column mean?" or "where does this number come from?", the answer is in the docs, not in an engineer's memory. Analysts can self-serve instead of filing a ticket.

Changes ship safely through CI/CD. Analytics engineers work in a development environment, test their changes, get them reviewed, and only then promote them to production — the same safe pipeline you'd use for application code. A bad change to the revenue definition gets caught in review, not in a board meeting.

Put it together and the two-dashboards problem is structurally solved: there's one definition of revenue, it's version-controlled, it's tested, it's documented, and both dashboards are built from the same model. They can't disagree, because they're the same source of truth.

Why this is worth understanding (even if you're not in data)

Maybe you're a web developer thinking "cool, but this isn't my world." Here's why it's worth having in your head anyway.

First, it's one of the fastest-growing, most in-demand roles in tech right now, and a lot of developers have vaguely heard "dbt" and "analytics engineer" without being able to explain either. Now you can.

Second — and more interesting — it's a beautiful case study in a pattern that shows up everywhere: a messy, error-prone, tribal-knowledge process gets transformed the moment someone applies basic engineering discipline to it. Version control, tests, documentation, modularity, review. These aren't exotic ideas. They're the boring fundamentals of software engineering. And the entire analytics-engineering movement is, at its core, just those boring fundamentals finally arriving in a place that didn't have them.

That's a useful lens to carry around. Any time you find a critical process that runs on undocumented scripts, copy-paste, no tests, and "ask the person who wrote it" — you've found the same gap dbt filled for data. The fix is almost always the same fundamentals.

The takeaway

Two dashboards disagree on revenue because, for most of software's history, the code that turns raw data into business numbers was written with none of the discipline we demand everywhere else — no version control, no tests, no documentation, no single source of truth. The disagreement wasn't a bug. It was the predictable result of an un-engineered process.

Analytics engineering is the discipline that closes that gap, and dbt is the tool that made it mainstream, by doing something almost unglamorous: treating data transformation like real software. Version-control it. Test it. Document it. Make one definition and reuse it.

The next time two dashboards disagree, you'll know exactly why — and exactly what fixes it. It's not a better chart. It's engineering discipline, finally applied to the last place that didn't have it.


A question for the room: what's the most painful version of "two numbers that should match but don't" you've run into at work? And if you've adopted dbt (or something like it) — what actually changed once you did? I'm curious whether the "single source of truth" promise holds up in practice or just moves the arguments somewhere new.

Top comments (0)