DEV Community

The Database Is a Detail. The Data Is Not.

Clean Architecture was right to push the database to the edge. The mistake was pushing everything the product needs to know about itself along with it.


The diagram filled the entire screen. It showed the new order flow: creation, payment, inventory, cancellation, and refund, each handled by its own service. Queues were named, retries configured, there was a circuit breaker between payment and inventory, and even the authentication flow had been drawn arrow by arrow. It was good work.

In the bottom-right corner, there was a gray box, smaller than the others, labeled "Analytics." A dotted arrow pointed to it.

Nobody mentioned it during the presentation.

Nobody asked, for example, how the business would know why orders were being cancelled. The cancellation service changed the order status to CANCELLED and moved on. The reason lived in a log, if it was recorded at all.

A reconstruction of the scene: everything was detailed except the box where the business would eventually discover why orders were being cancelled.

I have seen this scene two or three times, across different teams, working with very good software engineers. And as a data engineer, it bothers me for a specific reason.

That gray box is where the product finds out whether it is actually working.

It is where conversion metrics come from. It is where the risk model gets its inputs and where the regulatory report gets its numbers. Increasingly, it is also where an AI agent gets the context it needs to make a decision.

The question behind this article is simple: if data is so central to the product, why does it appear as a footnote in the architecture?

What Uncle Bob Said, and Why He Was Right

In Clean Architecture, Robert C. Martin dedicates an entire chapter to a provocative statement: "The Database Is a Detail."

The argument is solid.

Business rules, the entities and use cases at the center of the circles, should not know whether their data lives in Oracle, MySQL, or a file. The database belongs in the outermost ring, alongside frameworks, drivers, and the web interface.

Dependencies always point inward: the database knows about the domain, but the domain does not know about the database.

I agree with this completely.

An architectural decision should not be held hostage by the choice between Postgres and DynamoDB. Anyone who has migrated a system tightly coupled to an ORM knows the cost of ignoring that advice.

But there is a detail in that chapter that is easy to overlook.

Martin distinguishes between two things: the data model, which he considers architecturally significant, and the technology used to store and retrieve that data, which is the detail.

The database is a detail.

The structure of the information is not.

In the order flow from the beginning of this article, the distinction is easy to see. Storing a cancellation in a Postgres table or a MongoDB document is a detail. But the rule that a cancellation has a reason, a refunded amount, and a timestamp, and that this fact matters to the business, belongs to the model.

That was exactly the part missing from the diagram.

The problem starts when this distinction gets lost in practice.

Where the Interpretation Goes Wrong

In practice, "the database is a detail" has often turned into "data is a detail."

They sound similar, but they are not.

The transactional database is where the system stores its state so it can operate. But the product needs to know much more about its own data than its current state:

  • what happened, not just what the current state is. The order was cancelled, but when, by whom, and after how many attempts?
  • reliable history for auditing, regulation, and learning;
  • whether a feature actually produced the expected outcome;
  • the information required by models, recommendation systems, and now agents that make decisions autonomously based on that context.

None of these are infrastructure details. They are business requirements.

The decision that a "payment declined" event needs to include the reason for the decline does not belong to the DBA or the data team. It comes from the business rules.

When architecture pushes all of this into the frameworks & drivers ring, something predictable happens.

The use case does not produce the information the business actually needs.

So someone comes along later and tries to reconstruct it from the final state of the transactional database.

That is how we end up with pipelines reverse-engineering application tables, updated_at columns being treated as sources of truth, and dashboards that "sometimes don't match."

The gray box in the corner is not small because its job is simple.

It is small because its work was postponed.

And postponing it makes it more expensive.

Data Engineering Has Already Become Software Engineering

Part of the problem is perception.

Many software engineers still picture data engineering as it existed fifteen years ago.

In 2004, Google published the MapReduce paper. Hadoop followed, with HDFS and batch jobs running overnight. Back then, "data" really was a separate world: scripts, nightly loads, specialized tools, and a separate team receiving a production database dump and figuring out what to do with it.

That world has changed.

Modern data engineering uses version control, automated testing, CI/CD, infrastructure as code, schema contracts, observability, and real-time streaming. Tools like dbt brought pull requests and code reviews into data transformation workflows. Kafka, created at LinkedIn and open-sourced in 2011, became a common component in microservice architectures.

And here is the irony: software engineers already use data engineering practices every day.

Topics, queues, events, event sourcing, CDC (change data capture, capturing database changes as a stream of events), and the outbox pattern are part of the same vocabulary.

The boundary between the two disciplines has almost disappeared in the code.

But it is still there in the diagram.

With AI agents, that boundary becomes even harder to defend.

An agent making decisions based on customer history depends on the quality, semantics, and freshness of that data. If the data is bad, the agent does not simply become "less accurate."

It starts making the wrong decisions autonomously and at scale.

So Should the Database Move to the Center?

No.

This is important because it is the easy, and wrong, conclusion.

Moving the database, or the data platform, to the center of the circles would repeat exactly the mistake Uncle Bob was trying to prevent: coupling business rules to technology.

No use case should know whether an event is going to Kafka, Pub/Sub, or a bucket.

Something else needs to move to the center: data requirements.

My proposal is to treat as part of the domain what is usually left implicit today:

  • Domain events as use case outputs. "Order cancelled" is a business fact, with meaning and fields defined by the business. It should be modeled alongside the entity, not discovered later by comparing database snapshots.
  • Data contracts as interfaces. The schema of what a service exposes to the analytical world is a public API. It deserves versioning, review, and compatibility guarantees just like any endpoint.
  • Ports for publishing, adapters at the edge. The use case declares "I need to publish this fact" through an interface, a port in hexagonal architecture terminology. The technology implementing that interface remains a detail.

In other words, the dependency rule still holds.

What changes is that information is intentionally produced by the center instead of being hastily extracted from the edge.

Data requirements, shown in orange, originate at the center. Technology, shown in gray, remains at the edge, and dependencies still point inward.

What This Looks Like in Code

A minimal example makes the idea more concrete.

Below is a use case for cancelling an order. Notice that it knows nothing about Kafka, databases, or warehouses.

But as a business rule, it does know that a cancellation is a fact that must be published, and it knows which information that fact contains.

from dataclasses import dataclass
from datetime import datetime
from typing import Protocol


@dataclass(frozen=True)
class OrderCancelled:
    order_id: str
    customer_id: str
    reason: str             # defined by the business, not the data team
    refunded_amount: int    # in cents
    cancelled_at: datetime


class EventPublisher(Protocol):
    def publish(self, event: OrderCancelled) -> None: ...
Enter fullscreen mode Exit fullscreen mode

The event and the port live at the center.

The use case simply brings them together:

class CancelOrder:
    def __init__(self, orders, events: EventPublisher):
        self.orders = orders
        self.events = events

    def execute(self, order_id: str, reason: str) -> None:
        order = self.orders.get(order_id)
        order.cancel(reason)
        self.orders.save(order)

        self.events.publish(OrderCancelled(
            order_id=order.id,
            customer_id=order.customer_id,
            reason=reason,
            refunded_amount=order.amount_paid,
            cancelled_at=datetime.now(),
        ))
Enter fullscreen mode Exit fullscreen mode

Three things are worth noticing.

OrderCancelled is a domain structure, with a name and fields that anyone working on the product can understand.

The use case depends only on the EventPublisher interface, not on any specific technology.

And the concrete implementation, perhaps a transactional outbox that eventually publishes to Kafka, lives in the outer ring, exactly where it belongs.

The difference from the common scenario is that nobody needs to come along months later and figure out how to infer the cancellation reason from an overwritten status column.

The use case decides what to publish. Outbox, Kafka, and consumers can change without touching the domain.

What Changes in the Architecture Diagram

If I could ask for one change in architecture reviews, it would be this:

The gray box stops being a box and becomes a set of questions.

Before approving the design, the team should answer:

  1. What business facts does this solution produce, and who needs them?
  2. What is the contract for those facts, including fields, meaning, and version, and who owns it?
  3. How will we use data to determine whether this feature actually worked?
  4. What history do we need to preserve for auditing, regulation, or models?
  5. If an agent or model consumes this data, what happens when the data is wrong or delayed?

None of these questions requires choosing a technology.

They require business decisions.

That is why they belong at the center of the conversation, not in a footnote.

Where I Might Be Wrong

Not every system needs this.

An internal CRUD application with ten users does not need versioned event contracts. Requiring them would just create bureaucracy.

There is also a real risk of polluting the domain with analytical requirements that change every week. If every dashboard request turns into a new field on an entity, the solution becomes another problem.

The filter I use is simple:

Would this fact still make sense to the business if no dashboard existed?

"Order cancelled because the item was out of stock" passes the test.

"Helper column for the quarterly report" does not.

There is also an organizational argument I respect. In many companies, the boundary in the architecture diagram reflects the boundary between teams. Changing the diagram without changing ownership does not solve anything.

Approaches such as data mesh and "data as a product" are, in part, attempts to address exactly this problem.

Data Is Not a Detail

Uncle Bob was right about the database.

Oracle, MySQL, and Kafka are details, and the architecture should remain independent of them.

But the data a product generates about itself is not a detail.

It is how the business knows what happened, proves what it did, and decides what to do next, increasingly without a human in the loop.

If you design software architecture, try three things in your next proposal:

  • model domain events alongside entities, as explicit outputs of use cases;
  • treat event schemas as public contracts, with ownership and versioning;
  • bring someone from data into the architecture review before the final design, not after the first broken dashboard.

The gray box in the corner was never small.

It was just drawn in the wrong place.

If you have experienced the scene from the beginning of this article, from either side of the table, share it in the comments. I would like to hear how other teams have approached this problem.

Top comments (0)