What separation buys us, and what the running system pays for it.
Something I keep noticing in software discussions is how easily we talk about boundaries as if they were inherently beneficial. Separate the responsibilities. Decouple the components. Extract the service. Put a queue between them. Give each domain its own database. The direction is almost always toward more separation, and the separation itself is often presented as evidence that the system is becoming better designed.
But a boundary can change the physical constraints of the problem. Work that once happened inside a process can now require data to move between machines. We introduce latency, communication overhead, and coordination requirements into the same business operation. And depending on where we put that boundary, those changes can be surprisingly expensive.
In my previous post, I wrote about the distinction between physics, engineering, and architecture, and how I think the industry often approaches them in the wrong order. Architecture becomes the starting point, while the underlying constraints become something we discover later, usually when the system is already difficult to change. I think the way we discuss boundaries is one of the clearest examples of that inversion.
Because on a diagram, a boundary is almost free. You draw another box, move a responsibility into it, and connect it to the rest of the system with an arrow. The diagram becomes more organized. The responsibilities become more explicit. Everything looks a little more intentional.
But in the running system, that arrow has to become something.
Sometimes it becomes a function call. Sometimes it becomes a network request. Sometimes it becomes a message that sits in a queue until another process is available to handle it. Those are very different mechanisms, with very different costs, even when the diagrams look almost identical.
And I think we often spend much more time discussing the boxes than understanding what we have placed inside the arrows.
Imagine a fairly ordinary application where a customer places an order. The system checks inventory, calculates the price, processes payment, records the order, and eventually sends a confirmation email. There are business rules involved, there is state to coordinate, and there are external dependencies. It is already an engineering problem before we introduce any particular architectural style.
Now imagine that inventory, pricing, payments, and orders all become separate services.
Maybe that is the correct decision. But the customer still wants to perform the same operation. We have not reduced the amount of agreement the business requires between those responsibilities. We have changed the mechanisms through which that agreement has to happen.
Checking inventory now involves serialization, communication, a timeout, and a decision about what to do when the answer does not arrive. Reserving inventory introduces another problem: the reservation might have succeeded even if the caller never received the response. Retrying the request now requires us to understand whether repeating the operation is safe. Recording the order in one database while updating inventory in another requires us to define what happens when one succeeds and the other fails.
That is how an architectural decision reaches down into the physics of the system. We have introduced communication and data movement that the operation previously avoided. Engineering now has to account for those costs and the uncertainty they introduce. The business operation still needs to make sense. The architecture does not get to negotiate that away.
And that is where I think a lot of the language around decoupling becomes misleading. We separate the code, the processes, the repositories, and the databases, and then talk as if we have separated the underlying responsibilities of the business itself. But if an order requires an inventory reservation, that relationship still exists. If checkout requires a response from the pricing service, that dependency still exists. It has simply taken a different form.
Two components can be physically separated while remaining deeply dependent on each other.
In some cases, that dependency becomes harder to manage precisely because of the separation. What used to be enforced through a local transaction now requires a recovery process. What used to be checked by the compiler now depends on compatibility between independently deployed applications. What used to be a straightforward call stack now requires enough observability to reconstruct what happened across several machines.
These are all solvable problems. But solving them is work. And the work exists because of a decision we made.
I think that distinction matters because the cost of a boundary tends to become invisible once the boundary is accepted as part of the architecture. The retries, the tracing, the message handlers, the reconciliation jobs, the deployment coordination, and the operational procedures become things the system simply “needs.” We stop asking which of those needs came from the business and which came from the way we chose to organize its implementation.
At some point, a considerable portion of the system can exist to maintain the conditions required by the system itself.
That does not make the architecture wrong. But it should make us curious about what those conditions are buying us.
Sometimes they buy something extremely valuable. A workload might need to scale independently. A component might need stronger resource isolation. A team might need to deploy changes without coordinating every release with several other teams. A security requirement might demand a separate execution environment. These are real constraints, and a service boundary can be a very reasonable response to them.
But the benefit should be concrete enough to discuss alongside the cost.
“Inventory is a separate business concept” tells me something about how the code might be organized. It does not, by itself, tell me why checking inventory should require a network request. A concept can have a clear interface, strong encapsulation, and explicit ownership inside the same process. Logical separation and physical separation are different decisions, even though architectural conversations often collapse them into one.
And this is also where I think the argument needs some balance. The machine is not the only thing whose constraints matter.
People have limits too. Teams have limited attention. Understanding a large codebase takes time. Coordinating changes across several groups can become expensive. A system that minimizes every runtime cost while making ordinary development painfully difficult is not automatically a well-engineered system.
Sometimes paying more at runtime is a reasonable way to reduce the cost of changing the software. Sometimes operational complexity is worth accepting because a team needs a degree of independence that the current structure cannot provide. Good engineering includes those trade-offs.
But human benefits deserve the same scrutiny as technical ones. Giving each team a service does not automatically give each team autonomy. If every useful change still requires coordinated updates across those services, we may have preserved the organizational dependency while adding a distributed system around it.
The names of the repositories do not determine how independently people can work.
Queues create a similar kind of confusion. They can be extremely useful when an operation does not need to complete immediately. Sending a confirmation email or updating analytics probably does not need to hold up the customer’s request. Moving that work into the background can be a sensible decision.
But accepting work and completing work are different promises. A queue gives us a place to hold work. We still need enough processing capacity to finish it, a way to understand how far behind it is, and a policy for what happens when processing repeatedly fails. The customer’s experience depends on whether the operation actually finishes within an acceptable amount of time.
A queue changes when the system does the work, but the work still has to be done. Physics still constrains how quickly consumers can process it. Engineering has to decide how much delay the business can tolerate and what capacity is needed to keep up. Calling something asynchronous does not make its consequences disappear.
The more I think about these decisions, the more I feel that the useful question is what a particular boundary allows us to do that we could not do adequately without it. Which constraint does it address? Which failure does it contain? Which coordination problem does it reduce? And what new problems do we accept in exchange?
The answer can change over time. A boundary that would be unnecessary today might become valuable as the workload grows or the organization changes. A boundary that once solved an important problem might eventually become an expensive historical artifact. Architecture has to remain open to that possibility, because the conditions that justified it are rarely permanent.
I think that is part of what makes engineering difficult. Adding a service is visible. Introducing a messaging platform is visible. A diagram with more components can look like progress. Understanding that a particular separation has no useful purpose requires a different kind of work, and removing it often requires more knowledge of the system than introducing it did.
The best decisions connect all three layers. We understand the physical costs, make an engineering judgment about which costs are worth paying, and let the architecture reflect that judgment. The boundary becomes a consequence of understanding the problem, and we can explain both the responsibilities it separates and the cooperation it still requires.
Every boundary has a cost.
Sometimes that cost buys us something we need.
Good engineering is being able to explain what that something is.
Top comments (0)