A team picks a read replica to fix database latency under load. It works, latency drops, everyone moves on. Eighteen months later, a new engineer joins, looks at the replication setup, and asks why reads are eventually consistent instead of strongly consistent, since that's clearly adding complexity somewhere downstream. Nobody still on the team remembers the original reasoning. The engineer, reasonably, proposes removing it to simplify things. It gets removed. Three weeks later, the database is saturated again during peak hours, and the team is back to solving a problem they already solved once, except this time with no memory of how.
Nothing about this story involves a bad technology choice. The read replica was the right call both times it got considered. The actual failure was that the first decision never got written down anywhere a future engineer could find it, so it looked, eighteen months later, like an arbitrary complexity someone had added for no clear reason. That's the problem architecture, as a discipline, actually exists to solve, and it's a narrower claim than it sounds like at first.
Architecture is a set of decisions, not a diagram
It's tempting to picture "architecture" as the boxes-and-arrows diagram that gets drawn on a whiteboard or dropped into a design doc. The diagram is a byproduct. What architecture actually is: the set of decisions that are expensive to change, the constraints a team deliberately chooses to live with so the system can meet its real requirements under real-world conditions.
Those decisions get shaped by three things pulling against each other. Requirements are what the system actually has to do, both the functional behavior and the non-functional expectations around it. Constraints are the hard limits already in place, budget, team size, a deadline, a compliance requirement, things that aren't negotiable no matter how elegant an alternative looks. Quality attributes describe how the system has to behave once it's running, its latency, its availability, its durability, its cost, how operable it is for the people on call for it. A decision that ignores any one of these three tends to look fine in a design review and then quietly becomes the wrong choice the moment it meets production.
Why undocumented decisions decay into mysteries
Most architectural failures aren't the result of picking the wrong technology. They're the result of a decision that got made implicitly, without anyone writing down the requirement it was satisfying, the alternatives that were considered and rejected, or the specific condition under which it should eventually get revisited.
Once that documentation doesn't exist, every new engineer who touches the system has to re-evaluate the decision from scratch, and they frequently end up re-selecting an option the team already tried and rejected, for reasons that are no longer visible to anyone. Explicit reasoning, written down somewhere, is what lets that shared understanding survive team turnover and the organization simply growing past the people who made the original call. The read replica story at the top is exactly this pattern: the decision was fine, the absence of a record was the actual bug.
The reasoning loop underneath a good decision
Stripped down, the architectural reasoning process is a short loop, and skipping a step in it tends to be where things go wrong later, not in any single step itself.
Start by identifying the requirements: what the system must actually do, and which quality attributes, availability, latency, consistency, cost, matter most for this particular decision. Then identify the constraints, team size, budget, deadline, what already exists, any compliance obligations, the hard limits that rule certain options out before you've even evaluated them. Next, enumerate every viable option before judging any of them, collapsing the option space too early is one of the most common ways teams end up with a worse decision than they needed to make. Evaluate the trade-offs of each option against the criteria that actually came from the requirements and constraints, not against whichever option happens to be trendiest. Decide, and write down why, the decision itself, which alternatives were rejected and why, and a specific trigger condition for when this decision should be revisited. And finally, acknowledge the consequences explicitly, every decision introduces new constraints of its own, and naming them up front is what lets a future team actually plan around them instead of discovering them by accident.
The artifact that captures most of this is usually called an Architecture Decision Record, an ADR, and the single most valuable section in one is almost always the rejected options. That section is specifically what stops a future engineer from re-opening a question the team already closed, for reasons that made sense at the time and are otherwise invisible to anyone who wasn't in the room.
The question experienced engineers actually ask
The best architecture isn't the most sophisticated one available. It's the one that satisfies the current requirements with the least unnecessary complexity, while still being genuinely evolvable once requirements change, which they reliably will.
Principal-level engineers tend to ask one specific question more than any other: "what would change my decision?" That's the mark of conditional reasoning rather than a fixed, context-free preference for one architecture over another. A cache is the right call precisely when some amount of stale data is acceptable, not on principle. A queue is the right call when asynchronous processing is tolerable for that specific workflow. Multi-region deployment is the right call when the cost of downtime genuinely exceeds the cost of the operational complexity multi-region adds, not just because the company is now large enough that it feels like the next step. Architecture isn't really about knowing the correct answer in the abstract. It's about knowing which questions to ask, and which of the constraints in front of you actually matter most for this particular decision.
Watching one system evolve through three stages
A small team launches with the simplest possible setup, one application server, one PostgreSQL database, deploys running straight from someone's laptop. Response times sit comfortably under 50ms. Nothing about this is wrong, it's exactly the right amount of architecture for where the system actually is.
Users grow by a factor of ten. The database's CPU starts saturating during peak hours, and read latency climbs. The team adds a read replica, writes still go to the primary, and reads now come with replication lag, eventual consistency instead of strong consistency. That's a real trade-off accepted deliberately, not a flaw slipping in unnoticed.
Traffic becomes globally distributed. Users in Asia are seeing 200ms-plus latency that a read replica does nothing for, since it's a geography problem, not a database-load problem. The team adds a CDN for static assets and starts seriously considering a multi-region deployment. But multi-region brings real data-consistency challenges, meaningfully more operational complexity, and real cost. The honest question at this stage isn't "should we go multi-region," it's what specific requirement would actually make multi-region necessary right now, and which constraints make it premature at this particular size. Different teams can legitimately land on different answers here, the point isn't that there's one correct stage to make this jump, it's that the decision should be made against an actual requirement, not against a vague sense that the company has gotten big enough that it's time.
The trade-offs that show up in nearly every decision
A handful of tensions recur across almost any architectural decision you'll actually face. Simplicity versus scalability: a monolith is simpler to operate day to day but harder to scale independently piece by piece, while microservices scale more flexibly at the cost of a real multiplication in operational burden. Consistency versus availability: strong consistency requires coordination, which costs latency and can reduce availability during a network partition, while eventual consistency buys higher availability at the cost of application logic that now has to account for staleness. Performance versus cost: more replicas, more caching layers, more regions all improve performance, and each one adds real infrastructure cost and real operational complexity that someone has to carry. Flexibility versus commitment: delaying a decision keeps options open longer, but it can also quietly accumulate technical debt, while committing early enables real optimization at the cost of reduced room to adapt later if the assumption turns out wrong.
None of these trade-offs resolve to a universal right answer. They resolve differently depending on the actual requirements and constraints in front of a given team at a given stage, which is exactly why the reasoning loop above matters more than memorizing which side of each trade-off is "correct."
The failure modes that show up again and again
A specific, recurring handful of failure patterns account for a disproportionate share of real architectural pain. Implicit decisions, a choice made with no documentation, so six months later nobody remembers why it was made, and it gets reversed, reintroducing the exact problem it was originally meant to solve, is the read-replica story from the opening, and it's extremely common. Solution-first framing starts from "we need Kafka" instead of starting from the actual requirement, collapsing the option space before any real evaluation happens, and frequently building the wrong fit as a result. No change trigger means a decision that was genuinely correct back at an earlier stage is still in place at a much later stage, with nobody having recorded when it should be revisited, so the old architecture quietly becomes a constraint on everything built after it. Over-engineering early looks like choosing microservices for a three-person team because "we'll need it eventually," and watching the resulting operational complexity consume engineering capacity that could have gone toward the product itself. Ignoring constraints means picking the technically superior option on paper, one that happens to require expertise the team doesn't actually have, and discovering that the correct choice in theory is the wrong choice in practice. And architecture by committee, requiring full consensus on every decision regardless of how reversible it actually is, treats cheap, easily-undone choices with the same ceremony as genuinely irreversible ones, and velocity collapses under the weight of that mismatch.
Not every decision deserves the same amount of rigor
Applying the full reasoning loop, options, trade-offs, a written ADR, to every single decision is its own kind of failure, the architecture-by-committee pattern above. The actual skill is matching the rigor to how reversible the decision is.
Full rigor earns its cost when choosing a primary datastore or data model, defining a public API contract, committing to a platform-level technology, making a call that affects multiple teams at once, or any decision where the cost of changing course later is genuinely high. Moving fast is the right call when choosing an internal library or a local caching strategy, deciding the shape of an internal API within a single service, making something that's reversible within hours or days if it turns out wrong, when the blast radius stays bounded to one team, or when the choice can just be validated with a quick spike instead of a formal process. The mismatch in either direction, heavy process on a trivial, reversible choice, or an implicit, undocumented call on something genuinely hard to undo, is where most of the pain in the failure modes above actually comes from.
The sentence worth remembering
If a startup's monolith is slowing feature delivery and a senior engineer proposes a full microservices rewrite, the most important question isn't which framework to use or how many services to split into. It's what specific constraint the monolith is actually creating right now, since that's the requirement any proposed solution, rewrite included, actually needs to satisfy, and skipping straight to the solution is the exact solution-first framing mistake from above.
The single most useful sentence in any architectural decision, the one worth writing down every time, is "revisit this if...". It's what turns evolution into something intentional, a planned transition triggered by a condition someone actually named in advance, rather than a surprise rework triggered by nobody remembering why the original decision was made in the first place.
Explore It Visually
I put together an interactive walkthrough of this entire reasoning cycle, requirements, constraints, options, trade-offs, the decision and its ADR, and then watching a system evolve through an actual change trigger, on SeeItFlow, if you'd like to see the cycle play out scene by scene rather than read through it linearly.
Top comments (0)