<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bogdan Nechyporenko</title>
    <description>The latest articles on DEV Community by Bogdan Nechyporenko (@bogdan_nechyporenko).</description>
    <link>https://dev.to/bogdan_nechyporenko</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122424%2Fb6b1c3fd-e6b8-4257-a336-936eb7c09eab.jpg</url>
      <title>DEV Community: Bogdan Nechyporenko</title>
      <link>https://dev.to/bogdan_nechyporenko</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bogdan_nechyporenko"/>
    <language>en</language>
    <item>
      <title>Four Ways to Survive a Network Split</title>
      <dc:creator>Bogdan Nechyporenko</dc:creator>
      <pubDate>Sun, 13 Sep 2026 08:45:10 +0000</pubDate>
      <link>https://dev.to/bogdan_nechyporenko/four-ways-to-survive-a-network-split-olk</link>
      <guid>https://dev.to/bogdan_nechyporenko/four-ways-to-survive-a-network-split-olk</guid>
      <description>&lt;p&gt;&lt;em&gt;What Paxos, VR, Zab, and Raft teach us.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;You may never implement consensus yourself. You still need its mental models to choose databases, coordination services, and infrastructure that fail the way you expect.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;At 2:13 a.m., your monitoring says the service is healthy.&lt;/p&gt;

&lt;p&gt;The API is answering. The database has replicas. Every dashboard is green enough to let you hope the pager was a mistake.&lt;/p&gt;

&lt;p&gt;But two nodes disagree about who the leader is.&lt;/p&gt;

&lt;p&gt;One accepted a write just before the network split. Another is about to accept a conflicting write. Both are alive. Both are behaving rationally with the information they have. And somewhere between them sits a question that a health check cannot answer:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When several machines remember different versions of reality, which one becomes the truth?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is the problem consensus algorithms solve.&lt;/p&gt;

&lt;p&gt;Most developers will never implement one—and probably should not. Yet many of us choose and operate systems built on them: databases, configuration stores, service-discovery platforms, schedulers, control planes, and distributed locks. If you understand the ideas underneath those tools, their tradeoffs stop looking like mysterious product limitations. They become predictable consequences of design.&lt;/p&gt;

&lt;p&gt;This article looks at four protocols: Multi-Paxos, Viewstamped Replication, Zab, and Raft. The goal is not to reproduce their proofs or prepare you to write one over a weekend. It is to learn four ways of thinking about the same hard problem—and then use those mental models when evaluating real infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this is for
&lt;/h2&gt;

&lt;p&gt;This is for backend developers, platform engineers, SREs, architects, and technical leads who choose or operate distributed systems but do not build consensus protocols themselves.&lt;/p&gt;

&lt;p&gt;You should leave with enough intuition to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ask better questions during a database or coordination-tool evaluation.&lt;/li&gt;
&lt;li&gt;Understand why a system may reject traffic while several machines are still running.&lt;/li&gt;
&lt;li&gt;Reason about leader changes, stale reads, write latency, quorum size, and recovery time.&lt;/li&gt;
&lt;li&gt;Recognize which complexity a tool has absorbed—and which complexity it has handed back to you.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frqc5lgmhyr7hgdg6ozib.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frqc5lgmhyr7hgdg6ozib.png" alt="A quorum commits a replicated log entry when majorities overlap" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Suppose three replicas maintain the same state. A client sends &lt;code&gt;SET plan = pro&lt;/code&gt;. The replicas must agree not only that the command happened, but where it belongs in the history. If command number 42 is &lt;code&gt;SET plan = pro&lt;/code&gt; on one replica, command number 42 cannot be &lt;code&gt;DELETE account&lt;/code&gt; on another.&lt;/p&gt;

&lt;p&gt;The usual mechanism is a replicated log. Each entry is an ordered command. Once a quorum—typically a majority—has accepted an entry under the protocol's rules, the cluster can treat it as committed and apply it to a state machine.&lt;/p&gt;

&lt;p&gt;A majority matters because any two majorities overlap. In a five-node cluster, a quorum of three can tolerate two crashed or unreachable nodes. Two separate groups cannot both form a majority, so a network partition cannot safely produce two independent committed histories.&lt;/p&gt;

&lt;p&gt;That safety has a price. If no majority can communicate, a correctly designed cluster may stop accepting writes. The machines are up; the service is unavailable by choice. This is not necessarily a bug. It may be the system refusing to invent two truths.&lt;/p&gt;

&lt;p&gt;One more boundary: the four protocols here are designed primarily for crash faults—nodes fail, restart, or become unreachable—not Byzantine behavior where a participant lies or sends maliciously contradictory messages.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why exactly these four?
&lt;/h2&gt;

&lt;p&gt;These four are useful together because they form a compact tour through the major design ideas behind crash-tolerant replicated state machines.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Protocol&lt;/th&gt;
&lt;th&gt;The design question it highlights&lt;/th&gt;
&lt;th&gt;The lesson worth keeping&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Multi-Paxos&lt;/td&gt;
&lt;td&gt;How can repeated decisions reuse stable leadership?&lt;/td&gt;
&lt;td&gt;Separate safety from the optimizations that make a protocol fast.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Viewstamped Replication (VR)&lt;/td&gt;
&lt;td&gt;How does a primary-backup system change leaders without losing committed work?&lt;/td&gt;
&lt;td&gt;Recovery and view change are part of the protocol, not cleanup after it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zab&lt;/td&gt;
&lt;td&gt;How do we preserve a strict prefix of history for a coordination service?&lt;/td&gt;
&lt;td&gt;A workload-specific ordering contract can shape the entire protocol.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Raft&lt;/td&gt;
&lt;td&gt;Can the same safety goals be expressed as understandable, enforceable rules?&lt;/td&gt;
&lt;td&gt;Comprehensibility is an operational feature.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;They are not the only important consensus algorithms, nor four entries in a leaderboard. They overlap substantially. The value of comparing them is seeing where each one places structure: in ballots, views, epochs, terms, logs, leaders, and recovery rules.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F32i1dair73c3opceykfa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F32i1dair73c3opceykfa.png" alt="Multi-Paxos reuses stable leadership to commit many log entries" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-Paxos
&lt;/h2&gt;

&lt;p&gt;Basic Paxos reaches agreement on one value. Roughly, a proposer first asks acceptors to promise not to accept older proposals, then asks them to accept a value under a numbered ballot. The subtle rules ensure that once a value can be chosen, a later ballot cannot replace it with a conflicting value.&lt;/p&gt;

&lt;p&gt;Doing that full exchange for every log position would be expensive. Multi-Paxos adds the practical move: establish a stable leader for a ballot, then let that leader drive many log positions without repeating the prepare phase every time.&lt;/p&gt;

&lt;p&gt;In steady state, the shape becomes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A leader receives a command.&lt;/li&gt;
&lt;li&gt;It proposes the command for the next log position.&lt;/li&gt;
&lt;li&gt;A quorum of acceptors accepts it.&lt;/li&gt;
&lt;li&gt;The value becomes chosen and can be learned or applied.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The important insight is not “Paxos is complicated.” It is that the protocol's safety core is more general than the leader-based system we usually run. Stable leadership is an optimization for progress and efficiency. When leadership becomes unstable, the system falls back into the harder work of establishing a newer ballot and discovering what may already have been chosen.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where this mental model helps
&lt;/h3&gt;

&lt;p&gt;Multi-Paxos is a good lens for highly optimized or custom replicated services. Implementations can pipeline, batch, use flexible quorum arrangements, and make other choices around the safety core. That flexibility is powerful, but it increases the amount of protocol detail an implementation team must get right.&lt;/p&gt;

&lt;p&gt;When a vendor says its system is “Paxos-based,” the label is only the beginning. Ask what kind of Paxos, how leadership works, how log gaps are repaired, how membership changes are handled, whether reads use quorum or lease mechanisms, and what operators can observe during recovery.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What to remember:&lt;/strong&gt; A general safety foundation can support many optimized implementations, so two “Paxos-based” products may behave very differently in production.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F416ccqbfloekl2n509qy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F416ccqbfloekl2n509qy.png" alt="Viewstamped Replication moves leadership and safe history together" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Viewstamped Replication (VR)
&lt;/h2&gt;

&lt;p&gt;Viewstamped Replication starts from a primary-backup picture. At any moment, the replicas operate in a numbered view, and one replica is the primary for that view.&lt;/p&gt;

&lt;p&gt;During normal operation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The client sends a request to the primary.&lt;/li&gt;
&lt;li&gt;The primary gives it the next operation number and sends a prepare message to backups.&lt;/li&gt;
&lt;li&gt;Once enough replicas acknowledge it, the operation commits.&lt;/li&gt;
&lt;li&gt;Replicas execute committed operations in order, and the primary replies to the client.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The interesting part begins when the primary appears to fail. Replicas move to a higher view, exchange information about their logs, and the new primary constructs a state that preserves committed operations. In the revisited protocol, the primary for a view is selected deterministically from the view number and group membership.&lt;/p&gt;

&lt;p&gt;VR makes an architectural point that is easy to miss: a leader change is a state-transfer protocol. Electing a name is not enough. The new primary must know which history it is allowed to continue.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where this mental model helps
&lt;/h3&gt;

&lt;p&gt;VR is especially useful for understanding primary-backup storage and replicated services. Even when a product does not literally implement VR, its vocabulary—views, operation numbers, commit numbers, normal processing, view change, recovery—gives you a clean way to interrogate failover.&lt;/p&gt;

&lt;p&gt;Ask: What state is transferred before a promoted replica serves writes? Can an out-of-date replica become primary? When does a client retry become a duplicate operation? How does the system distinguish a recovering replica from a participant in the current view?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What to remember:&lt;/strong&gt; Failover is safe only when leadership and history move together.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flu39grpmx8qw2g2lcyak.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flu39grpmx8qw2g2lcyak.png" alt="Zab synchronizes a safe history before broadcasting ordered transactions" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Zab
&lt;/h2&gt;

&lt;p&gt;Zab—ZooKeeper Atomic Broadcast—was built for Apache ZooKeeper's primary-backup architecture. ZooKeeper is not merely storing independent keys. It provides a coordination namespace where the order of changes matters: create a membership node, update configuration, delete a lock contender, trigger watchers.&lt;/p&gt;

&lt;p&gt;Zab therefore emphasizes a totally ordered stream of state changes and divides the protocol into two broad modes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Recovery&lt;/strong&gt;, where a leader is established and replicas synchronize on a valid history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Broadcast&lt;/strong&gt;, where the leader proposes transactions, followers acknowledge them, and committed transactions are delivered in order.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Transactions carry a &lt;code&gt;zxid&lt;/code&gt;, a monotonically ordered identifier containing an epoch-related component and a counter. A prospective leader cannot simply start appending after an election. It must complete the synchronization work that gives the ensemble a safe common prefix.&lt;/p&gt;

&lt;p&gt;That is Zab's central lesson: sometimes the product's data model tells you which consensus property deserves the spotlight. ZooKeeper's hierarchical namespace and coordination semantics make ordered broadcast—not isolated agreement on unrelated values—the natural abstraction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where this mental model helps
&lt;/h3&gt;

&lt;p&gt;You encounter Zab when using ZooKeeper directly or operating platforms that rely on it for metadata and coordination. The practical questions are not “Is Zab better than Raft?” but “Do ZooKeeper's semantics fit this job?”&lt;/p&gt;

&lt;p&gt;ZooKeeper is well suited to small coordination data, configuration, membership, and synchronization primitives. It is not a general replacement for a high-volume application database. You should also examine read semantics separately from write ordering: a protocol can totally order updates while a product still offers local reads with freshness tradeoffs.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What to remember:&lt;/strong&gt; Choose a system whose ordering contract matches the meaning of your data, not one whose algorithm name sounds strongest.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qg32glli9leryy85aml.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qg32glli9leryy85aml.png" alt="Raft elects an up-to-date leader and repairs conflicting follower logs" width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Raft
&lt;/h2&gt;

&lt;p&gt;Raft was designed around understandability. It decomposes replicated-log consensus into leader election, log replication, and safety, then makes the leader-follower relationship deliberately asymmetric.&lt;/p&gt;

&lt;p&gt;In normal operation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A client sends a command to the leader.&lt;/li&gt;
&lt;li&gt;The leader appends it to its log and sends &lt;code&gt;AppendEntries&lt;/code&gt; RPCs to followers.&lt;/li&gt;
&lt;li&gt;After the entry is safely replicated according to Raft's commit rule, the leader applies it and replies.&lt;/li&gt;
&lt;li&gt;Followers learn the commit position and apply the same entries in order.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Terms act as logical eras of leadership. Each log entry records the term in which it was created. During an election, a voter rejects a candidate whose log is less up to date than its own. Once elected, the leader uses the previous log index and term in &lt;code&gt;AppendEntries&lt;/code&gt; to detect divergence, then repairs followers by replacing conflicting uncommitted suffixes.&lt;/p&gt;

&lt;p&gt;The nuance matters: a new leader does not win because it has every byte any replica has ever seen. It wins under rules designed to ensure that committed entries cannot be lost. Uncommitted entries may disappear, which is why a client timeout does not always tell you whether an operation committed.&lt;/p&gt;

&lt;p&gt;Raft's greatest contribution may be sociotechnical: a protocol that engineers can explain, review, test, and debug is less likely to be implemented incorrectly. Understandability is not cosmetic when the code decides which data survives a failure.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where this mental model helps
&lt;/h3&gt;

&lt;p&gt;Raft appears under widely used infrastructure such as etcd and Consul; CockroachDB uses many Raft groups to replicate ranges of data. Knowing the shared foundation helps, but it does not make these products interchangeable. Their data models, transaction layers, read paths, placement, snapshots, reconfiguration, and operational tooling differ enormously.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What to remember:&lt;/strong&gt; An understandable consensus core reduces one category of risk, but the surrounding distributed system still determines the user experience.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The four protocols in one operational picture
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Lens&lt;/th&gt;
&lt;th&gt;Multi-Paxos&lt;/th&gt;
&lt;th&gt;Viewstamped Replication&lt;/th&gt;
&lt;th&gt;Zab&lt;/th&gt;
&lt;th&gt;Raft&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Leadership era&lt;/td&gt;
&lt;td&gt;Ballot&lt;/td&gt;
&lt;td&gt;View&lt;/td&gt;
&lt;td&gt;Epoch&lt;/td&gt;
&lt;td&gt;Term&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Normal-path coordinator&lt;/td&gt;
&lt;td&gt;Stable proposer/leader in practical deployments&lt;/td&gt;
&lt;td&gt;Primary&lt;/td&gt;
&lt;td&gt;Leader&lt;/td&gt;
&lt;td&gt;Leader&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Replicated object&lt;/td&gt;
&lt;td&gt;Sequence of consensus instances, commonly used as a log&lt;/td&gt;
&lt;td&gt;Ordered operation log&lt;/td&gt;
&lt;td&gt;Totally ordered transaction stream&lt;/td&gt;
&lt;td&gt;Ordered log&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recovery emphasis&lt;/td&gt;
&lt;td&gt;Discover and preserve already chosen values&lt;/td&gt;
&lt;td&gt;Build a safe new view from replica state&lt;/td&gt;
&lt;td&gt;Synchronize a safe history before broadcast&lt;/td&gt;
&lt;td&gt;Elect an up-to-date candidate; repair follower suffixes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Design personality&lt;/td&gt;
&lt;td&gt;General and optimization-friendly&lt;/td&gt;
&lt;td&gt;Primary-backup made explicit&lt;/td&gt;
&lt;td&gt;Coordination-workload-specific&lt;/td&gt;
&lt;td&gt;Structured for understandability&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This table is a map, not a benchmark. Performance depends on implementation, batching, storage, network topology, durability settings, read mode, workload, and cluster health. “Raft versus Paxos” is rarely a useful procurement question by itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to apply this knowledge without implementing anything
&lt;/h2&gt;

&lt;p&gt;Imagine you are choosing a distributed database for an order service. Product A advertises Raft. Product B advertises Paxos. The tempting conclusion is that the protocol name settles the decision. It does not.&lt;/p&gt;

&lt;p&gt;Use the algorithms to generate better questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. What is the unit of consensus?
&lt;/h3&gt;

&lt;p&gt;Is there one log for the whole cluster, one group per shard, one group per data range, or a metadata consensus group plus separate data replication? This determines where contention appears and how failure domains interact.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Which operations actually pass through consensus?
&lt;/h3&gt;

&lt;p&gt;Writes probably do. What about reads? Are they linearizable, lease-based, quorum-based, or served locally and potentially stale? Does a transaction spanning several consensus groups require another coordination layer?&lt;/p&gt;

&lt;h3&gt;
  
  
  3. What happens without a quorum?
&lt;/h3&gt;

&lt;p&gt;Will the system stop writes, serve stale reads, fail over elsewhere, or expose a tunable consistency mode? A safe refusal can be preferable to silent divergence, but your product must be designed for that refusal.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. What does “acknowledged” mean?
&lt;/h3&gt;

&lt;p&gt;Was the operation accepted by memory, written to durable storage, replicated to a majority, committed, or applied to the state machine? Those are different milestones. Ask what survives a leader crash immediately after the client receives success.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. How is leadership moved?
&lt;/h3&gt;

&lt;p&gt;Look for election timeouts, planned leadership transfer, fencing, catch-up requirements, and behavior under asymmetric packet loss. Failover time is often a distribution, not a single number.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. How does a slow replica recover?
&lt;/h3&gt;

&lt;p&gt;Does it replay a log, install a snapshot, fetch a checkpoint, or rebuild from another store? Can recovery saturate the same network and disks serving production traffic?&lt;/p&gt;

&lt;h3&gt;
  
  
  7. How does membership change?
&lt;/h3&gt;

&lt;p&gt;Adding and removing voters is consensus about who participates in consensus. Safe reconfiguration is not equivalent to editing a host list. Understand the product's supported procedure.&lt;/p&gt;

&lt;h3&gt;
  
  
  8. Can operators see protocol state?
&lt;/h3&gt;

&lt;p&gt;You want metrics for leader changes, term/view/epoch changes, commit lag, proposal latency, quorum health, snapshot transfer, and rejected requests. A correct protocol hidden behind poor observability can still create a terrible incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  What these algorithms do not solve for you
&lt;/h2&gt;

&lt;p&gt;Consensus is a foundation, not a complete distributed system.&lt;/p&gt;

&lt;p&gt;It does not automatically give you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multi-key or cross-shard transactions.&lt;/li&gt;
&lt;li&gt;Exactly-once side effects.&lt;/li&gt;
&lt;li&gt;A correct retry strategy for ambiguous client timeouts.&lt;/li&gt;
&lt;li&gt;Good geographic latency.&lt;/li&gt;
&lt;li&gt;Elastic scaling.&lt;/li&gt;
&lt;li&gt;Protection against malicious replicas.&lt;/li&gt;
&lt;li&gt;Correct application invariants.&lt;/li&gt;
&lt;li&gt;Painless upgrades and disaster recovery.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, if a payment request times out during a leader change, consensus may preserve the command perfectly while the client cannot tell whether it committed. The application still needs an idempotency key and a safe retry contract.&lt;/p&gt;

&lt;p&gt;Likewise, consensus can order &lt;code&gt;reserve item&lt;/code&gt; and &lt;code&gt;charge card&lt;/code&gt; without guaranteeing that the business workflow across two external systems is atomic. That is an application-level problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small real-life decision exercise
&lt;/h2&gt;

&lt;p&gt;Suppose you need three capabilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Service discovery and a small amount of strongly consistent configuration.&lt;/li&gt;
&lt;li&gt;A globally scaled transactional database.&lt;/li&gt;
&lt;li&gt;A custom replicated control plane.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The protocol perspective changes how you evaluate them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;For coordination, you inspect ZooKeeper, etcd, or Consul semantics: watch behavior, session or lease models, read consistency, quorum loss, and operational maturity. Zab versus Raft is context, not the final score.&lt;/li&gt;
&lt;li&gt;For the database, you ask how many consensus groups exist, how transactions cross them, where leaders are placed, and what geographic topology does to commit latency.&lt;/li&gt;
&lt;li&gt;For the custom control plane, you should strongly prefer a mature library or service over implementing a paper. If custom behavior is truly necessary, Multi-Paxos shows the optimization freedom available; Raft shows the value of constrained, reviewable state transitions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In all three cases, knowing consensus helps you see the system you are actually buying.&lt;/p&gt;

&lt;h2&gt;
  
  
  The deeper lesson: failure behavior is part of the API
&lt;/h2&gt;

&lt;p&gt;We often evaluate infrastructure by its happy-path interface: SQL syntax, key-value operations, SDK quality, or throughput in a benchmark. Distributed systems reveal their real contract when messages are delayed, leaders restart, disks stall, and clients retry.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multi-Paxos teaches us to distinguish a safety core from its performance optimizations.&lt;/li&gt;
&lt;li&gt;Viewstamped Replication teaches us that a new leader must inherit a safe history, not merely a title.&lt;/li&gt;
&lt;li&gt;Zab teaches us that ordering should match the workload's meaning.&lt;/li&gt;
&lt;li&gt;Raft teaches us that understandability can improve implementation and operations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You do not need to implement any of them to benefit. You need to recognize the questions they force every distributed tool to answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The questions worth asking:&lt;/strong&gt; Who may speak for the cluster? Which history survives? What can progress without a majority? And how will we know what happened after the network heals?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The next time a product page says “strongly consistent” or “powered by Raft,” do not stop at the label. Ask for the failure story.&lt;/p&gt;

&lt;p&gt;Because at 2:13 a.m., the failure story is the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://lamport.azurewebsites.net/pubs/paxos-simple.pdf" rel="noopener noreferrer"&gt;Leslie Lamport, &lt;em&gt;Paxos Made Simple&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dspace.mit.edu/handle/1721.1/71763" rel="noopener noreferrer"&gt;Barbara Liskov and James Cowling, &lt;em&gt;Viewstamped Replication Revisited&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://classpages.cselabs.umn.edu/Spring-2021/csci8980/papers/zab.pdf" rel="noopener noreferrer"&gt;Flavio P. Junqueira, Benjamin C. Reed, and Marco Serafini, &lt;em&gt;Zab: High-performance broadcast for primary-backup systems&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://raft.github.io/raft.pdf" rel="noopener noreferrer"&gt;Diego Ongaro and John Ousterhout, &lt;em&gt;In Search of an Understandable Consensus Algorithm&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://etcd.io/docs/v3.3/faq/" rel="noopener noreferrer"&gt;etcd FAQ: Raft and cluster behavior&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.hashicorp.com/consul/docs/architecture/consensus" rel="noopener noreferrer"&gt;Consul documentation: Consensus protocol&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>architecture</category>
      <category>database</category>
      <category>backend</category>
      <category>programming</category>
    </item>
    <item>
      <title>Why we moved our Backstage platform from Yarn to pnpm</title>
      <dc:creator>Bogdan Nechyporenko</dc:creator>
      <pubDate>Sat, 12 Sep 2026 19:44:23 +0000</pubDate>
      <link>https://dev.to/bogdan_nechyporenko/why-we-moved-our-backstage-platform-from-yarn-to-pnpm-21ap</link>
      <guid>https://dev.to/bogdan_nechyporenko/why-we-moved-our-backstage-platform-from-yarn-to-pnpm-21ap</guid>
      <description>&lt;p&gt;&lt;em&gt;What Git worktrees, parallel coding agents, and a cache pointing at the wrong directory taught us about the assumptions hiding inside &lt;code&gt;node_modules&lt;/code&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjgz61flcj32iv3dqwxi6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjgz61flcj32iv3dqwxi6.png" alt="Parallel coding agents sharing packages through the pnpm store" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I’m a developer at bol, where we have been building and running our internal developer platform on Backstage for more than five years.&lt;/p&gt;

&lt;p&gt;After that much time, a Backstage platform becomes more than a collection of plugins. It accumulates its own ways of working: CI conventions, local startup tooling, container images, Git worktree helpers, and shortcuts that make sense inside one engineering organization.&lt;/p&gt;

&lt;p&gt;One of our most productive choices has been using Git worktrees for parallel development with coding agents. An agent can fix a bug in one worktree while another handles a small improvement elsewhere, without branches fighting over the same working directory. When several small tasks are moving at once, this can multiply the amount of useful work we finish.&lt;/p&gt;

&lt;p&gt;That workflow also made our package manager part of the architecture. We wanted a more forward-looking setup, and pnpm looked like a better fit: a shared content-addressable store, efficient reuse across checkouts, and a dependency layout managed by the package manager itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Our goal was simple:&lt;/strong&gt; let independent worktrees move in parallel, while package content is reused safely from one shared store.&lt;/p&gt;

&lt;p&gt;So we moved from Yarn 4 to pnpm. I expected the lockfile and commands to be the main work. They were only the entrance to the migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The migration looked finished until CI ran
&lt;/h2&gt;

&lt;p&gt;The visible switch was straightforward: pin pnpm 12.3.0, commit &lt;code&gt;pnpm-lock.yaml&lt;/code&gt;, replace the Yarn commands, and update the contributor guides. Then CI ran and downloaded 4,226 packages from scratch.&lt;/p&gt;

&lt;p&gt;The next run did the same. The log read &lt;code&gt;reused: 0, downloaded: 4226&lt;/code&gt;. The download itself took about 18 seconds, but the native &lt;code&gt;node-gyp&lt;/code&gt; builds added several minutes on top. We had changed package managers, and the pipeline was behaving as if no cache existed.&lt;/p&gt;

&lt;p&gt;It turned out that the cache did exist. It was looking in the wrong place.&lt;/p&gt;

&lt;h3&gt;
  
  
  The store had two addresses
&lt;/h3&gt;

&lt;p&gt;Our container image carried a global pnpm configuration that sent the store to &lt;code&gt;/builds/.pnpm-store&lt;/code&gt;. GitLab, meanwhile, was trying to archive &lt;code&gt;.pnpm-store&lt;/code&gt; inside the project checkout. pnpm wrote outside that directory, the archiver found no matching files, and every job that followed started cold.&lt;/p&gt;

&lt;p&gt;The fix was small: set the location explicitly inside CI, and print the effective value so the job log can prove it.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;.gitlab-ci.yml&lt;/code&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;variables&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;npm_config_store_dir&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;.pnpm-store'&lt;/span&gt;

&lt;span class="na"&gt;before_script&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;pnpm config set store-dir "$CI_PROJECT_DIR/.pnpm-store"&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;pnpm store path&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That was the first real lesson of this migration: never reason about a package-manager cache from configuration files alone. Ask the running job where its store actually is.&lt;/p&gt;

&lt;p&gt;We also stopped retrying script failures automatically. A dropped connection may deserve a retry. A &lt;code&gt;--frozen-lockfile&lt;/code&gt; mismatch will fail six times for exactly the same reason, and bill you for six runners on the way.&lt;/p&gt;

&lt;h2&gt;
  
  
  When a working cache became the slow part
&lt;/h2&gt;

&lt;p&gt;Once the store was finally cached, the next problem surfaced. The archive now held the pnpm store and every &lt;code&gt;node_modules&lt;/code&gt; tree: 1.42 GB spread across roughly 601,000 files.&lt;/p&gt;

&lt;p&gt;Four jobs pulled that archive. Saving, transferring, and extracting it became a substantial part of the pipeline. Worse, our installs still rebuilt the dependency layout and the native modules anyway, so those cached &lt;code&gt;node_modules&lt;/code&gt; trees were not buying the time we expected.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh3fzmj6r22e30not7j7z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh3fzmj6r22e30not7j7z.png" alt="Archive size and entry count after caching only the pnpm store" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;After keeping the pnpm store and dropping &lt;code&gt;node_modules&lt;/code&gt; from this cache, its size fell by about 72% and its file count by about 59%. Figures are approximate observations taken during the migration.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We removed &lt;code&gt;node_modules&lt;/code&gt;, &lt;code&gt;plugins/*/node_modules&lt;/code&gt;, and &lt;code&gt;packages/*/node_modules&lt;/code&gt; from the cache paths. The archive fell to roughly 400 MB and 244,000 files—about six minutes saved per pipeline, by our estimate.&lt;/p&gt;

&lt;p&gt;The useful lesson was to measure the dependency phase we actually had: restore, install, native builds, save. pnpm’s own CI documentation warns that caching its store is not guaranteed to make installation faster; the right policy depends on your runner, your network, and your workload. [1]&lt;/p&gt;

&lt;h3&gt;
  
  
  The fastest retry is the one you do not schedule
&lt;/h3&gt;

&lt;p&gt;We switched installs to &lt;code&gt;--frozen-lockfile --prefer-offline&lt;/code&gt;, made transfer progress visible in the logs, and chose faster cache compression. Each change removed a little uncertainty. That mattered, because the migration was no longer one problem: it was a chain in which every fix exposed the next bottleneck.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then pnpm met our custom worktree machinery
&lt;/h2&gt;

&lt;p&gt;The hardest surprise was local development. Our worktree bootstrap contained about 350 lines of custom symlink-farm logic, all of it built around the &lt;code&gt;node_modules&lt;/code&gt; structure we had used with Yarn.&lt;/p&gt;

&lt;p&gt;The helper mirrored third-party dependencies from the main checkout and redirected workspace packages to each worktree’s own source. It was clever, fast when its assumptions held, and deeply coupled to a filesystem layout that the package manager was free to change.&lt;/p&gt;

&lt;p&gt;pnpm uses a virtual store and links packages into the dependency graph. [2] In our setup, the old mirroring code met pnpm links where it expected ordinary directories, and we began seeing &lt;code&gt;ENOTDIR&lt;/code&gt; failures. The optimization that made worktrees convenient had become the thing preventing worktrees from working.&lt;/p&gt;

&lt;h3&gt;
  
  
  The best fix deleted the clever part
&lt;/h3&gt;

&lt;p&gt;Instead of teaching our symlink farm every detail of pnpm’s layout, I deleted it. The central helper went from 370 lines to 31, and the whole bootstrap became one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pnpm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--frozen-lockfile&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbuny0857iezvpuioyl0e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbuny0857iezvpuioyl0e.png" alt="Replacing custom worktree symlink logic with pnpm install" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Each checkout gets an installation created by pnpm, while package content can still be reused through the shared store.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This was the point where the migration began to feel successful. We had not recreated the previous mechanism under a new name. We had returned ownership of dependency layout to the package manager.&lt;/p&gt;

&lt;p&gt;There was a tradeoff. Our replacement verifier is much lighter: it checks for &lt;code&gt;node_modules/.pnpm&lt;/code&gt; instead of walking dangling links and proving that every workspace package resolves correctly. Less maintenance code is valuable, but I still want an integration check that edits a plugin in a secondary worktree and confirms the running application uses that source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Existing laptops remembered the old world
&lt;/h2&gt;

&lt;p&gt;A clean checkout worked. Some existing developer checkouts did not.&lt;/p&gt;

&lt;p&gt;They still contained the project-local &lt;code&gt;.pnpm-store&lt;/code&gt; created by our earlier configuration. Once local development moved to the global store, pnpm detected the mismatch and relinked roughly 4,800 packages on repeated starts—about four minutes, for a change that should not have required a dependency install at all.&lt;/p&gt;

&lt;p&gt;We added a one-time migration to startup: detect the obsolete local store, remove it together with the root &lt;code&gt;node_modules&lt;/code&gt; and the dependency hash, then run one clean install. Later starts with an unchanged lockfile take the existing hash-skip path and return almost immediately.&lt;/p&gt;

&lt;p&gt;That distinction matters. pnpm did not install 4,800 packages in zero seconds. Our startup code learned when no install was needed.&lt;/p&gt;

&lt;h3&gt;
  
  
  A migration must include yesterday’s state
&lt;/h3&gt;

&lt;p&gt;This changed how I think about developer-tool migrations. Testing a fresh clone is necessary, but your colleagues do not all have fresh clones. They have old caches, generated files, global configuration, half-finished branches, and worktrees created before the migration.&lt;/p&gt;

&lt;p&gt;The cleanup also had to be environment-specific. CI intentionally kept a project-local store so GitLab could archive it. Local development treated that same directory as stale. The name of a folder is not enough context to decide whether it should be deleted.&lt;/p&gt;

&lt;h3&gt;
  
  
  Slow networks exposed another edge
&lt;/h3&gt;

&lt;p&gt;On slower office and VPN connections, pnpm’s default request concurrency could lead to timeouts. We time a request for a small artifact in our registry; when a successful request takes more than three seconds, bootstrap adds &lt;code&gt;--network-concurrency=1&lt;/code&gt;, five fetch retries, and a longer maximum retry timeout.&lt;/p&gt;

&lt;p&gt;The idea helped, but the implementation taught us something as well: a failed probe never reaches the slow-success branch. A production version of this pattern should decide explicitly what a timeout, an authentication failure, and a slow success each mean.&lt;/p&gt;

&lt;h2&gt;
  
  
  The journey improved more than installation
&lt;/h2&gt;

&lt;p&gt;Once we were looking closely at the pipeline, we found work that was only loosely related to pnpm but still affected the experience of the migration.&lt;/p&gt;

&lt;p&gt;Jest coverage was one example. Its cache had grown to 5.2 GB. We switched to the V8 coverage provider and enabled inline source maps for &lt;code&gt;@swc/jest&lt;/code&gt;, which brought the cache back to no more than 2 GB. [3] A later change removed Jest-cache transfer entirely and revisited test selection, worker limits, and when coverage runs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnlabfmt2mgl40kt4szyd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnlabfmt2mgl40kt4szyd.png" alt="CI improvements beyond the package manager migration" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;These are reported outcomes from the wider CI work. The unit-test change includes test selection, coverage policy, worker, and cache changes, so it is not a pnpm-only benchmark.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Unit-test wall time moved from 17 minutes to roughly 10. That result is useful, but I would not present it as “pnpm made our tests 41% faster.” The amount and the kind of test work changed too.&lt;/p&gt;

&lt;p&gt;We enabled incremental TypeScript compilation and baked the pinned pnpm version into the runtime image. We also had to make our native-build policy explicit: one image gained &lt;code&gt;pkg-config&lt;/code&gt; and the &lt;code&gt;libsecret&lt;/code&gt; development headers, while our pnpm workspace policy later disabled &lt;code&gt;keytar&lt;/code&gt;’s build.&lt;/p&gt;

&lt;p&gt;Installing native prerequisites and disabling a build are different choices. If the application needs the native module at runtime, a green install is not sufficient evidence that the decision is safe.&lt;/p&gt;

&lt;h3&gt;
  
  
  The results I would confidently share
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;The dependency cache shrank from 1.42 GB to about 400 MB, and from roughly 601,000 files to 244,000.&lt;/li&gt;
&lt;li&gt;A worktree bootstrap helper shrank from 370 lines to 31, by replacing custom filesystem logic with &lt;code&gt;pnpm install&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Warm installs in the reported CI run reused 4,226 packages and downloaded none, once the store path was corrected.&lt;/li&gt;
&lt;li&gt;Unchanged local startup can skip dependency installation entirely, after the stale state is cleaned once.&lt;/li&gt;
&lt;li&gt;Jest cache size and unit-test duration both came down—though several of those causes sit outside the package-manager switch.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frwgizayn14tdh90e76gg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frwgizayn14tdh90e76gg.png" alt="Summary of the Backstage migration results" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell the next team
&lt;/h2&gt;

&lt;p&gt;Moving a mature Backstage monorepo to pnpm is not mainly a search-and-replace exercise. The package manager sits underneath a web of assumptions about directories, caches, startup order, native builds, and developer habits.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Start with the workflow you want to enable. For us, that was dependable parallel development with Git worktrees and coding agents.&lt;/li&gt;
&lt;li&gt;Map every place that knows about dependencies: CI jobs, images, bootstrap scripts, worktree tools, docs, generators, and assistant instructions.&lt;/li&gt;
&lt;li&gt;Print the effective store path inside the real job image. Run the same lockfile twice and capture restore, install, native-build, and save time separately.&lt;/li&gt;
&lt;li&gt;Benchmark no cache, store-only caching, and your existing strategy. Measure elapsed pipeline time separately from the sum of concurrent job durations.&lt;/li&gt;
&lt;li&gt;Exercise a clean clone and an old checkout. Test stale stores, unchanged restarts, dependency changes, slow networks, and authentication failures.&lt;/li&gt;
&lt;li&gt;Change source code inside a secondary worktree and prove the running Backstage instance uses it. Directory existence alone is a weak health check.&lt;/li&gt;
&lt;li&gt;Prefer deleting compatibility machinery when the new package manager can own the same responsibility.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What happened on our journey
&lt;/h3&gt;

&lt;p&gt;We began with a productivity idea: parallel agents working safely in separate Git worktrees. We chose pnpm because its shared store and installation model fit that direction. Then the migration forced us to confront every place where our platform still depended on the old world.&lt;/p&gt;

&lt;p&gt;We fixed a cache that pointed to the wrong directory. We removed hundreds of thousands of unnecessary files from that cache. We deleted a custom symlink farm. We cleaned stale state from existing laptops, adapted to slow connections, and made native build policy explicit.&lt;/p&gt;

&lt;p&gt;The biggest gain was not one benchmark. Our setup became easier to explain: pnpm manages dependencies, CI caches a deliberate store, and each worktree is an independent place for an engineer or a coding agent to work. That clarity is what lets the productivity benefit survive after the migration project is over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Appendix: the code changes, end to end
&lt;/h2&gt;

&lt;p&gt;Everything above is the story. This is the diff. Below are the concrete changes we made, in the order that mattered, so you can lift them into your own repository without repeating the week we spent finding them. Versions and paths are ours—adjust them to your setup.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do step 5 first.&lt;/strong&gt; If the store lands in the wrong directory, everything else here is wasted: the cache misses on every run, and every number you measure afterward is the wrong number.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  1. &lt;code&gt;package.json&lt;/code&gt;—declare the package manager
&lt;/h3&gt;

&lt;p&gt;Pin the version once, so CI, laptops, and containers cannot disagree about it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"packageManager"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pnpm@12.3.0"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. &lt;code&gt;.npmrc&lt;/code&gt;—use the global store, prefer offline
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-store-dir=.pnpm-store
&lt;/span&gt;&lt;span class="gi"&gt;+prefer-offline=true
+side-effects-cache=true
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Removing &lt;code&gt;store-dir&lt;/code&gt; is the line that matters. A project-local store conflicts with the global one, and it is exactly the leftover that step 11 has to clean up later.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. &lt;code&gt;tsconfig.json&lt;/code&gt;—enable incremental compilation
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;   "compilerOptions": {
&lt;span class="gi"&gt;+    "incremental": true,
+    "tsBuildInfoFile": "build/.tsbuildinfo"
&lt;/span&gt;   }
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add &lt;code&gt;build/.tsbuildinfo&lt;/code&gt; to &lt;code&gt;.gitignore&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Jest—switch the coverage provider to V8
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;jest.config.js&lt;/code&gt; (or the Jest block in &lt;code&gt;package.json&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-"sourceMaps": false
&lt;/span&gt;&lt;span class="gi"&gt;+"sourceMaps": "inline"
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Inline source maps are what V8 needs to map coverage back to your sources. Then change the CI test command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-pnpm test:all --coverage
&lt;/span&gt;&lt;span class="gi"&gt;+pnpm test:all --coverage --coverageProvider=v8
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why: the Babel provider writes duplicate transform-cache entries—one instrumented, one not—which is what took our Jest cache from around 2 GB to over 5 GB. [3]&lt;/p&gt;

&lt;h3&gt;
  
  
  5. GitLab CI—make the store path win
&lt;/h3&gt;

&lt;p&gt;The Docker image’s global pnpm configuration overrides &lt;code&gt;.npmrc&lt;/code&gt;, so configuration alone will not win this argument. An environment variable will. Log the effective path in the same breath, so the job can prove where its store is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;variables&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;npm_config_store_dir&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;.pnpm-store'&lt;/span&gt; &lt;span class="c1"&gt;# env var beats global config&lt;/span&gt;

&lt;span class="na"&gt;before_script&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;corepack enable pnpm&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;pnpm config set store-dir "$CI_PROJECT_DIR/.pnpm-store"&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;pnpm store path&lt;/span&gt; &lt;span class="c1"&gt;# log it, so you can verify it&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  6. GitLab CI—cache the store, not &lt;code&gt;node_modules&lt;/code&gt;
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt; cache:
   paths:
     - '.pnpm-store/'
&lt;span class="gd"&gt;-    - 'node_modules/'
-    - 'plugins/*/node_modules/'
-    - 'packages/*/node_modules/'
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;pnpm re-links from the store on every install, so caching &lt;code&gt;node_modules&lt;/code&gt; adds archive overhead with no install-time benefit. While you are in there, switch cache compression to fast—a large store is network-bound, not CPU-bound:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;variables&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;CACHE_COMPRESSION_LEVEL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;fast'&lt;/span&gt;
  &lt;span class="na"&gt;TRANSFER_METER_FREQUENCY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;2s'&lt;/span&gt; &lt;span class="c1"&gt;# shows transfer times in the job log&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  7. GitLab CI—pass &lt;code&gt;--prefer-offline&lt;/code&gt; in every job
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-pnpm install --frozen-lockfile
&lt;/span&gt;&lt;span class="gi"&gt;+pnpm install --frozen-lockfile --prefer-offline
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In every job that installs—&lt;code&gt;install_deps&lt;/code&gt;, &lt;code&gt;build&lt;/code&gt;, &lt;code&gt;test-and-lint&lt;/code&gt;—not only the install job, so the pipeline still behaves when the cache is cold.&lt;/p&gt;

&lt;h3&gt;
  
  
  8. GitLab CI—stop retrying script failures
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt; retry:
   when:
     - runner_system_failure
     - stuck_or_timeout_failure
     - api_failure
&lt;span class="gd"&gt;-    - script_failure # six retries for one deterministic lockfile failure
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  9. CI image—give &lt;code&gt;keytar&lt;/code&gt; what it needs, or turn its build off
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt; RUN apt-get install -y \
     python3 g++ build-essential \
&lt;span class="gi"&gt;+    pkg-config libsecret-1-dev
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The alternative is adding &lt;code&gt;keytar&lt;/code&gt; to &lt;code&gt;allowBuilds: false&lt;/code&gt;. These are genuinely different decisions: if the application needs the native module at runtime, a green install is not evidence that skipping the build was safe.&lt;/p&gt;

&lt;h3&gt;
  
  
  10. Runtime image—bake Corepack in once
&lt;/h3&gt;

&lt;p&gt;Dockerfile (base image):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;ARG&lt;/span&gt;&lt;span class="s"&gt; PNPM_VERSION=12.3.0&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;corepack &lt;span class="nb"&gt;enable&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; corepack prepare pnpm@&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;PNPM_VERSION&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt; &lt;span class="nt"&gt;--activate&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Dockerfile (application image):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;&lt;span class="gd"&gt;-RUN corepack enable &amp;amp;&amp;amp; corepack prepare pnpm@12.3.0
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Caveat: if your build environment cannot reach a private registry from &lt;code&gt;RUN&lt;/code&gt; layers—Cloud Build, in our case—point this one step at the public npm registry.&lt;/p&gt;

&lt;h3&gt;
  
  
  11. Startup script—clean up the stale local store
&lt;/h3&gt;

&lt;p&gt;If &lt;code&gt;store-dir=.pnpm-store&lt;/code&gt; was ever in your &lt;code&gt;.npmrc&lt;/code&gt;, every existing checkout still has a project-local store. pnpm sees the &lt;code&gt;storeDir&lt;/code&gt; mismatch and re-links every package on every install—roughly four minutes, for us. One guard in the start script fixes it for the whole team, once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;staleLocalStore&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;root&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;.pnpm-store&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;existsSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;staleLocalStore&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;rmSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;staleLocalStore&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;recursive&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;force&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="nf"&gt;rmSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;root&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;node_modules&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;recursive&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;force&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;existsSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;hashFile&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="nf"&gt;rmSync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;hashFile&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep this environment-specific. CI deliberately keeps a project-local store so GitLab can archive it; only local development should treat that directory as stale. The name of a folder is not enough context to decide whether it should be deleted.&lt;/p&gt;

&lt;h3&gt;
  
  
  12. Install scripts—a slow-network guard
&lt;/h3&gt;

&lt;p&gt;Time one small request to the registry, and back off the concurrency only when a successful request is slow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;_pnpm_extra_flags&lt;/span&gt;&lt;span class="o"&gt;=()&lt;/span&gt;
&lt;span class="nv"&gt;_start_ms&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%s%3N&lt;span class="si"&gt;)&lt;/span&gt;
curl &lt;span class="nt"&gt;-fsSo&lt;/span&gt; /dev/null &lt;span class="nt"&gt;--max-time&lt;/span&gt; 5 &lt;span class="se"&gt;\&lt;/span&gt;
  https://registry.npmjs.org/is-odd/-/is-odd-3.0.1.tgz 2&amp;gt;/dev/null &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nv"&gt;_elapsed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%s%3N&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; _start_ms &lt;span class="k"&gt;))&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;((&lt;/span&gt; _elapsed &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; 3000 &lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  _pnpm_extra_flags+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="nt"&gt;--network-concurrency&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nt"&gt;--fetch-retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;5 &lt;span class="se"&gt;\&lt;/span&gt;
                      &lt;span class="nt"&gt;--fetch-retry-maxtimeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;120000&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;true

&lt;/span&gt;pnpm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--frozen-lockfile&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;_pnpm_extra_flags&lt;/span&gt;&lt;span class="p"&gt;[@]+&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;_pnpm_extra_flags&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note the edge in our version: a probe that fails outright falls straight through to the normal path. Decide deliberately what a timeout, an authentication failure, and a slow success should each do.&lt;/p&gt;

&lt;h3&gt;
  
  
  13. Worktrees—do not install in secondary checkouts
&lt;/h3&gt;

&lt;p&gt;If you have worktrees or slot-mode checkouts that share the main &lt;code&gt;node_modules&lt;/code&gt;, do not run &lt;code&gt;pnpm install&lt;/code&gt; in the secondary checkout. In hoisted mode, pnpm 12 hits &lt;code&gt;ENOTDIR&lt;/code&gt; when it tries to create symlinks over the real directories the shared farm has already placed there.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the numbers looked like
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;install_deps&lt;/code&gt; on a new lockfile&lt;/td&gt;
&lt;td&gt;~3.5 min (&lt;code&gt;reused 0&lt;/code&gt;, &lt;code&gt;downloaded 4226&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Seconds (&lt;code&gt;reused 4226&lt;/code&gt;, &lt;code&gt;downloaded 0&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CI cache size&lt;/td&gt;
&lt;td&gt;1.42 GB / 601k files&lt;/td&gt;
&lt;td&gt;~400 MB / 244k files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-pipeline cache overhead (4 jobs)&lt;/td&gt;
&lt;td&gt;+~12 min&lt;/td&gt;
&lt;td&gt;Baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jest cache size&lt;/td&gt;
&lt;td&gt;5.2 GB&lt;/td&gt;
&lt;td&gt;≤ 2 GB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unit-test wall time&lt;/td&gt;
&lt;td&gt;17 min&lt;/td&gt;
&lt;td&gt;~10 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local startup (executor JIT)&lt;/td&gt;
&lt;td&gt;22–26 s silent gap&lt;/td&gt;
&lt;td&gt;~2–3 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local install after a 1-line change&lt;/td&gt;
&lt;td&gt;~4 min (stale store)&lt;/td&gt;
&lt;td&gt;~0 s (hash skip)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CI Corepack setup&lt;/td&gt;
&lt;td&gt;Every application build&lt;/td&gt;
&lt;td&gt;Once per runtime image rebuild&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;These are observations from our pipeline, not a controlled benchmark. Several of them—the test numbers in particular—include changes that have nothing to do with the package manager. Measure your own before and after.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;a href="https://pnpm.io/continuous-integration" rel="noopener noreferrer"&gt;pnpm—Continuous Integration&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pnpm.io/symlinked-node-modules-structure" rel="noopener noreferrer"&gt;pnpm—Symlinked &lt;code&gt;node_modules&lt;/code&gt; structure&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jestjs.io/docs/configuration#coverageprovider-string" rel="noopener noreferrer"&gt;Jest—&lt;code&gt;coverageProvider&lt;/code&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>backstage</category>
      <category>npm</category>
      <category>devops</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
