<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alex Georgiev</title>
    <description>The latest articles on DEV Community by Alex Georgiev (@alexgeorgiev17).</description>
    <link>https://dev.to/alexgeorgiev17</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F374570%2Ffd187417-a674-43c9-951a-81c6fb965471.png</url>
      <title>DEV Community: Alex Georgiev</title>
      <link>https://dev.to/alexgeorgiev17</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alexgeorgiev17"/>
    <language>en</language>
    <item>
      <title>Bun 1.4's bun test --parallel cuts a 4-second suite to about 1 second</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Tue, 06 Oct 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/bun-14s-bun-test-parallel-cuts-a-4-second-suite-to-about-1-second-4481</link>
      <guid>https://dev.to/alexgeorgiev17/bun-14s-bun-test-parallel-cuts-a-4-second-suite-to-about-1-second-4481</guid>
      <description>&lt;p&gt;I had a twenty-file test suite that took just over four seconds to run. I added one flag, &lt;code&gt;--parallel&lt;/code&gt;, and it took one second. Then I tried to push it further and found the exact point where the flag stops helping, which turned out to depend on what kind of work the tests were actually doing, not on how many files there were.&lt;/p&gt;

&lt;p&gt;Bun 1.4 shipped on 20 August 2026 and brought &lt;code&gt;bun test --parallel&lt;/code&gt;, which distributes test files across worker processes instead of running them one after another in a single process. I'm running Bun 1.4.2 on a 4-core Intel Xeon VM (KVM, 4 logical CPUs, no cgroup limit), and I wanted to know what the flag actually buys you, where it tops out, and what the "implies &lt;code&gt;--isolate&lt;/code&gt;" line in the docs means in practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  The headline number
&lt;/h2&gt;

&lt;p&gt;I wrote twenty test files, two tests each, where every test awaits a 100ms &lt;code&gt;setTimeout&lt;/code&gt; before asserting — a stand-in for the API calls and debounces that make a lot of real integration tests slow. Three runs each, fastest and slowest within 15ms of each other:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Wall time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;bun test&lt;/code&gt; (serial)&lt;/td&gt;
&lt;td&gt;4.04s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--parallel=1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4.06s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--parallel=2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2.05s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;--parallel=4&lt;/code&gt; (default, = CPU count)&lt;/td&gt;
&lt;td&gt;1.03s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--parallel=8&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.63s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With no argument, &lt;code&gt;--parallel&lt;/code&gt; defaults to the number of CPU cores, which on this box is 4. That alone took the suite from 4.04s to 1.03s, a 3.9x speedup that lines up almost exactly with four workers splitting twenty files. What surprised me is that &lt;code&gt;--parallel=8&lt;/code&gt; kept improving, down to 0.63s, on a machine with only four logical cores. I'll get to why in a minute, because it's the most important part of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The opposite reading: CPU-bound tests hit a wall at the core count
&lt;/h2&gt;

&lt;p&gt;A test that awaits &lt;code&gt;setTimeout&lt;/code&gt; isn't actually using the CPU while it waits. It's just occupying a slot. So I built a second suite, same shape, twenty files and two tests each, but this time each test runs a real computation: a prime sieve up to 20 million, which takes genuine CPU time with nothing to wait on.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Wall time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;bun test&lt;/code&gt; (serial)&lt;/td&gt;
&lt;td&gt;6.1s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--parallel=1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;6.2s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--parallel=2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3.3s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--parallel=4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.8s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--parallel=6&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.9s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--parallel=8&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1.8s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is the number that argues against the headline. Scaling is close to linear up to four workers, matching the four physical cores, and then it stops. Going to six or eight workers bought nothing, and in a couple of runs was a touch slower than four, presumably from context-switch overhead as processes fight for the same cores. &lt;code&gt;--parallel&lt;/code&gt; schedules OS processes; it doesn't make more CPU appear. For a suite that's actually CPU-bound, the number that matters isn't how many files you have, it's &lt;code&gt;nproc&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Put the two tables together and the real rule is simpler than either one alone: &lt;code&gt;--parallel&lt;/code&gt; can scale past your core count for tests that spend most of their time waiting, and it cannot for tests that spend most of their time computing. Most real suites are a mix of both, so the honest advice is to benchmark your own suite rather than assume either curve.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "implies --isolate" actually isolates
&lt;/h2&gt;

&lt;p&gt;The docs say &lt;code&gt;--parallel&lt;/code&gt; implies &lt;code&gt;--isolate&lt;/code&gt;, and describe it as giving "each file... a fresh global object even when two files land on the same worker." I wanted to know exactly what that resets, because the phrasing leaves open whether it's per-file or something coarser or finer.&lt;/p&gt;

&lt;p&gt;I wrote a module with a plain top-level counter and four test files that import it and log what they see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// shared.ts&lt;/span&gt;
&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;bump&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;count&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;count&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under plain &lt;code&gt;bun test&lt;/code&gt;, the whole run shares one process and one module cache, so the counter climbs across files in whatever order they happen to run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;file-4 sees counter = 1
file-1 sees counter = 2
file-2 sees counter = 3
file-3 sees counter = 4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under &lt;code&gt;--parallel&lt;/code&gt;, every file saw &lt;code&gt;1&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;file-1 sees counter = 1   worker = 1
file-2 sees counter = 1   worker = 1
file-3 sees counter = 1   worker = 1
file-4 sees counter = 1   worker = 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All four ran in the same worker process (same PID), so this isn't about separate OS processes — it's a fresh JS global per file inside one process, exactly as documented. Adding &lt;code&gt;--no-isolate&lt;/code&gt; to &lt;code&gt;--parallel&lt;/code&gt; brought the leak straight back: &lt;code&gt;1, 2, 3, 4&lt;/code&gt; again, same as serial.&lt;/p&gt;

&lt;p&gt;The part the one-line doc summary doesn't spell out: isolation is per file, not per test. Three tests inside a single file, run under &lt;code&gt;--parallel&lt;/code&gt;, saw the counter climb &lt;code&gt;1, 2, 3&lt;/code&gt; with no reset between them. If your tests rely on module-level setup running fresh for every single test, not just every file, &lt;code&gt;--isolate&lt;/code&gt; won't give you that, parallel or not.&lt;/p&gt;

&lt;p&gt;Bun also sets &lt;code&gt;BUN_TEST_WORKER_ID&lt;/code&gt; and &lt;code&gt;JEST_WORKER_ID&lt;/code&gt; to the worker's 1-based index, which I confirmed by printing both in the test above — worth knowing if you're porting a Jest setup that keys a database name or port off that variable.&lt;/p&gt;

&lt;h2&gt;
  
  
  --bail stops starting files, not running ones
&lt;/h2&gt;

&lt;p&gt;The docs describe it precisely: "the coordinator handles &lt;code&gt;--bail&lt;/code&gt; at file granularity: once the failure threshold is reached it starts no new files, but files already running finish." I tested this against twenty files where the first one fails a 300ms-long test.&lt;/p&gt;

&lt;p&gt;Without &lt;code&gt;--bail&lt;/code&gt;, under &lt;code&gt;--parallel&lt;/code&gt;, all twenty files ran regardless of the failure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;19 pass
1 fail
Ran 20 tests across 20 files. [1.53s]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;--bail&lt;/code&gt; added:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bailed out after 1 failure
3 pass
1 fail
Ran 4 tests across 4 files. [320.00ms]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four files ran, not one. That matches four workers already mid-flight when the failure landed; the other sixteen files never started. This is a real time saving in CI (1.53s down to 0.32s here), but it also means &lt;code&gt;--bail&lt;/code&gt; under &lt;code&gt;--parallel&lt;/code&gt; won't tell you about any other failures in those sixteen untouched files, including ones with nothing to do with whatever broke the first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coverage merges correctly, refusals are clean, one bad file doesn't take down the run
&lt;/h2&gt;

&lt;p&gt;Three more things I checked, more briefly:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Coverage merging.&lt;/strong&gt; I split three functions of one module across three test files, each covering one function, and compared &lt;code&gt;--coverage&lt;/code&gt; output. Serial and &lt;code&gt;--parallel --coverage&lt;/code&gt; reported identical numbers — 75% functions, 60% lines, same uncovered line range — so the claimed coverage merge across workers held up exactly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Invalid input is rejected before anything runs.&lt;/strong&gt; &lt;code&gt;--parallel=0&lt;/code&gt;, &lt;code&gt;--parallel=-1&lt;/code&gt;, and &lt;code&gt;--parallel=abc&lt;/code&gt; all produced the same clean message and exit code 1, with nothing executed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;error: --parallel expects a positive integer, received "0"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;A crash in one file doesn't sink the run.&lt;/strong&gt; I had one test file throw at module load time, independent of any test inside it. Under &lt;code&gt;--parallel&lt;/code&gt;, the other three files in that batch still ran and passed, and the crash was reported as a distinct "error" rather than lumped in with assertion failures:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;bun test v1.4.2 (744846f84) 4x PARALLEL

tests-crash/cr2.test.ts:

# Unhandled error between tests
-------------------------------
error: boom at module load time
-------------------------------

 3 pass
 1 fail
 1 error
Ran 4 tests across 4 files. [11.00ms]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;4x PARALLEL&lt;/code&gt; line in the banner is also the easiest way to confirm from a CI log how many workers actually ran, without adding anything to your own test code.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first attempt at a "CPU-bound" suite used &lt;code&gt;while (Date.now() - start &amp;lt; 100) {}&lt;/code&gt; as a stand-in for real work, expecting it to behave like the prime-sieve suite. It didn't: &lt;code&gt;--parallel=8&lt;/code&gt; kept getting faster on a 4-core box well past where real CPU work should have plateaued, the same shape as the IO-bound suite. I spent a while assuming Bun was doing something clever with scheduling before I worked out the actual cause: a spin loop checking the wall clock only needs to be scheduled once after its own deadline to notice time has passed and exit. It behaves like a sleep, not like computation, because the OS can preempt it for as long as it likes without the loop caring, so long as real time has moved on by the time it next gets to check. That's why I rebuilt the CPU-bound suite around an actual prime sieve with no clock check in it at all, which is the version in the table above.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This is a trimmed version of what produced the first table. It needs Bun 1.4 or later.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://bun.sh/install | bash
&lt;span class="nb"&gt;mkdir &lt;/span&gt;bun-parallel-demo &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;bun-parallel-demo

&lt;span class="k"&gt;for &lt;/span&gt;i &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;seq&lt;/span&gt; &lt;span class="nt"&gt;-w&lt;/span&gt; 1 20&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
&lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"t&lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="s2"&gt;.test.ts"&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;
import { test, expect } from "bun:test";
test("io-&lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="sh"&gt;", async () =&amp;gt; {
  await new Promise((r) =&amp;gt; setTimeout(r, 100));
  expect(1 + 1).toBe(2);
});
&lt;/span&gt;&lt;span class="no"&gt;EOF
&lt;/span&gt;&lt;span class="k"&gt;done

&lt;/span&gt;bun &lt;span class="nb"&gt;test&lt;/span&gt;                 &lt;span class="c"&gt;# serial baseline&lt;/span&gt;
bun &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="nt"&gt;--parallel&lt;/span&gt;      &lt;span class="c"&gt;# defaults to your CPU count&lt;/span&gt;
bun &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="nt"&gt;--parallel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;8    &lt;span class="c"&gt;# push past it, since these tests just wait&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Swap the body of the generated test for a real computation (a sieve, a sort, anything with no timer in it) to see the plateau instead of continued scaling, and compare against &lt;code&gt;nproc&lt;/code&gt; on your own machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this
&lt;/h2&gt;

&lt;p&gt;Before turning &lt;code&gt;--parallel&lt;/code&gt; on for an existing suite, work out which kind of suite you actually have. If most of your tests wait on timers, sockets, or a database round trip, raising &lt;code&gt;--parallel&lt;/code&gt; past your core count is free speed and worth trying immediately. If your tests are doing real computation in-process, there is no benefit past &lt;code&gt;nproc&lt;/code&gt;, and you should set &lt;code&gt;--parallel&lt;/code&gt; to that number rather than guessing higher. Either way, check what your tests actually share at module scope before you lean on &lt;code&gt;--isolate&lt;/code&gt; for correctness: it resets state between files, not between tests in the same file, and that gap is exactly where flaky suites tend to live.&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>testing</category>
      <category>performance</category>
      <category>webdev</category>
    </item>
    <item>
      <title>nginx 1.29.6 moves cookie-based session affinity out of nginx Plus</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/nginx-1296-moves-cookie-based-session-affinity-out-of-nginx-plus-26c</link>
      <guid>https://dev.to/alexgeorgiev17/nginx-1296-moves-cookie-based-session-affinity-out-of-nginx-plus-26c</guid>
      <description>&lt;p&gt;For as long as I've configured nginx upstream blocks, cookie-based session affinity has been something you paid for. The free build had &lt;code&gt;ip_hash&lt;/code&gt; and the &lt;code&gt;hash&lt;/code&gt; directive, both of which pin a client to a backend using properties of the connection rather than anything the application controls. Proper sticky sessions, the kind that follow a browser cookie, lived behind an nginx Plus subscription.&lt;/p&gt;

&lt;p&gt;nginx 1.29.6, released in March 2026, moves the &lt;code&gt;sticky&lt;/code&gt; directive into the open-source build. I pulled it, broke my usual &lt;code&gt;ip_hash&lt;/code&gt; setup on purpose, and measured what changes when session affinity stops depending on where a request comes from and starts depending on a cookie.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;Three backend containers, each a plain Python HTTP server that returns its own container hostname so I could tell which one answered. One nginx container in front, reconfigured between runs: default round robin, &lt;code&gt;sticky cookie&lt;/code&gt;, and &lt;code&gt;ip_hash&lt;/code&gt;, all pointed at the same three backends over a Docker bridge network. Every backend was reachable and healthy throughout; nothing here depends on failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The headline number
&lt;/h2&gt;

&lt;p&gt;One client, 30 requests, three configurations:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;upstream config&lt;/th&gt;
&lt;th&gt;backend 1&lt;/th&gt;
&lt;th&gt;backend 2&lt;/th&gt;
&lt;th&gt;backend 3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;default (round robin)&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sticky cookie srv_id&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With the cookie set, every one of 30 requests from the same client landed on the backend nginx picked on request one. Without it, the same client's requests rotated evenly across all three. That's the entire pitch of the feature, and it held up exactly as advertised on the first test.&lt;/p&gt;

&lt;h2&gt;
  
  
  What ip_hash actually does to five different users
&lt;/h2&gt;

&lt;p&gt;The reason I'd been using &lt;code&gt;ip_hash&lt;/code&gt; for years is that it needed no application cooperation: no cookie, no header, just the client's address. The problem only shows up when several distinct users share one address, which is normal behind NAT, a corporate proxy, or a mobile carrier. I simulated that by running five separate clients, each with its own cookie jar, all issuing requests from the same source container, so nginx saw one IP address for all of them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;upstream config&lt;/th&gt;
&lt;th&gt;client 1&lt;/th&gt;
&lt;th&gt;client 2&lt;/th&gt;
&lt;th&gt;client 3&lt;/th&gt;
&lt;th&gt;client 4&lt;/th&gt;
&lt;th&gt;client 5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ip_hash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;backend 1&lt;/td&gt;
&lt;td&gt;backend 1&lt;/td&gt;
&lt;td&gt;backend 1&lt;/td&gt;
&lt;td&gt;backend 1&lt;/td&gt;
&lt;td&gt;backend 1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sticky cookie&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;backend 2&lt;/td&gt;
&lt;td&gt;backend 3&lt;/td&gt;
&lt;td&gt;backend 1&lt;/td&gt;
&lt;td&gt;backend 2&lt;/td&gt;
&lt;td&gt;backend 3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Under &lt;code&gt;ip_hash&lt;/code&gt;, all five clients, despite being genuinely different sessions, landed on the identical backend, every single request, 30 hits out of 30. Under &lt;code&gt;sticky cookie&lt;/code&gt;, the same five clients, from the same shared address, spread themselves across all three backends and each one stayed put once assigned. This is the actual argument for the feature. It isn't really about stickiness at all; &lt;code&gt;ip_hash&lt;/code&gt; was already sticky. It's about stickiness that doesn't quietly merge unrelated users the moment they share a network path.&lt;/p&gt;

&lt;p&gt;I re-ran this twice more to be sure it wasn't an artefact of request order, and both numbers reproduced exactly: five-for-five collapse under &lt;code&gt;ip_hash&lt;/code&gt;, correct separation under &lt;code&gt;sticky cookie&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  It combines with least_conn, which surprised me
&lt;/h2&gt;

&lt;p&gt;I assumed &lt;code&gt;sticky&lt;/code&gt; would be mutually exclusive with another balancing method, the way &lt;code&gt;ip_hash&lt;/code&gt; and &lt;code&gt;least_conn&lt;/code&gt; can't be used together (nginx rejects that combination outright). I was wrong. &lt;code&gt;sticky cookie&lt;/code&gt; layered on top of &lt;code&gt;least_conn&lt;/code&gt; passed &lt;code&gt;nginx -t&lt;/code&gt; cleanly and behaved correctly at runtime: new sessions were assigned using &lt;code&gt;least_conn&lt;/code&gt;'s logic, and once assigned, every client stayed on its backend for the rest of the test, five clients, ten requests each, zero migration. The two aren't the same kind of directive. &lt;code&gt;ip_hash&lt;/code&gt; and &lt;code&gt;least_conn&lt;/code&gt; are both load-balancing &lt;em&gt;methods&lt;/em&gt;, and nginx only allows one. &lt;code&gt;sticky&lt;/code&gt; is a binding layer that sits on top of whichever method you choose for new sessions. Worth knowing if, like me, you assumed otherwise from the &lt;code&gt;ip_hash&lt;/code&gt; precedent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's still gated behind a subscription
&lt;/h2&gt;

&lt;p&gt;The open-sourcing isn't complete. &lt;code&gt;sticky learn&lt;/code&gt;, which watches upstream responses for an application-issued session cookie rather than inventing its own, now works in the free build:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ nginx -t
nginx: configuration file /etc/nginx/nginx.conf test is successful
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But its &lt;code&gt;sync&lt;/code&gt; parameter, which replicates the learned-session table across a cluster of nginx instances, is still commercial-only, and says so plainly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ nginx -t
nginx: [emerg] unknown parameter "sync" in /etc/nginx/nginx.conf:7
nginx: configuration file /etc/nginx/nginx.conf test failed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the boundary is specific: &lt;code&gt;cookie&lt;/code&gt;, &lt;code&gt;route&lt;/code&gt;, plain &lt;code&gt;learn&lt;/code&gt;, and the new &lt;code&gt;drain&lt;/code&gt;/&lt;code&gt;route&lt;/code&gt; server parameters are free as of 1.29.6. Cross-instance session replication is not. If your actual requirement is a cluster of load balancers that agree on where a session lives, this release doesn't get you there by itself.&lt;/p&gt;

&lt;p&gt;And none of it exists before 1.29.6 at all. Running the identical config against 1.29.5:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ nginx -t
nginx: [emerg] unknown directive "sticky" in /etc/nginx/nginx.conf:7
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not a deprecated-but-working directive, not a warning. The whole block fails to parse.&lt;/p&gt;

&lt;h2&gt;
  
  
  Draining a backend without losing bound sessions
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;drain&lt;/code&gt; server parameter is the other half of this release, also previously commercial. The documented behaviour is that a draining server keeps serving clients already bound to it via &lt;code&gt;sticky&lt;/code&gt;, while refusing any brand-new session. I tested both halves. Across 40 fresh, cookie-less requests against two separately configured draining backends, zero landed on the drained server. Against a session cookie established before drain was switched on, eight replayed requests landed on the drained server eight times out of eight, whether the binding was a plain cookie hash or an explicit &lt;code&gt;route&lt;/code&gt; value. The documentation's claim held exactly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs on every response
&lt;/h2&gt;

&lt;p&gt;The sticky cookie isn't set once and forgotten. nginx resends &lt;code&gt;Set-Cookie&lt;/code&gt; on every single response, not just the one that creates the session. Measuring raw response headers with and without it on an otherwise identical request:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;headers, bytes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;plain upstream&lt;/td&gt;
&lt;td&gt;156&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;sticky cookie&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;268&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;112 extra bytes, on every request, forever, for as long as the session exists. That's nothing on a typical API response, but it's not zero, and nobody mentions it because the feature announcement is naturally about the backend behaviour, not the wire cost of the mechanism that provides it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first attempt at testing &lt;code&gt;drain&lt;/code&gt; produced a result that looked like a documentation bug: a session already bound to the draining backend got silently redirected elsewhere, which is the opposite of what nginx's docs promise. I was about to write that up as the most interesting finding in this post.&lt;/p&gt;

&lt;p&gt;The actual cause was my test harness, not nginx. I'd established the session against one nginx container and replayed its cookie against a second, separately started container, using Python's &lt;code&gt;http.cookiejar&lt;/code&gt;, which enforces domain matching the way a real browser would: a cookie set while talking to host A doesn't get sent when the code then talks to host B, even if both containers run identical config. The cookie I thought I was replaying never left the client. Switching to an explicit &lt;code&gt;Cookie:&lt;/code&gt; header, bypassing the jar's domain logic entirely, showed the real behaviour, which matched the documentation.&lt;/p&gt;

&lt;p&gt;A second, unrelated harness bug showed up when I tried to reproduce the same test with curl's &lt;code&gt;-b&lt;/code&gt;/&lt;code&gt;-c&lt;/code&gt; cookie-jar file instead of Python: curl refused to load a saved cookie back for a bare, dotless hostname like &lt;code&gt;nx-sticky&lt;/code&gt;, with the error &lt;code&gt;cookie 'srv_id' dropped, domain '[file]' must not set cookies for 'nx-sticky'&lt;/code&gt;, even though it had just set that exact cookie from a live response moments earlier. That's a curl quirk specific to single-label Docker Compose-style hostnames, not an nginx issue, but it's exactly the kind of thing that will quietly break a local reproduction if you reach for curl's file-based cookie jar rather than an explicit &lt;code&gt;-b "name=value"&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Both mistakes map to the same lesson: a tool designed to behave like a cautious real browser will sometimes be more cautious than your test setup expects, and the resulting "failure" is the harness, not the thing you're testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This reproduces the headline comparison. It needs Docker and nothing else.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker network create stickynet

&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; backend.py &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;PY&lt;/span&gt;&lt;span class="sh"&gt;'
import http.server, socket
class H(http.server.BaseHTTPRequestHandler):
    def do_GET(self):
        self.send_response(200); self.end_headers()
        self.wfile.write(f"{socket.gethostname()}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;".encode())
    def log_message(self, *a): pass
http.server.ThreadingHTTPServer(('0.0.0.0', 8080), H).serve_forever()
&lt;/span&gt;&lt;span class="no"&gt;PY

&lt;/span&gt;&lt;span class="k"&gt;for &lt;/span&gt;n &lt;span class="k"&gt;in &lt;/span&gt;b1 b2 b3&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="nv"&gt;$n&lt;/span&gt; &lt;span class="nt"&gt;--network&lt;/span&gt; stickynet &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/backend.py:/backend.py:ro"&lt;/span&gt; python:3.12-slim python3 /backend.py
&lt;span class="k"&gt;done

&lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; sticky.conf &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;'
events {}
http {
    upstream backend {
        server b1:8080;
        server b2:8080;
        server b3:8080;
        sticky cookie srv_id expires=1h path=/;
    }
    server { listen 80; location / { proxy_pass http://backend; } }
}
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; nx-sticky &lt;span class="nt"&gt;--network&lt;/span&gt; stickynet &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/sticky.conf:/etc/nginx/nginx.conf:ro"&lt;/span&gt; nginx:1.29.8

&lt;span class="c"&gt;# one client, repeated requests, explicit cookie header (not a jar file --&lt;/span&gt;
&lt;span class="c"&gt;# curl's jar file rejects bare Docker hostnames, see below)&lt;/span&gt;
docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;--network&lt;/span&gt; stickynet curlimages/curl:latest sh &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'
cookie=""
for i in $(seq 1 10); do
  curl -s -D /tmp/h.txt -o /tmp/b.txt -b "$cookie" http://nx-sticky/
  cookie=$(grep -oE "srv_id=[a-f0-9]+" /tmp/h.txt | head -1)
  cat /tmp/b.txt
done'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I ran that exact block before putting it in this post; it produced the same backend hostname ten times in a row.&lt;/p&gt;

&lt;p&gt;What I measured here applies directly to any upstream block that uses &lt;code&gt;ip_hash&lt;/code&gt; for a reason that was never really about load distribution. The failure mode is specific: a shared VPN exit, a corporate NAT gateway, or a mobile carrier's address pool puts several real sessions behind one IP, and &lt;code&gt;ip_hash&lt;/code&gt; cannot tell them apart, because it was never given anything to tell them apart with. The fix is a single directive, &lt;code&gt;sticky cookie srv_id expires=1h;&lt;/code&gt;, swapped in where &lt;code&gt;ip_hash;&lt;/code&gt; used to sit, and as of 1.29.6 it compiles into the open-source binary with no licence file involved. Whether it belongs in a given config is a question about that config's own traffic, not one this post can answer from three containers on one machine.&lt;/p&gt;

</description>
      <category>nginx</category>
      <category>devops</category>
      <category>networking</category>
      <category>performance</category>
    </item>
    <item>
      <title>Valkey 9.2's forkless BGSAVE cuts my memory spike from 350MB to 10MB</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Sun, 04 Oct 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/valkey-92s-forkless-bgsave-cuts-my-memory-spike-from-350mb-to-10mb-23km</link>
      <guid>https://dev.to/alexgeorgiev17/valkey-92s-forkless-bgsave-cuts-my-memory-spike-from-350mb-to-10mb-23km</guid>
      <description>&lt;p&gt;Every BGSAVE I have ever watched on a busy Redis or Valkey instance does the same thing: the process forks, and for a few seconds the host's free memory drops like someone pulled a plug. Valkey 9.2.0-rc1, tagged on Docker Hub on 16 September, adds a way to skip the fork entirely. I wanted to know what that actually buys you, not what the release notes say it buys you, so I ran both paths against the same dataset and measured what happened.&lt;/p&gt;

&lt;p&gt;The short version: the memory spike nearly disappeared, and the save got noticeably slower doing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;I pulled &lt;code&gt;valkey/valkey:9.2.0-rc1&lt;/code&gt; and ran two containers from the identical image, differing only in one setting. The first used the default: &lt;code&gt;bgsave-default-method fork&lt;/code&gt;. The second started with &lt;code&gt;--forkless-infrastructure-enabled yes --bgsave-default-method forkless&lt;/code&gt;, which is the only way to turn it on — more on that below.&lt;/p&gt;

&lt;p&gt;Into each I loaded 3,000,000 keys at 300 bytes each, a little over 1.1GB of data (&lt;code&gt;used_memory&lt;/code&gt; reported 1.04GiB). Then, against each container, I ran eight threads hammering &lt;code&gt;SET&lt;/code&gt; on random existing keys as fast as they could go, and from a ninth connection I sent a &lt;code&gt;PING&lt;/code&gt; every 2ms and timed the round trip. Two seconds into each 12-second run I fired &lt;code&gt;BGSAVE&lt;/code&gt;. I repeated this seven times per mode.&lt;/p&gt;

&lt;p&gt;This setup matters for one reason: copy-on-write only costs you anything if pages are actually being written while the fork holds them. A BGSAVE against an idle dataset tells you almost nothing. Mine wasn't idle — the eight writer threads pushed roughly 27,000 operations a second into whichever container was running fork mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  The memory spike, measured twice
&lt;/h2&gt;

&lt;p&gt;I tracked two numbers: the RDB-reported copy-on-write size (&lt;code&gt;rdb_last_cow_size&lt;/code&gt; in &lt;code&gt;INFO persistence&lt;/code&gt;), and the container's own cgroup memory usage sampled every 150ms through the save.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;fork (default)&lt;/th&gt;
&lt;th&gt;forkless&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;rdb_last_cow_size&lt;/code&gt;, mean of 7 runs&lt;/td&gt;
&lt;td&gt;354.7MB&lt;/td&gt;
&lt;td&gt;0MB (always)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cgroup memory delta during save, mean of 4 runs&lt;/td&gt;
&lt;td&gt;354.4MB&lt;/td&gt;
&lt;td&gt;9.9MB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;save wall time (&lt;code&gt;rdb_last_bgsave_time_sec&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;3s, every run&lt;/td&gt;
&lt;td&gt;5-6s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two memory measurements agree with each other, which is the point of taking both: the kernel's own copy-on-write accounting and an independent cgroup memory sample converge on the same number through two unrelated instruments. On a 1.1GB working set under sustained writes, forking cost roughly a third of the dataset's size in extra memory, every single time, with a tight range of 319.6MB to 367.5MB across the seven runs. Forkless never went above 10.0MB.&lt;/p&gt;

&lt;p&gt;That is the headline, and it held up every time I ran it. But it's not the whole story.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the smoother memory curve cost
&lt;/h2&gt;

&lt;p&gt;The forkless save took 5 to 6 seconds against fork's steady 3. That's not a rounding difference — it's roughly 70% longer for the same data, every time I measured it, over seven runs each.&lt;/p&gt;

&lt;p&gt;It also cost write throughput. Across the 12-second window surrounding each save, the eight writer threads delivered a mean of 27,202 operations a second under fork mode and 23,034 under forkless — about 15% less. The background thread doing the serialising in forkless mode isn't free; it competes with the main thread for the same lock and the same CPU core budget, for longer, and the client-facing throughput shows it.&lt;/p&gt;

&lt;p&gt;So the trade is real on both sides: fork spikes memory hard and briefly; forkless keeps memory nearly flat but takes longer and leans on your write throughput the whole time it's doing it. Neither is free. If your host is memory-constrained, forkless is the one you want. If your host has memory to spare and your write path cares more about throughput than about a short memory bump, the fork default is doing less damage than its memory graph suggests.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pause the release notes mention
&lt;/h2&gt;

&lt;p&gt;The one documented caveat I could find for this feature was that if the main thread writes to a key the background serialising thread hasn't reached yet, that client's request briefly stalls while the key gets moved to the front of the queue. I went looking for that stall in my latency samples.&lt;/p&gt;

&lt;p&gt;Across seven forkless runs, the worst single round-trip during the 12-second window ranged from 5.1ms to 14.2ms, with a mean of 7.2ms. Fork's worst case ranged from 11.5ms to 25.7ms, mean 20.1ms — consistently about three times rougher. But forkless wasn't perfectly smooth either: one run spiked to 14.2ms, almost twice its own average worst case, which is consistent with hitting that documented stall on an unlucky key. The feature narrows the tail. It doesn't flatten it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it refuses, and what else still forks
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;forkless-infrastructure-enabled&lt;/code&gt; cannot be turned on with &lt;code&gt;CONFIG SET&lt;/code&gt; against a running server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;valkey-cli config &lt;span class="nb"&gt;set &lt;/span&gt;forkless-infrastructure-enabled &lt;span class="nb"&gt;yes&lt;/span&gt;
&lt;span class="go"&gt;(error) ERR CONFIG SET failed (possibly related to argument
'forkless-infrastructure-enabled') - can't set immutable config
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It has to go on the command line or in the config file at startup, which means adopting this on an existing fleet needs a restart, not a hot config push. Trying to select the method without that flag fails too, with a message that at least tells you why:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;valkey-cli config &lt;span class="nb"&gt;set &lt;/span&gt;bgsave-default-method forkless
&lt;span class="go"&gt;(error) ERR CONFIG SET failed (possibly related to argument
'bgsave-default-method') - 'forkless' can only be selected when the
server was started with 'forkless-infrastructure-enabled yes'
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A second &lt;code&gt;BGSAVE&lt;/code&gt; while one is already running is refused outright, in both modes, with &lt;code&gt;ERR Background save already in progress&lt;/code&gt; — unsurprising, but worth confirming it isn't silently queued.&lt;/p&gt;

&lt;p&gt;The more useful limitation is this: enabling &lt;code&gt;forkless-infrastructure-enabled&lt;/code&gt; only changes how RDB snapshots are taken. I turned on &lt;code&gt;appendonly&lt;/code&gt; on the forkless-enabled instance, which triggers Valkey's automatic initial AOF rewrite, and checked &lt;code&gt;aof_last_cow_size&lt;/code&gt; afterwards. It came back at roughly 10MB, not zero. AOF rewrites still fork, forkless RDB setting or not. If your actual pain is AOF rewrite stalls rather than BGSAVE stalls, this feature does not touch that path at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching it happen
&lt;/h2&gt;

&lt;p&gt;The INFO fields the release notes promised for tracking progress are genuinely there and genuinely populated mid-save, which I didn't expect to hold up as cleanly as it did:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;valkey-cli info persistence | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'save_keys|remaining'&lt;/span&gt;
current_save_keys_processed:555092
current_save_keys_total:3000000
forkless_estimated_seconds_remaining:3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Polled a second and a half after triggering a save on the same 3,000,000-key dataset, that's a believable in-flight number with a plausible estimate attached, not a placeholder. That's one documented claim that checked out exactly as described, which is worth saying plainly since the other numbers in this post are all pushing back on some part of the pitch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first version of the latency probe recorded each sample's offset by subtracting &lt;code&gt;time.time()&lt;/code&gt; (wall clock) from a &lt;code&gt;time.perf_counter()&lt;/code&gt; reading (a separate monotonic clock with an arbitrary zero point). The two aren't comparable, so every timestamp in my first batch of CSVs came out as a nonsensical negative number in the billions. The aggregate percentiles I'd already printed to stdout were unaffected, because those only used the latency values, not the timestamps — but the per-sample files were useless for lining a latency spike up against the exact moment &lt;code&gt;BGSAVE&lt;/code&gt; started. I fixed it by recording one wall-clock epoch at the start of the run and adding monotonic deltas to it, then re-ran the affected trials. The lesson, again: check that your instrumentation's own clock is the one you think it is before trusting a timing result, even one that looks plausible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This is the full loop for one side of the comparison; swap the container flags to switch modes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; vk-forkless &lt;span class="nt"&gt;-p&lt;/span&gt; 16379:6379 valkey/valkey:9.2.0-rc1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--forkless-infrastructure-enabled&lt;/span&gt; &lt;span class="nb"&gt;yes&lt;/span&gt; &lt;span class="nt"&gt;--bgsave-default-method&lt;/span&gt; forkless &lt;span class="nt"&gt;--save&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt;

docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; vk-fork &lt;span class="nt"&gt;-p&lt;/span&gt; 16380:6379 valkey/valkey:9.2.0-rc1 &lt;span class="nt"&gt;--save&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt;

python3 - &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;'
import socket, sys
def load(host, port, n=3_000_000, vsize=300):
    s = socket.create_connection((host, port))
    val = b'x' * vsize
    buf, batch = bytearray(), 2000
    for i in range(n):
        key = f'key:{i}'.encode()
        buf += b'*3&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="nv"&gt;$3&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;SET&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="nv"&gt;$%&lt;/span&gt;&lt;span class="sh"&gt;d&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;%s&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="nv"&gt;$%&lt;/span&gt;&lt;span class="sh"&gt;d&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;%s&lt;/span&gt;&lt;span class="se"&gt;\r\n&lt;/span&gt;&lt;span class="sh"&gt;' % (len(key), key, len(val), val)
        if (i + 1) % batch == 0:
            s.sendall(buf); buf = bytearray()
            got = 0
            while got &amp;lt; batch:
                got += s.recv(65536).count(b'+OK')
load('127.0.0.1', 16379); load('127.0.0.1', 16380)
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;docker &lt;span class="nb"&gt;exec &lt;/span&gt;vk-fork valkey-cli bgsave
docker &lt;span class="nb"&gt;exec &lt;/span&gt;vk-forkless valkey-cli bgsave
docker &lt;span class="nb"&gt;exec &lt;/span&gt;vk-fork valkey-cli info persistence | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'rdb_last_cow_size|rdb_last_bgsave_time_sec'&lt;/span&gt;
docker &lt;span class="nb"&gt;exec &lt;/span&gt;vk-forkless valkey-cli info persistence | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'rdb_last_cow_size|rdb_last_bgsave_time_sec'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Running just this much against an idle dataset gave me an &lt;code&gt;rdb_last_cow_size&lt;/code&gt; of about 12MB, nowhere near the 350MB in the table above. That gap is the whole point: copy-on-write only costs you something when pages are being written while the save runs, so to see the real effect you need the eight-thread write load from the full test, not this stripped-down version.&lt;/p&gt;

&lt;p&gt;If you're running Valkey on a host where memory headroom during BGSAVE has actually bitten you — not hypothetically, but a real OOM or a real eviction storm triggered by a save — this is worth testing against your own dataset shape before 9.2 ships as stable. If your problem has instead been a save that takes too long and steals too much write throughput while it runs, measure that before switching the default, because this release doesn't make BGSAVE cheaper. It moves where the cost shows up.&lt;/p&gt;

</description>
      <category>valkey</category>
      <category>database</category>
      <category>performance</category>
      <category>devops</category>
    </item>
    <item>
      <title>Docker Engine 29.7's overlay networking breaks every Swarm task without IPv6</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Sat, 03 Oct 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/docker-engine-297s-overlay-networking-breaks-every-swarm-task-without-ipv6-72l</link>
      <guid>https://dev.to/alexgeorgiev17/docker-engine-297s-overlay-networking-breaks-every-swarm-task-without-ipv6-72l</guid>
      <description>&lt;p&gt;I asked a Swarm service for twenty replicas on an overlay network and got zero. Not slow, not partially up. Zero, forever, with the exact same command that gives me twenty on the engine version one release behind it.&lt;/p&gt;

&lt;p&gt;I'd gone looking at the Docker Engine 29 release notes for something smaller: a line in the 29.8.0 changelog about spreading out the daemon's periodic overlay-network gossip so it doesn't burn CPU in bursts. That seemed like a nice, modest thing to measure. While setting up a test rig for it, before I'd measured anything about gossip at all, every overlay-attached service I created came up with 0 running tasks instead of the number I'd asked for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;I had Docker Engine 29.6.2 already running as this machine's system install. I downloaded the static builds for 29.7.0, 29.7.2 and 29.8.2 from &lt;code&gt;download.docker.com&lt;/code&gt; and ran each as its own &lt;code&gt;dockerd&lt;/code&gt;, pointed at its own data directory and its own Unix socket, so I could run two or more versions side by side on one host and compare them directly rather than trusting memory of "what it used to do."&lt;/p&gt;

&lt;p&gt;Each one got &lt;code&gt;docker swarm init&lt;/code&gt; on a single node, an overlay network, and a service with a handful of replicas running &lt;code&gt;sleep 600&lt;/code&gt;, nothing fancier.&lt;/p&gt;

&lt;p&gt;On 29.6.2:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;docker service create &lt;span class="nt"&gt;--detach&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; gossip-svc &lt;span class="nt"&gt;--replicas&lt;/span&gt; 20 &lt;span class="nt"&gt;--network&lt;/span&gt; gossip-net alpine:3.20 &lt;span class="nb"&gt;sleep &lt;/span&gt;600
&lt;span class="nv"&gt;$ &lt;/span&gt;docker service &lt;span class="nb"&gt;ls
&lt;/span&gt;ID             NAME             MODE         REPLICAS   IMAGE         PORTS
e8v48kmvwvya   gossip-svc-old   replicated   20/20      alpine:3.20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Twenty requested, twenty running. I exec'd into two of the containers and pinged across the overlay network to be sure it wasn't just lying about the count:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;docker &lt;span class="nb"&gt;exec &lt;/span&gt;17a672fb04a7 sh &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"ping -c3 -W2 10.0.1.16"&lt;/span&gt;
PING 10.0.1.16 &lt;span class="o"&gt;(&lt;/span&gt;10.0.1.16&lt;span class="o"&gt;)&lt;/span&gt;: 56 data bytes
64 bytes from 10.0.1.16: &lt;span class="nb"&gt;seq&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0 &lt;span class="nv"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;64 &lt;span class="nb"&gt;time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1.051 ms
64 bytes from 10.0.1.16: &lt;span class="nb"&gt;seq&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1 &lt;span class="nv"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;64 &lt;span class="nb"&gt;time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.066 ms
64 bytes from 10.0.1.16: &lt;span class="nb"&gt;seq&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2 &lt;span class="nv"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;64 &lt;span class="nb"&gt;time&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.075 ms

&lt;span class="nt"&gt;---&lt;/span&gt; 10.0.1.16 ping statistics &lt;span class="nt"&gt;---&lt;/span&gt;
3 packets transmitted, 3 packets received, 0% packet loss
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Real traffic, real VXLAN tunnel, both containers reachable. Good baseline.&lt;/p&gt;

&lt;p&gt;On 29.8.2, the identical sequence of commands, same host, same image:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;docker service create &lt;span class="nt"&gt;--detach&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; gossip-svc3 &lt;span class="nt"&gt;--replicas&lt;/span&gt; 5 &lt;span class="nt"&gt;--network&lt;/span&gt; gossip-net3 alpine:3.20 &lt;span class="nb"&gt;sleep &lt;/span&gt;600
&lt;span class="nv"&gt;$ &lt;/span&gt;docker service &lt;span class="nb"&gt;ls
&lt;/span&gt;ID            NAME          MODE         REPLICAS   IMAGE         PORTS
iaki8w2lsuwr  gossip-svc3   replicated   0/5        alpine:3.20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;docker service ps&lt;/code&gt; named the reason:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;network sandbox join failed: subnet sandbox join failed for "10.0.1.0/24":
overlay: cannot determine address family of transport: the local data-plane
address is not currently known
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every one of the five tasks hit this. Not one came up.&lt;/p&gt;

&lt;h2&gt;
  
  
  It isn't a fluke or a port clash
&lt;/h2&gt;

&lt;p&gt;My first instinct was that I'd done something wrong with the test harness itself — I'd given the two engines different VXLAN data-path ports to avoid a bind conflict, so I reran 29.8.2 on the default port 4789 with the older engine's overlay network torn down first, to rule that out. Same error, same 0/5.&lt;/p&gt;

&lt;p&gt;Then I widened the test backwards through the point releases to find where it actually started, since I only had 29.6.2 and 29.8.2 at first:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Engine&lt;/th&gt;
&lt;th&gt;Replicas requested&lt;/th&gt;
&lt;th&gt;Replicas running&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;29.6.2&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;overlay traffic confirmed with &lt;code&gt;ping&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;29.7.0&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;identical "address family" error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;29.7.2&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;identical "address family" error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;29.8.0–29.8.2&lt;/td&gt;
&lt;td&gt;5 and 20&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;identical "address family" error&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It's not version-specific to 29.8. It was already broken in 29.7.0, the release immediately after the one that works, and it's still broken in 29.8.2, the newest static build on &lt;code&gt;download.docker.com&lt;/code&gt; as I write this. I didn't bisect down to a single commit — I don't have the moby source checked out here — but the official 29.7.0 and 29.8.0 release notes both list several Swarm networking changes in that stretch, including one about published ports on the service mesh sharing infrastructure with locally published ports, which touches exactly the code path that's failing here.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it isn't
&lt;/h2&gt;

&lt;p&gt;Plain containers on this same 29.8.2 install work fine. A &lt;code&gt;docker run&lt;/code&gt; with a published port, no Swarm involved:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; 18080:80 &lt;span class="nt"&gt;--name&lt;/span&gt; plainweb nginx:alpine
&lt;span class="nv"&gt;$ &lt;/span&gt;docker &lt;span class="nb"&gt;exec &lt;/span&gt;plainweb wget &lt;span class="nt"&gt;-qO-&lt;/span&gt; http://127.0.0.1:80 | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-3&lt;/span&gt;
&amp;lt;&lt;span class="o"&gt;!&lt;/span&gt;DOCTYPE html&amp;gt;
&amp;lt;html&amp;gt;
&amp;lt;&lt;span class="nb"&gt;head&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That rules out something broad like the bridge driver or the whole networking stack being down. This is specific to the Swarm overlay driver's VXLAN data plane.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one thing I could find that correlates
&lt;/h2&gt;

&lt;p&gt;This host has no IPv6 stack at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/sys/net/ipv6/conf/all/disable_ipv6
&lt;span class="go"&gt;cat: /proc/sys/net/ipv6/conf/all/disable_ipv6: No such file or directory
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/net/if_inet6
&lt;span class="go"&gt;cat: /proc/net/if_inet6: No such file or directory
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not disabled — absent. No IPv6 sysctls, no IPv6 socket table, nothing. Both 29.6.2 and every later engine log the same ipv6-related read failures at startup (&lt;code&gt;failed to read ipv6 net.ipv6.conf.&amp;lt;bridge&amp;gt;.accept_ra&lt;/code&gt;), so the daemon itself already knows this host has none. The difference is what each version does with that fact once a container actually tries to join an overlay sandbox: 29.6.2 carries on and completes the join over IPv4; 29.7.0 onward asks something that comes back unanswered and calls it fatal.&lt;/p&gt;

&lt;p&gt;I want to be careful here: I didn't read the overlay driver's source to confirm this is the actual cause rather than a correlated symptom. What I can say is that the error text names exactly this — "cannot determine address family of transport" — on the one property of this host that changed nothing between engine versions and that I can independently confirm is true.&lt;/p&gt;

&lt;p&gt;I checked whether Docker's own documentation sets an expectation either way. Its swarm networking page lists the ports that need to be open between hosts and says nothing about IPv6 either requiring it or ruling it out:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Port 2377 TCP for communication with and between manager nodes&lt;br&gt;
Port 7946 TCP/UDP for overlay network node discovery&lt;br&gt;
Port 4789 UDP (configurable) for overlay network traffic&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Three protocols, all IPv4-shaped, no mention of IPv6 as a prerequisite anywhere on that page. By that documentation, what I ran should have worked on 29.8.2 exactly as it did on 29.6.2.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workaround that didn't work
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;docker network create&lt;/code&gt; takes an explicit &lt;code&gt;--ipv6&lt;/code&gt; flag, so I tried forcing it off on the overlay network itself, on the theory that the daemon might be tripping over IPv6 address assignment specifically and that turning it off per-network would route around that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;docker network create &lt;span class="nt"&gt;-d&lt;/span&gt; overlay &lt;span class="nt"&gt;--ipv6&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false &lt;/span&gt;gossip-net3
&lt;span class="nv"&gt;$ &lt;/span&gt;docker service create &lt;span class="nt"&gt;--detach&lt;/span&gt; &lt;span class="nt"&gt;--replicas&lt;/span&gt; 5 &lt;span class="nt"&gt;--network&lt;/span&gt; gossip-net3 alpine:3.20 &lt;span class="nb"&gt;sleep &lt;/span&gt;600
&lt;span class="nv"&gt;$ &lt;/span&gt;docker service &lt;span class="nb"&gt;ls
&lt;/span&gt;ID            NAME           MODE         REPLICAS   IMAGE
zwuxi4opkw5  gossip-svc3     replicated   0/5        alpine:3.20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Identical failure, identical error text. Whatever is going wrong, it isn't happening at the per-network IPv6-enable flag, so there's no documented flag I found that gets you out of this on an IPv6-less host.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually costs, which surprised me
&lt;/h2&gt;

&lt;p&gt;I expected a service stuck retrying forever to be quietly expensive — some background loop hammering the scheduler. I sampled &lt;code&gt;dockerd&lt;/code&gt;'s own CPU time once a second for thirty seconds under two conditions: 29.6.2 running twenty real, working replicas, and 29.8.2 sitting on its broken five-replica service.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;Mean CPU&lt;/th&gt;
&lt;th&gt;Max CPU (1s sample)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;29.6.2, 20/20 replicas running&lt;/td&gt;
&lt;td&gt;0.52%&lt;/td&gt;
&lt;td&gt;0.98%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;29.8.2, 0/5 replicas, stuck&lt;/td&gt;
&lt;td&gt;0.30%&lt;/td&gt;
&lt;td&gt;1.97%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The broken service was cheaper on average, not more expensive. Watching &lt;code&gt;docker service ps&lt;/code&gt; over time explained why: each task gets retried three times, each attempt a few seconds apart, and then Swarm stops trying and leaves it sitting in &lt;code&gt;Assigned&lt;/code&gt; forever. No more log lines, no more CPU, no more anything. It fails hard once and then goes completely quiet. That's worse for anyone watching dashboards rather than logs, because nothing about resource usage tells you it's broken.&lt;/p&gt;

&lt;h2&gt;
  
  
  How you'd actually notice
&lt;/h2&gt;

&lt;p&gt;One command shows it plainly, which is the only genuinely good news in this post:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ docker service ls
ID            NAME          MODE         REPLICAS   IMAGE
iaki8w2lsuwr  gossip-svc3   replicated   0/5        alpine:3.20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;0/5&lt;/code&gt;, sitting there indefinitely, next to a service that's genuinely fine. If your deploy tooling checks &lt;code&gt;docker service ls&lt;/code&gt; or the equivalent API field for convergence before calling a rollout successful, you'll catch this immediately. If it only checks that &lt;code&gt;docker service create&lt;/code&gt; returned exit code 0 — which it does, every time, failure included — you won't.&lt;/p&gt;

&lt;h2&gt;
  
  
  A smaller thing I checked and it didn't hold up
&lt;/h2&gt;

&lt;p&gt;While I was in there I also tried to reproduce a specific fix listed in the 29.8.0 notes, about service creation failing when an automatically generated name collides with an existing one. I fired thirty concurrent &lt;code&gt;docker service create&lt;/code&gt; calls with no &lt;code&gt;--name&lt;/code&gt; at both 29.6.2 and 29.8.2, hoping to force a collision in the random name generator. Zero collisions on either version, in thirty concurrent attempts each. The namespace is evidently large enough that this doesn't show up at a scale I can produce by hand in an hour. I'm noting it rather than dropping it, because a finding of "I couldn't reproduce this" is still useful if you were about to spend time worrying about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first run of the concurrent service-creation test hung indefinitely. I'd used &lt;code&gt;docker service create&lt;/code&gt; without &lt;code&gt;--detach&lt;/code&gt;, and with &lt;code&gt;--restart-condition none&lt;/code&gt; the task is supposed to exit after &lt;code&gt;sleep 1&lt;/code&gt; rather than stay running — which meant the CLI's default wait for "the service has converged to its desired running count" could never be satisfied, since the desired count for a task designed to exit is never going to match "currently running." It wasn't a Docker bug, it was me asking the CLI to wait for a state I'd specifically engineered never to arrive. Adding &lt;code&gt;--detach&lt;/code&gt; fixed it in about ten seconds once I noticed what the command was actually doing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This needs root and a Linux host with Docker's static builds reachable from &lt;code&gt;download.docker.com&lt;/code&gt;. It stands up two engines side by side on their own sockets so you can compare without touching your real Docker install.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# grab a second engine build to compare against your system one&lt;/span&gt;
curl &lt;span class="nt"&gt;-sA&lt;/span&gt; &lt;span class="s1"&gt;'Mozilla/5.0'&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; docker-29.8.2.tgz &lt;span class="se"&gt;\&lt;/span&gt;
  https://download.docker.com/linux/static/stable/x86_64/docker-29.8.2.tgz
&lt;span class="nb"&gt;mkdir &lt;/span&gt;new &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;tar &lt;/span&gt;xzf docker-29.8.2.tgz &lt;span class="nt"&gt;-C&lt;/span&gt; new
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; /var/lib/docker-new /run/docker-new

&lt;span class="nb"&gt;nohup &lt;/span&gt;new/docker/dockerd &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--data-root&lt;/span&gt; /var/lib/docker-new &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--exec-root&lt;/span&gt; /run/docker-new/exec &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--host&lt;/span&gt; unix:///run/docker-new/docker.sock &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--exec-opt&lt;/span&gt; native.cgroupdriver&lt;span class="o"&gt;=&lt;/span&gt;cgroupfs &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; new.log 2&amp;gt;&amp;amp;1 &amp;amp;
&lt;span class="nb"&gt;sleep &lt;/span&gt;5

&lt;span class="nv"&gt;ADDR&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;hostname&lt;/span&gt; &lt;span class="nt"&gt;-I&lt;/span&gt; | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'{print $1}'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DOCKER_HOST&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;unix:///run/docker-new/docker.sock
new/docker/docker swarm init &lt;span class="nt"&gt;--advertise-addr&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ADDR&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
new/docker/docker network create &lt;span class="nt"&gt;-d&lt;/span&gt; overlay testnet
new/docker/docker service create &lt;span class="nt"&gt;--detach&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; testsvc &lt;span class="nt"&gt;--replicas&lt;/span&gt; 5 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--network&lt;/span&gt; testnet alpine:3.20 &lt;span class="nb"&gt;sleep &lt;/span&gt;600

&lt;span class="nb"&gt;sleep &lt;/span&gt;15
new/docker/docker service &lt;span class="nb"&gt;ls
&lt;/span&gt;new/docker/docker service ps testsvc &lt;span class="nt"&gt;--no-trunc&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-5&lt;/span&gt;

&lt;span class="c"&gt;# check whether this host even has IPv6 at all&lt;/span&gt;
&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/net/if_inet6 2&amp;gt;&amp;amp;1 &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"no IPv6 stack on this host"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;service ls&lt;/code&gt; shows &lt;code&gt;0/5&lt;/code&gt; and your host has no &lt;code&gt;/proc/net/if_inet6&lt;/code&gt;, you've reproduced this. If it shows &lt;code&gt;5/5&lt;/code&gt;, either this has been fixed since 29.8.2 or your host has IPv6, and either way that's worth knowing before you plan around it.&lt;/p&gt;

&lt;p&gt;If you run Swarm in production on a host or VM image that has IPv6 turned off — which is a common, deliberate choice on plenty of minimal server images and sandboxed CI runners — test an upgrade past 29.6.x on a throwaway node before you roll it out, and check &lt;code&gt;docker service ls&lt;/code&gt; for the replica count rather than trusting a clean exit code from &lt;code&gt;service create&lt;/code&gt; or &lt;code&gt;service update&lt;/code&gt;.&lt;/p&gt;

</description>
      <category>docker</category>
      <category>devops</category>
      <category>containers</category>
      <category>networking</category>
    </item>
    <item>
      <title>The BMW manual was off-limits, so I built my friend something better</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Fri, 02 Oct 2026 20:07:29 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/i-couldnt-legally-use-the-repair-manual-so-i-built-my-friend-something-better-84</link>
      <guid>https://dev.to/alexgeorgiev17/i-couldnt-legally-use-the-repair-manual-so-i-built-my-friend-something-better-84</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;My friend repairs his own BMW E90 325i. Not professionally, not as a hobby exactly, just the way someone does when they'd rather understand their car than hand it to someone who won't explain it back.&lt;/p&gt;

&lt;p&gt;Watching him work, the actual bottleneck is never the wrench. It's that the information is scattered. A torque spec lives in a forum post from 2013. A part number lives in a different tab. Half the "guides" are a video where someone talks for four minutes before touching the car. He's lying under a car on jack stands with dirty hands trying to pinch-zoom a phone screen.&lt;/p&gt;

&lt;p&gt;So I built him &lt;strong&gt;&lt;a href="https://georgievalex.github.io/bmw-repair-assistant/" rel="noopener noreferrer"&gt;BMW Repair Workshop&lt;/a&gt;&lt;/strong&gt;: a fast, mobile-first reference for his exact car. Twenty repair procedures, seven reference guides, every torque value and part number laid out in a table instead of buried in prose. And an Ask box where he can describe a problem in his own words and get an answer built only from those procedures.&lt;/p&gt;

&lt;p&gt;The interesting part isn't the site. It's what I wasn't allowed to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  The constraint that shaped everything
&lt;/h2&gt;

&lt;p&gt;My first plan was obvious: feed the service manual into a RAG pipeline and let an agent answer questions from it.&lt;/p&gt;

&lt;p&gt;That plan died about an hour in. The real BMW service documentation (Bentley's manual, BMW's own TIS system) is commercial, copyrighted, and sold for money. It's also mirrored absolutely everywhere, a chapter-split copy is one search away, and there are entire sites dedicated to serving it for free.&lt;/p&gt;

&lt;p&gt;Wide availability isn't permission. So I didn't use any of it. Not the PDFs, not the mirrors, not "just extracting the facts" from them, which is the same thing wearing a different hat.&lt;/p&gt;

&lt;p&gt;That left me with a harder problem and, it turns out, a better project.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where the content actually came from:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Procedures and specs written from general public automotive knowledge, with a handful of figures cross-checked against publicly published sources (owner forums, retailer install guides)&lt;/li&gt;
&lt;li&gt;Every single entry states its own provenance in a Source field, visible on the page&lt;/li&gt;
&lt;li&gt;Every torque figure and part number carries an explicit "verify this against RealOEM or your dealer" caveat, because I am not a mechanic and these came from general knowledge, not from his car's documentation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point matters more than it sounds. A repair reference that quietly presents an unverified number as fact is worse than no reference, because someone torques a bolt to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Live:&lt;/strong&gt; &lt;a href="https://georgievalex.github.io/bmw-repair-assistant/" rel="noopener noreferrer"&gt;https://georgievalex.github.io/bmw-repair-assistant/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Try the Ask box with something vague, the way you'd actually say it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;em&gt;"my coolant keeps disappearing but there is no puddle"&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;"clunk from the front over bumps, what should I check first"&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;"how do I rebuild the automatic transmission?"&lt;/em&gt; ← watch it refuse&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/GeorgievAlex/bmw-repair-assistant" rel="noopener noreferrer"&gt;https://github.com/GeorgievAlex/bmw-repair-assistant&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Open-source AI at the core:&lt;/strong&gt; &lt;a href="https://deepmind.google/models/gemma/" rel="noopener noreferrer"&gt;Gemma&lt;/a&gt; (&lt;code&gt;gemma-4-31B-it&lt;/code&gt;), Google's open-weight model, served through &lt;strong&gt;DigitalOcean Serverless Inference&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;DigitalOcean is doing the actual model serving here. I send an OpenAI-compatible chat completion to &lt;code&gt;inference.do-ai.run/v1/chat/completions&lt;/code&gt;, authenticated with a DigitalOcean model access key, and Gemma runs on their infrastructure. No GPU to provision, no model to host, no container to keep warm. The spend ceiling is their prepaid Inference &amp;amp; Agents balance with auto-reload off, which means the absolute worst case for a runaway is a capped loss and the Ask box going quiet, rather than a surprise bill.&lt;/p&gt;

&lt;p&gt;That combination, an open-weight model served by a provider I'm not locked into, is most of the "open innovation" argument below in one line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Architecture, deliberately boring:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;browser → GitHub Pages (static site, no backend)
            ↓ POST { question }
          Cloudflare Worker  ← holds the model access key
            ↓
          DigitalOcean Serverless Inference → Gemma
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The site is 100% static. Content lives as Markdown in the repo, a dependency-free Node script compiles it into a search index plus one static HTML page per procedure. Search runs client-side. There is no database, no framework, no build toolchain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The design decision I'm happiest with: there is no RAG pipeline.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I started to build one, then measured the corpus. The entire knowledge base is about 11,000 tokens. That fits in a single prompt with room to spare.&lt;/p&gt;

&lt;p&gt;So every question sends &lt;em&gt;all twenty-seven entries&lt;/em&gt; to the model. No embeddings, no vector store, no chunking, no retrieval step that can silently fetch the wrong chunk and answer confidently from it. The model sees the complete corpus every time.&lt;/p&gt;

&lt;p&gt;This is faster to build, free of an entire category of bug, &lt;em&gt;and&lt;/em&gt; more accurate than chunked retrieval at this scale. Costs about $0.002 per question. The reflex to reach for a vector database is strong, and measuring first saved me from it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keeping it honest.&lt;/strong&gt; The system prompt does real work here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Answer only from the provided entries&lt;/li&gt;
&lt;li&gt;Never invent a torque value, part number, or step&lt;/li&gt;
&lt;li&gt;Carry each entry's verification caveats through into the answer&lt;/li&gt;
&lt;li&gt;Keep safety warnings (spring compressors, hot cooling systems, torquing bushings at ride height) rather than trimming them for brevity&lt;/li&gt;
&lt;li&gt;Name the procedure the answer came from&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Asked to rebuild an automatic transmission, it says it has no such procedure and points at the nearest relevant entry. That refusal is the single most important behaviour in the whole project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Four reasons, in increasing order of how much I actually believe them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It's cheap.&lt;/strong&gt; $0.002 a question. The whole thing runs on free static hosting with a free worker in front of an open-weight model. A closed frontier model would work too, and cost more for a job that doesn't need it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model choice is a swap, not a migration.&lt;/strong&gt; Gemma is one line of config. If something better ships, or Gemma gets cheaper elsewhere, I change a string. Nothing else in the project knows or cares which model answers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It's inspectable.&lt;/strong&gt; When the model does something strange, I can read the entire input that produced it: 27 Markdown files in a public repo and a system prompt in a 160-line worker. No hidden retrieval step deciding what the model sees. For something that tells people how tight to torque a brake caliper, I want to be able to explain any answer it gives.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And the one that actually drove the project:&lt;/strong&gt; the open approach was the only one I could build honestly.&lt;/p&gt;

&lt;p&gt;The closed path here isn't a proprietary model. It's the proprietary &lt;em&gt;manual&lt;/em&gt;, and the ecosystem of unauthorised mirrors around it. Those mirrors work right up until they don't. They vanish, they get taken down, they're of unknown provenance, and nothing in them tells you where a number came from.&lt;/p&gt;

&lt;p&gt;What I built instead is small, honest about its limits, and entirely inspectable. Every entry says where it came from. Every number says to verify it. The model can only speak from content that's sitting in a public repo with its sources stated.&lt;/p&gt;

&lt;p&gt;That's a worse product than BMW's actual service manual. It's a far better one than a pirated copy, because you can see exactly what you're trusting.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;I guessed a model ID.&lt;/strong&gt; &lt;code&gt;gemma-4&lt;/code&gt; seemed reasonable. The real one is &lt;code&gt;gemma-4-31B-it&lt;/code&gt;, and a wrong ID fails as a bare &lt;code&gt;404&lt;/code&gt; from the inference API, nothing that says "check your config." Twenty minutes gone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I wrote an XSS bug into the answer renderer.&lt;/strong&gt; My escaping helper used the &lt;code&gt;textContent&lt;/code&gt; → &lt;code&gt;innerHTML&lt;/code&gt; trick, which escapes &lt;code&gt;&amp;lt;&lt;/code&gt;, &lt;code&gt;&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;amp;&lt;/code&gt; but &lt;em&gt;not&lt;/em&gt; quote characters, because quotes aren't special in text nodes. I then dropped that "escaped" output straight into an &lt;code&gt;alt="..."&lt;/code&gt; attribute. One stray &lt;code&gt;"&lt;/code&gt; in a caption, even my own typo, would have broken out of the attribute. Fixed with proper attribute-context escaping.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I told my friend CORS would protect the API budget.&lt;/strong&gt; It won't. CORS is enforced by browsers, so it stops another &lt;em&gt;website&lt;/em&gt; calling the worker from page JavaScript, and does precisely nothing against &lt;code&gt;curl&lt;/code&gt;. Rate limiting and a capped prepaid balance are the actual controls.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I shipped reference pages that rendered &lt;code&gt;## Heading&lt;/code&gt; as literal text&lt;/strong&gt; for a day, because my first reference docs were plain prose and never exercised the heading path.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I didn't do
&lt;/h2&gt;

&lt;p&gt;No VIN decoding yet, so the site can't narrow specs to an exact build, which matters on a chassis spanning several engine variants. No photos of the actual parts, because the only correctly-licensed one I found was a single public-domain engine shot, and a generic stock photo of &lt;em&gt;someone else's&lt;/em&gt; brake caliper on a BMW brake page is the kind of small dishonesty this project is supposed to avoid. No testing against a second chassis, though the data structure already carries chassis and engine code per entry so adding one is content work, not a rewrite.&lt;/p&gt;

&lt;p&gt;And the big one: &lt;strong&gt;every technical figure on that site needs a mechanic's eye.&lt;/strong&gt; The provenance lines say so plainly. Which brings me to the only part of this I couldn't do myself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Handing it over
&lt;/h2&gt;

&lt;p&gt;I sent it to him while he was at the garage, with four questions. The most important was the first: &lt;em&gt;is anything on here actually wrong?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;He went through the jobs he's done himself and nothing contradicted what he knows. I want to be precise about what that is and isn't: it's a working mechanic reading it and nothing jumping out, which is genuine signal. It is not a line-by-line audit against documentation. Those caveats on every number stay exactly where they are.&lt;/p&gt;

&lt;p&gt;On the AI, he was more measured than I expected, in a way I liked:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The answers it gives are valid, you still need to check them of course, like diagnosis and etc"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;He arrived at the project's own position without being told it. That's about the best outcome available for a tool like this: it was useful, and it didn't make him credulous.&lt;/p&gt;

&lt;p&gt;Then he corrected an assumption I'd built the whole input design around. I'd been thinking about voice input, on the theory that dirty hands and phone screens don't mix. He doesn't want it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I always prefer text as long answers can be quickly forgotten and I work with gloves so taking them off to ask is not a problem"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two things I hadn't considered. Gloves come off anyway. And more interesting: a spoken answer &lt;em&gt;evaporates&lt;/em&gt;, while text stays on screen while you're under the car with your hands busy. The persistence is the feature. I'd have built the wrong thing.&lt;/p&gt;

&lt;p&gt;On the look, which I'd worried was too plain:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I like the idea it looks simple and old style - no weird images, adds and stuff"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the one real feature request, which he raised and then talked himself out of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Perhaps it misses like video tutorial itself but this is not possible I guess, like to search in youtube if someone performed this repair on my car/engine"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is possible, just not the way he assumed. I can't host or embed video. But I can hand off a search already narrowed to his exact chassis and engine, so he gets &lt;code&gt;BMW E90 N52B25 brake pads and rotors change&lt;/code&gt; rather than generic results for a different car. Every procedure now has a "See it done" link doing that. It shipped before I finished writing this post.&lt;/p&gt;

&lt;p&gt;He also asked for exhaust and gearbox procedures. Exhaust is a straightforward driveway job and it's added. The gearbox I partly declined: the card covers fluid and pan service, and then says plainly that internal rebuild is a specialist bench job with no honest driveway version. Writing a rebuild procedure from general knowledge, for a job where a wrong clearance is a destroyed transmission, is exactly the failure mode this whole project was built to avoid. Saying "this isn't something I should write" is part of the same discipline as the model refusing the same question.&lt;/p&gt;

&lt;p&gt;The line I didn't expect:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"You're my go to AI guy and also good mechanic yourself so I would love if you can work on this in the future so me, you and other petrol heads can use it"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I'm not a good mechanic. I just wrote down what I could verify and was careful about what I couldn't. But "so me, you and other petrol heads can use it" is a better description of why this is worth continuing than anything in my own notes.&lt;/p&gt;

&lt;h3&gt;
  
  
  What his feedback actually changed
&lt;/h3&gt;

&lt;p&gt;All of this is live on the site now, shipped between his message and this post going up:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;He said&lt;/th&gt;
&lt;th&gt;What changed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wanted to find video of the job on &lt;em&gt;his&lt;/em&gt; engine, assumed impossible&lt;/td&gt;
&lt;td&gt;Every procedure has a "See it done" link, a search pre-narrowed to &lt;code&gt;BMW E90 N52B25 &amp;lt;job&amp;gt;&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Asked for an exhaust procedure&lt;/td&gt;
&lt;td&gt;Added, full job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Asked to "dissamble the gearbox"&lt;/td&gt;
&lt;td&gt;Added fluid and pan service; internal rebuild explicitly declined as a specialist bench job&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefers text over voice, with reasons&lt;/td&gt;
&lt;td&gt;Voice input dropped from the roadmap entirely&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the one I'd have got wrong on my own. The feature I was about to build is the feature he actively didn't want, and I'd never have found that out by thinking harder about it.&lt;/p&gt;

&lt;p&gt;He had one more request, which I'm not going to be able to ship:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"it will be good if the site can handle those repairs and changes for myself but AI is not there yet :D"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Correct on both counts. And I think that joke is a decent summary of where this sits. The useful version of AI here isn't the one that does the job. It's the one that tells him the caliper guide bolts are 30 Nm, points at the procedure it got that from, reminds him to check it, and then gets out of the way while he does the work he already knows how to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Best Use of DigitalOcean&lt;/strong&gt; — DigitalOcean Serverless Inference serves the open-weight model behind the Ask feature, via &lt;code&gt;inference.do-ai.run&lt;/code&gt;, with spend capped by the prepaid Inference &amp;amp; Agents balance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best Use of Gemma&lt;/strong&gt; — &lt;code&gt;gemma-4-31B-it&lt;/code&gt; is the open-weight model doing the reasoning, constrained to the site's own content and instructed to refuse rather than invent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Credits
&lt;/h2&gt;

&lt;p&gt;The project has no dependencies, the code is all hand-written, and the content was written for this project rather than taken from anywhere. Two things worth naming anyway:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The GitHub mark in the site navigation is GitHub's logo, used only to link to the repository.&lt;/li&gt;
&lt;li&gt;The post and the project were built with AI assistance, declared in the front matter. The specs in the site came from general knowledge, and every one of them says on the page that you should verify it yourself.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>ai</category>
    </item>
    <item>
      <title>Git 2.56's fetch.followRemoteHEAD setting fixes stale origin/HEAD for every remote</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Fri, 02 Oct 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/git-256s-fetchfollowremotehead-setting-fixes-stale-originhead-for-every-remote-4hc6</link>
      <guid>https://dev.to/alexgeorgiev17/git-256s-fetchfollowremotehead-setting-fixes-stale-originhead-for-every-remote-4hc6</guid>
      <description>&lt;p&gt;I have an &lt;code&gt;origin/HEAD&lt;/code&gt; pointing at &lt;code&gt;origin/master&lt;/code&gt; in a repository I cloned two years ago, long after the project moved its default branch to &lt;code&gt;main&lt;/code&gt;. It has never fixed itself. Every &lt;code&gt;git fetch&lt;/code&gt; since then has pulled the right commits and left that one symbolic ref exactly where it was.&lt;/p&gt;

&lt;p&gt;Git 2.56, released in the past few weeks, adds a config variable called &lt;code&gt;fetch.followRemoteHEAD&lt;/code&gt; that is supposed to address this, by giving you one global switch instead of a separate setting per remote. I built three versions of Git from source — the Ubuntu-packaged 2.43.0, a 2.55.0 I compiled myself, and 2.56.0 — and used them against a handful of local bare repositories to see what the setting actually does, what it doesn't, and what happens when you reach for the wrong value.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproducing the stale ref
&lt;/h2&gt;

&lt;p&gt;First I confirmed the bug I already own a copy of. I cloned a bare repo while its &lt;code&gt;HEAD&lt;/code&gt; pointed at &lt;code&gt;main&lt;/code&gt;, then renamed the remote's default branch to &lt;code&gt;trunk&lt;/code&gt;, the way GitHub does when you change the default branch in the settings page:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git clone remote2 clone-create &lt;span class="nt"&gt;-q&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;-C&lt;/span&gt; clone-create symbolic-ref refs/remotes/origin/HEAD
refs/remotes/origin/main
&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;--git-dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;remote2 branch &lt;span class="nt"&gt;-m&lt;/span&gt; main trunk
&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;-C&lt;/span&gt; clone-create fetch
From /tmp/gittest/remote2
 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;new branch]      trunk      -&amp;gt; origin/trunk
&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;-C&lt;/span&gt; clone-create symbolic-ref refs/remotes/origin/HEAD
refs/remotes/origin/main
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fetch worked. The new branch arrived. &lt;code&gt;origin/HEAD&lt;/code&gt; still points at &lt;code&gt;origin/main&lt;/code&gt;, a ref to a branch that no longer exists on the remote. That's &lt;code&gt;fetch.followRemoteHEAD&lt;/code&gt;'s default value, &lt;code&gt;create&lt;/code&gt;: it will only ever create the symref if none exists. Once it exists, nothing updates it, forever, on every Git version that has the setting at all.&lt;/p&gt;

&lt;p&gt;Setting the per-remote value to &lt;code&gt;warn&lt;/code&gt; doesn't fix it either, it just tells you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;hint: Run 'git remote set-head origin trunk' to follow the change, or modify
hint: either of the 'remote.origin.followRemoteHEAD' or 'fetch.followRemoteHEAD'
hint: configuration variables to handle the situation differently.
hint:
hint: Using this specific setting
hint:
hint:     git config set remote.origin.followRemoteHEAD warn-if-not-trunk
hint:
hint: will suppress the warning until the remote changes HEAD to something else.
'HEAD' at 'origin' is 'trunk', but we have 'main' locally.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;always&lt;/code&gt;, set per-remote, does what you'd expect: the next fetch silently repoints the symref to &lt;code&gt;trunk&lt;/code&gt;. All three of these modes (&lt;code&gt;create&lt;/code&gt;, &lt;code&gt;warn&lt;/code&gt;, &lt;code&gt;always&lt;/code&gt;, plus &lt;code&gt;never&lt;/code&gt;) existed before 2.56. What's new in 2.56 is that you can set the default for all of them at once, with &lt;code&gt;fetch.followRemoteHEAD&lt;/code&gt;, instead of writing &lt;code&gt;remote.&amp;lt;name&amp;gt;.followRemoteHEAD&lt;/code&gt; into every remote's config by hand.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does the global setting actually save you per-remote config
&lt;/h2&gt;

&lt;p&gt;This is the part worth measuring rather than describing. I set &lt;code&gt;fetch.followRemoteHEAD always&lt;/code&gt; once, in a config file pointed to by &lt;code&gt;GIT_CONFIG_GLOBAL&lt;/code&gt; so I wasn't touching anything outside my test directory, then added three remotes to a fresh repository and fetched all of them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git config &lt;span class="nt"&gt;--global&lt;/span&gt; fetch.followRemoteHEAD always
&lt;span class="nv"&gt;$ &lt;/span&gt;git remote add r1 ../remote1   &lt;span class="c"&gt;# HEAD -&amp;gt; trunk&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;git remote add r2 ../remote2   &lt;span class="c"&gt;# HEAD -&amp;gt; trunk&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;git remote add r3 ../remote3   &lt;span class="c"&gt;# HEAD -&amp;gt; main&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;git fetch &lt;span class="nt"&gt;--all&lt;/span&gt;
...
&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r1/HEAD
refs/remotes/r1/trunk
&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r2/HEAD
refs/remotes/r2/trunk
&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r3/HEAD
refs/remotes/r3/main
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; followRemoteHEAD .git/config
0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero lines of per-remote config across three remotes, all three correctly tracking. But that much is also true of the old default behaviour, &lt;code&gt;create&lt;/code&gt;, for a brand new remote with no existing symref — I'll come back to that, because it's where I fooled myself for a while.&lt;/p&gt;

&lt;p&gt;The real test is what happens later, when a remote I already have changes its default branch again, with no new config written anywhere:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r1/HEAD
refs/remotes/r1/trunk
&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;--git-dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;../remote1 branch &lt;span class="nt"&gt;-m&lt;/span&gt; trunk develop
&lt;span class="nv"&gt;$ &lt;/span&gt;git fetch r1
From ../remote1
 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;new branch]      develop    -&amp;gt; r1/develop
&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r1/HEAD
refs/remotes/r1/develop
&lt;span class="nv"&gt;$ &lt;/span&gt;git config &lt;span class="nt"&gt;--get-regexp&lt;/span&gt; &lt;span class="s1"&gt;'remote\.r1\..*'&lt;/span&gt;
remote.r1.url ../remote1
remote.r1.fetch +refs/heads/&lt;span class="k"&gt;*&lt;/span&gt;:refs/remotes/r1/&lt;span class="k"&gt;*&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the actual improvement: one global line, set once, keeps every remote's &lt;code&gt;HEAD&lt;/code&gt; current indefinitely, without ever touching that remote's own config block. Before 2.56 the equivalent was one &lt;code&gt;remote.&amp;lt;name&amp;gt;.followRemoteHEAD always&lt;/code&gt; line added for every single remote you cared about, by hand, after noticing the problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  It did nothing on last month's Git
&lt;/h2&gt;

&lt;p&gt;To be sure the global key is genuinely new and not just newly documented, I pointed the same &lt;code&gt;GIT_CONFIG_GLOBAL&lt;/code&gt; file at Git 2.55.0, which I'd compiled from the official &lt;code&gt;v2.55.0&lt;/code&gt; tag:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;--version&lt;/span&gt;
git version 2.55.0
&lt;span class="nv"&gt;$ &lt;/span&gt;git fetch r1          &lt;span class="c"&gt;# first fetch, symref doesn't exist yet, 'create' kicks in&lt;/span&gt;
 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;new branch]      develop    -&amp;gt; r1/develop
&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r1/HEAD
refs/remotes/r1/develop
&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;--git-dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;../remote1 branch &lt;span class="nt"&gt;-m&lt;/span&gt; develop final
&lt;span class="nv"&gt;$ &lt;/span&gt;git fetch r1          &lt;span class="c"&gt;# second fetch, remote renamed again&lt;/span&gt;
 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;new branch]      final      -&amp;gt; r1/final
&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r1/HEAD
refs/remotes/r1/develop
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same global config file, same rename, same remote. On 2.55 the second rename leaves &lt;code&gt;HEAD&lt;/code&gt; stale at &lt;code&gt;develop&lt;/code&gt;, because 2.55 only understands the per-remote key. It doesn't error on the unrecognised global setting, it just never consults it. That first fetch succeeding was the default &lt;code&gt;create&lt;/code&gt; behaviour doing its ordinary job on a never-before-seen remote, not the global setting working — which is exactly the mistake I made on my first pass through this test, below.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it refuses
&lt;/h2&gt;

&lt;p&gt;The per-remote setting accepts one value the global one doesn't: &lt;code&gt;warn-if-not-$branch&lt;/code&gt;, which behaves like &lt;code&gt;warn&lt;/code&gt; but keeps quiet as long as the remote's &lt;code&gt;HEAD&lt;/code&gt; still matches the branch you name. I tried setting that at the global level, since the documentation for &lt;code&gt;remote.&amp;lt;name&amp;gt;.followRemoteHEAD&lt;/code&gt; says it accepts "the values supported by &lt;code&gt;fetch.followRemoteHEAD&lt;/code&gt;" plus this one, which reads as though it might go either way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git config &lt;span class="nt"&gt;--global&lt;/span&gt; fetch.followRemoteHEAD warn-if-not-main
&lt;span class="nv"&gt;$ &lt;/span&gt;git fetch r3
warning: unrecognized fetch.followRemoteHEAD value &lt;span class="s1"&gt;'warn-if-not-main'&lt;/span&gt; ignored
From ../remote3
 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;new branch]      newmain    -&amp;gt; r3/newmain
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nv"&gt;$?&lt;/span&gt;
0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No error, no non-zero exit, just a warning and a silent fall-back to &lt;code&gt;create&lt;/code&gt;. If you set this in your global config expecting it to suppress warnings for a specific branch name on every remote, it will quietly not do that, and the only sign is a line easy to miss in fetch output you weren't reading closely.&lt;/p&gt;

&lt;p&gt;I also checked precedence between the two settings, since the docs state the per-remote value overrides the global one but I wanted to see it fail in both directions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git config &lt;span class="nt"&gt;--global&lt;/span&gt; fetch.followRemoteHEAD never
&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;--git-dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;../remote2 branch &lt;span class="nt"&gt;-m&lt;/span&gt; trunk r2final
&lt;span class="nv"&gt;$ &lt;/span&gt;git fetch r2                                    &lt;span class="c"&gt;# no per-remote override: stays stale&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r2/HEAD
refs/remotes/r2/trunk
&lt;span class="nv"&gt;$ &lt;/span&gt;git config remote.r2.followRemoteHEAD always
&lt;span class="nv"&gt;$ &lt;/span&gt;git &lt;span class="nt"&gt;--git-dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;../remote2 branch &lt;span class="nt"&gt;-m&lt;/span&gt; r2final r2final2
&lt;span class="nv"&gt;$ &lt;/span&gt;git fetch r2                                    &lt;span class="c"&gt;# per-remote override: updates&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;git symbolic-ref refs/remotes/r2/HEAD
refs/remotes/r2/r2final2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That matched the documentation exactly, both ways.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs, and what it doesn't tell you
&lt;/h2&gt;

&lt;p&gt;I timed five fetches each, against a local filesystem remote, with the setting on &lt;code&gt;create&lt;/code&gt; and then on &lt;code&gt;always&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;mode&lt;/th&gt;
&lt;th&gt;run 1&lt;/th&gt;
&lt;th&gt;run 2&lt;/th&gt;
&lt;th&gt;run 3&lt;/th&gt;
&lt;th&gt;run 4&lt;/th&gt;
&lt;th&gt;run 5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;create&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;always&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;td&gt;12ms&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;td&gt;11ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;No measurable difference. That's a local filesystem transport, so it only tells you the extra ref write costs nothing on top of a fetch that's already happening; it says nothing about network latency, which this setup can't produce.&lt;/p&gt;

&lt;p&gt;The more interesting cost is that &lt;code&gt;always&lt;/code&gt; mode gives you no feedback when it does something. Compare the two fetches I ran against &lt;code&gt;r3&lt;/code&gt; with &lt;code&gt;always&lt;/code&gt; configured:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git fetch r3
From ../remote3
 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt;new branch]      finalbranch -&amp;gt; r3/finalbranch
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nv"&gt;$?&lt;/span&gt;
0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That output is identical whether or not &lt;code&gt;origin/HEAD&lt;/code&gt; just got silently repointed. If you want to know it happened, you have to go and check with &lt;code&gt;git symbolic-ref&lt;/code&gt; or &lt;code&gt;git remote show &amp;lt;remote&amp;gt;&lt;/code&gt; afterwards; nothing in a normal &lt;code&gt;fetch&lt;/code&gt; run says so. For a setting whose entire purpose is to change state you didn't ask this invocation to change, that's worth knowing before you turn it on for every clone a CI job makes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first attempt at testing the global default was to add three brand new remotes, set &lt;code&gt;fetch.followRemoteHEAD always&lt;/code&gt; beforehand, fetch all three, and check that all three &lt;code&gt;HEAD&lt;/code&gt; refs came out correct with no per-remote config. They did, and for a few minutes I treated that as proof the feature worked.&lt;/p&gt;

&lt;p&gt;It wasn't proof of anything. &lt;code&gt;create&lt;/code&gt;, the old default, already creates a correct &lt;code&gt;HEAD&lt;/code&gt; symref for any remote that doesn't have one yet, with no config at all, on every Git version back to whenever the per-remote setting first existed. Three fresh remotes with no prior &lt;code&gt;HEAD&lt;/code&gt; are exactly the case where &lt;code&gt;create&lt;/code&gt; and &lt;code&gt;always&lt;/code&gt; produce the same result. I only caught this rereading the &lt;code&gt;fetch.followRemoteHEAD&lt;/code&gt; documentation a second time, which is what sent me back to re-run the test against a remote that had already been fetched once and then changed its default branch again — the case the old default genuinely can't handle and the new global setting can.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This needs Git built from source, since Ubuntu's packaged Git (2.43.0 here) predates the per-remote setting entirely. Building takes a few minutes on four cores.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# build 2.56.0 and 2.55.0 for comparison&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;apt-get &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; build-essential libssl-dev libcurl4-openssl-dev libexpat1-dev gettext zlib1g-dev
git clone &lt;span class="nt"&gt;--depth&lt;/span&gt; 1 &lt;span class="nt"&gt;--branch&lt;/span&gt; v2.56.0 https://github.com/git/git /tmp/git-256
&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/git-256 &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make configure &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; ./configure &lt;span class="nt"&gt;--prefix&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/opt/git-2.56 &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make &lt;span class="nt"&gt;-j4&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; make &lt;span class="nb"&gt;install

mkdir&lt;/span&gt; /tmp/gittest &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/gittest
git init &lt;span class="nt"&gt;--bare&lt;/span&gt; &lt;span class="nt"&gt;-b&lt;/span&gt; main remote-a
git clone remote-a seed &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;seed &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git commit &lt;span class="nt"&gt;--allow-empty&lt;/span&gt; &lt;span class="nt"&gt;-qm&lt;/span&gt; init &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; git push &lt;span class="nt"&gt;-q&lt;/span&gt; origin main
&lt;span class="nb"&gt;cd&lt;/span&gt; .. &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; seed

&lt;span class="nv"&gt;GIT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/opt/git-2.56/bin/git
&lt;span class="nv"&gt;$GIT&lt;/span&gt; clone remote-a work &lt;span class="nt"&gt;-q&lt;/span&gt;
git &lt;span class="nt"&gt;--git-dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;remote-a branch &lt;span class="nt"&gt;-m&lt;/span&gt; main trunk   &lt;span class="c"&gt;# simulate a default-branch rename&lt;/span&gt;

&lt;span class="nb"&gt;cd &lt;/span&gt;work
&lt;span class="nv"&gt;$GIT&lt;/span&gt; fetch                                     &lt;span class="c"&gt;# pulls trunk, origin/HEAD stays on main&lt;/span&gt;
&lt;span class="nv"&gt;$GIT&lt;/span&gt; symbolic-ref refs/remotes/origin/HEAD     &lt;span class="c"&gt;# -&amp;gt; refs/remotes/origin/main, stale&lt;/span&gt;

&lt;span class="nv"&gt;$GIT&lt;/span&gt; config remote.origin.followRemoteHEAD always
git &lt;span class="nt"&gt;--git-dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;../remote-a branch &lt;span class="nt"&gt;-m&lt;/span&gt; trunk final
&lt;span class="nv"&gt;$GIT&lt;/span&gt; fetch                                     &lt;span class="c"&gt;# now it follows&lt;/span&gt;
&lt;span class="nv"&gt;$GIT&lt;/span&gt; symbolic-ref refs/remotes/origin/HEAD     &lt;span class="c"&gt;# -&amp;gt; refs/remotes/origin/final&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you maintain more than a couple of remotes, or you're responsible for a CI image that clones repositories whose owners might rename &lt;code&gt;master&lt;/code&gt; to &lt;code&gt;main&lt;/code&gt; out from under you, set &lt;code&gt;fetch.followRemoteHEAD&lt;/code&gt; to &lt;code&gt;always&lt;/code&gt; once in your global config and stop thinking about it per-remote. Just don't expect &lt;code&gt;warn-if-not-$branch&lt;/code&gt; to work there, and don't expect any output when it quietly fixes something for you.&lt;/p&gt;

</description>
      <category>git</category>
      <category>devops</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>SQLite 3.53's SET NOT NULL replaces a full table rebuild with six writes</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Thu, 01 Oct 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/sqlite-353s-set-not-null-replaces-a-full-table-rebuild-with-six-writes-110e</link>
      <guid>https://dev.to/alexgeorgiev17/sqlite-353s-set-not-null-replaces-a-full-table-rebuild-with-six-writes-110e</guid>
      <description>&lt;p&gt;For as long as I've used SQLite, adding a NOT NULL constraint to an existing column meant the twelve-step dance: create a new table with the constraint you want, copy every row across, drop the old table, rename the new one, and rebuild whatever indexes you just lost in the process. SQLite 3.53.0, released in April 2026, adds &lt;code&gt;ALTER TABLE ... ALTER COLUMN ... SET NOT NULL&lt;/code&gt; and &lt;code&gt;DROP NOT NULL&lt;/code&gt;, so I built a few test tables and measured what actually changes.&lt;/p&gt;

&lt;p&gt;I ran this against the real 3.53.4 CLI, not the one that ships in Ubuntu's &lt;code&gt;apt&lt;/code&gt; repositories, which still tops out at 3.45.1 and doesn't understand the new syntax at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rebuild, timed
&lt;/h2&gt;

&lt;p&gt;I generated a 2 million row table with an index on the column I was about to constrain, dropped the OS page cache before each run (&lt;code&gt;echo 3 &amp;gt; /proc/sys/vm/drop_caches&lt;/code&gt;) so I wasn't measuring the filesystem cache instead of SQLite, and timed both approaches three times each.&lt;/p&gt;

&lt;p&gt;The old way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;t_new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="nb"&gt;INTEGER&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;INSERT&lt;/span&gt; &lt;span class="k"&gt;INTO&lt;/span&gt; &lt;span class="n"&gt;t_new&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;DROP&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;t_new&lt;/span&gt; &lt;span class="k"&gt;RENAME&lt;/span&gt; &lt;span class="k"&gt;TO&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;INDEX&lt;/span&gt; &lt;span class="n"&gt;idx_email&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The new way:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Run 1&lt;/th&gt;
&lt;th&gt;Run 2&lt;/th&gt;
&lt;th&gt;Run 3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Rebuild (old way)&lt;/td&gt;
&lt;td&gt;1.788s&lt;/td&gt;
&lt;td&gt;1.495s&lt;/td&gt;
&lt;td&gt;1.594s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;SET NOT NULL&lt;/code&gt; (indexed column)&lt;/td&gt;
&lt;td&gt;0.0093s&lt;/td&gt;
&lt;td&gt;0.0113s&lt;/td&gt;
&lt;td&gt;0.0097s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's roughly 150 to 180 times faster, and it isn't a rounding trick: the rebuild also has to recreate &lt;code&gt;idx_email&lt;/code&gt; from scratch, which the old approach needs you to remember to do by hand. Forget that step and your queries silently fall back to a table scan. &lt;code&gt;SET NOT NULL&lt;/code&gt; leaves the index exactly where it was.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's that fast: it barely touches the disk
&lt;/h2&gt;

&lt;p&gt;The speed difference made me suspicious, so I ran both under &lt;code&gt;strace -c&lt;/code&gt; counting &lt;code&gt;pwrite64&lt;/code&gt; calls on a fresh copy of the same database.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;=== OLD WAY ===
% time     seconds  usecs/call     calls    errors syscall
100.00    1.001682          14     69843           pwrite64

=== NEW WAY ===
% time     seconds  usecs/call     calls    errors syscall
  0.00    0.000000           0         6           pwrite64
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six writes totalling 8,716 bytes, against 69,843 writes for the rebuild. The rebuild touches every row twice — once to write it into the new table, once again when the index is rebuilt. &lt;code&gt;SET NOT NULL&lt;/code&gt; only rewrites the schema record in &lt;code&gt;sqlite_master&lt;/code&gt;; the row data pages never move.&lt;/p&gt;

&lt;h2&gt;
  
  
  The catch: it still has to read every row, unless there's an index
&lt;/h2&gt;

&lt;p&gt;Here's the finding that argues against the headline number. SQLite's own release notes say &lt;code&gt;ALTER TABLE&lt;/code&gt; execution time for a new NOT NULL constraint "is proportional to the amount of data in the table," because every existing row has to be checked. On my first table that wasn't visible at all — 0.01 seconds for 2 million rows looked too good to be true, so I ran it again on a column with no index.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;=== indexed column, 8M rows: read syscalls ===
calls
   13   pread64

=== unindexed column, 8M rows: read syscalls ===
calls
25372   pread64
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With an index on the target column, SQLite checks for nulls by walking the index itself — NULLs sort first in SQLite's b-trees, so it only has to look at the left edge, 13 page reads regardless of table size. Without an index, it has to scan every data page: 25,372 reads for an 8 million row table.&lt;/p&gt;

&lt;p&gt;On a warm cache this difference nearly disappears in wall-clock terms (both finish in well under a second, because Linux is just serving pages out of RAM). Cold-cache numbers tell a clearer story:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Column has index?&lt;/th&gt;
&lt;th&gt;Cold-cache time (8M rows, created_at / email)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;~0.010s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;~0.11s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Still much faster than a rebuild either way, but the documentation's "proportional to table size" claim only really bites when there's no index backing the column you're constraining.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it refuses, and how unhelpfully
&lt;/h2&gt;

&lt;p&gt;Feed it a table with an existing null and it refuses, correctly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;sqlite3 t.db &lt;span class="s2"&gt;"ALTER TABLE t ALTER COLUMN b SET NOT NULL;"&lt;/span&gt;
Error &lt;span class="k"&gt;in &lt;/span&gt;2nd &lt;span class="nb"&gt;command &lt;/span&gt;line argument: constraint failed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare that with the error you get for the exact same violation through a normal &lt;code&gt;INSERT&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ sqlite3 t.db "INSERT INTO t VALUES (1, NULL);"
Error near line 1: NOT NULL constraint failed: t.b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;INSERT&lt;/code&gt; error names the table and the column. The &lt;code&gt;ALTER TABLE&lt;/code&gt; error says "constraint failed" and nothing else — not which column, not which constraint type, not how many rows fail or which row id. On a real table with several nullable columns you're trying to lock down one at a time, that message won't tell you which one is the problem. You have to go find the offending rows yourself, with something like &lt;code&gt;SELECT rowid FROM t WHERE email IS NULL LIMIT 5&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  A working feature nowhere in the docs
&lt;/h2&gt;

&lt;p&gt;While testing CHECK constraints I tried the ANSI-SQL-style syntax out of habit, expecting a parse error, because SQLite's own &lt;code&gt;lang_altertable.html&lt;/code&gt; page only documents &lt;code&gt;SET NOT NULL&lt;/code&gt; / &lt;code&gt;DROP NOT NULL&lt;/code&gt; and says CHECK constraints can only be added via &lt;code&gt;ADD COLUMN&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="k"&gt;ADD&lt;/span&gt; &lt;span class="k"&gt;CONSTRAINT&lt;/span&gt; &lt;span class="n"&gt;age_positive&lt;/span&gt; &lt;span class="k"&gt;CHECK&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;age&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It worked. It validated existing rows, it rejected a negative insert afterwards with a properly named error (&lt;code&gt;CHECK constraint failed: age_positive&lt;/code&gt;), and &lt;code&gt;ALTER TABLE t DROP CONSTRAINT age_positive;&lt;/code&gt; cleanly removed it again. I checked the raw HTML of the ALTER TABLE reference page for the literal strings "ADD CONSTRAINT" and "DROP CONSTRAINT" — zero matches, in either direction. This is a fully working, round-trippable feature that I couldn't find documented anywhere on sqlite.org.&lt;/p&gt;

&lt;h2&gt;
  
  
  Locking: reads pass through, writes don't
&lt;/h2&gt;

&lt;p&gt;I used a FIFO to keep one sqlite3 CLI session open with an uncommitted &lt;code&gt;BEGIN IMMEDIATE&lt;/code&gt; write transaction, then tried to run the ALTER from a second connection against the locked database.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;--- ALTER with busy_timeout=0, while a write lock is held ---
Error in 2nd command line argument: database is locked
immediate-fail attempt wall=0.004s

--- ALTER with busy_timeout=3000, same lock held for 0.3s then released ---
wait-then-succeed attempt wall=0.982s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With no timeout it fails immediately. With a timeout it queues and runs as soon as the lock is free, same as any other write. Unsurprising, but worth confirming, since it means you do need &lt;code&gt;busy_timeout&lt;/code&gt; set on whatever runs this migration, same as any other schema change.&lt;/p&gt;

&lt;p&gt;The more interesting case is a concurrent reader rather than a concurrent writer. I started the slow (unindexed, 8M row) &lt;code&gt;SET NOT NULL&lt;/code&gt; in the background, let it get 150ms into its scan, then fired a &lt;code&gt;SELECT count(*) FROM t&lt;/code&gt; from a second connection with &lt;code&gt;busy_timeout=0&lt;/code&gt;, in both the default rollback-journal mode and WAL mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DELETE mode: reader wall=0.097s, returned 8000000, ALTER total wall=0.658s
WAL mode:    reader wall=0.111s, returned 8000000, ALTER total wall=0.322s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The reader succeeded immediately in both modes, without waiting for the ALTER to finish. The NOT NULL scan only needs a shared lock for the read phase; it doesn't take an exclusive lock until the very end, for that tiny schema write. If your application is read-heavy, running this migration won't stall your readers even on a large table.&lt;/p&gt;

&lt;h2&gt;
  
  
  The no-op claim checks out
&lt;/h2&gt;

&lt;p&gt;The docs say calling &lt;code&gt;SET NOT NULL&lt;/code&gt; on a column that's already &lt;code&gt;NOT NULL&lt;/code&gt; is a no-op. I ran it twice in a row:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="c1"&gt;-- succeeds, adds the constraint&lt;/span&gt;
&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;COLUMN&lt;/span&gt; &lt;span class="n"&gt;email&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="c1"&gt;-- succeeds again, exit code 0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The schema was identical before and after the second call, and the second call didn't trigger the pread64-heavy validation scan a fresh constraint would. &lt;code&gt;DROP NOT NULL&lt;/code&gt; behaved correctly too: after dropping it, an &lt;code&gt;INSERT&lt;/code&gt; with a null value in that column succeeded again.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first attempt at the concurrency test used a background shell loop that checked for the existence of a flag file before writing anything. The loop and the &lt;code&gt;touch&lt;/code&gt; of that flag file were two separate commands, and when I ran them back to back, the loop occasionally started, checked for the file, found it missing, and exited — before the &lt;code&gt;touch&lt;/code&gt; had actually run. The result was a test that reported "0 writer attempts logged" and looked, misleadingly, like no concurrent activity had happened at all. I rebuilt it using a named pipe holding one persistent sqlite3 session with an explicit &lt;code&gt;BEGIN IMMEDIATE&lt;/code&gt;, which gave me a lock I controlled directly instead of a lock I was hoping a loop would acquire in time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;Get the real CLI, since distro packages lag badly behind:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-L&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; sqlite-tools.zip &lt;span class="se"&gt;\&lt;/span&gt;
  https://www.sqlite.org/2026/sqlite-tools-linux-x64-3530400.zip
unzip &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; sqlite-tools.zip
./sqlite3 &lt;span class="nt"&gt;--version&lt;/span&gt;   &lt;span class="c"&gt;# should print 3.53.4 or later&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Build a test table and compare the two approaches:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./sqlite3 bench.db &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;'
CREATE TABLE t(id INTEGER PRIMARY KEY, email TEXT, created_at TEXT);
WITH RECURSIVE seq(x) AS (
  SELECT 1 UNION ALL SELECT x+1 FROM seq WHERE x &amp;lt; 2000000
)
INSERT INTO t(id, email, created_at)
  SELECT x, 'user'||x||'@example.com', datetime('now') FROM seq;
CREATE INDEX idx_email ON t(email);
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;&lt;span class="nb"&gt;cp &lt;/span&gt;bench.db newway.db
&lt;span class="nb"&gt;time&lt;/span&gt; ./sqlite3 newway.db &lt;span class="s2"&gt;"ALTER TABLE t ALTER COLUMN email SET NOT NULL;"&lt;/span&gt;

&lt;span class="nb"&gt;cp &lt;/span&gt;bench.db oldway.db
&lt;span class="nb"&gt;time&lt;/span&gt; ./sqlite3 oldway.db &lt;span class="s2"&gt;"
  CREATE TABLE t_new(id INTEGER PRIMARY KEY, email TEXT NOT NULL, created_at TEXT);
  INSERT INTO t_new SELECT * FROM t;
  DROP TABLE t;
  ALTER TABLE t_new RENAME TO t;
  CREATE INDEX idx_email ON t(email);
"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you want the write-count comparison, run the same two commands under &lt;code&gt;strace -f -e trace=pwrite64 -c&lt;/code&gt; instead of &lt;code&gt;time&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this
&lt;/h2&gt;

&lt;p&gt;If you're on SQLite 3.53 or later and need to tighten a schema that's already in production, use &lt;code&gt;ALTER TABLE ... ALTER COLUMN ... SET NOT NULL&lt;/code&gt; without hesitation — it's faster, it doesn't disturb your indexes, and it won't block concurrent readers. But check your target column for nulls yourself first, with an explicit &lt;code&gt;SELECT ... WHERE col IS NULL&lt;/code&gt;, before you run it on anything you can't immediately re-run: the error message won't tell you where the violation is. And if you need a CHECK constraint added or removed on a live table, &lt;code&gt;ADD CONSTRAINT&lt;/code&gt; and &lt;code&gt;DROP CONSTRAINT&lt;/code&gt; already work, even though the manual doesn't mention them yet — just don't be surprised if that's a documentation gap rather than a permanent feature, and keep an eye on future release notes in case the syntax changes before it's formally written up.&lt;/p&gt;

</description>
      <category>sqlite</category>
      <category>database</category>
      <category>performance</category>
      <category>devops</category>
    </item>
    <item>
      <title>MongoDB Community Edition's new search index stayed pending above 89% disk use</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/mongodb-community-editions-new-search-index-stayed-pending-above-89-disk-use-5bng</link>
      <guid>https://dev.to/alexgeorgiev17/mongodb-community-editions-new-search-index-stayed-pending-above-89-disk-use-5bng</guid>
      <description>&lt;p&gt;MongoDB Search and Vector Search for Community Edition went GA on 30 June 2026. Self-managed deployments get the same &lt;code&gt;mongot&lt;/code&gt; search process that used to be Atlas-only, running alongside &lt;code&gt;mongod&lt;/code&gt; instead of bolted onto a separate engine. I followed MongoDB's own Docker installation guide word for word, and the search index never left &lt;code&gt;PENDING&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting it up
&lt;/h2&gt;

&lt;p&gt;The docs describe two containers on one network: &lt;code&gt;mongod&lt;/code&gt; (MongoDB Community Server 9.0.2) and &lt;code&gt;mongot-community&lt;/code&gt; (&lt;code&gt;mongot&lt;/code&gt; 1.70.5), joined by a gRPC connection and a SCRAM user with the &lt;code&gt;searchCoordinator&lt;/code&gt; role. I ran it exactly as documented: a single-member replica set, a &lt;code&gt;mongod.conf&lt;/code&gt; pointing at the search host, a &lt;code&gt;mongot.conf&lt;/code&gt; pointing back at &lt;code&gt;mongod&lt;/code&gt;, then &lt;code&gt;createSearchIndex&lt;/code&gt; on a 50,000-document test collection.&lt;/p&gt;

&lt;p&gt;The index came back immediately, as expected:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;default&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;PENDING&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;queryable&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;PENDING&lt;/code&gt; is normal for the first few seconds. It was not normal five minutes later, still retrying every 30 seconds with the same message in the &lt;code&gt;mongot&lt;/code&gt; log:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"msg"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Transient error while trying to sync ... Retrying in 30000 milliseconds"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="nl"&gt;"stack_trace"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"... PauseInitialSyncException: Initial syncs are paused ..."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Finding the actual cause
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;mongot&lt;/code&gt; exposes Prometheus metrics on port 9946, and one of them settled it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight prometheus"&gt;&lt;code&gt;&lt;span class="n"&gt;mongot_system_disk_space_data_path_free_bytes&lt;/span&gt; &lt;span class="mf"&gt;2.8686217216&lt;/span&gt;&lt;span class="n"&gt;E10&lt;/span&gt;
&lt;span class="n"&gt;mongot_system_disk_space_data_path_total_bytes&lt;/span&gt; &lt;span class="mf"&gt;2.70553174016&lt;/span&gt;&lt;span class="n"&gt;E11&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;28.7GB free out of 270.6GB, which is 89.4% used. MongoDB's own self-managed troubleshooting page states the rule directly: "Replication stops when disk usage exceeds roughly 90% and resumes after usage drops below roughly 85%. For a new index or rebuild, expect the definition to be accepted but the build to stay stuck if disk pressure is already above the protective threshold." My host had drifted into that band and the index was never going to move.&lt;/p&gt;

&lt;p&gt;Nothing about this showed up on the client side. &lt;code&gt;createSearchIndex&lt;/code&gt; returned normally, &lt;code&gt;$listSearchIndexes&lt;/code&gt; just said &lt;code&gt;PENDING&lt;/code&gt;, and &lt;code&gt;df -h&lt;/code&gt; on the same path reported 28% used, not 89%, because &lt;code&gt;df&lt;/code&gt; was measuring against a quota-limited allowance while &lt;code&gt;mongot&lt;/code&gt; was dividing by the underlying device's full reported capacity. Two tools looking at the same disk gave answers 60 points apart.&lt;/p&gt;

&lt;p&gt;I moved &lt;code&gt;mongot&lt;/code&gt;'s data directory onto &lt;code&gt;tmpfs&lt;/code&gt; to get a mount with headroom, and the same index went from &lt;code&gt;PENDING&lt;/code&gt; to serving queries in under five seconds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"msg":"Finished a collection scan phase.","attr":{"numDocumentsIndexed":50000}
"msg":"Completed initial sync. Beginning first commit.","attr":{"duration":"4.850 s", ...}
"msg":"Transitioning from INITIAL_SYNC to STEADY_STATE."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The block was real and the fix was disk headroom, nothing else.&lt;/p&gt;

&lt;h2&gt;
  
  
  A status field that keeps lying
&lt;/h2&gt;

&lt;p&gt;After I recreated the &lt;code&gt;mongot&lt;/code&gt; container, &lt;code&gt;$listSearchIndexes&lt;/code&gt; kept reporting &lt;code&gt;default&lt;/code&gt; as &lt;code&gt;PENDING&lt;/code&gt; even though queries against it worked. The full &lt;code&gt;statusDetail&lt;/code&gt; array explained why: it still listed the old, now-dead &lt;code&gt;mongot&lt;/code&gt; host as &lt;code&gt;PENDING&lt;/code&gt; alongside the new host, which was &lt;code&gt;READY&lt;/code&gt;. The aggregated top-level status is the worst of all hosts it has ever seen, including ones that no longer exist. A vector index I created fresh, with no dead host in its history, reported &lt;code&gt;READY&lt;/code&gt; cleanly. If you restart &lt;code&gt;mongot&lt;/code&gt; in place, don't trust the summary status; read &lt;code&gt;statusDetail&lt;/code&gt; per host.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it refuses, and what it doesn't
&lt;/h2&gt;

&lt;p&gt;Querying an index that exists but isn't ready yet is a hard, specific error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OperationFailure: cannot query search index ... while in state NOT_STARTED (code 8)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Querying an index name that doesn't exist at all is not an error. It returns zero results silently, with only a &lt;code&gt;WARN "No index in catalog"&lt;/code&gt; line in &lt;code&gt;mongot&lt;/code&gt;'s own log, invisible to the client. A typo in an index name and a genuinely empty result set look identical from the application side.&lt;/p&gt;

&lt;h2&gt;
  
  
  The speed comparison I expected to favour $search
&lt;/h2&gt;

&lt;p&gt;Once the index was live, I ran 15 repetitions each of five single-word and two-word queries against &lt;code&gt;$search&lt;/code&gt;, a classic MongoDB &lt;code&gt;$text&lt;/code&gt; index, and a plain case-insensitive regex, all on the same 50,000-document collection.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;min&lt;/th&gt;
&lt;th&gt;median&lt;/th&gt;
&lt;th&gt;p95&lt;/th&gt;
&lt;th&gt;max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;$search&lt;/code&gt; (mongot)&lt;/td&gt;
&lt;td&gt;6.25ms&lt;/td&gt;
&lt;td&gt;9.36ms&lt;/td&gt;
&lt;td&gt;15.43ms&lt;/td&gt;
&lt;td&gt;56.05ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;classic &lt;code&gt;$text&lt;/code&gt; index&lt;/td&gt;
&lt;td&gt;0.48ms&lt;/td&gt;
&lt;td&gt;0.67ms&lt;/td&gt;
&lt;td&gt;0.89ms&lt;/td&gt;
&lt;td&gt;1.57ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;regex scan&lt;/td&gt;
&lt;td&gt;0.48ms&lt;/td&gt;
&lt;td&gt;0.79ms&lt;/td&gt;
&lt;td&gt;2.86ms&lt;/td&gt;
&lt;td&gt;3.95ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a plain term lookup, the new search stack was 10 to 14 times slower than the mechanism it's positioned to replace. That's the gRPC round trip to a separate process rather than an in-process index scan, and it's not a reason to avoid &lt;code&gt;mongot&lt;/code&gt; — it's a reason not to switch to it for cases the old &lt;code&gt;$text&lt;/code&gt; index already handled well.&lt;/p&gt;

&lt;p&gt;Where &lt;code&gt;$search&lt;/code&gt; earns the extra hop is fuzzy matching. I queried all four methods with &lt;code&gt;backpak&lt;/code&gt;, a one-letter typo for &lt;code&gt;backpack&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;results for "backpak"&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;$search&lt;/code&gt; with &lt;code&gt;fuzzy: {maxEdits: 1}&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;$search&lt;/code&gt;, no fuzzy option&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;classic &lt;code&gt;$text&lt;/code&gt; index&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;regex&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Only the fuzzy option found anything. That's the actual capability being sold here, and it worked as documented.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vector search against doing it by hand
&lt;/h2&gt;

&lt;p&gt;I ran &lt;code&gt;$vectorSearch&lt;/code&gt; against a 32-dimension cosine-similarity index over the same 50,000 documents, and compared it to scoring every vector in Python after fetching them once:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;min&lt;/th&gt;
&lt;th&gt;median&lt;/th&gt;
&lt;th&gt;max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;$vectorSearch&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10.38ms&lt;/td&gt;
&lt;td&gt;15.52ms&lt;/td&gt;
&lt;td&gt;93.93ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;brute-force Python cosine scan&lt;/td&gt;
&lt;td&gt;123.23ms&lt;/td&gt;
&lt;td&gt;134.89ms&lt;/td&gt;
&lt;td&gt;243.87ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;$vectorSearch&lt;/code&gt; was roughly 8.7 times faster at the median, and that's generous to the brute-force side: I excluded the time to fetch all 50,000 embeddings over the wire, which a real application doing this in-process would also have to pay unless it already held every vector in memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Concurrency and cost
&lt;/h2&gt;

&lt;p&gt;Ten concurrent clients issuing &lt;code&gt;$search&lt;/code&gt; queries finished 20 total queries in 98.6ms of wall time, against 180.7ms for one client doing the same 20 sequentially — a real throughput gain. But per-query latency degraded: the median went from 8.41ms to 28.77ms, and the worst case rose from 13.07ms to 83.40ms. &lt;code&gt;mongot&lt;/code&gt; is a shared, single-threaded-ish gRPC service from the client's point of view, and concurrency buys you overall throughput at the cost of tail latency.&lt;/p&gt;

&lt;p&gt;I also checked &lt;code&gt;docker stats&lt;/code&gt; while both containers sat idle, with nothing but two small indexes over 50,000 short documents behind them. &lt;code&gt;mongod&lt;/code&gt; held 229.5MiB resident. &lt;code&gt;mongot&lt;/code&gt; held 1.033GiB, over four times as much, and it hadn't answered a single query yet. Anyone budgeting a self-managed box for &lt;code&gt;mongod&lt;/code&gt; alone needs to add a second, JVM-sized line item for &lt;code&gt;mongot&lt;/code&gt;, and that line item doesn't shrink if the collection is small.&lt;/p&gt;

&lt;p&gt;MongoDB's launch post says &lt;code&gt;$search&lt;/code&gt; runs on "the same query model you already know." I wanted to see if that survived contact with a real pipeline rather than a one-line example, so I chained &lt;code&gt;$match&lt;/code&gt;, &lt;code&gt;$project&lt;/code&gt;, &lt;code&gt;$sort&lt;/code&gt;, &lt;code&gt;{$meta: "searchScore"}&lt;/code&gt; and &lt;code&gt;$limit&lt;/code&gt; after a &lt;code&gt;$search&lt;/code&gt; stage. It worked without any special-casing, filtering on price after the text match and sorting on the filtered set, which is a normal thing to want and not something the docs show directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;The sync config in the docs sets &lt;code&gt;readPreference: secondaryPreferred&lt;/code&gt;, and when the index sat in &lt;code&gt;PENDING&lt;/code&gt; for the first few minutes I decided that was the problem: a single-node replica set with no secondary to prefer. So I built one. I added a second &lt;code&gt;mongod&lt;/code&gt; container, ran &lt;code&gt;rs.add()&lt;/code&gt;, watched &lt;code&gt;rs.status()&lt;/code&gt; report a healthy &lt;code&gt;SECONDARY&lt;/code&gt;, and waited for the index to move. It didn't. Another five minutes of the same 30-second retry loop passed before I gave up on that theory and went looking at &lt;code&gt;mongot&lt;/code&gt;'s metrics instead, where the disk numbers had been sitting the entire time. &lt;code&gt;df -h&lt;/code&gt; on the same host never once suggested there was a problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This needs Docker with a working daemon, not just the CLI.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker network create search-community

&lt;span class="c"&gt;# mongod.conf: net.bindIpAll true, replication.replSetName rs0,&lt;/span&gt;
&lt;span class="c"&gt;# plus the setParameter block pointing at mongot-community:27028&lt;/span&gt;
docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; mongod &lt;span class="nt"&gt;--network&lt;/span&gt; search-community &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--network-alias&lt;/span&gt; mongod.search-community &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/mongod.conf:/etc/mongod.conf:ro"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/data/db:/data/db"&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; 27017:27017 &lt;span class="se"&gt;\&lt;/span&gt;
  mongodb/mongodb-community-server:latest &lt;span class="nt"&gt;--config&lt;/span&gt; /etc/mongod.conf

docker &lt;span class="nb"&gt;exec &lt;/span&gt;mongod mongosh &lt;span class="nt"&gt;--port&lt;/span&gt; 27017 &lt;span class="nt"&gt;--quiet&lt;/span&gt; &lt;span class="nt"&gt;--eval&lt;/span&gt; &lt;span class="s1"&gt;'
  rs.initiate({_id:"rs0", members:[{_id:0, host:"mongod.search-community:27017"}]})'&lt;/span&gt;
docker &lt;span class="nb"&gt;exec &lt;/span&gt;mongod mongosh &lt;span class="nt"&gt;--port&lt;/span&gt; 27017 &lt;span class="nt"&gt;--quiet&lt;/span&gt; &lt;span class="nt"&gt;--eval&lt;/span&gt; &lt;span class="s1"&gt;'
  db.getSiblingDB("admin").createUser({user:"mongotUser", pwd:"testpass123", roles:["searchCoordinator"]})'&lt;/span&gt;

&lt;span class="c"&gt;# mongot.conf: syncSource pointing at mongod, storage.dataPath /data/mongot&lt;/span&gt;
docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; mongot-community &lt;span class="nt"&gt;--network&lt;/span&gt; search-community &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--network-alias&lt;/span&gt; mongot-community.search-community &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/data/mongot:/data/mongot"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/mongot.conf:/mongot-community/config.default.yml:ro"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PWD&lt;/span&gt;&lt;span class="s2"&gt;/passwordFile:/passwordFile:ro"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-p&lt;/span&gt; 8080:8080 &lt;span class="nt"&gt;-p&lt;/span&gt; 9946:9946 &lt;span class="se"&gt;\&lt;/span&gt;
  mongodb/mongodb-community-search:latest

curl localhost:8080/health          &lt;span class="c"&gt;# expect {"status":"SERVING"}&lt;/span&gt;
curl localhost:9946/metrics | &lt;span class="nb"&gt;grep &lt;/span&gt;disk_space_data_path   &lt;span class="c"&gt;# check your own headroom first&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check your disk metric before you file a bug: &lt;code&gt;(total - free) / total&lt;/code&gt; from &lt;code&gt;mongot_system_disk_space_data_path_total_bytes&lt;/code&gt; and &lt;code&gt;_free_bytes&lt;/code&gt;, not &lt;code&gt;df -h&lt;/code&gt;. If you're above roughly 85%, free space before you touch the index definition.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this
&lt;/h2&gt;

&lt;p&gt;Before you create a first index on a self-managed deployment, pull &lt;code&gt;mongot_system_disk_space_data_path_free_bytes&lt;/code&gt; and do the arithmetic yourself rather than trusting &lt;code&gt;df&lt;/code&gt; on the same box; the two disagreed by 60 percentage points on my host, and only one of them is the number &lt;code&gt;mongot&lt;/code&gt; actually acts on. Don't rip out a working &lt;code&gt;$text&lt;/code&gt; index or a regex query just because &lt;code&gt;$search&lt;/code&gt; is newer — on plain term lookups mine was an order of magnitude slower. Reach for it when you need fuzzy matching or relevance scoring you can't already get, and reach for &lt;code&gt;$vectorSearch&lt;/code&gt; once scoring vectors in application code starts to hurt, which on my 50,000-row test happened well before the collection felt large. If you restart &lt;code&gt;mongot&lt;/code&gt;, open &lt;code&gt;statusDetail&lt;/code&gt; before you believe whatever the top-level &lt;code&gt;queryable&lt;/code&gt; field tells you.&lt;/p&gt;

</description>
      <category>mongodb</category>
      <category>database</category>
      <category>docker</category>
      <category>performance</category>
    </item>
    <item>
      <title>Docker Compose's daily pull policy cuts registry checks by 80% during up</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Tue, 29 Sep 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/docker-composes-daily-pull-policy-cuts-registry-checks-by-80-during-up-468b</link>
      <guid>https://dev.to/alexgeorgiev17/docker-composes-daily-pull-policy-cuts-registry-checks-by-80-during-up-468b</guid>
      <description>&lt;p&gt;Five &lt;code&gt;docker compose up&lt;/code&gt; cycles against the same project produced five HEAD requests to my registry when the service used &lt;code&gt;pull_policy: always&lt;/code&gt;, one when it used &lt;code&gt;pull_policy: daily&lt;/code&gt;. That part worked exactly as documented. The surprise came when I ran the same comparison against &lt;code&gt;docker compose pull&lt;/code&gt; directly: the registry got hit every single time, regardless of the policy or how long I waited between calls.&lt;/p&gt;

&lt;p&gt;Docker Compose's &lt;code&gt;pull_policy&lt;/code&gt; field has supported time-based refresh windows — &lt;code&gt;daily&lt;/code&gt;, &lt;code&gt;weekly&lt;/code&gt;, and &lt;code&gt;every_&amp;lt;duration&amp;gt;&lt;/code&gt; — since v2.34.0. The idea is straightforward: instead of choosing between "always re-check the registry" and "never re-check unless the image is missing locally", you tell Compose how stale an image is allowed to get before it bothers asking the registry again. The v5.5.0 changelog, released on 17 August 2026, adds this line: "&lt;code&gt;compose pull&lt;/code&gt; now honors &lt;code&gt;pull_policy&lt;/code&gt; refresh windows (&lt;code&gt;daily&lt;/code&gt;, &lt;code&gt;weekly&lt;/code&gt;, &lt;code&gt;every_N&lt;/code&gt;)." I wanted to know what that was worth in practice, and whether it actually did what the changelog said.&lt;/p&gt;

&lt;p&gt;The Docker Engine install I was working with (29.3.1) shipped the Compose plugin at v5.1.1, already older than the release that claims to fix this. So I downloaded the v5.1.1 and v5.5.1 binaries directly from GitHub and ran both against the same setup: a local &lt;code&gt;registry:2&lt;/code&gt; container holding a single small image, with its access log tailed so I could count exactly when a manifest got checked.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the window actually saves under &lt;code&gt;up&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;With a fresh project — local image removed with &lt;code&gt;docker rmi -f&lt;/code&gt; so there was no trace of it left on the daemon — and &lt;code&gt;pull_policy: daily&lt;/code&gt;, I ran &lt;code&gt;up -d&lt;/code&gt; then &lt;code&gt;down&lt;/code&gt; five times in a row, immediately after each other. Compose only checked the registry on the first cycle. The other four reused the image it already had, without a single request leaving the box:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Policy&lt;/th&gt;
&lt;th&gt;Cycles&lt;/th&gt;
&lt;th&gt;Manifest checks&lt;/th&gt;
&lt;th&gt;Image re-pulled&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;always&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;every cycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;daily&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;first cycle only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's the headline number: four out of five checks avoided, on a workload that's realistic for anyone who runs &lt;code&gt;docker compose up&lt;/code&gt; repeatedly during a working session, or as part of a CI job that spins the same stack up and down. Locally, against a registry on the same machine, the wall-clock difference across five cycles was noise — both runs took about 53 seconds, dominated by container and network setup rather than the registry call. That call is not free everywhere, though. Against Docker Hub, resolving an anonymous pull token and checking one manifest took between 0.40 and 1.01 seconds across three tries from this machine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;trial 1: token+HEAD roundtrip 1.014s
trial 2: token+HEAD roundtrip 0.431s
trial 3: token+HEAD roundtrip 0.404s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the cost &lt;code&gt;daily&lt;/code&gt; is built to avoid, and on a slow or rate-limited registry it adds up over a day of repeated &lt;code&gt;up&lt;/code&gt; calls.&lt;/p&gt;

&lt;p&gt;I also checked that the window genuinely expires rather than just remembering "already pulled once, never again". With &lt;code&gt;pull_policy: every_3s&lt;/code&gt;, an immediate second &lt;code&gt;up&lt;/code&gt; skipped the check as expected, but after sleeping six seconds a third &lt;code&gt;up&lt;/code&gt; checked the registry again and re-pulled. The mechanism is time-based, not a permanent cache flag.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the changelog claim didn't hold up
&lt;/h2&gt;

&lt;p&gt;The part I couldn't reproduce is the specific one the release notes highlight. I built a project with &lt;code&gt;pull_policy: every_3s&lt;/code&gt; and ran &lt;code&gt;docker compose pull&lt;/code&gt; — the explicit subcommand, not &lt;code&gt;up&lt;/code&gt; — three times: cold, immediately again, and a third time after sleeping five seconds so the window had clearly elapsed. Both v5.1.1 (before the fix) and v5.5.1 (the release that claims to have shipped it) hit the registry all three times:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pull #1 (cold): 1 check
pull #2 (immediate): 1 check
pull #3 (after 5s, window elapsed): 1 check
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same pattern with &lt;code&gt;daily&lt;/code&gt; in place of &lt;code&gt;every_3s&lt;/code&gt;. &lt;code&gt;docker compose pull --help&lt;/code&gt; shows no new flag between the two versions, and &lt;code&gt;docker compose up --help&lt;/code&gt; is identical too, so whatever changed in v5.5.0 isn't something a flag exposes. I can't tell from the outside what the fix actually touched, but on the exact scenario the changelog names — plain &lt;code&gt;daily&lt;/code&gt;/&lt;code&gt;weekly&lt;/code&gt;/&lt;code&gt;every_N&lt;/code&gt; values, explicit &lt;code&gt;pull&lt;/code&gt;, nothing else in play — I saw no difference in behaviour between the pre-fix and post-fix binary. If you were hoping the fix meant your CI pipeline's &lt;code&gt;docker compose pull&lt;/code&gt; step would start skipping unnecessary registry calls, it doesn't, at least not in this scenario.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tradeoff nobody puts on the label
&lt;/h2&gt;

&lt;p&gt;A refresh window that skips checks also skips seeing real changes. I pushed a different image to the same tag while a &lt;code&gt;daily&lt;/code&gt;-policied service was still inside its window, then reset the local tag back to the old image ID to rule out Compose picking up the change through the local Docker image store rather than the registry. Running &lt;code&gt;up&lt;/code&gt; again produced zero manifest checks and started the container with the old image, silently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;manifest hits: 0
NAME="Alpine Linux"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the documented tradeoff, working as intended rather than as a bug, but it's worth being explicit about: any push to that tag during the window is invisible to &lt;code&gt;up&lt;/code&gt; until the window closes. If you deploy hotfixes by re-pushing a mutable tag, &lt;code&gt;pull_policy: daily&lt;/code&gt; will hide them from you for up to a day.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it refuses
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;pull_policy&lt;/code&gt; field is validated against a fairly narrow pattern. &lt;code&gt;hourly&lt;/code&gt; and &lt;code&gt;monthly&lt;/code&gt; are both rejected outright:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;services.app.pull_policy 'monthly' does not match pattern
'^(always|never|build|if_not_present|missing|refresh|daily|weekly|every_([0-9]+[wdhms])+)+$'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;every_3x&lt;/code&gt; fails the same way — the suffix after the number has to be &lt;code&gt;w&lt;/code&gt;, &lt;code&gt;d&lt;/code&gt;, &lt;code&gt;h&lt;/code&gt;, &lt;code&gt;m&lt;/code&gt; or &lt;code&gt;s&lt;/code&gt;, nothing else. Neither rejection surprised me once I'd read the regex. What did surprise me was a sixth word sitting in that same pattern, one the docs page never names: &lt;code&gt;refresh&lt;/code&gt;, on its own, with no number attached. I set &lt;code&gt;pull_policy: refresh&lt;/code&gt; and ran &lt;code&gt;up&lt;/code&gt; twice in a row. It checked the registry both times — same as &lt;code&gt;always&lt;/code&gt;, no window at all. My guess is that &lt;code&gt;daily&lt;/code&gt;, &lt;code&gt;weekly&lt;/code&gt; and &lt;code&gt;every_N&lt;/code&gt; all compile down to this same internal policy with a duration attached, and a duration of zero just means always trip it. That's a guess, not something I read in source; what I did confirm is that dropping the duration entirely is accepted by the schema and does something, and none of the three docs pages I checked mention it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first attempt at the staleness test looked like it disproved the whole feature: I pushed a new image under the same tag, ran &lt;code&gt;up&lt;/code&gt; again inside the window, and the container came back running the new image with zero manifest checks logged. That looked like &lt;code&gt;daily&lt;/code&gt; was somehow psychic. It wasn't. Pushing the new image from the same Docker daemon that was also running the compose project updates that daemon's local tag mapping immediately, with no registry round trip involved — &lt;code&gt;docker tag&lt;/code&gt; and &lt;code&gt;docker push&lt;/code&gt; together mean the local image store already knows about the new image before Compose ever looks at it. Compose was comparing local state, not talking to the registry at all, and getting the right answer for the wrong reason. Resetting the local tag back to the old image ID before the second &lt;code&gt;up&lt;/code&gt; — so the only place the new content existed was the registry — is what actually isolated the registry-check behaviour I was trying to measure.&lt;/p&gt;

&lt;p&gt;The same mechanism bit me a second time while writing the reproduction script below. I tagged and pushed the test image, then went straight into the five-cycle loop without removing the local image first. Every single cycle came back with zero registry checks, including the first one, which made it look like &lt;code&gt;daily&lt;/code&gt; was broken in the opposite direction. It wasn't that either: &lt;code&gt;docker tag&lt;/code&gt; stamps the image's &lt;code&gt;LastTagTime&lt;/code&gt; the moment it runs, and Compose's freshness window reads that same timestamp. As far as Compose was concerned, the image had already been "checked" a second earlier, by an operation that never went near the registry. The fix was mechanical — &lt;code&gt;docker rmi -f&lt;/code&gt; the tag before the loop, so the only honest timestamp left is the one Compose itself sets on a real pull — but it means the window isn't tracking "the last time Compose checked", it's tracking "the last time anything touched this tag locally", which is worth knowing if you also build or &lt;code&gt;docker tag&lt;/code&gt; images under the same name your compose file uses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This needs Docker with a running daemon; no cloud account or paid registry involved.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# start a local registry and push a tiny test image&lt;/span&gt;
docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; 5500:5000 &lt;span class="nt"&gt;--name&lt;/span&gt; test-registry registry:2
docker pull alpine:3.20
docker tag alpine:3.20 localhost:5500/pulltest:latest
docker push localhost:5500/pulltest:latest

&lt;span class="c"&gt;# tagging and pushing just now already stamped this image's local&lt;/span&gt;
&lt;span class="c"&gt;# LastTagTime, which is what the refresh window reads -- remove the&lt;/span&gt;
&lt;span class="c"&gt;# local tag so the first cycle below is a genuinely cold check&lt;/span&gt;
docker rmi &lt;span class="nt"&gt;-f&lt;/span&gt; localhost:5500/pulltest:latest

&lt;span class="c"&gt;# tail the registry's access log in another terminal&lt;/span&gt;
docker logs &lt;span class="nt"&gt;-f&lt;/span&gt; test-registry

&lt;span class="c"&gt;# a minimal project using the refresh window&lt;/span&gt;
&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; compose.yaml &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;'
services:
  app:
    image: localhost:5500/pulltest:latest
    pull_policy: daily
    command: ["sleep", "3600"]
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;&lt;span class="c"&gt;# five cycles: only the first should touch the registry&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;i &lt;span class="k"&gt;in &lt;/span&gt;1 2 3 4 5&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;docker compose up &lt;span class="nt"&gt;-d&lt;/span&gt;
  docker compose down
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Switch &lt;code&gt;pull_policy&lt;/code&gt; to &lt;code&gt;always&lt;/code&gt; and repeat to see the difference in the registry log. Swap &lt;code&gt;daily&lt;/code&gt; for &lt;code&gt;every_3s&lt;/code&gt; and add a &lt;code&gt;sleep 6&lt;/code&gt; between two &lt;code&gt;up&lt;/code&gt; calls to see the window actually expire. And to reproduce the part that didn't work as advertised, replace the &lt;code&gt;for&lt;/code&gt; loop with two back-to-back &lt;code&gt;docker compose pull&lt;/code&gt; calls and watch the registry log show a check both times regardless.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this
&lt;/h2&gt;

&lt;p&gt;If you're using &lt;code&gt;pull_policy: daily&lt;/code&gt; or &lt;code&gt;weekly&lt;/code&gt; anywhere and relying on it to make &lt;code&gt;docker compose pull&lt;/code&gt; in a CI step cheaper, check your registry's own access logs during that step. On the evidence here, it isn't skipping anything, on v5.5.1 or on the older v5.1.1 it's supposed to have fixed. The saving is real, but it's currently only real under &lt;code&gt;up&lt;/code&gt;. If your workflow calls &lt;code&gt;pull&lt;/code&gt; directly, and you want fewer registry round trips, you'd need to drop the explicit &lt;code&gt;pull&lt;/code&gt; step and let &lt;code&gt;up&lt;/code&gt; do the pulling instead — with the tradeoff that any mutable-tag push you make will stay invisible to your stack until the window you chose runs out.&lt;/p&gt;

</description>
      <category>docker</category>
      <category>devops</category>
      <category>containers</category>
      <category>performance</category>
    </item>
    <item>
      <title>PostgreSQL 19's WAIT FOR LSN command replaces the read-your-writes polling loop</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Mon, 28 Sep 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/postgresql-19s-wait-for-lsn-command-replaces-the-read-your-writes-polling-loop-1e5m</link>
      <guid>https://dev.to/alexgeorgiev17/postgresql-19s-wait-for-lsn-command-replaces-the-read-your-writes-polling-loop-1e5m</guid>
      <description>&lt;p&gt;I had ten clients hammering a PostgreSQL standby with the same query in a loop: has my last write landed yet? Standby CPU sat at a mean of 57% across two runs. I swapped the loop for a single blocking call, same ten clients, same workload, and CPU dropped to 38% while throughput went up.&lt;/p&gt;

&lt;p&gt;That blocking call is &lt;code&gt;WAIT FOR LSN&lt;/code&gt;, new in PostgreSQL 19, which shipped its fourth beta on 24 September 2026. It gives an application a way to ask the database directly: block this connection until a specific write-ahead-log position has reached a named state on this server, then return. Read-your-writes consistency against an async replica has always been possible by hand, by capturing the WAL position after a write and polling for it on the replica. &lt;code&gt;WAIT FOR&lt;/code&gt; turns that into one statement.&lt;/p&gt;

&lt;p&gt;I built a primary and a streaming-replication standby with &lt;code&gt;postgres:19beta4&lt;/code&gt; and measured the old way against the new one. The honest version of the result is more interesting than "faster": the raw latency is close to a tie, and the real difference only shows up once you put load on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting it up
&lt;/h2&gt;

&lt;p&gt;Two Docker containers, same image, one network:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker network create pgwaitnet
docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; pg-primary &lt;span class="nt"&gt;--network&lt;/span&gt; pgwaitnet &lt;span class="nt"&gt;-p&lt;/span&gt; 15432:5432 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nv"&gt;POSTGRES_PASSWORD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;postgres &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nv"&gt;POSTGRES_HOST_AUTH_METHOD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;trust &lt;span class="se"&gt;\&lt;/span&gt;
  postgres:19beta4 &lt;span class="se"&gt;\&lt;/span&gt;
  postgres &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="nv"&gt;wal_level&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;replica &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="nv"&gt;max_wal_senders&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;10 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="nv"&gt;hot_standby&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;on
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I created a replication role and a slot on the primary, opened &lt;code&gt;pg_hba.conf&lt;/code&gt; for replication and normal connections from the container network, then took a base backup straight into the standby's data directory with &lt;code&gt;pg_basebackup -R&lt;/code&gt;, which writes &lt;code&gt;standby.signal&lt;/code&gt; and &lt;code&gt;primary_conninfo&lt;/code&gt; for me:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker &lt;span class="nb"&gt;exec &lt;/span&gt;pg-primary psql &lt;span class="nt"&gt;-U&lt;/span&gt; postgres &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"CREATE ROLE replicator WITH REPLICATION LOGIN PASSWORD 'replpass';"&lt;/span&gt;
docker &lt;span class="nb"&gt;exec &lt;/span&gt;pg-primary psql &lt;span class="nt"&gt;-U&lt;/span&gt; postgres &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"SELECT pg_create_physical_replication_slot('standby_slot');"&lt;/span&gt;
docker &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; postgres pg-standby bash &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"PGPASSWORD=replpass pg_basebackup -h pg-primary -U replicator &lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;&lt;span class="s2"&gt;
   -D /var/lib/postgresql/19/docker -Fp -Xs -P -R -S standby_slot"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;pg_stat_replication&lt;/code&gt; on the primary showed the standby streaming within a couple of seconds, replay lag around 0.1ms on this one machine. Everything below ran against that pair, over TCP from the host, not through &lt;code&gt;docker exec&lt;/code&gt; (I mention this because it mattered — see the mistake section further down).&lt;/p&gt;

&lt;h2&gt;
  
  
  The headline number, and why I don't fully trust it
&lt;/h2&gt;

&lt;p&gt;For each of 20 trials, I inserted a row on the primary, captured the resulting LSN, and timed how long it took the standby to satisfy that LSN, three ways: a polling loop with a 1ms sleep between checks, a busy-polling loop with no sleep at all, and one &lt;code&gt;WAIT FOR LSN ... WITH (MODE 'standby_replay')&lt;/code&gt; call.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;method&lt;/th&gt;
&lt;th&gt;median&lt;/th&gt;
&lt;th&gt;mean&lt;/th&gt;
&lt;th&gt;max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;poll, 1ms sleep&lt;/td&gt;
&lt;td&gt;0.49–1.65ms&lt;/td&gt;
&lt;td&gt;0.96–1.42ms&lt;/td&gt;
&lt;td&gt;1.97–2.02ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;busy-poll, no sleep&lt;/td&gt;
&lt;td&gt;0.35–0.41ms&lt;/td&gt;
&lt;td&gt;0.39–0.40ms&lt;/td&gt;
&lt;td&gt;0.55–1.04ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;WAIT FOR LSN&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.28–0.36ms&lt;/td&gt;
&lt;td&gt;0.31–0.38ms&lt;/td&gt;
&lt;td&gt;0.56–0.77ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;(Ranges are across three repeated runs of 20 trials each for the sleeping poll, two for the others.)&lt;/p&gt;

&lt;p&gt;Against the sleeping poll, &lt;code&gt;WAIT FOR&lt;/code&gt; looks like a clean 3–5x win. Against the busy-poll, the gap almost disappears — &lt;code&gt;WAIT FOR&lt;/code&gt; is still slightly ahead on every run, but by tenths of a millisecond, not multiples. A busy loop that never sleeps will notice a replayed LSN about as fast as a purpose-built wait primitive will, on one connection, on an idle standby. The 3–5x number is really measuring my arbitrary 1ms sleep, not the feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it actually wins: ten clients, continuously
&lt;/h2&gt;

&lt;p&gt;I ran ten worker threads, each looping insert-on-primary, then wait-on-standby, for five seconds straight, and sampled standby CPU with &lt;code&gt;docker stats&lt;/code&gt; during the run. Averaged over two runs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;method&lt;/th&gt;
&lt;th&gt;completions/sec&lt;/th&gt;
&lt;th&gt;standby CPU (mean)&lt;/th&gt;
&lt;th&gt;standby CPU (peak)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;busy-poll, 10 workers&lt;/td&gt;
&lt;td&gt;3,182&lt;/td&gt;
&lt;td&gt;57%&lt;/td&gt;
&lt;td&gt;67%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;WAIT FOR LSN&lt;/code&gt;, 10 workers&lt;/td&gt;
&lt;td&gt;3,552&lt;/td&gt;
&lt;td&gt;38%&lt;/td&gt;
&lt;td&gt;52%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;More completed work, at roughly two-thirds the mean CPU. This is the number I'd actually plan around: a busy-poll loop is fine for one client checking one write, but it doesn't idle — every waiting connection is a backend spinning a query in a tight cycle, and ten of those add up on the machine serving your replica reads. &lt;code&gt;WAIT FOR&lt;/code&gt; parks the backend instead of spinning it. I only got two or three CPU samples per five-second run because spawning &lt;code&gt;docker stats --no-stream&lt;/code&gt; has its own overhead, so treat the CPU figures as directional rather than to the percentage point — the completions-per-second numbers, counted directly rather than sampled, I trust more.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it refuses
&lt;/h2&gt;

&lt;p&gt;I fed it every wrong input I could think of.&lt;/p&gt;

&lt;p&gt;Garbage LSN:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR:  invalid input syntax for type pg_lsn: "not-a-lsn"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;standby_replay&lt;/code&gt; mode run against the primary itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR:  recovery is not in progress
HINT:  Waiting for the standby_replay LSN can only be executed during recovery.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;primary_flush&lt;/code&gt; mode run against the standby:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR:  recovery is in progress
HINT:  Waiting for primary_flush can only be done on a primary server. Use standby_flush mode on a standby server.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A made-up mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;ERROR&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="n"&gt;unrecognized&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;WAIT&lt;/span&gt; &lt;span class="k"&gt;option&lt;/span&gt; &lt;span class="nv"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;"bogus_mode"&lt;/span&gt;
&lt;span class="n"&gt;LINE&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;WAIT&lt;/span&gt; &lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="n"&gt;LSN&lt;/span&gt; &lt;span class="s1"&gt;'0/3000000'&lt;/span&gt; &lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;MODE&lt;/span&gt; &lt;span class="s1"&gt;'bogus_mode'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;TIMEOUT&lt;/span&gt; &lt;span class="s1"&gt;'1...
                                       ^
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A negative timeout:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR:  timeout cannot be negative
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All four of the role and mode errors carry a HINT that names the fix, which is more than the average Postgres error gives you. &lt;code&gt;primary_flush&lt;/code&gt; mode itself works fine when run in the right place — I confirmed it on the primary directly, waiting on a just-issued LSN, and it returned &lt;code&gt;success&lt;/code&gt; immediately once the WAL was flushed locally.&lt;/p&gt;

&lt;h2&gt;
  
  
  NO_THROW turns the timeout into data
&lt;/h2&gt;

&lt;p&gt;By default, a &lt;code&gt;WAIT FOR&lt;/code&gt; that times out raises an error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ERROR:  timed out while waiting for target LSN FFFFFFFF/FFFFFFFF to be replayed;
current standby_replay LSN 0/03F18920
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add &lt;code&gt;NO_THROW&lt;/code&gt; and the same situation returns a row instead of raising:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; status
---------
 timeout
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful wait returns &lt;code&gt;status = success&lt;/code&gt; the same way. For application code, that's the difference between wrapping every read-your-writes check in exception handling and just branching on a column, which matters if you're calling this from a connection pool path that runs on every request.&lt;/p&gt;

&lt;h2&gt;
  
  
  The documented claim I could check
&lt;/h2&gt;

&lt;p&gt;The reference page says a &lt;code&gt;TIMEOUT&lt;/code&gt; of zero or an omitted &lt;code&gt;TIMEOUT&lt;/code&gt; both mean wait indefinitely, not "don't wait." I set up a standby waiting on an LSN that didn't exist yet, with no &lt;code&gt;TIMEOUT&lt;/code&gt; clause, checked that it was still running after two seconds, then generated enough WAL on the primary to pass that LSN. The wait returned &lt;code&gt;success&lt;/code&gt; right after, in step with the insert, not before it and not with an error. I repeated the same check with &lt;code&gt;TIMEOUT '0'&lt;/code&gt; explicitly and got the same behaviour. Both hold as documented, and neither reads as "zero means immediate timeout," which is the more common convention elsewhere and the one I half-expected to trip over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching it happen
&lt;/h2&gt;

&lt;p&gt;While a session is blocked in &lt;code&gt;WAIT FOR&lt;/code&gt;, another connection can see exactly what it's doing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt; &lt;span class="n"&gt;pid&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="k"&gt;state&lt;/span&gt;  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;wait_event_type&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt;    &lt;span class="n"&gt;wait_event&lt;/span&gt;    &lt;span class="o"&gt;|&lt;/span&gt;                  &lt;span class="n"&gt;query&lt;/span&gt;
&lt;span class="c1"&gt;-----+--------+------------------+-------------------+-------------------------------------------&lt;/span&gt;
 &lt;span class="mi"&gt;110&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;active&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;IPC&lt;/span&gt;              &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;WaitForWalReplay&lt;/span&gt;  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;WAIT&lt;/span&gt; &lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="n"&gt;LSN&lt;/span&gt; &lt;span class="s1"&gt;'0/03F7A450'&lt;/span&gt; &lt;span class="k"&gt;WITH&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;MODE&lt;/span&gt;&lt;span class="p"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;wait_event_type = IPC&lt;/code&gt;, &lt;code&gt;wait_event = WaitForWalReplay&lt;/code&gt;, sitting right there in &lt;code&gt;pg_stat_activity&lt;/code&gt;. A polling loop shows up in that view only for the instant each individual poll query runs; between polls, from the database's point of view, the connection was never doing anything in particular. &lt;code&gt;WAIT FOR&lt;/code&gt; gives you one row that tells you, unambiguously, "this connection is blocked waiting for replication to catch up," which is the kind of thing you want in a slow-query or blocked-session dashboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Concurrent waiters don't queue behind each other
&lt;/h2&gt;

&lt;p&gt;I was slightly worried that ten or more processes all waiting on the same not-yet-issued LSN might be serialised — woken up one at a time as some internal list gets walked. I started 15 connections waiting on a single future LSN, inserted enough rows on the primary to satisfy it, and measured how spread out their wake-up times were.&lt;/p&gt;

&lt;p&gt;Across two runs, the gap between the first and last waiter to return was under 10ms in both, with a standard deviation of 0.5ms and 3.3ms respectively, against absolute wait times in the low hundreds of milliseconds. They wake together, not in a queue.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first version of this post would have led with "&lt;code&gt;WAIT FOR&lt;/code&gt; is 3 to 5 times faster than polling," because that's what the first benchmark showed, and it looked like a clean, presentable number. I only added the busy-poll variant afterwards, as a sanity check on whether I was measuring the mechanism or my own choice of sleep interval. It was the latter. If I hadn't gone back and questioned the first result, the whole post would have overstated the case for a feature that, on the evidence I actually have, is worth adopting for a completely different reason than the one I started out expecting to write about.&lt;/p&gt;

&lt;p&gt;I also lost the first primary container to my own &lt;code&gt;docker rm -f&lt;/code&gt; while trying to add a port mapping I'd forgotten at setup, which meant rebuilding the base backup from scratch. Nothing in that was interesting, but it's why every container in the commands above gets its ports mapped on the very first &lt;code&gt;docker run&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This is the core of the latency comparison, trimmed to what you need; the full concurrency and CPU-sampling scripts are longer but follow the same shape.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;psycopg2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;

&lt;span class="n"&gt;PRIM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;localhost&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;15432&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;postgres&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;password&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;postgres&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dbname&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;postgres&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;STBY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;localhost&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;15433&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;postgres&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;password&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;postgres&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dbname&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;postgres&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;psycopg2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;d&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;autocommit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;

&lt;span class="n"&gt;pc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PRIM&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nf"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;STBY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;pcur&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scur&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;sc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;pcur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;INSERT INTO t(v) VALUES (&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;x&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;) RETURNING id;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;pcur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT pg_current_wal_lsn();&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;lsn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pcur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetchone&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="n"&gt;t0&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;perf_counter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;scur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;WAIT FOR LSN %s WITH (MODE &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;standby_replay&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, TIMEOUT &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;5s&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;);&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lsn&lt;/span&gt;&lt;span class="p"&gt;,)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scur&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetchone&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;perf_counter&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;t0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You need &lt;code&gt;postgres:19beta4&lt;/code&gt; or later on both ends of a streaming-replication pair, &lt;code&gt;psycopg2-binary&lt;/code&gt; on the client, and a table called &lt;code&gt;t&lt;/code&gt; that exists on both sides. I ran this on one machine, so treat the millisecond figures as belonging to that machine. The CPU ratio under ten concurrent workers is the number I'd actually carry over to a different box, because it comes from comparing two ways of using the CPU, not from network distance.&lt;/p&gt;

&lt;p&gt;Postgres 19 is still in beta, due for a release candidate in early October according to the project's own announcement, so nobody should put this on a production standby yet. What I'd do instead is find the polling loop already sitting in a read-replica code path — most teams doing async reads have one somewhere, even if it's disguised as a retry-with-backoff around a stale read — and check how many connections hit it concurrently at peak. If it's one or two, this change buys almost nothing. Past that, the CPU freed up is CPU your standby was supposed to be spending on the reads it exists to serve, not on asking itself the same question forty times a second.&lt;/p&gt;

</description>
      <category>postgres</category>
      <category>database</category>
      <category>devops</category>
      <category>performance</category>
    </item>
    <item>
      <title>containerd 2.2's mount manager panics on a one-mount mkfs chain</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Sun, 27 Sep 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/containerd-22s-mount-manager-panics-on-a-one-mount-mkfs-chain-545n</link>
      <guid>https://dev.to/alexgeorgiev17/containerd-22s-mount-manager-panics-on-a-one-mount-mkfs-chain-545n</guid>
      <description>&lt;p&gt;containerd 2.2 shipped a mount manager: a service that can format a file as ext4 or xfs, attach it as a loopback device, and hand the result to a runtime, all from a single &lt;code&gt;Activate&lt;/code&gt; call instead of the usual &lt;code&gt;truncate&lt;/code&gt;, &lt;code&gt;mkfs&lt;/code&gt;, &lt;code&gt;losetup&lt;/code&gt;, &lt;code&gt;mount&lt;/code&gt; sequence. I wanted to know whether that call is actually faster than doing it by hand, so I wrote a small Go program against the manager's package directly and ran both paths three times each on the same machine.&lt;/p&gt;

&lt;p&gt;The manual sequence came out at 23.5 to 25.9 milliseconds. The mount manager's &lt;code&gt;Activate&lt;/code&gt; came out at 29.6 to 47.2 milliseconds. It was not faster. Along the way it also panicked once, leaked a raw BoltDB error once, and left a loop device attached with nothing able to find it again.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the mount manager actually is
&lt;/h2&gt;

&lt;p&gt;There's no &lt;code&gt;ctr&lt;/code&gt; subcommand for any of this. The manager lives at &lt;code&gt;github.com/containerd/containerd/v2/core/mount/manager&lt;/code&gt; and is meant to be embedded by a snapshotter or a runtime shim, not driven from a terminal. Its job is to let a mount type be built out of steps: a "transformer" can create a file, format it, and format directories, and a "handler" can attach it as a loopback device, before the result gets handed off as an ordinary system mount.&lt;/p&gt;

&lt;p&gt;The container running this test was Docker Engine 29.3.1, whose bundled containerd reports itself as &lt;code&gt;v2.2.2&lt;/code&gt;. I confirmed the mount manager package is at that exact version by pinning it in &lt;code&gt;go.mod&lt;/code&gt; and building against it, not by trusting the daemon's version string.&lt;/p&gt;

&lt;p&gt;A working activation for a 200MiB ext4 image looks like this, once I had the templating right:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;mounts&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="n"&gt;mount&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Mount&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;   &lt;span class="s"&gt;"mkfs/loop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Source&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;imgPath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Options&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="s"&gt;"X-containerd.mkfs.size=200MiB"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="s"&gt;"X-containerd.mkfs.fs=ext4"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;   &lt;span class="s"&gt;"format/ext4"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Source&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"{{ mount 0 }}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;mgr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Activate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"demo1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mounts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Activate&lt;/code&gt; returns two sets of mounts: &lt;code&gt;info.Active&lt;/code&gt;, the ones it handled itself (the loopback attach), and &lt;code&gt;info.System&lt;/code&gt;, the ones it expects the caller to mount with the ordinary &lt;code&gt;mount(2)&lt;/code&gt; syscall (the ext4 filesystem on that loop device). The manager does not mount your rootfs for you. You still call &lt;code&gt;Mount()&lt;/code&gt; on whatever comes back in &lt;code&gt;info.System&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The speed comparison
&lt;/h2&gt;

&lt;p&gt;Both paths format a 200MiB ext4 image, attach it as a loopback device, and mount it. I timed the manager's &lt;code&gt;Activate&lt;/code&gt; call and, separately, the same steps run as four shell commands: &lt;code&gt;truncate -s 200M&lt;/code&gt;, &lt;code&gt;mkfs.ext4 -q -F&lt;/code&gt;, &lt;code&gt;losetup -f --show&lt;/code&gt;, &lt;code&gt;mount&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;th&gt;Run 1&lt;/th&gt;
&lt;th&gt;Run 2&lt;/th&gt;
&lt;th&gt;Run 3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Mount manager &lt;code&gt;Activate&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;29.6ms&lt;/td&gt;
&lt;td&gt;39.0ms&lt;/td&gt;
&lt;td&gt;47.2ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Manual (truncate + mkfs.ext4 + losetup + mount)&lt;/td&gt;
&lt;td&gt;24.4ms&lt;/td&gt;
&lt;td&gt;23.6ms&lt;/td&gt;
&lt;td&gt;25.9ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The manual path was faster on every run. That is not a criticism of the manager's design so much as an observation about what it's for: I checked the source of the &lt;code&gt;mkfs&lt;/code&gt; transformer (&lt;code&gt;core/mount/manager/mkfs.go&lt;/code&gt;) and it shells out to the real &lt;code&gt;mkfs.ext4&lt;/code&gt; and &lt;code&gt;mkfs.xfs&lt;/code&gt; binaries with &lt;code&gt;exec.CommandContext&lt;/code&gt;, exactly what the manual path calls directly. There's no custom fast-path formatter underneath. The manager's job is composability across snapshotters, not raw speed, and on this measurement it costs a small constant overhead (BoltDB writes, symlink creation, bookkeeping) rather than removing any.&lt;/p&gt;

&lt;p&gt;Cleanup showed the same pattern: &lt;code&gt;Deactivate&lt;/code&gt; plus &lt;code&gt;umount&lt;/code&gt; took 32 to 45ms; &lt;code&gt;umount&lt;/code&gt; plus &lt;code&gt;losetup -d&lt;/code&gt; by hand took 11 to 13ms.&lt;/p&gt;

&lt;p&gt;Disk usage was identical either way. &lt;code&gt;du --apparent-size&lt;/code&gt; on both images read 200M; actual usage was 17M on both, since neither &lt;code&gt;mkfs.ext4&lt;/code&gt; nor &lt;code&gt;truncate&lt;/code&gt; writes real data to most of the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first version of the manual comparison used &lt;code&gt;dd if=/dev/zero of=disk.img bs=1M count=200&lt;/code&gt; to create the image before formatting it, because that's the version of this recipe I'd actually run in the past. It came out at 358ms to 1.6 seconds; against that, the mount manager looked ten to fifty times faster, which would have been the headline of this post.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;dd&lt;/code&gt; was writing 200MB of real zero bytes to disk before &lt;code&gt;mkfs.ext4&lt;/code&gt; ever ran. &lt;code&gt;truncate -s 200M&lt;/code&gt; creates the same size file as a sparse hole in under a millisecond, and &lt;code&gt;mkfs.ext4&lt;/code&gt; doesn't need the data pre-zeroed, it only writes its own metadata. The mount manager's &lt;code&gt;mkfs&lt;/code&gt; transformer already does the equivalent of &lt;code&gt;truncate&lt;/code&gt;, via &lt;code&gt;os.OpenFile&lt;/code&gt; plus &lt;code&gt;f.Truncate(size)&lt;/code&gt;, then calls the real &lt;code&gt;mkfs.ext4&lt;/code&gt; binary on the result. Once I made the manual comparison do the same thing, the ten-times "win" disappeared and mildly reversed. The lesson: when a new API bundles three steps into one call, benchmark it against the shortest correct version of those three steps, not the version you happen to type from muscle memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Under concurrency
&lt;/h2&gt;

&lt;p&gt;Ten goroutines calling &lt;code&gt;Activate&lt;/code&gt; in parallel against the same manager instance, each formatting its own 50MiB image, completed in 70.1ms of wall time, with individual calls ranging from 30.5 to 69.8ms. All ten succeeded. If activations were serialized behind a lock, ten of them would have taken close to 300-400ms; they didn't, so the manager's internal locking (&lt;code&gt;RLock&lt;/code&gt; during normal activation, held exclusively only during garbage collection) does allow real concurrent formatting.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it refuses
&lt;/h2&gt;

&lt;p&gt;An unsupported filesystem type is rejected before any file gets created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;unsupported filesystem "btrfs": invalid argument
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A missing size option is also rejected cleanly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mkfs requires mkfs.size option: invalid argument
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A path outside the manager's configured root is rejected too, but with a misleading error class:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;no root "/tmp/not-the-root/disk.img" configured for mkfs: not implemented
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That comes back as &lt;code&gt;errdefs.ErrNotImplemented&lt;/code&gt;, the same error category containerd uses for "this operation genuinely doesn't exist here." Code that checks &lt;code&gt;errdefs.IsNotImplemented()&lt;/code&gt; to decide whether to fall back to a different mount path would treat "you forgot to allow this directory" the same as "this feature isn't built." I read the source (&lt;code&gt;core/mount/manager/mkfs.go&lt;/code&gt;) to confirm this isn't a formatting quirk on my end; the transformer returns exactly that wrapped error whenever the source path doesn't match any configured root.&lt;/p&gt;

&lt;p&gt;Activating a second time under the same name, without deactivating the first, doesn't get a clean "already exists" either:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;bucket already exists
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's a raw &lt;code&gt;bbolt&lt;/code&gt; error surfacing straight from the metadata store, with no containerd-level wrapping. It's accurate, but it tells you about the manager's storage engine rather than about your mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  The panic
&lt;/h2&gt;

&lt;p&gt;The documentation's own examples always chain at least two mounts: something that produces a loopback device, and something that mounts a filesystem on it. I tried activating a single &lt;code&gt;mkfs/loop&lt;/code&gt; mount on its own, with nothing consuming its output, expecting either a successful format-only activation or a clean validation error.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="nb"&gt;panic&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;runtime&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="n"&gt;of&lt;/span&gt; &lt;span class="k"&gt;range&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="n"&gt;with&lt;/span&gt; &lt;span class="n"&gt;length&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt;

&lt;span class="n"&gt;goroutine&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;running&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="n"&gt;github&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;com&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;containerd&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;containerd&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;core&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;mount&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;manager&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;mountManager&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Activate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;...&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="o"&gt;.../&lt;/span&gt;&lt;span class="n"&gt;core&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;mount&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;manager&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;manager&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="k"&gt;go&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="m"&gt;421&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="m"&gt;0x2030&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I read &lt;code&gt;manager.go&lt;/code&gt; to find out why. The function tracks &lt;code&gt;firstSystemMount&lt;/code&gt;, the index of the first mount it expects the caller to handle. When a mount is only ever a transform target (my single &lt;code&gt;mkfs/loop&lt;/code&gt; mount, with no second mount to hand the loop device to), that index gets set to &lt;code&gt;i+1&lt;/code&gt;, which in a one-mount list equals &lt;code&gt;len(mounts)&lt;/code&gt;. A later loop indexes into &lt;code&gt;mountConv[firstSystemMount]&lt;/code&gt; to apply any pending format templating, and &lt;code&gt;mountConv&lt;/code&gt; was allocated with &lt;code&gt;len(mounts)&lt;/code&gt; elements. Index &lt;code&gt;1&lt;/code&gt; into a slice of length &lt;code&gt;1&lt;/code&gt; panics.&lt;/p&gt;

&lt;p&gt;This matters beyond the crash itself. The loopback device gets attached to the backing file before the panic point, and because &lt;code&gt;Activate&lt;/code&gt; never returns successfully, nothing gets written to the manager's BoltDB. There is no record of this activation to recover.&lt;/p&gt;

&lt;h2&gt;
  
  
  Crash recovery, and where it doesn't reach
&lt;/h2&gt;

&lt;p&gt;The manager persists activation state in BoltDB specifically so a restarted process can find and clean up mounts from before a crash. I tested that separately from the panic: one process called &lt;code&gt;Activate&lt;/code&gt; and then exited hard with &lt;code&gt;os.Exit(0)&lt;/code&gt;, skipping &lt;code&gt;Deactivate&lt;/code&gt; entirely, to simulate a daemon that died mid-operation.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;./mmdemo &lt;span class="nt"&gt;-mode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;crash-activate
&lt;span class="gp"&gt;activated in 30.7ms, err=&amp;lt;nil&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="go"&gt;
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;./mmdemo &lt;span class="nt"&gt;-mode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;crash-recover
&lt;span class="gp"&gt;List() after simulated crash: 1 activations, err=&amp;lt;nil&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="go"&gt;  recovered activation: crashy active=[{loop ... /tmp/mm-crash/targets/1/1}]
&lt;/span&gt;&lt;span class="gp"&gt;cleanup via Deactivate: &amp;lt;nil&amp;gt;&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A second process, pointed at the same database and target directory, listed the orphaned activation and deactivated it cleanly, and the loop device it had been holding was released. That documented claim held up.&lt;/p&gt;

&lt;p&gt;But it only works because the first &lt;code&gt;Activate&lt;/code&gt; call returned successfully and got committed. Compare that against the one-mount panic above: I left that test running separately, and the loop device it opened is still attached to a deleted backing file with no BoltDB record anywhere pointing at it. &lt;code&gt;losetup -a&lt;/code&gt; still shows it. No amount of restarting a mount manager pointed at any database will find it, because it was never written down. The crash-recovery mechanism protects against a daemon dying after an activation completes. It has no way to protect against the daemon crashing during one.&lt;/p&gt;

&lt;h2&gt;
  
  
  xfs, and a size limit that isn't the manager's
&lt;/h2&gt;

&lt;p&gt;I ran the same &lt;code&gt;mkfs/loop&lt;/code&gt; chain with &lt;code&gt;X-containerd.mkfs.fs=xfs&lt;/code&gt; at 200MiB and got a full &lt;code&gt;mkfs.xfs&lt;/code&gt; usage message back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mkfs.xfs failed: Filesystem must be larger than 300MB.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's &lt;code&gt;mkfs.xfs&lt;/code&gt; itself refusing, not the manager. At 400MiB, formatting succeeded in 48 to 72ms. Mounting the result failed in this environment specifically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mount source: ".../targets/1/1", target: ".../rootfs", fstype: xfs, flags: 0, data: "", err: no such device
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;/proc/filesystems&lt;/code&gt; on this container's kernel has no &lt;code&gt;xfs&lt;/code&gt; entry and there's no &lt;code&gt;modprobe&lt;/code&gt; to load one. That's a property of the sandbox this test ran in, not of containerd, and I'm noting it rather than counting it as a finding against the feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This needs Go 1.24+, root (for loopback devices and mounts), and &lt;code&gt;mkfs.ext4&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir &lt;/span&gt;mounttest &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;mounttest
go mod init mounttest
go get github.com/containerd/containerd/v2@v2.2.2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Save as &lt;code&gt;main.go&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;package&lt;/span&gt; &lt;span class="n"&gt;main&lt;/span&gt;

&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="s"&gt;"context"&lt;/span&gt;
    &lt;span class="s"&gt;"fmt"&lt;/span&gt;
    &lt;span class="s"&gt;"os"&lt;/span&gt;

    &lt;span class="s"&gt;"github.com/containerd/containerd/v2/core/mount"&lt;/span&gt;
    &lt;span class="s"&gt;"github.com/containerd/containerd/v2/core/mount/manager"&lt;/span&gt;
    &lt;span class="s"&gt;"github.com/containerd/containerd/v2/pkg/namespaces"&lt;/span&gt;
    &lt;span class="n"&gt;bolt&lt;/span&gt; &lt;span class="s"&gt;"go.etcd.io/bbolt"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="s"&gt;"/tmp/mm-demo"&lt;/span&gt;
    &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RemoveAll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MkdirAll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="s"&gt;"/targets"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="m"&gt;0755&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;bolt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="s"&gt;"/meta.db"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="m"&gt;0644&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;defer&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Close&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;mgr&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;manager&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NewManager&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="s"&gt;"/targets"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;manager&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;WithMountHandler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"loop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mount&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;LoopbackHandler&lt;/span&gt;&lt;span class="p"&gt;()),&lt;/span&gt;
        &lt;span class="n"&gt;manager&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;WithAllowedRoot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nb"&gt;panic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;namespaces&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;WithNamespace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Background&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="s"&gt;"default"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;mounts&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="n"&gt;mount&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Mount&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"mkfs/loop"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Source&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;"/disk.img"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Options&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="s"&gt;"X-containerd.mkfs.size=200MiB"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="s"&gt;"X-containerd.mkfs.fs=ext4"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;}},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;Type&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"format/ext4"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Source&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"{{ mount 0 }}"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;mgr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Activate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"demo1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mounts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Printf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"info=%+v err=%v&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MkdirAll&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="s"&gt;"/rootfs"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="m"&gt;0755&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sm&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="k"&gt;range&lt;/span&gt; &lt;span class="n"&gt;info&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;System&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;sm&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Mount&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="s"&gt;"/rootfs"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Println&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"mount failed:"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c"&gt;// activating the same name again without deactivating first:&lt;/span&gt;
    &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;mgr&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Activate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"demo1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mounts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Println&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"second activate, same name:"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;go build &lt;span class="nt"&gt;-o&lt;/span&gt; mmdemo &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;span class="nb"&gt;sudo&lt;/span&gt; ./mmdemo
mount | &lt;span class="nb"&gt;grep &lt;/span&gt;mm-demo
&lt;span class="nb"&gt;sudo &lt;/span&gt;umount /tmp/mm-demo/rootfs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I verified every command in this section on a fresh checkout before publishing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this
&lt;/h2&gt;

&lt;p&gt;If you're building on the mount manager today, treat it as a composability primitive, not a performance one; it won't beat a shell one-liner on latency. Never construct a mount list that ends in a transform-only type like &lt;code&gt;mkfs/loop&lt;/code&gt; without a mount that actually consumes its output, until this specific crash is fixed upstream. Don't branch on &lt;code&gt;errdefs.IsNotImplemented()&lt;/code&gt; from this package without also checking the error text, because it currently covers both "unsupported" and "not configured." And if you're relying on its crash recovery for anything in production, test it against a process that dies mid-&lt;code&gt;Activate&lt;/code&gt;, not just one that dies after — those are different guarantees, and only one of them is covered right now.&lt;/p&gt;

</description>
      <category>containerd</category>
      <category>docker</category>
      <category>devops</category>
      <category>go</category>
    </item>
    <item>
      <title>systemd's mstack tool mounts a single-layer directory writable by default</title>
      <dc:creator>Alex Georgiev</dc:creator>
      <pubDate>Sat, 26 Sep 2026 10:00:00 +0000</pubDate>
      <link>https://dev.to/alexgeorgiev17/systemds-mstack-tool-mounts-a-single-layer-directory-writable-by-default-48p</link>
      <guid>https://dev.to/alexgeorgiev17/systemds-mstack-tool-mounts-a-single-layer-directory-writable-by-default-48p</guid>
      <description>&lt;p&gt;I built a two-file directory, mounted it with systemd's new &lt;code&gt;mstack&lt;/code&gt; tool, wrote one line into it, and watched that line permanently overwrite a file I'd marked read-only. No error, no warning, exit code 0.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;mstack&lt;/code&gt; is one of the features systemd picked up this year: a way to describe an overlayfs mount — several read-only layers plus one writable layer — as an ordinary directory full of symlinks, instead of a long &lt;code&gt;mount -t overlay -o lowerdir=...&lt;/code&gt; command line. It landed in systemd 260 in March and &lt;code&gt;systemd-nspawn&lt;/code&gt; grew more support for it in 261, released in June and now the current stable release (261.3 is what ships in today's Arch Linux image). I ran everything below against that build, in Docker, since my sandbox has no systemd running as PID 1.&lt;/p&gt;

&lt;p&gt;The idea is genuinely useful: put symlinks named &lt;code&gt;layer@0&lt;/code&gt;, &lt;code&gt;layer@1&lt;/code&gt; and so on in a directory suffixed &lt;code&gt;.mstack&lt;/code&gt;, add a &lt;code&gt;rw/&lt;/code&gt; subdirectory for the writable top, and &lt;code&gt;systemd-mstack --mount that.mstack /somewhere&lt;/code&gt; assembles the whole overlay in one call. &lt;code&gt;systemd-nspawn --mstack=&lt;/code&gt;, and a service's &lt;code&gt;RootMStack=&lt;/code&gt; directive, can point at the same directory. It's meant to make mount stacks shareable and inspectable rather than baked into a shell script.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building one that works
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; base app app.mstack
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"base content"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; base/f.txt
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"app content"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; app/g.txt
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;ln&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; ../base app.mstack/layer@0
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;ln&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; ../app app.mstack/layer@1
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;mkdir &lt;/span&gt;app.mstack/rw
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack app.mstack
&lt;span class="go"&gt;TYPE  NAME    IMAGE     WHAT       WHERE SORT
layer layer@0 directory /work/base /     0
layer layer@1 directory /work/app  /     1
rw    rw      directory /work/app.mstack/rw /     -
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack &lt;span class="nt"&gt;--mount&lt;/span&gt; app.mstack /mnt/target
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; /mnt/target
&lt;span class="go"&gt;f.txt  g.txt
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That worked first try, on a tmpfs-backed directory. Both layers' files showed up merged, as expected.&lt;/p&gt;

&lt;h2&gt;
  
  
  The overlay-on-overlay refusal, and it isn't mstack's fault
&lt;/h2&gt;

&lt;p&gt;My first attempt wasn't on tmpfs, it was in a plain Docker container's own filesystem, which uses the &lt;code&gt;overlay2&lt;/code&gt; storage driver by default. There, the exact same &lt;code&gt;app.mstack&lt;/code&gt; directory produced this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack &lt;span class="nt"&gt;--mount&lt;/span&gt; app.mstack /mnt/target
&lt;span class="go"&gt;'(layerfd)' failed with exit status 1.
Failed to apply .mstack/ directory '/work/app.mstack': Invalid argument
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That told me nothing useful, and for a while I assumed &lt;code&gt;mstack&lt;/code&gt; itself was broken or that I'd built the directory wrong. It wasn't. I ran the equivalent hand-written command against the same two directories:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;mount &lt;span class="nt"&gt;-t&lt;/span&gt; overlay overlay &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;lowerdir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;app:base,upperdir&lt;span class="o"&gt;=&lt;/span&gt;upper,workdir&lt;span class="o"&gt;=&lt;/span&gt;workdir merged
&lt;span class="go"&gt;mount: /work/merged: fsconfig() failed: overlay: filesystem on upper not supported as upperdir.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same failure, plainer error. The Linux kernel won't let you stack a new overlayfs mount on top of directories that are themselves already on an overlayfs — which is exactly what a Docker container's root filesystem is under the default &lt;code&gt;overlay2&lt;/code&gt; driver. This isn't new in 261 and it isn't &lt;code&gt;mstack&lt;/code&gt;'s doing; the manual, decade-old mount command hits the identical wall. &lt;code&gt;mstack&lt;/code&gt; just explains it worse: it names an internal function, &lt;code&gt;(layerfd)&lt;/code&gt;, instead of naming the actual constraint. Anyone testing this feature inside a plain Docker container, which is most people's first instinct, will hit it and get the confusing version.&lt;/p&gt;

&lt;p&gt;The workaround is the same for both: put the layers on a filesystem that isn't itself overlayfs. A tmpfs works. So does a bind-mounted host directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;docker run &lt;span class="nt"&gt;--privileged&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; /some/host/dir:/work archlinux bash
&lt;span class="gp"&gt;#&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;inside&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack &lt;span class="nt"&gt;--mount&lt;/span&gt; app.mstack /mnt/target
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo &lt;/span&gt;EXIT:&lt;span class="nv"&gt;$?&lt;/span&gt;
&lt;span class="go"&gt;EXIT:0
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I confirmed this against an ext-family host directory and it mounted cleanly, no tmpfs required.&lt;/p&gt;

&lt;h2&gt;
  
  
  The version-sort claim held up
&lt;/h2&gt;

&lt;p&gt;The documentation says layer ordering is decided by a version sort, not a plain string sort, specifically so &lt;code&gt;layer@10&lt;/code&gt; doesn't get treated as coming before &lt;code&gt;layer@2&lt;/code&gt;. I built layers &lt;code&gt;layer@1&lt;/code&gt; through &lt;code&gt;layer@3&lt;/code&gt; and &lt;code&gt;layer@10&lt;/code&gt; through &lt;code&gt;layer@12&lt;/code&gt; and asked for the JSON view:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;systemd-mstack&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;vsort.mstack&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;--json=short&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"layer@1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"sort"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;},{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"layer@2"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"sort"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"2"&lt;/span&gt;&lt;span class="p"&gt;},{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"layer@3"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"sort"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"3"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"layer@10"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"sort"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"10"&lt;/span&gt;&lt;span class="p"&gt;},{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"layer@11"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"sort"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"11"&lt;/span&gt;&lt;span class="p"&gt;},{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"layer@12"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"sort"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"12"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A plain lexicographic sort of those strings would put &lt;code&gt;layer@10&lt;/code&gt; right after &lt;code&gt;layer@1&lt;/code&gt;. It didn't. The claim checks out.&lt;/p&gt;

&lt;h2&gt;
  
  
  No speed advantage, and no extra disk cost
&lt;/h2&gt;

&lt;p&gt;I timed five mounts each of the same two-layer stack, &lt;code&gt;mstack&lt;/code&gt; against the hand-written &lt;code&gt;mount -t overlay&lt;/code&gt; command, both on tmpfs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;method&lt;/th&gt;
&lt;th&gt;run 1&lt;/th&gt;
&lt;th&gt;run 2&lt;/th&gt;
&lt;th&gt;run 3&lt;/th&gt;
&lt;th&gt;run 4&lt;/th&gt;
&lt;th&gt;run 5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;systemd-mstack --mount&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5ms&lt;/td&gt;
&lt;td&gt;4ms&lt;/td&gt;
&lt;td&gt;4ms&lt;/td&gt;
&lt;td&gt;4ms&lt;/td&gt;
&lt;td&gt;5ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;mount -t overlay&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4ms&lt;/td&gt;
&lt;td&gt;3ms&lt;/td&gt;
&lt;td&gt;3ms&lt;/td&gt;
&lt;td&gt;3ms&lt;/td&gt;
&lt;td&gt;3ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Roughly a millisecond slower, consistently, which is the extra work of resolving symlinks and building the option string. At this scale it's noise, not a cost anyone would notice. The value of &lt;code&gt;mstack&lt;/code&gt; isn't speed. Anyone expecting a performance win from the new tool won't find one here — it's an ergonomics and inspectability feature, not a faster overlay.&lt;/p&gt;

&lt;p&gt;I also put a 100MB file inside one layer and mounted it, to check whether &lt;code&gt;mstack&lt;/code&gt; copies layer content anywhere before mounting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;dd &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/dev/zero &lt;span class="nv"&gt;of&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;biglayer/bigfile &lt;span class="nv"&gt;bs&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1M &lt;span class="nv"&gt;count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;100
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack &lt;span class="nt"&gt;--mount&lt;/span&gt; biglayer.mstack /mnt/bigmount
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;du&lt;/span&gt; &lt;span class="nt"&gt;-sh&lt;/span&gt; /tmp/mstack-temporary-&lt;span class="k"&gt;*&lt;/span&gt;
&lt;span class="go"&gt;40      /tmp/mstack-temporary-focKQP
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;stat&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'%s bytes'&lt;/span&gt; /mnt/bigmount/bigfile
&lt;span class="go"&gt;104857600 bytes
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The temporary staging directory &lt;code&gt;mstack&lt;/code&gt; creates is 40 bytes, not 100 megabytes. The file is fully visible through the mount, so nothing was lost, and nothing was duplicated. That's a fair result and worth stating plainly since a "self-describing" abstraction is exactly the kind of thing I'd expect to add a copy step somewhere.&lt;/p&gt;

&lt;p&gt;One thing I couldn't explain: &lt;code&gt;systemd-mstack --mount&lt;/code&gt; on a genuine two-layer stack shows two identical &lt;code&gt;lowerdir+=&lt;/code&gt; entries pointing at the same temporary path in &lt;code&gt;/proc/self/mountinfo&lt;/code&gt;, rather than two distinct paths for the two distinct source layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;overlay /work/app.mstack rw,lowerdir+=/tmp/mstack-temporary-focKQP,lowerdir+=/tmp/mstack-temporary-focKQP,upperdir=/tmp/mstack-temporary-focKQP/data,...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mount worked and both layers' files were present in the result, so it isn't a functional bug I could demonstrate. I just don't know why the kernel reports it that way, and I'd rather say that than guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it silently gives you nothing
&lt;/h2&gt;

&lt;p&gt;Two failure modes produced no error text at all, which is worse than the opaque &lt;code&gt;(layerfd)&lt;/code&gt; message above.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;.mstack&lt;/code&gt; directory with a symlink pointing nowhere:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;ln&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; ../does-not-exist bad.mstack/layer@0
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack bad.mstack
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo &lt;/span&gt;EXIT:&lt;span class="nv"&gt;$?&lt;/span&gt;
&lt;span class="go"&gt;EXIT:1
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole output. &lt;code&gt;--show&lt;/code&gt; gives you an exit code and nothing to act on. Only &lt;code&gt;--mount&lt;/code&gt; on the same broken directory names the problem, and only because mounting forces it to actually try to open the target:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack &lt;span class="nt"&gt;--mount&lt;/span&gt; bad.mstack /mnt/bad
&lt;span class="go"&gt;Failed to apply .mstack/ directory '/work/bad.mstack': No such file or directory
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A stray file or directory inside a &lt;code&gt;.mstack/&lt;/code&gt; that doesn't match &lt;code&gt;layer@*&lt;/code&gt; or &lt;code&gt;rw&lt;/code&gt; does the same thing: both the plain and &lt;code&gt;--json&lt;/code&gt; show commands exit 1 with no message. If you're building these directories by a script and something goes wrong upstream — a template that leaves a stray file behind, say — you get a bare failure and have to guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that would actually catch someone out
&lt;/h2&gt;

&lt;p&gt;This is the finding I'd act on. A &lt;code&gt;.mstack&lt;/code&gt; directory with exactly one layer, no &lt;code&gt;rw/&lt;/code&gt; subdirectory, and no &lt;code&gt;--read-only&lt;/code&gt; flag does not refuse to mount and does not fall back to a safe, discardable overlay. It bind-mounts the source directory directly, writable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;ln&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; ../base noRw.mstack/layer@0
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack &lt;span class="nt"&gt;--mount&lt;/span&gt; noRw.mstack /mnt/norw
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;mount | &lt;span class="nb"&gt;grep &lt;/span&gt;norw
&lt;span class="go"&gt;tmpfs on /mnt/norw type tmpfs (rw,relatime)
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"MODIFIED via mount"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /mnt/norw/f.txt
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"written via mount"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /mnt/norw/new-from-mount.txt
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack &lt;span class="nt"&gt;--umount&lt;/span&gt; /mnt/norw
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;cat &lt;/span&gt;base/f.txt
&lt;span class="go"&gt;MODIFIED via mount
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;ls &lt;/span&gt;base/
&lt;span class="go"&gt;f.txt  new-from-mount.txt
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;base/f.txt&lt;/code&gt; on disk now says "MODIFIED via mount", permanently, and the new file is sitting there too, after unmount. The mount type even reports as &lt;code&gt;tmpfs&lt;/code&gt;, not &lt;code&gt;overlay&lt;/code&gt;, which is the tell: with a single layer and no upper directory, &lt;code&gt;mstack&lt;/code&gt; skips building an overlay at all and just bind-mounts the layer as-is. Nothing about the directory name, &lt;code&gt;layer@0&lt;/code&gt;, or the general framing of "layers" as something you stack read-only content underneath, warns you that this specific shape writes straight through.&lt;/p&gt;

&lt;p&gt;I checked whether the same thing happens with two layers and no &lt;code&gt;rw/&lt;/code&gt;, since that felt like the more likely real-world mistake — someone forgetting the writable directory on a proper multi-layer stack:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;systemd-mstack &lt;span class="nt"&gt;--mount&lt;/span&gt; twolayer.mstack /mnt/two
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;mount | &lt;span class="nb"&gt;grep&lt;/span&gt; /mnt/two
&lt;span class="go"&gt;overlay ... (ro,relatime,...)
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"written"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /mnt/two/newfile.txt
&lt;span class="go"&gt;bash: /mnt/two/newfile.txt: Read-only file system
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With two or more layers, the same omission correctly falls back to a read-only overlay. The unsafe default is specific to the single-layer case, where &lt;code&gt;mstack&lt;/code&gt; decides an overlay isn't necessary and hands you the raw directory instead. That's a narrow trigger, but a single read-only "base" layer referenced from several application stacks — the exact use case the tool's docs describe — is precisely the shape where this would fire.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I got wrong on the way
&lt;/h2&gt;

&lt;p&gt;My first read of the &lt;code&gt;(layerfd)&lt;/code&gt; error was that &lt;code&gt;mstack&lt;/code&gt; had a bug, or that I'd built my test directory wrong, and I spent time double-checking symlink targets and directory names before it occurred to me to try the plain &lt;code&gt;mount -t overlay&lt;/code&gt; command against the same two directories. Once that failed with the same underlying cause and a clearer message, it was obvious the fault was the container's storage driver, not the tool. I should have reached for that control test first rather than last.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;This needs systemd 260 or newer for the &lt;code&gt;systemd-mstack&lt;/code&gt; binary. Arch Linux's current image ships 261.3, which is what I used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;--privileged&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;pwd&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;/work:/work"&lt;/span&gt; archlinux:latest bash &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'
  cd /work
  mkdir -p base app app.mstack
  echo "base content" &amp;gt; base/f.txt
  echo "app content" &amp;gt; app/g.txt
  ln -s ../base app.mstack/layer@0
  ln -s ../app app.mstack/layer@1
  mkdir app.mstack/rw
  mkdir -p /mnt/target
  systemd-mstack --mount app.mstack /mnt/target
  ls /mnt/target
  systemd-mstack --umount /mnt/target
'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Bind-mounting a real host directory as &lt;code&gt;/work&lt;/code&gt;, as above, sidesteps the overlay-on-overlay refusal. Drop the &lt;code&gt;-v&lt;/code&gt; and it will fail with the &lt;code&gt;(layerfd)&lt;/code&gt; error on most default Docker setups, which is itself worth seeing once.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this
&lt;/h2&gt;

&lt;p&gt;If you're packaging application images with a single, shared read-only base and building &lt;code&gt;.mstack&lt;/code&gt; directories by script, add the &lt;code&gt;rw/&lt;/code&gt; subdirectory even when you intend the mount to be read-only, and pass &lt;code&gt;--read-only&lt;/code&gt; explicitly rather than relying on the absence of a writable layer to protect anything. Don't test this feature for the first time inside a stock Docker container without bind-mounting a real directory in — you'll spend time chasing a kernel limitation that has nothing to do with the tool. And if a &lt;code&gt;.mstack&lt;/code&gt; directory fails to show or mount with no message at all, check for a broken symlink or a stray file before assuming the tool is the problem; in my testing that combination produced silence rather than a diagnostic every time.&lt;/p&gt;

&lt;p&gt;I didn't get as far as testing &lt;code&gt;RootMStack=&lt;/code&gt; from an actual service unit or &lt;code&gt;systemd-nspawn --mstack=&lt;/code&gt; end to end, because that needs systemd running as PID 1, which my sandbox doesn't have. That's the natural next thing to check.&lt;/p&gt;

</description>
      <category>systemd</category>
      <category>linux</category>
      <category>docker</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
