<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mustafa ERBAY</title>
    <description>The latest articles on DEV Community by Mustafa ERBAY (@merbayerp).</description>
    <link>https://dev.to/merbayerp</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3921203%2Fe3a198a1-49a0-466f-99e6-74bdf202a867.png</url>
      <title>DEV Community: Mustafa ERBAY</title>
      <link>https://dev.to/merbayerp</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/merbayerp"/>
    <language>en</language>
    <item>
      <title>CPU Quota Does Not Look at Your Average</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Sun, 04 Oct 2026 15:00:33 +0000</pubDate>
      <link>https://dev.to/merbayerp/cpu-quota-does-not-look-at-your-average-1n5j</link>
      <guid>https://dev.to/merbayerp/cpu-quota-does-not-look-at-your-average-1n5j</guid>
      <description>&lt;p&gt;It is very easy to look at a container's CPU graph and conclude the thing is loafing. Average&lt;br&gt;
consumption: a fifth of its quota. Flat line. No alerts.&lt;/p&gt;

&lt;p&gt;That same container was stopped 509,912 times over the past ten days.&lt;/p&gt;

&lt;p&gt;These two sentences do not contradict each other. The contradiction is in what I was looking at.&lt;br&gt;
Linux's CFS quota mechanism has no interest whatsoever in your average; it cares about 100&lt;br&gt;
millisecond windows. And for years, every time I typed &lt;code&gt;--cpus&lt;/code&gt; on a container, I was not saying&lt;br&gt;
"half a core per second" — I was saying "50 ms in every 100 ms, no matter what." Those are not&lt;br&gt;
the same thing. In the &lt;a href="https://mustafaerbay.com.tr/en/blog/tutorials/vpste-kaynak-canavar-container-tespit-edip-limitleme/" rel="noopener noreferrer"&gt;resource-hog container guide&lt;/a&gt;&lt;br&gt;
I wrote back in May, I handed out the &lt;code&gt;--cpus "0.5"&lt;/code&gt; example without a second thought; the advice&lt;br&gt;
in that guide is not wrong, but this measurement showed me it was incomplete.&lt;/p&gt;
&lt;h2&gt;
  
  
  I swept the fleet and found 16 quota'd cgroups
&lt;/h2&gt;

&lt;p&gt;On VPS3 (Ubuntu, &lt;code&gt;6.8.0-142-generic&lt;/code&gt;, 18 cores) I walked every cgroup that had a quota and read&lt;br&gt;
its &lt;code&gt;cpu.stat&lt;/code&gt;. In all 16 cgroups where I had set a quota, &lt;code&gt;cpu.max.burst&lt;/code&gt; is &lt;code&gt;0&lt;/code&gt; and the&lt;br&gt;
&lt;code&gt;nr_bursts&lt;/code&gt; counter is &lt;code&gt;0&lt;/code&gt;. I will get to why that surprised me shortly.&lt;/p&gt;

&lt;p&gt;The resulting table gave me pause, because the two most heavily throttled services were the two&lt;br&gt;
with the least tolerance for latency. Top five by throttle count:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;cgroup&lt;/th&gt;
&lt;th&gt;quota&lt;/th&gt;
&lt;th&gt;periods&lt;/th&gt;
&lt;th&gt;throttled&lt;/th&gt;
&lt;th&gt;throttle rate&lt;/th&gt;
&lt;th&gt;throttled time&lt;/th&gt;
&lt;th&gt;avg. quota use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;vpsman-etcd&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.5 CPU&lt;/td&gt;
&lt;td&gt;9,297,557&lt;/td&gt;
&lt;td&gt;509,912&lt;/td&gt;
&lt;td&gt;5.48%&lt;/td&gt;
&lt;td&gt;14.24 hours&lt;/td&gt;
&lt;td&gt;17.55%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;licman-etcd&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.5 CPU&lt;/td&gt;
&lt;td&gt;5,195,009&lt;/td&gt;
&lt;td&gt;308,177&lt;/td&gt;
&lt;td&gt;5.93%&lt;/td&gt;
&lt;td&gt;8.45 hours&lt;/td&gt;
&lt;td&gt;19.82%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;messageman-api-1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1 CPU&lt;/td&gt;
&lt;td&gt;3,336,709&lt;/td&gt;
&lt;td&gt;7,270&lt;/td&gt;
&lt;td&gt;0.22%&lt;/td&gt;
&lt;td&gt;18 minutes&lt;/td&gt;
&lt;td&gt;3.46%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;burncpu-app&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2 CPU&lt;/td&gt;
&lt;td&gt;125,783&lt;/td&gt;
&lt;td&gt;1,725&lt;/td&gt;
&lt;td&gt;1.37%&lt;/td&gt;
&lt;td&gt;22 minutes&lt;/td&gt;
&lt;td&gt;12.24%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;messageman-pg-1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1 CPU&lt;/td&gt;
&lt;td&gt;201,165&lt;/td&gt;
&lt;td&gt;813&lt;/td&gt;
&lt;td&gt;0.40%&lt;/td&gt;
&lt;td&gt;41 seconds&lt;/td&gt;
&lt;td&gt;5.16%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;etcd. The consensus ledger. The thing that triggers a leader election when heartbeats are late,&lt;br&gt;
the thing every write waits on. Because its average load is 17% of quota, I had assumed for years&lt;br&gt;
that "half a core is more than enough." Technically true. In practice it stopped etcd half a&lt;br&gt;
million times over ten days.&lt;/p&gt;

&lt;p&gt;Cumulative counters tell you about the past, not about today. So I read &lt;code&gt;cpu.stat&lt;/code&gt; twice, 120&lt;br&gt;
seconds apart:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;vpsman-etcd, 120-second window
  periods     : 1191   (= 119.1 s — the timer almost never goes idle)
  throttled   : 107 times (8.98% of periods)
  throttled time : 12.42 s
  CPU used       : 13.60 s = 22.84% of its quota
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One hundred and seven times in two minutes. While using a quarter of its quota. No average-based&lt;br&gt;
dashboard will show you this, because the average genuinely is low — the problem is that the&lt;br&gt;
average is the wrong question.&lt;/p&gt;
&lt;h2&gt;
  
  
  The window is 100 ms, and there is no average in there
&lt;/h2&gt;

&lt;p&gt;The mechanism is stated plainly in the kernel docs: within each period the group is allocated&lt;br&gt;
&lt;code&gt;quota&lt;/code&gt; microseconds of CPU, and once that is exhausted its threads are throttled until the&lt;br&gt;
period refreshes. The default period is 100 ms, the default quota is unconstrained, the minimum&lt;br&gt;
quota and period are 1 ms, and the maximum period is 1 second.&lt;/p&gt;

&lt;p&gt;Wondering whether quota carries across periods, I went to the "Caveats" section of the docs —&lt;br&gt;
because I was about to write "it never carries," and I would have been wrong. What it says is&lt;br&gt;
that once a slice is assigned to a CPU it does &lt;strong&gt;not&lt;/strong&gt; expire; if every thread on that CPU&lt;br&gt;
becomes unrunnable, all but 1 ms of the slice may be returned to the global pool. So there is a&lt;br&gt;
small carryover — "typically at most 1 ms per cpu or as defined by &lt;code&gt;min_cfs_rq_runtime&lt;/code&gt;." In the&lt;br&gt;
kernel source that constant sits there as &lt;code&gt;1 * NSEC_PER_MSEC&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The nuance matters, and so does its scale: on an 18-core machine that is a few milliseconds of&lt;br&gt;
residue at best. Rounding error next to an 80 ms peak demand. The same paragraph also names what&lt;br&gt;
this carryover is meant to eliminate: "the propensity to throttle these applications while&lt;br&gt;
simultaneously using less than quota amounts of cpu." The exact phenomenon these counters record,&lt;br&gt;
named in the kernel's own documentation — and still happening on my server.&lt;/p&gt;

&lt;p&gt;It is worth doing the arithmetic once. A half-core quota grants 50 ms per period. If your job&lt;br&gt;
needs 80 ms of CPU in one go, it must be split no matter what your average consumption is —&lt;br&gt;
because what is missing is not the budget but its &lt;em&gt;distribution&lt;/em&gt;. &lt;code&gt;vpsman-etcd&lt;/code&gt; uses 17% of its&lt;br&gt;
quota on average; across 10.76 days of periods it never spent 82% of its entitlement, close to&lt;br&gt;
nine days' worth.&lt;/p&gt;

&lt;p&gt;Which is exactly what happened to etcd. For a &lt;code&gt;read-only range&lt;/code&gt; request — and for a request this&lt;br&gt;
small, at that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{"level":"warn","caller":"txn/util.go:93","msg":"apply request took too long",
 "took":"220.804807ms","expected-duration":"100ms",
 "prefix":"read-only range ","request":"key:\"health\" "}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reading the &lt;code&gt;health&lt;/code&gt; key took 220 ms. The work itself takes microseconds. The rest is waiting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four ways to misread cpu.stat
&lt;/h2&gt;

&lt;p&gt;I tripped over these counters four times. I could not be confident until I opened the kernel&lt;br&gt;
source (&lt;code&gt;kernel/sched/fair.c&lt;/code&gt;, v6.10) and found the lines — and all four run against intuition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;nr_periods&lt;/code&gt; is not wall-clock time.&lt;/strong&gt; The period timer is deactivated when the group goes&lt;br&gt;
idle; the comment at the head of &lt;code&gt;do_sched_cfs_period_timer&lt;/code&gt; says so explicitly. So&lt;br&gt;
&lt;code&gt;nr_periods × 100 ms&lt;/code&gt; is not elapsed time, it is "time during which there was work." In a lab run&lt;br&gt;
that lasted 124 seconds, the counter showed 750 periods (75 seconds). Inverted, it makes a nice&lt;br&gt;
diagnostic: for &lt;code&gt;vpsman-etcd&lt;/code&gt;, 9,297,557 periods works out to 10.76 days against an uptime of&lt;br&gt;
10.80 days — across ten days there was only about an hour in which it had no runnable work.&lt;br&gt;
&lt;code&gt;licman-etcd&lt;/code&gt; passes the same test independently: 6.01 days against 6.04.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;throttled_usec&lt;/code&gt; can exceed wall-clock time.&lt;/strong&gt; The counter is kept per cgroup but accumulated&lt;br&gt;
separately as each CPU's queue is released: &lt;code&gt;cfs_b-&amp;gt;throttled_time += rq_clock(rq) -&lt;br&gt;
cfs_rq-&amp;gt;throttled_clock&lt;/code&gt;. If the group is throttled on four cores simultaneously, 100 ms of real&lt;br&gt;
waiting adds 400 ms to the counter. The measurement itself shows this: dividing 51,261 seconds by&lt;br&gt;
509,912 throttle events gives 100.5 ms per event, which is above the period length — and since no&lt;br&gt;
single period can hold more waiting than that, the excess comes from summing across cores. So the&lt;br&gt;
"14.24 hours" in the table is not real waiting, it is an &lt;strong&gt;upper bound&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The counters are not hierarchical.&lt;/strong&gt; The docs cover this in a single sentence, but the&lt;br&gt;
consequence is heavy: those five CFS fields account only for throttling caused by that cgroup's&lt;br&gt;
&lt;em&gt;own&lt;/em&gt; bandwidth limit. If the quota lives on a parent slice (&lt;code&gt;kubepods.slice&lt;/code&gt; under Kubernetes, a&lt;br&gt;
parent unit under systemd), the container's own &lt;code&gt;cpu.stat&lt;/code&gt; looks spotless. In September, when&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/technology/cgroup-pressure-0-yazdim-86-ms-durma-deftere-girmedi/" rel="noopener noreferrer"&gt;I wrote 0 to cgroup.pressure&lt;/a&gt;&lt;br&gt;
the counter went quiet; this is the same blindness wearing a different file. There is a remedy,&lt;br&gt;
and it was news to me: &lt;code&gt;cpu.stat.local&lt;/code&gt; also reports throttling inherited from ancestors, and it&lt;br&gt;
has been there since Linux 6.6. I found the file on VPS3 (6.8); for my containers the two files&lt;br&gt;
hold identical values, meaning the quota really does sit at the leaf. But in a setup that puts&lt;br&gt;
the quota on a parent unit, that is the file to read. Note that &lt;code&gt;cpu.stat.local&lt;/code&gt; is also summed&lt;br&gt;
across CPUs, so it does not solve the second trap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The threshold in an application log is the application's, not yours.&lt;/strong&gt; etcd's 24-hour log holds&lt;br&gt;
891 "apply request took too long" lines for &lt;code&gt;vpsman-etcd&lt;/code&gt; and 12,395 for &lt;code&gt;licman-etcd&lt;/code&gt;. For a&lt;br&gt;
moment I got excited that the distribution never dipped below 100 ms — "look, the exact period&lt;br&gt;
boundary!" No. etcd's &lt;code&gt;expected-duration&lt;/code&gt; threshold is 100 ms; it never logs anything below that.&lt;br&gt;
The floor belongs to the logger, not the scheduler. For &lt;code&gt;licman-etcd&lt;/code&gt; the actual distribution is&lt;br&gt;
a smooth tail from 100 ms out to a second: p50 141 ms, p90 278 ms, p99 601 ms, worst case 996 ms.&lt;/p&gt;
&lt;h2&gt;
  
  
  The fix has been in the kernel for four years: cpu.max.burst
&lt;/h2&gt;

&lt;p&gt;Banking unused quota and spending it later is not a new idea. Commit &lt;code&gt;f4183717b370&lt;/code&gt;&lt;br&gt;
("sched/fair: Introduce the burstable CFS controller") is dated 21 June 2021 and landed in Linux&lt;br&gt;
5.14; the &lt;code&gt;cpu.max.burst&lt;/code&gt; interface arrived in the same release. In v5.13 there is no such thing&lt;br&gt;
as &lt;code&gt;cfs_burst&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The logic is ten lines inside &lt;code&gt;__refill_cfs_bandwidth_runtime&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVsiUGVyaW9kIHJlZnJlc2hlcyJdIC0tPiBCWyJydW50aW1lID0gcnVudGltZSArIHF1b3RhIl0KICBCIC0tPiBDeyJydW50aW1lX3NuYXAgLSBydW50aW1lID4gMCA_In0KICBDIC0tPnwieWVzOiBpdCBzcGVudCBhYm92ZSBxdW90YSJ8IERbIm5yX2J1cnN0Kys8YnIvPmJ1cnN0X3RpbWUgKz0gZGVsdGEiXQogIEMgLS0-fG5vfCBFWyJjb3VudGVycyBzdGF5IHB1dCJdCiAgRCAtLT4gRlsicnVudGltZSA9IG1pbihydW50aW1lLCBxdW90YSArIGJ1cnN0KSJdCiAgRSAtLT4gRgogIEYgLS0-IEd7IlF1b3RhIGV4Y2VlZGVkIHRoaXMgcGVyaW9kPyJ9CiAgRyAtLT58eWVzfCBIWyJjZnNfcnEgaXMgdGhyb3R0bGVkPGJyLz5ucl90aHJvdHRsZWQrKyJdCiAgRyAtLT58bm98IElbImtlZXBzIHJ1bm5pbmciXQ%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVsiUGVyaW9kIHJlZnJlc2hlcyJdIC0tPiBCWyJydW50aW1lID0gcnVudGltZSArIHF1b3RhIl0KICBCIC0tPiBDeyJydW50aW1lX3NuYXAgLSBydW50aW1lID4gMCA_In0KICBDIC0tPnwieWVzOiBpdCBzcGVudCBhYm92ZSBxdW90YSJ8IERbIm5yX2J1cnN0Kys8YnIvPmJ1cnN0X3RpbWUgKz0gZGVsdGEiXQogIEMgLS0-fG5vfCBFWyJjb3VudGVycyBzdGF5IHB1dCJdCiAgRCAtLT4gRlsicnVudGltZSA9IG1pbihydW50aW1lLCBxdW90YSArIGJ1cnN0KSJdCiAgRSAtLT4gRgogIEYgLS0-IEd7IlF1b3RhIGV4Y2VlZGVkIHRoaXMgcGVyaW9kPyJ9CiAgRyAtLT58eWVzfCBIWyJjZnNfcnEgaXMgdGhyb3R0bGVkPGJyLz5ucl90aHJvdHRsZWQrKyJdCiAgRyAtLT58bm98IElbImtlZXBzIHJ1bm5pbmciXQ%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="445" height="1262"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(The &lt;code&gt;nr_burst&lt;/code&gt; and &lt;code&gt;burst_time&lt;/code&gt; in the diagram are the kernel's internal field names; in the&lt;br&gt;
cgroup v2 interface you see the same things as &lt;code&gt;nr_bursts&lt;/code&gt; and &lt;code&gt;burst_usec&lt;/code&gt;.)&lt;/p&gt;

&lt;p&gt;The critical line is the ceiling at the bottom: &lt;code&gt;runtime = min(runtime, quota + burst)&lt;/code&gt;. The most&lt;br&gt;
you can spend in a single period is quota plus burst. The docs give the permitted range for&lt;br&gt;
&lt;code&gt;cpu.max.burst&lt;/code&gt; as &lt;code&gt;[0, $MAX]&lt;/code&gt;, default &lt;code&gt;0&lt;/code&gt;. So if you set burst equal to quota, you can spend at&lt;br&gt;
most twice the quota in one period, and not a microsecond more.&lt;/p&gt;

&lt;p&gt;That ceiling helped me design the experiment — because it tells you in advance what should happen.&lt;/p&gt;
&lt;h2&gt;
  
  
  The experiment: half a burst buys you nothing
&lt;/h2&gt;

&lt;p&gt;In Docker Desktop's linuxkit VM (&lt;code&gt;6.10.14-linuxkit&lt;/code&gt;, cgroup v2) I created my own cgroup and wrote&lt;br&gt;
a deliberately spiky workload: each round, idle for 400 ms, then burn 80 ms of CPU, 250 rounds.&lt;br&gt;
Quota &lt;code&gt;50000 100000&lt;/code&gt;. Average demand is only 33% of the quota. The quantity under measurement is&lt;br&gt;
the wall-clock duration of the work — since the burn loop spends exactly 80 ms of CPU as counted&lt;br&gt;
by &lt;code&gt;CLOCK_PROCESS_CPUTIME_ID&lt;/code&gt;, the difference is waiting.&lt;/p&gt;

&lt;p&gt;I ran three arms:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;quota 50000/100000, work: 400ms idle + 80ms CPU x 250 rounds

burst=0       p50  97.3 ms   throttled 250/250 rounds   throt 4.298 s   nr_bursts   0
burst=25000   p50  98.0 ms   throttled 250/250 rounds   throt 4.416 s   nr_bursts 250  burst 6.250 s
burst=50000   p50  80.0 ms   throttled   0/250 rounds   throt 0     s   nr_bursts 208  burst 3.765 s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The third row was what I expected: with burst equal to quota there is not a single throttle event&lt;br&gt;
and latency settles at 80.0 ms — the pure CPU cost of the work. A 22% latency tax, gone.&lt;/p&gt;

&lt;p&gt;The second row is what stopped me. With burst set to half the quota, &lt;code&gt;nr_bursts&lt;/code&gt; came out at 250:&lt;br&gt;
the mechanism fired on &lt;strong&gt;every single round&lt;/strong&gt;. &lt;code&gt;burst_usec&lt;/code&gt; is 6.250 seconds, that is exactly&lt;br&gt;
25,000 µs per round — the budget spent down to the last drop. And latency did not improve by one&lt;br&gt;
millisecond.&lt;/p&gt;

&lt;p&gt;The ceiling formula explains why: 50 ms quota + 25 ms burst = 75 ms, and the job needs 80 ms.&lt;br&gt;
Five milliseconds short. Waiting for a period boundary over five milliseconds produces the same&lt;br&gt;
delay as waiting over fifty. Here a half measure arrives at full cost and zero benefit — while&lt;br&gt;
the counters report back that "burst is working!" Green on the dashboard, waiting for the user.&lt;/p&gt;

&lt;p&gt;If it were me, I would take exactly one rule away from this: there is no such thing as enabling&lt;br&gt;
burst "a little." If you do not know the peak your workload demands in one go, you do not know&lt;br&gt;
whether the burst value you picked will do anything at all.&lt;/p&gt;
&lt;h2&gt;
  
  
  The dial nobody turns: period
&lt;/h2&gt;

&lt;p&gt;Here is what gets lost in the burst conversation: &lt;code&gt;cpu.max&lt;/code&gt; is two numbers. Everyone tunes the&lt;br&gt;
first one and leaves the second at its default for four years. The same 0.5 CPU ratio, run at&lt;br&gt;
three different periods with burst disabled:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;same work (400ms idle + 80ms CPU), 0.5 CPU ratio throughout, burst=0

5000/10000      (10 ms period)   p50 151.0 ms   throttled 2892/3414 periods
50000/100000    (100 ms period)  p50  96.4 ms   throttled  197/600  periods
250000/500000   (500 ms period)  p50  80.0 ms   throttled    0/194  periods
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Shortening the period made latency &lt;strong&gt;57% worse&lt;/strong&gt;. That may read as counterintuitive — I had&lt;br&gt;
assumed "refreshing more often means waiting less" — but the opposite happens: as the window&lt;br&gt;
narrows, the slice you can take in one go shrinks, and an 80 ms job gets cut into sixteen pieces.&lt;br&gt;
Every wait is short, but there are a lot of waits.&lt;/p&gt;

&lt;p&gt;Lengthening the period gave the same result as burst: zero throttling, 80.0 ms. And this dial&lt;br&gt;
works &lt;em&gt;today&lt;/em&gt; — you set it with &lt;code&gt;docker run --cpu-period --cpu-quota&lt;/code&gt;, it is part of the&lt;br&gt;
container configuration so it survives a restart, and it needs no new tooling. The cost is&lt;br&gt;
symmetric: once you genuinely exhaust the quota, a single stall can now last up to 500 ms, and&lt;br&gt;
the group holds the machine for longer while it does. The documented ceiling is 1 second.&lt;/p&gt;

&lt;p&gt;So a solution that needs no burst at all was sitting there, and for two weeks I never looked at&lt;br&gt;
the second number in &lt;code&gt;cpu.max&lt;/code&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where the chain breaks
&lt;/h2&gt;

&lt;p&gt;As for burst: the reason it is off in all 16 of my cgroups is that nothing in my stack turns it&lt;br&gt;
on. I checked the chain link by link:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Kernel&lt;/strong&gt;: present since v5.14. ✓&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OCI runtime-spec&lt;/strong&gt;: a &lt;code&gt;burst&lt;/code&gt; field is defined, with the constraint that it must not exceed a
positive &lt;code&gt;quota&lt;/code&gt;. ✓&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;runc&lt;/strong&gt;: &lt;code&gt;specconv&lt;/code&gt; maps &lt;code&gt;r.CPU.Burst&lt;/code&gt; to &lt;code&gt;c.Resources.CpuBurst&lt;/code&gt;; the actual write happens in
&lt;code&gt;fs2/cpu.go&lt;/code&gt; inside &lt;code&gt;opencontainers/cgroups&lt;/code&gt;, to the &lt;code&gt;cpu.max.burst&lt;/code&gt; file. ✓&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Podman&lt;/strong&gt;: settable via &lt;a href="https://docs.podman.io/en/latest/markdown/podman-run.1.html" rel="noopener noreferrer"&gt;&lt;code&gt;--cgroup-conf=cpu.max.burst=50000&lt;/code&gt;&lt;/a&gt;
— documented behavior, writing to
an arbitrary cgroup v2 file as part of the container configuration. (I am taking this one from
the docs; I did not measure it, as podman is not installed on my machine.) ✓&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Docker CLI&lt;/strong&gt;: absent. On my 27.4.0 install, the word "burst" appears zero times in
&lt;code&gt;docker run --help&lt;/code&gt;; the current &lt;code&gt;docker run&lt;/code&gt; reference lists &lt;code&gt;--cpu-period&lt;/code&gt;, &lt;code&gt;--cpu-quota&lt;/code&gt; and
&lt;code&gt;--cpu-shares&lt;/code&gt;, and lists neither burst nor any &lt;code&gt;--cgroup-conf&lt;/code&gt;-style passthrough. ✗&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compose spec&lt;/strong&gt;: zero hits in &lt;code&gt;deploy.md&lt;/code&gt;. ✗&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;systemd&lt;/strong&gt;: &lt;code&gt;systemd.resource-control&lt;/code&gt; has &lt;code&gt;CPUQuota&lt;/code&gt; and &lt;code&gt;CPUQuotaPeriodSec&lt;/code&gt;, with no burst
equivalent (VPS3 runs systemd 255). ✗&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kubernetes&lt;/strong&gt;: no burst field in the Pod spec. On the CRI side there is a generic cgroup v2
passthrough map called &lt;code&gt;LinuxContainerResources.unified&lt;/code&gt; — so the pipe is not technically
closed, there is just no user-facing surface for the kubelet to fill it from. ✗&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The chain is not broken end to end; it breaks at &lt;strong&gt;the Docker CLI and the kubelet's Pod&lt;br&gt;
interface&lt;/strong&gt;. If you run podman, the switch is within reach today.&lt;/p&gt;

&lt;p&gt;"I will just write it by hand on Docker," I said, and tried. I started a container with&lt;br&gt;
&lt;code&gt;docker run --cpus=0.5&lt;/code&gt;, wrote &lt;code&gt;50000&lt;/code&gt; into its cgroup, verified it, then ran &lt;code&gt;docker restart&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;docker --cpus=0.5  -&amp;gt; cpu.max=[50000 100000]  burst=[0]
written by hand    -&amp;gt; burst=[50000]
after restart      -&amp;gt; cpu.max=[50000 100000]  burst=[0]   (same container id)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Docker rewrites &lt;code&gt;cpu.max&lt;/code&gt; because that is its configuration; it does not rewrite &lt;code&gt;burst&lt;/code&gt; because&lt;br&gt;
it knows no such thing exists. Any value you write by hand evaporates on the first restart. This&lt;br&gt;
is not configuration, it is a poke.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to ask about your own setup
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;First ask why the quota is there at all.&lt;/strong&gt; My etcd instances are squeezed into half a core
on an 18-core machine. The quota was put there to stop them starving neighboring projects,
but etcd is not the profile that starves anyone. The right answer here may be to drop the quota
and give them a share of the contention via &lt;code&gt;cpu.weight&lt;/code&gt; — but that is not free: a weight only
divides the spoils during contention, it sets no ceiling. A runaway compaction can still take
all 18 cores. A quota is a hard wall; a weight is a negotiation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consider the period; you have probably never touched it.&lt;/strong&gt; The measurement above shows that
lengthening the period at the same ratio achieves what burst does, and it works in Docker
today. Shortening it backfires. Limits: the documented maximum period is 1 second, and a longer
period means a longer single stall.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Raising the quota is boring, and it works.&lt;/strong&gt; The difference from burst is this: doubling the
quota also doubles your entitlement on average, whereas burst only gives back entitlement you
previously did &lt;em&gt;not&lt;/em&gt; spend.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Burst is not free — the docs say so outright.&lt;/strong&gt; The burst section of the kernel documentation
introduces the mechanism and its price in the same sentence: borrowing against a future
underrun happens "at the cost of increased interference against the other system users." The
consolation is that the tardiness is bounded, and that per the docs the interference stays
limited when there are many cgroups or the CPU is under-utilized. Still, do not read it as
"protects the neighbors": it protects the average and raises the instantaneous interference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you do choose burst, measure peak demand rather than guessing it&lt;/strong&gt; — and work out up front
how to make it stick. If it does not come through Docker, you need something that reapplies it
on every restart.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;At minimum, put the counter on a dashboard.&lt;/strong&gt; Even if you do none of the above, reading
&lt;code&gt;nr_throttled&lt;/code&gt; and &lt;code&gt;throttled_usec&lt;/code&gt; is free — or &lt;code&gt;cpu.stat.local&lt;/code&gt; if the quota lives on a
parent unit. An average CPU graph structurally cannot show you this event.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The last one is the cheapest and the most useful. A pause you do not measure has not been&lt;br&gt;
cancelled, only made invisible. And if you have an appetite for changing the scheduler more&lt;br&gt;
fundamentally, &lt;a href="https://mustafaerbay.com.tr/en/blog/technology/zamanlayiciyi-calisirken-degistirmek-sched-ext/" rel="noopener noreferrer"&gt;the sched_ext side&lt;/a&gt;&lt;br&gt;
is worth a look — though the answer to the quota problem is not there.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I did not prove
&lt;/h2&gt;

&lt;p&gt;I need to draw a line here, because the story looks too clean and for a while I believed it too&lt;br&gt;
much myself.&lt;/p&gt;

&lt;p&gt;In the lab there is causation: same work, same quota, burst the only variable, result 250 throttle&lt;br&gt;
events down to zero. But the lab workload is &lt;strong&gt;single-threaded&lt;/strong&gt;. The docs state that burst is not&lt;br&gt;
transferred between cores; in a multi-threaded group like etcd, per-CPU slice distribution enters&lt;br&gt;
the picture and the behavior may not be identical. What I proved is the single-threaded form of&lt;br&gt;
the mechanism.&lt;/p&gt;

&lt;p&gt;In production, all I have is a strong coincidence. I know etcd is slow, I know it is being&lt;br&gt;
throttled, I know the requests that slow down are things like &lt;code&gt;key:"health"&lt;/code&gt; with no CPU cost at&lt;br&gt;
all, and I know storage is not the culprit (against 891 and 12,395 apply warnings in 24 hours,&lt;br&gt;
&lt;code&gt;slow fdatasync&lt;/code&gt; appears only 1 and 3 times). I wanted to test it with a measurement: I took eight&lt;br&gt;
one-minute readings on &lt;code&gt;licman-etcd&lt;/code&gt; and lined up throttle counts against slow-request counts&lt;br&gt;
across the seven deltas.&lt;/p&gt;

&lt;p&gt;The result did not support me: Pearson correlation 0.437, n=7. That does not mean "no&lt;br&gt;
relationship" — it is moderately positive but not significant at this sample size, so the&lt;br&gt;
measurement neither supports nor refutes; it is underpowered. The reason is visible inside it:&lt;br&gt;
against an average of 61.2 throttle events per minute there are only 11.3 slow-request warnings.&lt;br&gt;
Most throttles never turn into a warning on the request path — they land on background threads,&lt;br&gt;
and one slow request can span several throttles besides. At one-minute resolution the signal&lt;br&gt;
drowns in noise, and the way to fix that is to enable burst in production and run an A/B, which I&lt;br&gt;
am not going to do on a live consensus ledger.&lt;/p&gt;

&lt;p&gt;So what I hold is this: exposure measured, mechanism proven in its single-threaded form,&lt;br&gt;
production causation unproven. Keeping those three apart matters as much as the measurement itself.&lt;/p&gt;

&lt;p&gt;What is certain, though, is that the question I had been asking since the day I set that quota was&lt;br&gt;
the wrong one. I was asking "how much of its quota is this service using?" and for ten days I&lt;br&gt;
proudly collected the answer "17%." What I should have asked was "how many times has this service&lt;br&gt;
hit its quota?" The first question measures an average. The second measures what the user is&lt;br&gt;
waiting for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.kernel.org/scheduler/sched-bwc.html" rel="noopener noreferrer"&gt;CFS Bandwidth Control — Linux kernel documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.kernel.org/admin-guide/cgroup-v2.html" rel="noopener noreferrer"&gt;Control Group v2 — CPU Interface Files&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/commit/f4183717b370ad28dd0c0d74760142b20e6e7931" rel="noopener noreferrer"&gt;sched/fair: Introduce the burstable CFS controller (f4183717b370)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/opencontainers/runtime-spec/blob/main/config-linux.md" rel="noopener noreferrer"&gt;OCI Runtime Specification — Linux CPU resources&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/opencontainers/cgroups/blob/main/fs2/cpu.go" rel="noopener noreferrer"&gt;opencontainers/cgroups — fs2/cpu.go&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.docker.com/reference/cli/docker/container/run/" rel="noopener noreferrer"&gt;docker container run — CLI reference&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>linux</category>
      <category>cgroup</category>
      <category>container</category>
      <category>olcum</category>
    </item>
    <item>
      <title>I Could Not Turn On Journal Sealing; Only the Debug Log Said Why</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Sun, 04 Oct 2026 12:46:48 +0000</pubDate>
      <link>https://dev.to/merbayerp/i-could-not-turn-on-journal-sealing-only-the-debug-log-said-why-1a4o</link>
      <guid>https://dev.to/merbayerp/i-could-not-turn-on-journal-sealing-only-the-debug-log-said-why-1a4o</guid>
      <description>&lt;p&gt;Logs are the most trusted and least questioned files on a machine. After an incident you open them and believe what they say. systemd has an answer to that trust: Forward Secure Sealing. journald seals the file cryptographically at regular intervals, and if somebody tampers with it afterwards, &lt;code&gt;journalctl --verify&lt;/code&gt; tells you.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;journald.conf&lt;/code&gt; documentation puts it this way for &lt;code&gt;Seal=&lt;/code&gt;: "If enabled (the default), &lt;strong&gt;and a sealing key is available&lt;/strong&gt; (as created by journalctl's &lt;code&gt;--setup-keys&lt;/code&gt; command), Forward Secure Sealing (FSS) for all persistent journal files is enabled." So what ships enabled by default is the mechanism, not the key. The key is yours to create. Creating it is what I sat down to do, and the first command hit a wall at the first step:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--setup-keys&lt;/span&gt; &lt;span class="nt"&gt;--interval&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1m
Generating seed...
Generating key pair...
Failed to generate key pair: Operation not supported
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The error message names nothing. Finding the cause took half an hour; fixing it took one package. Then sealing actually engaged, and the real half of this article began: what sealing catches, what it does not, and what that &lt;code&gt;PASS&lt;/code&gt; line really promises.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lab
&lt;/h2&gt;

&lt;p&gt;Every number below comes from two privileged containers with systemd genuinely running as PID 1:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--version&lt;/span&gt; | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;
systemd 257 &lt;span class="o"&gt;(&lt;/span&gt;257.13-1~deb13u1&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt; /etc/os-release&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PRETTY_NAME&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
Debian GNU/Linux 13 &lt;span class="o"&gt;(&lt;/span&gt;trixie&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;
6.10.14-linuxkit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On the comparison side sits Debian 12: &lt;code&gt;systemd 252 (252.39-1~deb12u2)&lt;/code&gt;. One disclosure up front: that &lt;code&gt;linuxkit&lt;/code&gt; in &lt;code&gt;uname -r&lt;/code&gt; is the kernel of a Docker Desktop virtual machine on a macOS host. Both Debians share the same non-Debian kernel. What gets measured here is Debian &lt;strong&gt;userspace&lt;/strong&gt;: journald, journalctl, packaging. No claim in this article rests on the kernel.&lt;/p&gt;

&lt;p&gt;Sealing needs a persistent journal. Without &lt;code&gt;/var/log/journal&lt;/code&gt;, journald keeps everything under &lt;code&gt;/run/log/journal&lt;/code&gt; in memory and there is no file to seal. The directory exists on both.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "Operation not supported" does not say
&lt;/h2&gt;

&lt;p&gt;The setting is right where it should be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;systemd-analyze cat-config systemd/journald.conf | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; seal
&lt;span class="c"&gt;#Seal=yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The man page still documents &lt;code&gt;--setup-keys&lt;/code&gt;. The file header shows no sign of a gap either, only the sealing flag is missing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--header&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s2"&gt;"File path|Compatible flags"&lt;/span&gt;
File path: /var/log/journal/a452a00b.../system.journal
Compatible flags: TAIL_ENTRY_BOOT_ID
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Raising the log level gave up the answer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ SYSTEMD_LOG_LEVEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;debug journalctl &lt;span class="nt"&gt;--setup-keys&lt;/span&gt; &lt;span class="nt"&gt;--force&lt;/span&gt; 2&amp;gt;&amp;amp;1 | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-3&lt;/span&gt;
libgcrypt.so.20 is not installed: libgcrypt.so.20: cannot open shared object
file: No such file or directory
Failed to generate key pair: Operation not supported
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That line is printed at &lt;code&gt;LOG_DEBUG&lt;/code&gt;, so it never shows up in a normal run. Sealing depends on libgcrypt, and since systemd v256 the library is no longer an ordinary shared-library dependency; it is loaded at runtime through &lt;code&gt;dlopen()&lt;/code&gt;. The upstream &lt;code&gt;NEWS&lt;/code&gt; file announced the change along with its warning: those libraries might no longer be pulled in automatically when ELF dependencies are resolved. On the Debian side the result looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;apt-cache show systemd | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-iE&lt;/span&gt; &lt;span class="s2"&gt;"^(Depends|Recommends|Suggests)"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="s1"&gt;','&lt;/span&gt; &lt;span class="s1"&gt;'\n'&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-ci&lt;/span&gt; gcrypt
0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;systemd&lt;/code&gt; package does not ask for &lt;code&gt;libgcrypt20&lt;/code&gt;, neither as a dependency nor as a recommendation. The fix is one line and it works instantly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;apt-get &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-y&lt;/span&gt; libgcrypt20
&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--setup-keys&lt;/span&gt; &lt;span class="nt"&gt;--force&lt;/span&gt; &lt;span class="nt"&gt;--interval&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1m &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"rc=&lt;/span&gt;&lt;span class="nv"&gt;$?&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;rc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scope of the trap deserves stating. On a desktop or a full server install, &lt;code&gt;libgcrypt20&lt;/code&gt; is usually already there because something else pulled it in; &lt;code&gt;apt-cache rdepends libgcrypt20&lt;/code&gt; lists &lt;code&gt;libxslt1.1&lt;/code&gt;, the webkit libraries, strongSwan plugins. This silent loss mostly hits minimal images, containers and thin servers built with &lt;code&gt;debootstrap&lt;/code&gt;. So the answer to "can sealing be turned on here" belongs to the machine, not to the distribution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Once sealing really engages
&lt;/h2&gt;

&lt;p&gt;With the key pair in place, a &lt;code&gt;/var/log/journal/&amp;lt;machine-id&amp;gt;/fss&lt;/code&gt; file appears: the sealing key, 482 bytes, owned by &lt;code&gt;root:systemd-journal&lt;/code&gt;. Watch the mode, because on two separate installs it came out as &lt;code&gt;0640&lt;/code&gt;. That means accounts in the &lt;code&gt;systemd-journal&lt;/code&gt; group can read the sealing key. The verification key is printed to the screen exactly once; keeping it off the machine is your job, because you will not see it a second time.&lt;/p&gt;

&lt;p&gt;Second point: journald cannot add a sealing flag to the header of a file it already has open. Generating the key and restarting the service was not enough, a new file was needed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--rotate&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--header&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"Compatible flags"&lt;/span&gt;
Compatible flags: SEALED SEALED_CONTINUOUS TAIL_ENTRY_BOOT_ID
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now verification says something meaningful:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;.../system.journal &lt;span class="nt"&gt;--verify&lt;/span&gt; &lt;span class="nt"&gt;--verify-key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$VK&lt;/span&gt;
PASS: /var/log/journal/a452a00b.../system.journal
&lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; Validated from Sun 2026-10-04 11:51:50 UTC to Sun 2026-10-04 11:53:00 UTC,
   final 2.758424s entries not sealed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is my favourite output in this article. The tool announces its own blind spot: the final 2.76 seconds of entries are unsealed. Sealing reaches as far as the next seal, and the window in between stays open. The man page says the same thing while explaining &lt;code&gt;--interval=&lt;/code&gt;: a shorter interval raises CPU consumption and shortens the time range of undetectable alterations. The default is 15 minutes. In the lab it was pulled down to one minute, which is why the window here is a matter of seconds.&lt;/p&gt;

&lt;p&gt;Verification without a key, on the other hand, did not behave the way I assumed. Expecting at least a structural check on a sealed file, this came back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;.../system.journal &lt;span class="nt"&gt;--verify&lt;/span&gt;
Journal file ... has sealing enabled but verification key has not been passed
using &lt;span class="nt"&gt;--verify-key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt;
FAIL: ... &lt;span class="o"&gt;(&lt;/span&gt;Required key not available&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The source is unambiguous about it: on a sealed file with no key, &lt;code&gt;journal_file_verify()&lt;/code&gt; returns &lt;code&gt;-ENOKEY&lt;/code&gt; before it even opens a temporary directory. Not a single structural check runs. The operational translation is just as blunt: lose the verification key and you cannot form even a structural opinion about your journal's integrity.&lt;/p&gt;

&lt;h3&gt;
  
  
  How wide does the tail get
&lt;/h3&gt;

&lt;p&gt;For five minutes nothing was written to the machine by hand. The container's own services kept writing, which matters, because a seal only lands while an entry is being appended. The result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;=&amp;gt; Validated from Sun 2026-10-04 12:18:58 UTC to Sun 2026-10-04 12:23:00 UTC,
   final 59.603160s entries not sealed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six seals had accumulated in the file, at epochs &lt;code&gt;0, 0, 1, 2, 3, 4&lt;/code&gt;. With a one-minute interval the tail reaches a full minute. At the 15-minute default that same window is a quarter of an hour.&lt;/p&gt;

&lt;p&gt;The source holds a more uncomfortable detail here. In v257 there are only two places that write a seal: &lt;code&gt;journal_file_append_first_tag()&lt;/code&gt; when the file is created, and &lt;code&gt;journal_file_maybe_append_tag()&lt;/code&gt; called from the entry-append path. No entry, no seal. On a quiet machine the tail is therefore not bounded by the interval at all; it stays unsealed indefinitely.&lt;/p&gt;

&lt;p&gt;My own assumption had been that running &lt;code&gt;journalctl --rotate&lt;/code&gt; after a critical event would seal the tail. In the v257 source there is no path that appends a seal while closing a file, and no run of mine showed rotation sealing the old file's tail either. Do not lean on it. What you can lean on is that &lt;code&gt;--verify&lt;/code&gt; prints the window on every run.&lt;/p&gt;

&lt;h2&gt;
  
  
  What sealing looks like inside the file
&lt;/h2&gt;

&lt;p&gt;Seeing where the seal lands is the shortest path to understanding its reach. The file-format document gives the TAG object's layout: an object header, a sequence number, an epoch number and a 32-byte SHA-256 HMAC, 64 bytes in total. I walked my own file raw, stepping through the objects from start to end and counting their types:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;object types: ENTRY 73, DATA 258, FIELD 40, ENTRY_ARRAY 90, TAG 3
TAG (seqnum, epoch): (1, 0)  (2, 0)  (3, 1)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three seals in a file spanning roughly 70 seconds, with the interval set to one minute. The first seal lands as the file is created: &lt;code&gt;journal_file_append_first_tag()&lt;/code&gt;, called from &lt;code&gt;journal_file_open()&lt;/code&gt; for newly created files only, folding the header plus the two hash-table objects into the HMAC. The second sits in the same epoch. Hold on to that detail, I will come back to it.&lt;/p&gt;

&lt;p&gt;Each seal is computed over the objects that arrived since the previous one. The document adds an important footnote: while computing the HMAC, &lt;strong&gt;volatile&lt;/strong&gt; header fields of objects are skipped, such as linked-list pointers added later. What validates the consistency of those fields is not the seal but the structural checks in &lt;code&gt;--verify&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I did not take the cost side, but it can be derived from the structure that was measured: a 64-byte TAG at the default 15-minute interval means 96 seals a day, roughly 6 KB. Pull it down to one minute and you get 1,440 seals, about 92 KB. The disk side is irrelevant. The cost the man page warns about is on the CPU, because the key is derived forward at every seal. That figure is not one I took, so I will not quote one; if you plan to shorten the interval, run it on your own hardware and look.&lt;/p&gt;

&lt;h2&gt;
  
  
  What &lt;code&gt;--verify&lt;/code&gt; checks independently of sealing
&lt;/h2&gt;

&lt;p&gt;Sealing is half the story. The &lt;code&gt;journal_file_verify()&lt;/code&gt; source holds a long checklist that runs even without a key: hashes of data and field objects, object sizes and alignment, consistency of compression flags, strictly increasing entry sequence numbers, timestamps ordered monotonically within the same boot, hash-table chains free of cycles, entry-array chains that never jump backwards, and header counters that reconcile with the actual content.&lt;/p&gt;

&lt;p&gt;That list catches corruption well. What it does not catch is a deliberate edit that leaves all of those fields consistent. The hash on data objects is keyed, and my file's header carries the &lt;code&gt;KEYED-HASH&lt;/code&gt; flag, but the key sits inside the file itself. It is protection against hash flooding, not authentication. Whoever holds the file can recompute the hash. Authentication needs a secret the attacker does not have, and sealing is exactly what brings one.&lt;/p&gt;

&lt;p&gt;There is one more wrinkle: a sealed file will not even open on a machine without sealing support. Copying the same sealed file to a Debian 13 host with no &lt;code&gt;libgcrypt20&lt;/code&gt; and trying to verify it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/root/t/muhurlu.journal &lt;span class="nt"&gt;--verify&lt;/span&gt;
Failed to open files: Operation not supported
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Verification does not fail; the file cannot be read at all. If you plan to open sealed archives on a thin recovery box later, &lt;code&gt;libgcrypt20&lt;/code&gt; has to be there too.&lt;/p&gt;

&lt;h2&gt;
  
  
  What sealing catches
&lt;/h2&gt;

&lt;p&gt;Three kinds of tampering, each on a copy of the file, at single-byte or single-bit granularity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A message inside the sealed region.&lt;/strong&gt; I changed one character inside a log line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;392be0: Invalid object contents: Bad message
File corruption detected at .../t13-mesaj.journal:3746784
  (of 8388608 bytes, 44%).
FAIL: ... (Bad message)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Caught, but sealing contributed nothing here. Every data object carries its own hash and changing the text breaks it. The same edit would be caught in an unsealed file. A corruption check and authentication are not the same thing, and the first one is what this test demonstrates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The seal itself.&lt;/strong&gt; To isolate sealing, I flipped a single bit in the HMAC field of a TAG object:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;3955d8: Tag failed verification
FAIL: ... (Bad message)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the proof that sealing works. According to the file-format document, every TAG carries a SHA-256 HMAC over the objects before it, with the key derived forward at each epoch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dropping the flag.&lt;/strong&gt; Can an attacker pretend the file was never sealed? I zeroed the compatibility flags in the header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;38f980: Tag object in file without sealing
FAIL: ... (Bad message)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No. With TAG objects still in the file, clearing the flag does not silence verification, it makes more noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I deleted the logs
&lt;/h2&gt;

&lt;p&gt;A real attacker does not flip bytes, they delete. The sealing key lives on the very machine it protects, so anybody with root can read it. I saved the harshest test for last: stop journald, delete every &lt;code&gt;.journal&lt;/code&gt; file, leave the &lt;code&gt;fss&lt;/code&gt; key alone, start the service, write a clean history.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--no-pager&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
1384
&lt;span class="nv"&gt;$ &lt;/span&gt;systemctl stop systemd-journald
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; /var/log/journal/&lt;span class="k"&gt;*&lt;/span&gt;/&lt;span class="k"&gt;*&lt;/span&gt;.journal
&lt;span class="nv"&gt;$ &lt;/span&gt;systemctl start systemd-journald
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="k"&gt;for &lt;/span&gt;i &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;seq &lt;/span&gt;1 6&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do &lt;/span&gt;logger &lt;span class="nt"&gt;-t&lt;/span&gt; audit &lt;span class="s2"&gt;"FORGE-&lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--no-pager&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
11
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One thousand three hundred and eighty-four lines gone, eleven left. The new file opened with &lt;code&gt;SEALED SEALED_CONTINUOUS&lt;/code&gt; flags, so on paper I have a sealed history. Here is what verification said:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;38f980: Tag failed verification
FAIL: ... (Bad message)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The wipe left a mark. The same result came back three times across three separate containers, and there is a control run to go with it: without any deletion, a new file produced by a plain rotation passes the same verification. So this is not a blanket false alarm.&lt;/p&gt;

&lt;p&gt;On Debian 12 the same move gives a different answer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PASS: /var/log/journal/41a11bc4.../system.journal
=&amp;gt; No sealing yet, 2.020730s of entries not sealed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There, 1,543 lines dropped to 19 and verification was not bothered in the slightest.&lt;/p&gt;

&lt;p&gt;This is where I owe you honesty. I did not trace the Debian 13 &lt;code&gt;FAIL&lt;/code&gt; down to which byte broke and why, so I am not pinning the mechanism on a patch or a flag. What I hold is a behaviour reproduced three times and a control run that does not refute it. I claim the behaviour, not the mechanism.&lt;/p&gt;

&lt;h2&gt;
  
  
  The patch that seals empty epochs
&lt;/h2&gt;

&lt;p&gt;Debian 13's files carry an extra flag, &lt;code&gt;SEALED_CONTINUOUS&lt;/code&gt;, that Debian 12's do not. The flag arrived with Felix Dörre's patch, merged on 8 November 2023 in the v255 cycle, and the patch states what it closes: "Currently empty epochs are not sealed. This allows an attacker to truncate a sealed log and continue it without any problems showing when verifying the log."&lt;/p&gt;

&lt;p&gt;The same gap is on record as CVE-2023-31438. The NVD description: a sealed log file can be truncated and log sealing resumed such that the integrity check shows no error, despite modifications. The record also notes the vendor replying that none of the findings was a security vulnerability, and the entry is tagged disputed.&lt;/p&gt;

&lt;p&gt;One sentence in the patch text sits where most summaries skip: "This &lt;strong&gt;partially&lt;/strong&gt; addresses CVE-2023-31438." It also says what closing it completely would take: verifying that there is exactly one seal per epoch, and not sealing before the epoch has ended. The author left the premature sealing in place, having found it deliberate but not understood its purpose.&lt;/p&gt;

&lt;p&gt;That open door is sitting in my own measurement. Look at the TAG dump above: &lt;code&gt;(1, 0)&lt;/code&gt; and &lt;code&gt;(2, 0)&lt;/code&gt;, two seals in epoch 0. The premature sealing the patch says it did not remove is exactly this. The verifier's continuity check accepts it too; in the source, a file's first tag and a second tag in the same epoch are explicitly exempt.&lt;/p&gt;

&lt;p&gt;The flag has one more practical side effect. Hand a sealed file produced by Debian 12 to Debian 13's &lt;code&gt;journalctl&lt;/code&gt; and you get this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;000000: This log file was sealed with an old journald version where the
sequence of seals might not be continuous. We cannot guarantee completeness.
PASS: /root/t/bw.journal
=&amp;gt; Validated from Sun 2026-10-04 11:45:41 UTC to Sun 2026-10-04 11:48:00 UTC,
   final 4.528312s entries not sealed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;PASS&lt;/code&gt;, but not a complete one. If you keep archives produced by older releases, make a habit of looking for that line in your verification report.&lt;/p&gt;

&lt;h2&gt;
  
  
  Delete the key too, and verification goes quiet
&lt;/h2&gt;

&lt;p&gt;So what if the attacker deletes the key file as well? Same scenario, one extra step:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;systemctl stop systemd-journald
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; /var/log/journal/&lt;span class="k"&gt;*&lt;/span&gt;/&lt;span class="k"&gt;*&lt;/span&gt;.journal /var/log/journal/&lt;span class="k"&gt;*&lt;/span&gt;/fss
&lt;span class="nv"&gt;$ &lt;/span&gt;systemctl start systemd-journald
&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--header&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"Compatible flags"&lt;/span&gt;
Compatible flags: TAIL_ENTRY_BOOT_ID
&lt;span class="nv"&gt;$ &lt;/span&gt;journalctl &lt;span class="nt"&gt;--verify&lt;/span&gt; &lt;span class="nt"&gt;--verify-key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$VK&lt;/span&gt;
PASS: /var/log/journal/a452a00b.../system.journal
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;PASS&lt;/code&gt;, even with the verification key handed over, and not a word of complaint. The logic is consistent: the file is not sealed, and an unsealed file has no seal to verify. From a security standpoint the consequence is this: &lt;code&gt;journalctl --verify&lt;/code&gt; never tells you "this file should have been sealed". You have to hold that expectation somewhere the machine cannot reach.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVsiam91cm5hbGN0bCAtLXZlcmlmeSAtLXZlcmlmeS1rZXkiXSAtLT4gQnsiRG9lcyB0aGUgZmlsZSBjYXJyeSBTRUFMRUQ_In0KICBCIC0tPnwiTm8ifCBDWyJQQVNTPGJyLz5idXQgbm90aGluZyB3YXMgcHJvdmVuIl0KICBCIC0tPnwiWWVzInwgRHsiV2FzIGEga2V5IHN1cHBsaWVkPyJ9CiAgRCAtLT58Ik5vInwgRVsiRkFJTDxici8-UmVxdWlyZWQga2V5IG5vdCBhdmFpbGFibGUiXQogIEQgLS0-fCJZZXMifCBGeyJEb2VzIHRoZSBUQUcgSE1BQyBjaGFpbiBob2xkPyJ9CiAgRiAtLT58Ik5vInwgR1siRkFJTDxici8-VGFnIGZhaWxlZCB2ZXJpZmljYXRpb24iXQogIEYgLS0-fCJZZXMifCBIWyJQQVNTPGJyLz5leGNlcHQgYWZ0ZXIgdGhlIGxhc3Qgc2VhbCJd%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVsiam91cm5hbGN0bCAtLXZlcmlmeSAtLXZlcmlmeS1rZXkiXSAtLT4gQnsiRG9lcyB0aGUgZmlsZSBjYXJyeSBTRUFMRUQ_In0KICBCIC0tPnwiTm8ifCBDWyJQQVNTPGJyLz5idXQgbm90aGluZyB3YXMgcHJvdmVuIl0KICBCIC0tPnwiWWVzInwgRHsiV2FzIGEga2V5IHN1cHBsaWVkPyJ9CiAgRCAtLT58Ik5vInwgRVsiRkFJTDxici8-UmVxdWlyZWQga2V5IG5vdCBhdmFpbGFibGUiXQogIEQgLS0-fCJZZXMifCBGeyJEb2VzIHRoZSBUQUcgSE1BQyBjaGFpbiBob2xkPyJ9CiAgRiAtLT58Ik5vInwgR1siRkFJTDxici8-VGFnIGZhaWxlZCB2ZXJpZmljYXRpb24iXQogIEYgLS0-fCJZZXMifCBIWyJQQVNTPGJyLz5leGNlcHQgYWZ0ZXIgdGhlIGxhc3Qgc2VhbCJd%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="801" height="1201"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The bottom-left corner of that diagram is the real lesson here. &lt;code&gt;PASS&lt;/code&gt; has two separate meanings and the output says both with the same word.&lt;/p&gt;

&lt;h2&gt;
  
  
  How v262 changes this picture
&lt;/h2&gt;

&lt;p&gt;systemd v262 shipped on 22 September 2026 and sealing moved from libgcrypt to OpenSSL. Debian 13's &lt;code&gt;systemd&lt;/code&gt; package already carries a &lt;code&gt;libssl3t64&lt;/code&gt; dependency, so the missing-package trap in this article does not change shape when 262 lands; it disappears.&lt;/p&gt;

&lt;p&gt;What replaces it is sneakier. In the new source, verifying a sealed file on a machine without sealing support no longer returns &lt;code&gt;-ENOKEY&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;journal_auth_supported&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;ENOKEY&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;
        &lt;span class="nf"&gt;log_notice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Journal file is sealed, but journal sealing support is "&lt;/span&gt;
                   &lt;span class="s"&gt;"disabled. Skipping seal verification."&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the situation that produces a hard &lt;code&gt;FAIL&lt;/code&gt; on 257 will, from 262 onward, leave a &lt;code&gt;notice&lt;/code&gt; line and carry on toward &lt;code&gt;PASS&lt;/code&gt;. It is another variant of the quiet &lt;code&gt;PASS&lt;/code&gt; this whole article is about, this time inside upstream itself. When your distribution moves to 262, that notice line is what your verification reports will need you to look for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I do in practice
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Check the flag, do not trust the default.&lt;/strong&gt; If &lt;code&gt;journalctl --header | grep "Compatible flags"&lt;/code&gt; shows no &lt;code&gt;SEALED&lt;/code&gt;, there is no sealing. On minimal images this is the first place to look.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not leave the check to a human.&lt;/strong&gt; Advice is not a control. A small timer that greps &lt;code&gt;journalctl --header&lt;/code&gt; for &lt;code&gt;SEALED&lt;/code&gt; and alerts when it is absent is the "hold the expectation off the machine" argument turned into something operational.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put &lt;code&gt;libgcrypt20&lt;/code&gt; on the install list.&lt;/strong&gt; When it is missing the error message hides the cause, and the diagnostic path is &lt;code&gt;SYSTEMD_LOG_LEVEL=debug&lt;/code&gt;. Recovery boxes that will open sealed archives need it too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Get the verification key off the machine.&lt;/strong&gt; A key kept on the same disk is a convenience for whoever takes that disk, not evidence. It is printed once.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Know the price of rotating keys.&lt;/strong&gt; &lt;code&gt;--setup-keys --force&lt;/code&gt; mints a new pair. The verification key belongs to a key pair, not to a machine, so the moment you rotate, your existing sealed files can no longer be verified with the key in your hand. Do not rotate without archiving the old one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remember that early boot sits outside the seal.&lt;/strong&gt; Only persistent files are sealed. Everything journald writes under &lt;code&gt;/run/log/journal&lt;/code&gt; before &lt;code&gt;systemd-journal-flush.service&lt;/code&gt; is unsealed by design, so the first seconds of boot are structurally out of scope.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pick the interval deliberately.&lt;/strong&gt; The default is 15 minutes, and so is the length of the undetectable window. The "entries not sealed" line in &lt;code&gt;--verify&lt;/code&gt; reports it on every run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not confuse sealing with staying put.&lt;/strong&gt; Sealing shows that existing bytes have not changed, not that the records still exist. Detecting a missing message is a different problem: signed syslog (RFC 5848) numbers messages so the receiver can see which ones are absent. A copy shipped to a central collector does the same job, and the setup in &lt;a href="https://mustafaerbay.com.tr/en/blog/tutorials/systemd-journal-remote-ile-merkezi-loglama-mtls-retention/" rel="noopener noreferrer"&gt;central logging with systemd-journal-remote&lt;/a&gt; exists for exactly that. What gets deleted locally survives remotely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remember what rotating logs costs.&lt;/strong&gt; How &lt;code&gt;copytruncate&lt;/code&gt; silently eats records when it truncates a file in place is something &lt;a href="https://mustafaerbay.com.tr/en/blog/tutorials/copytruncate-cok-kucuk-bir-an-33879-satir-etti/" rel="noopener noreferrer"&gt;I took the measurements for earlier&lt;/a&gt;; in a sealed world such truncations cost you not only data but evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Detection and repair are separate things.&lt;/strong&gt; The distinction &lt;a href="https://mustafaerbay.com.tr/en/blog/technology/dm-integrity-bozuk-blogu-bulur-duzeltmek-icin-bir-ayna-ister/" rel="noopener noreferrer"&gt;from the dm-integrity article&lt;/a&gt; holds here too: sealing tells you about damage, it does not undo it. What undoes it is a backup or a remote copy.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Sealing works wherever you can turn it on. What tripped me up was not its quality but the silent assumption that it existed. The &lt;code&gt;Seal=yes&lt;/code&gt; line sits in the documentation as a default, anyone skipping the condition next to it can believe their sealing is on, and nothing announces that it is not: no service log, no startup error, and certainly not the &lt;code&gt;PASS&lt;/code&gt; from &lt;code&gt;--verify&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Journal integrity is not a setting but a chain. The library has to be installed, the key generated, the file rotated, the verification key kept outside, and somebody has to hold the knowledge that "this file should have been sealed" off the machine. Break any link in that chain and &lt;code&gt;journalctl --verify&lt;/code&gt; still prints &lt;code&gt;PASS&lt;/code&gt;. A quiet &lt;code&gt;PASS&lt;/code&gt; is no different from an alarm that stays silent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/systemd/systemd/blob/v257/man/journalctl.xml" rel="noopener noreferrer"&gt;systemd v257 — journalctl man page source&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/systemd/systemd/blob/v257/man/journald.conf.xml" rel="noopener noreferrer"&gt;systemd v257 — journald.conf man page source (Seal=)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/systemd/systemd/blob/v257/src/libsystemd/sd-journal/journal-verify.c" rel="noopener noreferrer"&gt;systemd v257 — journal_file_verify() source&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/systemd/systemd/blob/v257/src/libsystemd/sd-journal/journal-authenticate.c" rel="noopener noreferrer"&gt;systemd v257 — sealing and TAG writing source&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/systemd/systemd/blob/main/docs/JOURNAL_FILE_FORMAT.md" rel="noopener noreferrer"&gt;systemd — Journal file format, TAG objects and sealing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/systemd/systemd/blob/main/NEWS" rel="noopener noreferrer"&gt;systemd NEWS — v256 dlopen and the v262 OpenSSL move&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/systemd/systemd/pull/28886" rel="noopener noreferrer"&gt;systemd PR #28886 — sealing empty epochs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2023-31438" rel="noopener noreferrer"&gt;NVD — CVE-2023-31438&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.rfc-editor.org/rfc/rfc5848.html" rel="noopener noreferrer"&gt;RFC 5848 — Signed Syslog Messages&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manpages.debian.org/trixie/systemd/journald.conf.5.en.html" rel="noopener noreferrer"&gt;Debian 13 — journald.conf man page&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>systemd</category>
      <category>journald</category>
      <category>linux</category>
      <category>security</category>
    </item>
    <item>
      <title>I Trust 158 Root Certificates. I Installed Three of Them</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Sun, 04 Oct 2026 06:06:09 +0000</pubDate>
      <link>https://dev.to/merbayerp/i-trust-158-root-certificates-i-installed-three-of-them-1jf6</link>
      <guid>https://dev.to/merbayerp/i-trust-158-root-certificates-i-installed-three-of-them-1jf6</guid>
      <description>&lt;p&gt;If someone asked me "how many certificate authorities do you trust?", my honest answer for years would have been "no idea, whatever Apple says." I see the padlock in the browser, I move on. I assume somebody back there has thought about it.&lt;/p&gt;

&lt;p&gt;This morning I sat down and counted that assumption. Trust isn't a feeling, after all; it's a list sitting on my machine. And despite two decades in this line of work, I had never once opened that list.&lt;/p&gt;

&lt;p&gt;What surprised me when I did wasn't how long it was. It was that the number turned out not to be a single number.&lt;/p&gt;

&lt;h2&gt;
  
  
  I asked the same question in four places and got four answers
&lt;/h2&gt;

&lt;p&gt;The macOS root store is a keychain file on disk. Counting what's inside is one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;security find-certificate &lt;span class="nt"&gt;-a&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  /System/Library/Keychains/SystemRootCertificates.keychain &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'BEGIN CERTIFICATE'&lt;/span&gt;
&lt;span class="c"&gt;# 158&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;A nice, clean number. Then I asked the trust settings the same question:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;security dump-trust-settings &lt;span class="nt"&gt;-s&lt;/span&gt; 2&amp;gt;/dev/null | &lt;span class="nb"&gt;head&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;
&lt;span class="c"&gt;# Number of trusted certs = 157&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;One short. A certificate that lives in the store but doesn't make it into the "trusted" count. I put the third question to the names: deduplicating the labels of those 158 certificates left me with 155. And the fourth question was this — how many of the 158 did I put there? Answer: none. The ones I added aren't even in this file; they live in entirely separate places, and there are three of them.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So: same question, and the answers are 158, 157, 155 and 3. None of them is wrong. Each measures something different. Ever since the day I &lt;a href="https://mustafaerbay.com.tr/en/blog/life/diskimde-ne-kadar-yer-var-bes-cevap-13-gib-fark/" rel="noopener noreferrer"&gt;asked five different tools how much free disk space I had and found a 13.5 GiB spread&lt;/a&gt;, this no longer surprises me — but when the subject is trust, it does make you wince a little.&lt;/p&gt;

&lt;h2&gt;
  
  
  That extra certificate in the store isn't a bug
&lt;/h2&gt;

&lt;p&gt;Tracking down the gap between 158 and 157 took some digging, but I liked the answer. I opened all 158 certificates in the store and compared the subject field against the issuer field. In 157 of them the two matched — self-signed. In exactly one, they didn't:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;subject= CN=Developer ID Certification Authority, O=Apple Inc., C=US
issuer = CN=Apple Root CA, O=Apple Inc., C=US
notAfter= Feb  1 22:12:15 2027 GMT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is Apple's developer signing intermediate. It lives in the root store but it isn't a root; it has a parent, and that parent — Apple Root CA — sits in the same store. Its name never appears in the &lt;code&gt;security dump-trust-settings -s&lt;/code&gt; output, because that command lists anchors, not residents of the store.&lt;/p&gt;

&lt;p&gt;There's a subtle but important distinction here. RFC 5914 defines a trust anchor as "an authoritative entity represented by a public key and associated data"; the anchor is an &lt;em&gt;input&lt;/em&gt; to path validation, not a link in the chain. An intermediate's trustworthiness doesn't come from itself, it comes from Apple Root CA. Sure enough, validation passes cleanly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;security verify-cert &lt;span class="nt"&gt;-c&lt;/span&gt; devid.pem &lt;span class="nt"&gt;-p&lt;/span&gt; basic
&lt;span class="c"&gt;# ...certificate verification successful.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's why 157 is the more honest number: it counts the points this machine accepts without question. The intermediate is parked there for convenience.&lt;/p&gt;

&lt;p&gt;The naming side has a similar trap. There are four separate certificates whose label is exactly &lt;code&gt;GlobalSign&lt;/code&gt;; the distinguishing information isn't in the CN but tucked away in the OU field (Root CA - R3, Root CA - R6, ECC Root CA - R5, ECC Root CA - R4). And if you search for certificates with "GlobalSign" anywhere in the CN, you get nine. Four or nine? Both are correct; one is an exact match, the other a substring search. All 158 certificates have distinct SHA-256 fingerprints — what repeats isn't the certificate, it's the name shown to the human.&lt;/p&gt;

&lt;p&gt;While comparing the two lists, by the way, I walked straight into a trap of my own making. Subtracting the trust list from the store list showed two "missing" names, and for a moment I was genuinely excited. The second one was real (the intermediate above); the first was entirely my own fault. NetLock Arany, a Hungarian root, has non-ASCII characters in its name, and the two commands print the same certificate two different ways — one as a hexadecimal blob, the other as proper UTF-8. Same certificate, two spellings, one phantom finding. When you're measuring, the most expensive mistakes always surface where the tools disagree with each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Counting certificates is not counting trust
&lt;/h2&gt;

&lt;p&gt;After finishing with the system store I looked at the other keychains on the machine, and that's where I got the actual lesson.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/Library/Keychains/System.keychain&lt;/code&gt; holds six certificates. Four are familiar: Apple's system identity, its Kerberos KDC, the developer relations certificate, and the AdGuard root we'll get to below. The fifth is a customer's internal root. The sixth one's name tells me nothing at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;subject= CN=Self-Signed-7C9F7AE3
notBefore= Dec  2 08:06:00 2020 GMT
notAfter = Nov 30 08:06:00 2030 GMT
X509v3 Basic Constraints:
    CA:FALSE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;December 2020. Nearly six years ago. It may predate this machine; it may have been carried over in a migration. I don't know what it is, and writing that I don't know beats inventing a guess.&lt;/p&gt;

&lt;p&gt;But here's the part that matters: &lt;code&gt;CA:FALSE&lt;/code&gt;. This isn't a root, it's a leaf certificate. Its name doesn't appear in any of the three trust domains — I checked all three, zero hits. So it sits on my machine and nothing trusts it. A harmless leftover.&lt;/p&gt;

&lt;p&gt;In my own user keychain things get more instructive still. There are 20 certificates there, and 12 of them are marked &lt;code&gt;CA:TRUE&lt;/code&gt; — structurally capable of signing certificates. But when I look at the trust settings, the user domain contains only &lt;strong&gt;two&lt;/strong&gt; roots.&lt;/p&gt;

&lt;p&gt;I have twelve CA certificates and I trust two of them. The other ten are merely being stored; somebody sent me an organization's root, I dropped it into the keychain, and never trusted it. That's the correct behaviour — a pleasant surprise, even.&lt;/p&gt;

&lt;p&gt;What follows from this is simple, and I had been assuming the opposite: &lt;strong&gt;you cannot measure trust by counting certificates.&lt;/strong&gt; The keychain is a drawer; the trust settings are a separate ledger. Something being in the drawer doesn't mean it's written in the ledger. The sentence "I have 184 certificates on my machine" tells you, on its own, nothing at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Of the 157 roots, zero carry a decision of mine
&lt;/h2&gt;

&lt;p&gt;Apple documents that it splits this list into three categories: trusted roots, ones that always ask, and blocked ones. Its wording for the blocked category is unambiguous — certificates believed to be compromised, which will never be trusted.&lt;/p&gt;

&lt;p&gt;Curious, I went looking for these categories on my own machine. Across all three trust domains (system, admin, user) I couldn't find a single record returning &lt;code&gt;Deny&lt;/code&gt; or an always-ask result. What I found instead was more interesting: &lt;strong&gt;every one&lt;/strong&gt; of the 157 anchors in the system domain looks like this.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cert 5: ISRG Root X1
   Number of trust settings : 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero. No explicit setting — implicit, blanket trust. Not one of the 157 roots carries a decision recorded by me or by any administrator of this machine. That isn't a bad thing in itself; Apple's store is managed through audited, public processes and can be updated remotely. But the subject of the sentence "I trust these" isn't me. I trust a list I inherited.&lt;/p&gt;

&lt;p&gt;So are there any records carrying an explicit decision? Three: two in the user domain, one in the admin domain. All three are mine.&lt;/p&gt;

&lt;p&gt;But the genuinely uncomfortable part surfaced right here. The AdGuard root, the protagonist of this post, also has an explicit setting count of &lt;strong&gt;zero&lt;/strong&gt;. I carry it exactly the way I carry Apple's 157 — without being asked anything. A root I added with my own hands sits in precisely the same silence as the list I inherited.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVsibWFjT1MgdHJ1c3QgZXZhbHVhdGlvbiJdIC0tPiBCWyJTeXN0ZW0gZG9tYWluPGJyLz4xNTcgYW5jaG9yczxici8-ZXhwbGljaXQgc2V0dGluZ3M6IDAiXQogIEEgLS0-IENbIkFkbWluIGRvbWFpbjxici8-MiByb290czxici8-YWRkZWQgYnkgaGFuZCJdCiAgQSAtLT4gRFsiVXNlciBkb21haW48YnIvPjIgcm9vdHM8YnIvPmFkZGVkIGJ5IGhhbmQiXQogIEIgLS0-IEVbIkF1ZGl0ZWQgYnkgQXBwbGU8YnIvPnVwZGF0ZWQgcmVtb3RlbHkiXQogIEMgLS0-IEZbIkF1ZGl0ZWQgYnkgbm9ib2R5Il0KICBEIC0tPiBGCiAgRiAtLT4gR1siVmFsaWQgdW50aWwgMjAzNiBhbmQgMjA0NSJd%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVsibWFjT1MgdHJ1c3QgZXZhbHVhdGlvbiJdIC0tPiBCWyJTeXN0ZW0gZG9tYWluPGJyLz4xNTcgYW5jaG9yczxici8-ZXhwbGljaXQgc2V0dGluZ3M6IDAiXQogIEEgLS0-IENbIkFkbWluIGRvbWFpbjxici8-MiByb290czxici8-YWRkZWQgYnkgaGFuZCJdCiAgQSAtLT4gRFsiVXNlciBkb21haW48YnIvPjIgcm9vdHM8YnIvPmFkZGVkIGJ5IGhhbmQiXQogIEIgLS0-IEVbIkF1ZGl0ZWQgYnkgQXBwbGU8YnIvPnVwZGF0ZWQgcmVtb3RlbHkiXQogIEMgLS0-IEZbIkF1ZGl0ZWQgYnkgbm9ib2R5Il0KICBEIC0tPiBGCiAgRiAtLT4gR1siVmFsaWQgdW50aWwgMjAzNiBhbmQgMjA0NSJd%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The three roots I added myself
&lt;/h2&gt;

&lt;p&gt;Dumping the admin and user domains turned up three distinct roots. I installed all three at some point, and I had forgotten all three.&lt;/p&gt;

&lt;p&gt;The first is my own work: &lt;code&gt;kopru-root-ca&lt;/code&gt;, the root of the köprü platform I use to manage customer servers. Self-signed, &lt;code&gt;CA:TRUE, pathlen:0&lt;/code&gt;, generated on 9 July 2026, valid until 6 July 2036. I put this there deliberately, I know what it does, and I know where the key behind it lives. No complaints.&lt;/p&gt;

&lt;p&gt;The second is a customer's internal root CA. I'm not naming it here — the name of an organization's internal infrastructure doesn't need to appear in my blog post. What makes it interesting is this: it's registered in both the admin &lt;em&gt;and&lt;/em&gt; the user domain. Same root, two places. I don't remember adding it twice. Most likely it went in once by hand and once with a tool's help during some setup.&lt;/p&gt;

&lt;p&gt;The third one actually stopped me: &lt;code&gt;Adguard Personal CA&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  I apparently promised an ad blocker until 2045
&lt;/h2&gt;

&lt;p&gt;I installed AdGuard to cut out ads. A reasonable request. During setup it told me I needed to install a certificate, and I said "sure, it has to filter HTTPS" and approved it. I haven't thought about it once since closing that dialog.&lt;/p&gt;

&lt;p&gt;Today I looked at that certificate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;subject= C=EN, O=AdGuard, CN=Adguard Personal CA
notBefore= Jan 19 22:27:46 2025 GMT
notAfter = Jan 14 22:27:46 2045 GMT
X509v3 Basic Constraints: critical
    CA:TRUE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Twenty years. 2045. The first thing that struck me reading that line was that this laptop will long since have been scrapped by then. The certificate is written to outlive the computer it sits on. I added this root thinking "for now"; the document says "for life."&lt;/p&gt;

&lt;p&gt;And this root is not dormant. AdGuard is running right now, with three processes open including a system extension:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/Applications/Adguard.app/Contents/MacOS/Adguard
com.adguard.mac.adguard.network-extension.systemextension
com.adguard.mac.adguard.helper
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A trusted root flagged &lt;code&gt;CA:TRUE&lt;/code&gt; can sign a certificate for any domain. Including my bank's. I'm not claiming AdGuard is malicious — I have no evidence of that. What I'm saying is structural: if that key is ever leaked or abused, it carries the same authority on my machine as DigiCert. The difference is that DigiCert is subject to public audits, transparency logs and a revocation mechanism. The only person auditing AdGuard's root is me, and I didn't check it once in twenty months.&lt;/p&gt;

&lt;p&gt;CISA's warning on this isn't new — it dates from 2017 and sits in their archive — but the mechanism hasn't changed. The gist is that products performing HTTPS inspection have to install a trusted certificate on the device in order to avoid showing the client warnings, and that many of these products don't properly verify the server's certificate chain. In other words, the layer in the middle may be telling you "this connection is secure" without looking as carefully as you would. From that point on, the padlock describes the intermediary, not the far end.&lt;/p&gt;

&lt;p&gt;This is not a call to delete AdGuard. Filtering ads and trackers is a real benefit and I wanted that benefit. But making the trade knowingly isn't the same as approving it on a setup screen and forgetting. I had done the second.&lt;/p&gt;

&lt;h2&gt;
  
  
  And then there are the expiry dates
&lt;/h2&gt;

&lt;p&gt;Once you start counting it's hard to stop. I pulled the end dates for all 158 certificates: none have expired as of today, and the furthest runs until 9 December 2054. But the nearest is &lt;strong&gt;27 November 2026&lt;/strong&gt; — less than eight weeks away. Three roots will lapse before 2028; the fourth certificate expiring in that window is the Apple intermediate we already met, in February 2027.&lt;/p&gt;

&lt;p&gt;Those dates aren't my problem; Apple handles them with updates. For the three roots I added myself, nobody is standing behind me. I'm the one who has to track what happens to those keys in 2036 and 2045, and I have hard data on my performance so far: twenty months, zero checks.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to read your own list
&lt;/h2&gt;

&lt;p&gt;These commands are all read-only; none of them changes anything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# How many certificates are in the store&lt;/span&gt;
security find-certificate &lt;span class="nt"&gt;-a&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  /System/Library/Keychains/SystemRootCertificates.keychain &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'BEGIN CERTIFICATE'&lt;/span&gt;

&lt;span class="c"&gt;# The anchors the system trusts without question&lt;/span&gt;
security dump-trust-settings &lt;span class="nt"&gt;-s&lt;/span&gt;

&lt;span class="c"&gt;# THE TWO COMMANDS THAT ACTUALLY MATTER: what was added by hand&lt;/span&gt;
security dump-trust-settings &lt;span class="nt"&gt;-d&lt;/span&gt;   &lt;span class="c"&gt;# admin domain&lt;/span&gt;
security dump-trust-settings      &lt;span class="c"&gt;# user domain&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The last two are the important ones. The first two show you Apple's work; the last two show you yours. For every name that comes up you should have an answer to three questions: why did I add this, is it still needed, and when does it expire? If you can't answer all three, that root hasn't earned its place on the list.&lt;/p&gt;

&lt;p&gt;The distinction that helped me decide what to do with those names was this. If there's somebody &lt;em&gt;other than you&lt;/em&gt; behind a root — an audited public CA, your organization's IT team, a process that lands in transparency logs — then there's a structure carrying the risk, and your job is just to know the list. But if you installed the root, you are its only caretaker, and that has three concrete consequences: you have to know where the key lives, you have to put its expiry in your calendar, and you have to delete the root the day you stop using the tool. The third is the most commonly skipped. Removing a tool often doesn't revoke the trust record; the application leaves, the root stays.&lt;/p&gt;

&lt;p&gt;The second distinction concerns intermediary layers. Ad blockers, corporate filters, debugging proxies — they all use the same mechanism and they all raise the same question: is this layer &lt;em&gt;actually&lt;/em&gt; validating the server's certificate on my behalf? That was the sharp edge of CISA's warning. If you don't know the answer, you need to walk back some of the meaning you've attached to that padlock.&lt;/p&gt;

&lt;p&gt;For me, this round came out like this: the köprü root stays, its rationale is clear. I'll deduplicate the customer root's twin records across the two domains. AdGuard's root stays too — but now knowingly; its expiry is going into my calendar.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trust is the sum of decisions you don't remember
&lt;/h2&gt;

&lt;p&gt;What I noticed while writing this is that the number 157 never bothered me. Those aren't my decisions; they're a list Apple audits, can revoke, and updates. There's risk, but there's an owner.&lt;/p&gt;

&lt;p&gt;The three bother me. Because I'm their owner and I was absent. Everyone who likes tinkering with infrastructure has a layer of sediment like this on their machine: a certificate, a proxy, a port opened for a test. Taking &lt;a href="https://mustafaerbay.com.tr/en/blog/technology/secure-boot-tpm-ile-guven-koku/" rel="noopener noreferrer"&gt;root of trust in server infrastructure&lt;/a&gt; seriously while not opening the root store on my own laptop for twenty months is a strange inconsistency. I'd made the same kind of slip while &lt;a href="https://mustafaerbay.com.tr/en/blog/technology/ip-listesi-kimlik-degil-authenticated-origin-pulls/" rel="noopener noreferrer"&gt;writing that identity doesn't come from network location&lt;/a&gt;: there I was insisting that "it came from Cloudflare" doesn't mean "it came for me," while here I handed out twenty years of authority because a setup screen asked for it.&lt;/p&gt;

&lt;p&gt;Most security writing tells you to &lt;em&gt;add&lt;/em&gt; something new. This morning's lesson was the reverse: reading the list once taught me more than installing another tool would have. Four commands, one morning, three forgotten roots.&lt;/p&gt;

&lt;p&gt;Go and look at yours. There's probably somebody on your list too, and you probably invited them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/en-us/103272" rel="noopener noreferrer"&gt;Apple — Available root certificates for Apple operating systems&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/en-us/103723" rel="noopener noreferrer"&gt;Apple — Lists of available trusted root certificates in macOS&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.rfc-editor.org/rfc/rfc5914" rel="noopener noreferrer"&gt;RFC 5914 — Trust Anchor Format&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.cisa.gov/news-events/alerts/2017/03/16/https-interception-weakens-tls-security" rel="noopener noreferrer"&gt;CISA — HTTPS Interception Weakens TLS Security (TA17-075A)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>security</category>
      <category>certificates</category>
      <category>macos</category>
      <category>privacy</category>
    </item>
    <item>
      <title>Forty-Four Posts in a Row From One Category, a Reader Caught It</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Sat, 03 Oct 2026 14:28:05 +0000</pubDate>
      <link>https://dev.to/merbayerp/forty-four-posts-in-a-row-from-one-category-a-reader-caught-it-4lc8</link>
      <guid>https://dev.to/merbayerp/forty-four-posts-in-a-row-from-one-category-a-reader-caught-it-4lc8</guid>
      <description>&lt;p&gt;This blog's publishing pipeline wakes up three times a day: 08:30, 13:30, 16:30.&lt;br&gt;
Each time it picks a topic, researches it, writes it in Turkish and English,&lt;br&gt;
renders a cover, pushes it through the gates and puts it live. Six months of&lt;br&gt;
that. Where it keeps breaking is something I have already&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/career/ajan-yazdi-ben-67-kez-tamir-ettim/" rel="noopener noreferrer"&gt;laid out in numbers&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In the middle of September the pipeline did not miss a single day. Pace on&lt;br&gt;
target, gates green, alarms quiet. And over that same stretch, the 44 posts it&lt;br&gt;
produced all sat in the same category.&lt;/p&gt;

&lt;p&gt;Once the count was in, the number was not the uncomfortable part. The&lt;br&gt;
uncomfortable part was that nothing had broken. Had it broken, it would have&lt;br&gt;
told me. It did not break; it just kept writing into the same shelf.&lt;/p&gt;
&lt;h2&gt;
  
  
  The queue ran out on 3 September
&lt;/h2&gt;

&lt;p&gt;The pipeline has a topic queue: &lt;code&gt;scripts/content-calendar.json&lt;/code&gt;. There are 514&lt;br&gt;
entries in it, 500 of them published. Each entry carries a &lt;code&gt;source&lt;/code&gt; note —&lt;br&gt;
recording where a topic came from is a habit that started embarrassingly late,&lt;br&gt;
but from 3 September onward the notes are explicit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"source": "nobetci-uretimi (kuyruk tukendi, taze konu)"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A second record agrees on the date. When the old generation workflow was retired&lt;br&gt;
on 9 September, the note left at the top of the file reads: "Bu cron ise 3 Eyl'den&lt;br&gt;
beri SIFIR makale uretti: takvim kuyrugu tukendi" — this cron has produced zero&lt;br&gt;
articles since 3 September, the calendar queue is empty. Two separate places,&lt;br&gt;
the same day.&lt;/p&gt;

&lt;p&gt;So no queued topic was left and the pipeline started finding its own. That part&lt;br&gt;
was no secret to me; I simply did not read it as a problem. I did not want&lt;br&gt;
production to stop when the queue emptied, and the fallback behaviour looked&lt;br&gt;
sound: find a mechanism the blog has never covered, verify it against the primary&lt;br&gt;
source, measure it in a lab, write it up.&lt;/p&gt;

&lt;p&gt;The behaviour was sound; the outcome was not.&lt;/p&gt;
&lt;h2&gt;
  
  
  The fingerprint of a reflex
&lt;/h2&gt;

&lt;p&gt;The window from 3 to 22 September — the day the queue emptied through the day the&lt;br&gt;
warning arrived — is 20 days. Counting from the frontmatter gives 65 posts:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Posts&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;technology&lt;/td&gt;
&lt;td&gt;59&lt;/td&gt;
&lt;td&gt;90.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tutorials&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;7.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;life&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;career&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On 15 of those 20 days every single post was &lt;code&gt;technology&lt;/code&gt;. The longest unbroken&lt;br&gt;
run is 8–21 September: &lt;strong&gt;14 days, 44 posts, one category, no interruption.&lt;/strong&gt; The&lt;br&gt;
last &lt;code&gt;career&lt;/code&gt; post before the gap went out on 1 September and the next one on 23&lt;br&gt;
September — 22 days apart. For &lt;code&gt;life&lt;/code&gt; the gap is 20 days, for &lt;code&gt;tutorials&lt;/code&gt; 17.&lt;br&gt;
&lt;code&gt;technology&lt;/code&gt;, meanwhile, never went quiet in that window at all: it published on&lt;br&gt;
all 20 of the 20 days.&lt;/p&gt;

&lt;p&gt;A table does not quite convey a run like that; a list does. The nine posts&lt;br&gt;
published between 13 and 15 September, in order: HSTS preload, Authenticated&lt;br&gt;
Origin Pulls, &lt;code&gt;fs.protected_regular&lt;/code&gt;, kTLS, &lt;code&gt;tcp_ecn&lt;/code&gt;, TCP keepalive, dm-verity,&lt;br&gt;
KSM, &lt;code&gt;ssh -J&lt;/code&gt;. Nine distinct mechanisms, nine distinct measurements — and all&lt;br&gt;
nine on the same shelf. For three days, a reader of this blog was offered exactly&lt;br&gt;
one kind of writing: a lab note that flips a kernel knob and shows the result&lt;br&gt;
differing from expectation. A good kind on its own. Three days running, it stops&lt;br&gt;
being a genre and becomes a tic.&lt;/p&gt;

&lt;p&gt;The tags point the same way. Of those 65 posts, 47 (72.3%) carry the &lt;code&gt;linux&lt;/code&gt; tag&lt;br&gt;
and 28 carry &lt;code&gt;kernel&lt;/code&gt;. Of those 59 &lt;code&gt;technology&lt;/code&gt; posts, 43 carry the &lt;code&gt;linux&lt;/code&gt; tag&lt;br&gt;
and 27 carry &lt;code&gt;kernel&lt;/code&gt; or &lt;code&gt;sysctl&lt;/code&gt;. None of them were bad posts; they&lt;br&gt;
were measured, sourced, tested in my own lab. They were simply all being filed&lt;br&gt;
in one place.&lt;/p&gt;
&lt;h2&gt;
  
  
  A control group: what came out while the queue was full
&lt;/h2&gt;

&lt;p&gt;A number like that is misleading on its own; it needs a comparison. The window&lt;br&gt;
immediately before, 14 August to 2 September, while the queue was still full:&lt;br&gt;
19 days, 59 posts.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Window&lt;/th&gt;
&lt;th&gt;Posts/day&lt;/th&gt;
&lt;th&gt;technology&lt;/th&gt;
&lt;th&gt;tutorials&lt;/th&gt;
&lt;th&gt;career&lt;/th&gt;
&lt;th&gt;life&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;linux&lt;/code&gt; tag&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;14 Aug – 2 Sep (queue full)&lt;/td&gt;
&lt;td&gt;3.11&lt;/td&gt;
&lt;td&gt;39.0%&lt;/td&gt;
&lt;td&gt;44.1%&lt;/td&gt;
&lt;td&gt;10.2%&lt;/td&gt;
&lt;td&gt;6.8%&lt;/td&gt;
&lt;td&gt;3.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3 – 22 Sep (queue empty)&lt;/td&gt;
&lt;td&gt;3.25&lt;/td&gt;
&lt;td&gt;90.8%&lt;/td&gt;
&lt;td&gt;7.7%&lt;/td&gt;
&lt;td&gt;0%&lt;/td&gt;
&lt;td&gt;1.5%&lt;/td&gt;
&lt;td&gt;72.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Exactly one column refuses to move: posts per day. It went from 3.11 to 3.25.&lt;br&gt;
Throughput held; composition collapsed.&lt;/p&gt;

&lt;p&gt;While the queue was full, the variety was not a virtue of mine. The balance was&lt;br&gt;
built months earlier, when the topics were written into the list by hand, and the&lt;br&gt;
list carried it. The list ran out, the balance went with it, and what was left was the reflex.&lt;/p&gt;
&lt;h2&gt;
  
  
  The volume indicator never turned red
&lt;/h2&gt;

&lt;p&gt;The pipeline has gates and they work: the source policy (at least three primary&lt;br&gt;
sources across two distinct hosts), the duplicate scan, word-count and&lt;br&gt;
reading-time consistency, MDX parsing, 93 policy tests, a publication-stalled&lt;br&gt;
alarm, an intraday pace check. Some of them do look at the individual post: the duplicate scan asks "has this&lt;br&gt;
topic run before?", the source policy asks "are the sources in order?" So the&lt;br&gt;
post-scale version of "what did it publish?" is covered. What is never asked is&lt;br&gt;
the scale above it: &lt;strong&gt;the ratio between posts.&lt;/strong&gt; No gate puts two posts side by&lt;br&gt;
side and asks whether they resemble each other.&lt;/p&gt;

&lt;p&gt;One of them deserves a closer look, because it is the watchman that counts the&lt;br&gt;
published output every hour: &lt;code&gt;daily-pace-check&lt;/code&gt;. The comment at the top of the&lt;br&gt;
file states its job in one line — it checks the maximum-four-posts-per-day rule&lt;br&gt;
once an hour. The step itself does exactly that: it counts the commits on &lt;code&gt;main&lt;/code&gt;&lt;br&gt;
since the start of the day whose message begins with &lt;code&gt;feat: yeni makale&lt;/code&gt;, and&lt;br&gt;
errors out only if the count &lt;strong&gt;exceeds&lt;/strong&gt; four. GitHub's documentation defines the&lt;br&gt;
&lt;code&gt;schedule&lt;/code&gt; event in a single sentence of its own: "The &lt;code&gt;schedule&lt;/code&gt; event allows you&lt;br&gt;
to trigger a workflow at a scheduled time." What fires it is not a judgement about&lt;br&gt;
content but a clock — and the same page noting that the trigger can be delayed&lt;br&gt;
under high load, with queued jobs sometimes dropped, is a reminder that it is a&lt;br&gt;
timer rather than an opinion. For two weeks that clock kept finding a number that&lt;br&gt;
did not exceed four — on 10 and 17 September exactly four posts went out, and&lt;br&gt;
neither day broke the rule. It was entirely right to stay silent; that was the&lt;br&gt;
question I had given it.&lt;/p&gt;

&lt;p&gt;Google's SRE book names two questions for a monitoring system: "what's broken,&lt;br&gt;
and why?" My indicators answered both correctly — nothing was broken. But the&lt;br&gt;
SLO chapter of the SRE workbook has a sharper sentence: if a disruption is not&lt;br&gt;
captured by any indicator, that is "a strong sign that your SLO lacks coverage".&lt;br&gt;
The gap in my coverage had a name, and the name was composition.&lt;/p&gt;

&lt;p&gt;The Prometheus alerting guide lands in the same place: alert on symptoms rather&lt;br&gt;
than causes, and aim for "as few alerts as possible". Sound advice. But defining&lt;br&gt;
the symptom is my job, and I had defined it as "publishing stopped". Publishing&lt;br&gt;
did not stop. The blog narrowed to one subject — a symptom that definition&lt;br&gt;
simply did not contain.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVtTY2hlZHVsZXIgd2FrZXMgdGhyZWUgdGltZXMgYSBkYXldIC0tPiBCe0FueSB0b3BpYyBxdWV1ZWR9CiAgQiAtLSBZZXMgLS0-IENbVGFrZSB0aGUgbmV4dCBvbmUgaW4gbGluZV0KICBCIC0tIE5vIC0tPiBEW1BpcGVsaW5lIGZpbmRzIGl0cyBvd24gdG9waWNdCiAgRCAtLT4gRVtDaGVhcGVzdCB2ZXJpZmlhYmxlIGFuZ2xlOjxici8-YSBrZXJuZWwgc2V0dGluZyBtZWFzdXJhYmxlPGJyLz5pbiBvbmUgc2Vzc2lvbl0KICBDIC0tPiBGW0dlbmVyYXRpb24gYW5kIGdhdGVzXQogIEUgLS0-IEYKICBGIC0tPiBHe1doYXQgZG8gdGhlIGdhdGVzIG1lYXN1cmV9CiAgRyAtLT4gSFtQdWJsaXNoZWQgwrcgc291cmNlIGNvdW50PGJyLz7CtyBkdXBsaWNhdGVzIMK3IHdvcmQgY291bnRdCiAgSCAtLT4gSVtWb2x1bWUgZ3JlZW4sIGFsYXJtcyBxdWlldF0KICBJIC0tPiBKW0NvbXBvc2l0aW9uIG5ldmVyIG1lYXN1cmVkXQ%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVtTY2hlZHVsZXIgd2FrZXMgdGhyZWUgdGltZXMgYSBkYXldIC0tPiBCe0FueSB0b3BpYyBxdWV1ZWR9CiAgQiAtLSBZZXMgLS0-IENbVGFrZSB0aGUgbmV4dCBvbmUgaW4gbGluZV0KICBCIC0tIE5vIC0tPiBEW1BpcGVsaW5lIGZpbmRzIGl0cyBvd24gdG9waWNdCiAgRCAtLT4gRVtDaGVhcGVzdCB2ZXJpZmlhYmxlIGFuZ2xlOjxici8-YSBrZXJuZWwgc2V0dGluZyBtZWFzdXJhYmxlPGJyLz5pbiBvbmUgc2Vzc2lvbl0KICBDIC0tPiBGW0dlbmVyYXRpb24gYW5kIGdhdGVzXQogIEUgLS0-IEYKICBGIC0tPiBHe1doYXQgZG8gdGhlIGdhdGVzIG1lYXN1cmV9CiAgRyAtLT4gSFtQdWJsaXNoZWQgwrcgc291cmNlIGNvdW50PGJyLz7CtyBkdXBsaWNhdGVzIMK3IHdvcmQgY291bnRdCiAgSCAtLT4gSVtWb2x1bWUgZ3JlZW4sIGFsYXJtcyBxdWlldF0KICBJIC0tPiBKW0NvbXBvc2l0aW9uIG5ldmVyIG1lYXN1cmVkXQ%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="555" height="1366"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  The warning came from a reader
&lt;/h2&gt;

&lt;p&gt;The best evidence of the pipeline's blindness to composition sits in the&lt;br&gt;
calendar's own record. The &lt;code&gt;source&lt;/code&gt; note on the topic pool added on 23 September&lt;br&gt;
begins like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"KONU HAVUZU (23 Eyl 2026, kullanici talebi: makaleler hep technology
cikiyordu, 7 Eyl'den beri kategori cesitliligi yok)"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Translated: topic pool, 23 September 2026, &lt;em&gt;user request&lt;/em&gt; — the articles kept&lt;br&gt;
coming out as technology, no category variety since 7 September. The note is&lt;br&gt;
unambiguous about the source of the signal. What flagged the broken balance was&lt;br&gt;
not an indicator, a test or a report; it was somebody reading the blog. From 7 to&lt;br&gt;
23 September is 16 days. That is the detection latency: sixteen days and zero&lt;br&gt;
automatic signals.&lt;/p&gt;

&lt;p&gt;There is no soft way to phrase this. The person who has spent six months&lt;br&gt;
counting and writing up every place this pipeline breaks failed to catch, by his&lt;br&gt;
own measurement, the crudest deviation in its most visible output. And what the&lt;br&gt;
side that did catch it said was not a subtle metric: the articles keep coming out&lt;br&gt;
as technology.&lt;/p&gt;

&lt;p&gt;One correction I owe here: "since 7 September" is the note's own phrasing, not my&lt;br&gt;
measurement. One of the four posts published on 7 September is &lt;code&gt;tutorials&lt;/code&gt;; the&lt;br&gt;
unbroken run begins the next day. What the record supports is this: 15 days&lt;br&gt;
counting from the start of the run, or 13 days counting from the day the detector&lt;br&gt;
I am about to build would have fired. Either way, the side that produced the&lt;br&gt;
signal does not change.&lt;/p&gt;
&lt;h2&gt;
  
  
  Back-testing the detector I did not have
&lt;/h2&gt;

&lt;p&gt;"Which indicator would have seen this?" is a question I did not want to leave&lt;br&gt;
hanging, because questions like that usually hang there and quietly die as good&lt;br&gt;
intentions. So here is a simple rule: slide a seven-day window, and if it holds at&lt;br&gt;
least seven posts, compute the category shares; if the dominant category crosses a&lt;br&gt;
threshold, warn. Then apply the rule backwards to every day from 1 April to 3&lt;br&gt;
October — a range covering 1,276 of the archive's 1,308 posts; the remaining 32&lt;br&gt;
carry older dates.&lt;/p&gt;

&lt;p&gt;At a 70% threshold: 22 warning days in six months, in three clusters — 4–7 August&lt;br&gt;
(dominant &lt;code&gt;tutorials&lt;/code&gt;, 19 of 27 posts), 26–27 August (12 of 17) and 9–24&lt;br&gt;
September (16 days; at first trigger, 19 of 24 posts &lt;code&gt;technology&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;At 80% a single cluster survives: 10–24 September. No other day fires at all.&lt;br&gt;
So a threshold exists that drives false alarms to zero, and that threshold would&lt;br&gt;
have caught the deviation on &lt;strong&gt;10 September&lt;/strong&gt;. The reader spoke on 23 September.&lt;br&gt;
Thirteen days apart, and 41 more posts went out in that gap.&lt;/p&gt;

&lt;p&gt;The weakness of this test belongs in the article too: the threshold was found in&lt;br&gt;
data where the deviation was already known. 80% fits this six-month archive well;&lt;br&gt;
a week whose topic distribution narrows by design — a series going deep on one&lt;br&gt;
product, a cluster of posts after an incident — would trip the same threshold&lt;br&gt;
unfairly. So the right shape is not "warn and halt" but "warn and ask": a&lt;br&gt;
question about whether the narrowing is a deliberate choice or a reflex. Still,&lt;br&gt;
to claim the threshold sits roughly in the right place, there is six months of&lt;br&gt;
real trigger history behind it, which beats having none.&lt;/p&gt;

&lt;p&gt;There is also a more honest route for anyone picking a threshold without the&lt;br&gt;
benefit of hindsight: instead of writing a fixed number, let your own history set&lt;br&gt;
it. Apply the same sliding window to the past, derive the distribution of the&lt;br&gt;
dominant category's share, and make an upper percentile — p95, say — the&lt;br&gt;
threshold; your normal gets defined by you, rather than borrowing my 80%. Window&lt;br&gt;
length follows the same logic: a window needs at least a dozen units in it, or a&lt;br&gt;
two-item day reads as "monoculture" by chance. For a pipeline producing three&lt;br&gt;
posts a day, seven days clears that bar; for somebody closing two tickets a week,&lt;br&gt;
the same rule needs a thirty-day window or it never fires at all.&lt;/p&gt;

&lt;p&gt;The count itself is about as hard as a shell loop. Category and date are already&lt;br&gt;
in every post's frontmatter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;for &lt;/span&gt;f &lt;span class="k"&gt;in &lt;/span&gt;src/content/blog/&lt;span class="k"&gt;*&lt;/span&gt;/&lt;span class="k"&gt;*&lt;/span&gt;.mdx&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  case&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt;.en.mdx&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;;&lt;/span&gt; &lt;span class="k"&gt;esac&lt;/span&gt;
  &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="nt"&gt;-F&lt;/span&gt;&lt;span class="s1"&gt;': *'&lt;/span&gt; &lt;span class="s1"&gt;'/^publishDate:/{d=$2} /^category:/{gsub(/"/,"",$2); c=$2} END{print d, c}'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;done&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'$1&amp;gt;="2026-09-03" &amp;amp;&amp;amp; $1&amp;lt;="2026-09-22" {n[$2]++; t++}
                   END{for (k in n) printf "%-11s %3d  %%%.1f\n", k, n[k], 100*n[k]/t}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;technology   59  %90.8
tutorials     5  %7.7
life          1  %1.5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That a query this small went unrun for sixteen days is the most expensive detail&lt;br&gt;
in this piece. Measuring was not hard; thinking of measuring was hard — because&lt;br&gt;
every counter was green, and a green counter does not provoke questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the reflex kept going to the same place
&lt;/h2&gt;

&lt;p&gt;This is the part that interests me most, because this part is mine, not the&lt;br&gt;
pipeline's.&lt;/p&gt;

&lt;p&gt;The topic-finding method goes like this: grep the archive, find a mechanism never&lt;br&gt;
covered, verify it against the primary source, measure it. That method behaves&lt;br&gt;
like a cost function — and kernel settings are by far its cheapest candidates.&lt;br&gt;
You flip a sysctl in one session, measure, compare, and put the evidence on&lt;br&gt;
screen. Cheap to verify, fresh evidence, low risk of being wrong.&lt;/p&gt;

&lt;p&gt;Career and life pieces demand the same evidentiary bar at a far higher price.&lt;br&gt;
The evidence for a kernel setting is on my own machine, now, in the output of one&lt;br&gt;
command. The evidence for a career piece is scattered across six months of git&lt;br&gt;
history, CI logs, the actual state of servers, or 1,300 published posts. It has&lt;br&gt;
to be searched, extracted, counted, and the count verified twice — the windows,&lt;br&gt;
the control group and the back-test in this article are exactly that expensive&lt;br&gt;
part. The method was picking the most easily verifiable topic; I had been reading&lt;br&gt;
that as the best topic.&lt;/p&gt;

&lt;p&gt;A distinction is worth drawing here: the reflex did not produce bad work. I still&lt;br&gt;
stand behind the technical content of those 44 posts; they were measured and&lt;br&gt;
their sources held. What the reflex broke was not individual posts but the ratio&lt;br&gt;
between them. And a ratio is a property no single post contains — it exists only&lt;br&gt;
when you look from above.&lt;/p&gt;

&lt;p&gt;The chapter on automation in the SRE book fits precisely here: automation is a&lt;br&gt;
"force multiplier, not a panacea", and much of its value comes from consistency —&lt;br&gt;
"very few of us will ever be as consistent as a machine". The book scopes that&lt;br&gt;
consistency narrowly: the execution of well-scoped, known procedures. What my&lt;br&gt;
pipeline was executing was not a procedure but a choice. The catch is that&lt;br&gt;
consistency never asks whether the decision was right. The pipeline applied my&lt;br&gt;
preference 44 times in a row with flawless fidelity. A human would have grown&lt;br&gt;
tired, grown bored, switched subjects. The machine did not get bored.&lt;/p&gt;

&lt;p&gt;I also do not think this failure is specific to automation. Looking at my own&lt;br&gt;
weeks, the same curve shows up: what I take up to learn is usually dictated not&lt;br&gt;
by curiosity but by how provable the thing is. I go where I can measure. Part of&lt;br&gt;
what I call expertise is the sediment of a reflex like that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing the rule into the system: compute, do not choose
&lt;/h2&gt;

&lt;p&gt;Two things happened on 23 September. A new topic pool went into the calendar,&lt;br&gt;
with category balance taken into account. And a step that runs &lt;strong&gt;before&lt;/strong&gt; topic&lt;br&gt;
selection was added to the production recipe: compute the target category from&lt;br&gt;
the last-written date, take the one that has gone longest without a post, and&lt;br&gt;
produce only there on that run.&lt;/p&gt;

&lt;p&gt;The detail that matters is that the rule is code rather than a sentence. Writing&lt;br&gt;
"mind the category variety" into the recipe would have achieved nothing, because&lt;br&gt;
minding runs at the same moment as the reflex that finds the cheapest topic, and&lt;br&gt;
it loses that race. The rule is now a step that computes the target category from&lt;br&gt;
the calendar and imposes it. Not a statement of intent; a gate.&lt;/p&gt;

&lt;p&gt;The price of that distinction has come due before: there was a line in the deploy&lt;br&gt;
workflow I took for a guard, and months later it turned out to be&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/career/doksan-saniyede-geri-aldim-bes-saat-bozuk-kaldi/" rel="noopener noreferrer"&gt;blocking nothing at all&lt;/a&gt;.&lt;br&gt;
A rule being written down does not mean it is enforced. That post closed on&lt;br&gt;
"being able to count the places the gates do not look"; the difference here is&lt;br&gt;
that this time the counting actually happened instead of the sentence being&lt;br&gt;
repeated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eleven days later — and the limits of this measurement
&lt;/h2&gt;

&lt;p&gt;From 23 September to 3 October: 11 days, 34 posts.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Posts&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;technology&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;29.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;tutorials&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;23.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;career&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;23.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;life&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;23.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This table does not count the post you are reading; with it the window holds 35&lt;br&gt;
posts, &lt;code&gt;career&lt;/code&gt; rises to 9 (25.7%) and posts per day to 3.18.&lt;/p&gt;

&lt;p&gt;On none of those 11 days did a day stay in a single category; in the previous&lt;br&gt;
window that ratio was 15 of 20. The average number of distinct categories per day&lt;br&gt;
went from 1.25 to 2.82, and the share of posts tagged &lt;code&gt;linux&lt;/code&gt; fell from 72.3% to&lt;br&gt;
44.1%.&lt;/p&gt;

&lt;p&gt;One more look at throughput, this time with the same denominator across all three&lt;br&gt;
windows — per calendar day, counting days with no publication too: 2.95 · 3.25 ·&lt;br&gt;
3.18. A number moving inside a ten-percent band. Composition collapsed and then&lt;br&gt;
recovered while volume did not budge; the balance was not bought by cutting&lt;br&gt;
output.&lt;/p&gt;

&lt;p&gt;Now, the way this table should not be read: "the rule worked, here is the proof."&lt;br&gt;
Two things changed on the same day; a topic pool built with category balance also&lt;br&gt;
went in. The pool alone could have produced this table. A measurement separating&lt;br&gt;
the two is not something I have, and not something I set up. Without a&lt;br&gt;
discriminating test there is no telling which intervention is doing the work —&lt;br&gt;
the same trap I fell into&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/career/dogru-hipotezi-iki-ay-erken-kurdum/" rel="noopener noreferrer"&gt;two months earlier with a LinkedIn hypothesis&lt;/a&gt;. And 11 days is a short window next to 20.&lt;/p&gt;

&lt;p&gt;The honest statement: composition recovered, and the cause cannot be assigned to&lt;br&gt;
a single intervention.&lt;/p&gt;

&lt;p&gt;And one more correction. "We will see when the pool runs dry again" is what I was&lt;br&gt;
about to write; the calendar says it has run dry &lt;strong&gt;already&lt;/strong&gt;. Of 514 entries, 500&lt;br&gt;
are generated and the remaining 14 are rejected — zero topics waiting. The 23&lt;br&gt;
September pool lasted eleven days. So the rule is already running on its own, and&lt;br&gt;
the post you are reading is the product of exactly such a run: rotation said&lt;br&gt;
"career", the career queue was empty, and the pipeline found the topic itself.&lt;br&gt;
This time what it found was not a kernel setting but its own publishing record.&lt;/p&gt;

&lt;p&gt;The real gain from this piece is not even the rotation rule&lt;br&gt;
— it is that a query which measures composition now exists, and noticing a&lt;br&gt;
fourteen-day run no longer requires waiting for a reader.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reflex moved up a level
&lt;/h2&gt;

&lt;p&gt;Running the same eye over the draft of this post produced a finding that stung,&lt;br&gt;
because it was right. All nine &lt;code&gt;career&lt;/code&gt; posts published since 23 September carry a&lt;br&gt;
counted number in the title: 67 repairs, seven servers, two months, ninety&lt;br&gt;
seconds, three scripts, twenty findings, seventy branches, sixty-eight backups —&lt;br&gt;
and the forty-four posts you are reading about. The category field recovered; the&lt;br&gt;
title shape settled into a single mould.&lt;/p&gt;

&lt;p&gt;The rotation rule looks at the &lt;code&gt;category&lt;/code&gt; field. It does not look at the shape of&lt;br&gt;
the title, the structure of the piece, or the kind of evidence used. The reflex&lt;br&gt;
did not disappear; it moved outside the field being measured. Put the constraint&lt;br&gt;
anywhere and repetition accumulates one notch above it — and my new composition&lt;br&gt;
query is blind to this by construction, because it too counts only categories.&lt;br&gt;
Starting to measure one monoculture does not mean you have started counting the&lt;br&gt;
dimensions you are not measuring.&lt;/p&gt;

&lt;p&gt;This paragraph is not here as a gesture of modesty. If the thesis holds, its first&lt;br&gt;
casualty should be the post itself: your coverage reaches exactly as far as the&lt;br&gt;
dimension you measure, and the next blind spot is waiting one level above wherever&lt;br&gt;
you put the gate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Applying it to your own work
&lt;/h2&gt;

&lt;p&gt;For anyone reading this without a blog pipeline, the same questions in short&lt;br&gt;
form:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What is the composition of your last 20 units?&lt;/strong&gt; The last 20 tickets you
closed, the last 20 reports you wrote, the last 20 proposals you sent. Split
them into categories and count. Not volume — distribution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Which of your indicators would have seen this drift?&lt;/strong&gt; If none would,
"everything is green" means no more than "the three things I measure are fine".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What is carrying your balance?&lt;/strong&gt; A list, a plan, a client portfolio — or
your appetite on the day? When the list ends, the balance ends with it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is cost making your choices?&lt;/strong&gt; Easily verified work is not the same as the
right work. Your reflex goes to the cheapest evidence, and the shape of your
expertise comes out of that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Did you write a rule or a gate?&lt;/strong&gt; "I will be careful" is not a rule. A rule
is computed, imposed, and visible when skipped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who is your detector?&lt;/strong&gt; If the answer is a person — a reader, a client, a
teammate — that person is your monitoring system, and it does not scale.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;A system looking healthy does not mean it is doing the right thing; it means it&lt;br&gt;
answers the question I gave it well. An indicator that measures volume cannot, by&lt;br&gt;
definition, see composition collapse: with 44 posts out of one category, the&lt;br&gt;
counter still writes the same three.&lt;/p&gt;

&lt;p&gt;Variety is not a by-product of productivity. What arrives as a by-product is&lt;br&gt;
repetition — because repetition is cheaper, faster, and its evidence is already&lt;br&gt;
at hand. If balance is wanted, it has to be written somewhere as a constraint,&lt;br&gt;
into the system or into the day. Otherwise the reflex does the choosing, and a&lt;br&gt;
reflex always reaches for the nearest shelf.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://sre.google/sre-book/monitoring-distributed-systems/" rel="noopener noreferrer"&gt;SRE Book — Monitoring Distributed Systems: the two questions monitoring must answer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sre.google/workbook/implementing-slos/" rel="noopener noreferrer"&gt;SRE Workbook — Implementing SLOs: a disruption no indicator captures is a coverage gap&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://sre.google/sre-book/automation-at-google/" rel="noopener noreferrer"&gt;SRE Book — Automation at Google: the value of consistency and how automation spreads mistakes at scale&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://prometheus.io/docs/practices/alerting/" rel="noopener noreferrer"&gt;Prometheus — Alerting: alert on symptoms and keep the number of alerts small&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/actions/reference/workflows-and-actions/events-that-trigger-workflows" rel="noopener noreferrer"&gt;GitHub Actions — the &lt;code&gt;schedule&lt;/code&gt; event: triggers a workflow at a scheduled time, and can be delayed under load&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>yazilimkariyeri</category>
      <category>otomasyon</category>
      <category>olcum</category>
      <category>process</category>
    </item>
    <item>
      <title>Who Inherits the Queue of a Closing Socket?</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Sat, 03 Oct 2026 13:28:57 +0000</pubDate>
      <link>https://dev.to/merbayerp/who-inherits-the-queue-of-a-closing-socket-42ji</link>
      <guid>https://dev.to/merbayerp/who-inherits-the-queue-of-a-closing-socket-42ji</guid>
      <description>&lt;p&gt;The favourite use of &lt;code&gt;SO_REUSEPORT&lt;/code&gt; goes like this: N worker processes listen on the same port, the kernel spreads incoming connections among them, and when you roll out a new version you replace the workers one at a time. The new worker comes up, the old one closes, the port is never idle. Wherever it is described, it sounds like zero downtime.&lt;/p&gt;

&lt;p&gt;It is not. The connections waiting in the closing worker's &lt;code&gt;accept()&lt;/code&gt; queue — connections that have &lt;strong&gt;finished&lt;/strong&gt; their handshake and that the application has not yet picked up — die with that socket. On the client side this is an RST; on the server side it is nothing at all. It never reaches the access log, because the application never saw that connection.&lt;/p&gt;

&lt;p&gt;The kernel has a switch for this: &lt;code&gt;net.ipv4.tcp_migrate_req&lt;/code&gt;. It has been there since Linux 5.14, it defaults to off, and its documentation drifts from the code in at least two places. So I sat down and poked at it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A note on the terminal output below: the measurement harness is mine and prints Turkish labels. &lt;code&gt;B kabul etti&lt;/code&gt; means "B accepted", &lt;code&gt;gocen&lt;/code&gt; means "migrated", &lt;code&gt;dagilim&lt;/code&gt; means "distribution", &lt;code&gt;istemci kaybi&lt;/code&gt; means "clients lost", and &lt;code&gt;sunucu tarafi&lt;/code&gt; means "server side". The numbers and labels are the harness's own; where a block is long I trimmed trailing fields, but nothing was retyped or translated.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The lab, and an honest disclaimer
&lt;/h2&gt;

&lt;p&gt;Let me say this first: this is not an outage story. I did not hit this problem on my own server, because there is no reuseport group there. Here is the reading:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;ssh vps3 &lt;span class="s1"&gt;'ss -ltnH | awk "{print \$4}" | sort | uniq -c | awk "\$1&amp;gt;1"'&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;ssh vps3 &lt;span class="s1"&gt;'ss -ltnH | wc -l'&lt;/span&gt;
90
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ninety listening sockets, and not a single address:port pair shared by two of them. nginx 1.30.0 is running, but neither &lt;code&gt;reuseport&lt;/code&gt; nor &lt;code&gt;backlog&lt;/code&gt; appears anywhere in its configuration — one shared socket, the default queue of 511. &lt;code&gt;tcp_migrate_req&lt;/code&gt; is off too, and both counters sit at zero:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;ssh vps3 &lt;span class="s1"&gt;'uname -r; cat /proc/sys/net/ipv4/tcp_migrate_req; nstat -az | grep -i migrate'&lt;/span&gt;
6.8.0-142-generic
0
TcpExtTCPMigrateReqSuccess      0                  0.0
TcpExtTCPMigrateReqFailure      0                  0.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rather than flipping this on a live server to see what happens, I took it to the lab. The environment is Docker Desktop's linuxkit VM:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;
6.10.14-linuxkit
&lt;span class="nv"&gt;$ &lt;/span&gt;python3 &lt;span class="nt"&gt;-V&lt;/span&gt;
Python 3.12.3
&lt;span class="nv"&gt;$ &lt;/span&gt;ss &lt;span class="nt"&gt;-V&lt;/span&gt;
ss utility, iproute2-6.1.0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using containers has a technical reason: &lt;code&gt;tcp_migrate_req&lt;/code&gt; is network-namespaced (&lt;code&gt;net-&amp;gt;ipv4.sysctl_tcp_migrate_req&lt;/code&gt;), so &lt;code&gt;docker run --sysctl&lt;/code&gt; sets it per container and leaves the host's value untouched. I could run one round off and one round on with nothing leaking between them. The harness is two Python scripts, 311 lines together: they open listeners with &lt;code&gt;SO_REUSEPORT&lt;/code&gt;, establish client connections, count who accepts what, and read counter deltas from &lt;code&gt;/proc/net/netstat&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  First measurement: is the loss real?
&lt;/h2&gt;

&lt;p&gt;The setup is simple. Listener A is up alone with a backlog of 128. Twenty clients connect, and each one sends a few bytes &lt;strong&gt;as soon as the connection is established&lt;/strong&gt;. A accepts none of them; all twenty sit in the queue. Then listener B joins the same port, and A closes. B tries to drain the queue.&lt;/p&gt;

&lt;p&gt;Before A closes, &lt;code&gt;ss&lt;/code&gt; shows the state clearly — the &lt;code&gt;Recv-Q&lt;/code&gt; column is the number of connections waiting in the queue:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LISTEN 20     128    127.0.0.1:18081 0.0.0.0:*
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here is the result with both settings side by side:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;tcp_migrate_req=0&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;tcp_migrate_req=1&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Accepted by B&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 / 20&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20 / 20&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clients that got their data back&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clients reset&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;TCPMigrateReqSuccess&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;TCPMigrateReqFailure&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A small caveat on that measurement: my harness drops both &lt;code&gt;ECONNRESET&lt;/code&gt; and a clean end-of-file read into the same "reset" bucket, so it does not prove an RST at packet level. The source makes clear it is an RST — the &lt;code&gt;inet_child_forget&lt;/code&gt; path leads to &lt;code&gt;tcp_send_active_reset&lt;/code&gt; — but I know that from the code, not from my own measurement.&lt;/p&gt;

&lt;p&gt;With the switch off, all twenty connections were gone. With it on, all twenty moved to B — and note this: the data the clients sent &lt;strong&gt;before&lt;/strong&gt; the migration moved with them, so B was able to answer all twenty requests. The socket is not handed over, it is &lt;strong&gt;moved&lt;/strong&gt;, receive buffer included.&lt;/p&gt;

&lt;p&gt;There is one more thing in that table, and it is the part that bothered me: in the "off" round, twenty connections died while both counters stayed at zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mechanism: when does the socket leave the group
&lt;/h2&gt;

&lt;p&gt;To understand the silent counter I had to read the close path. &lt;code&gt;__tcp_close&lt;/code&gt; runs a special branch for listening sockets, and the order is this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sk&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;sk_state&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;TCP_LISTEN&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;tcp_set_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sk&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;TCP_CLOSE&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="cm"&gt;/* Special case. */&lt;/span&gt;
        &lt;span class="n"&gt;inet_csk_listen_stop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sk&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the state change comes &lt;strong&gt;before&lt;/strong&gt; the queue is drained. For &lt;code&gt;TCP_CLOSE&lt;/code&gt;, &lt;code&gt;tcp_set_state&lt;/code&gt; removes the socket from the hash table (&lt;code&gt;sk-&amp;gt;sk_prot-&amp;gt;unhash(sk)&lt;/code&gt;), which calls &lt;code&gt;reuseport_stop_listen_sock()&lt;/code&gt; by way of &lt;code&gt;inet_unhash&lt;/code&gt;. That is where the fork in the road is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;READ_ONCE&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sock_net&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;ipv4&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sysctl_tcp_migrate_req&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prog&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;prog&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;expected_attach_type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;BPF_SK_REUSEPORT_SELECT_OR_MIGRATE&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="cm"&gt;/* Migration capable, move sk from the listening section
         * to the closed section.
         */&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the switch is on, the closing socket is not thrown out of the group; it moves from the group's "listening" section to its "closed" section, and the &lt;code&gt;sk-&amp;gt;sk_reuseport_cb&lt;/code&gt; pointer &lt;strong&gt;survives&lt;/strong&gt;. If the switch is off, it is, as the comment puts it, "detach immediately" — the pointer becomes &lt;code&gt;NULL&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Then &lt;code&gt;inet_csk_listen_stop&lt;/code&gt; calls &lt;code&gt;reuseport_migrate_sock()&lt;/code&gt; for every child in the queue. The first thing that function does is look up the group:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="n"&gt;reuse&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rcu_dereference&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sk&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;sk_reuseport_cb&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;reuse&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;goto&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;out&lt;/code&gt; label does not touch any counter. The label that increments is &lt;code&gt;failure&lt;/code&gt;, and you can only reach it after the group has been found. So with the switch off, the kernel does not even &lt;strong&gt;attempt&lt;/strong&gt; the migration — and because it does not attempt it, it does not count it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBBWyJjbG9zZSgpIG9yIHNodXRkb3duKCkiXSAtLT4gQlsidGNwX3NldF9zdGF0ZShzaywgVENQX0NMT1NFKSJdCiAgICBCIC0tPiBDWyJpbmV0X3VuaGFzaCB0aGVuIHJldXNlcG9ydF9zdG9wX2xpc3Rlbl9zb2NrKCkiXQogICAgQyAtLT4gRHsidGNwX21pZ3JhdGVfcmVxPTE8YnIvPm9yIGEgTUlHUkFURSBlQlBGIHByb2dyYW0gYXR0YWNoZWQ_In0KICAgIEQgLS0-fG5vfCBFWyJyZXVzZXBvcnRfZGV0YWNoX3NvY2soKTxici8-c2tfcmV1c2Vwb3J0X2NiID0gTlVMTCJdCiAgICBEIC0tPnx5ZXN8IEZbInNvY2tldCBtb3ZlcyB0byB0aGUgY2xvc2VkIHNlY3Rpb24sPGJyLz5ncm91cCBwb2ludGVyIHN1cnZpdmVzIl0KICAgIEUgLS0-IEdbImluZXRfY3NrX2xpc3Rlbl9zdG9wKCkiXQogICAgRiAtLT4gRwogICAgRyAtLT4gSFsiZm9yIGV2ZXJ5IGNoaWxkIGluIHRoZSBxdWV1ZTxici8-cmV1c2Vwb3J0X21pZ3JhdGVfc29jaygpIl0KICAgIEggLS0-IEl7ImRvZXMgdGhlIGdyb3VwIHBvaW50ZXIgc3Vydml2ZT8ifQogICAgSSAtLT58bm98IEpbInNpbGVudCBOVUxMPGJyLz5ubyBjb3VudGVyLCBjb25uZWN0aW9uIFJTVCJdCiAgICBJIC0tPnx5ZXN8IEt7ImFueSBsaXZlIGxpc3RlbmVyPyJ9CiAgICBLIC0tPnxub25lfCBMWyJUQ1BNaWdyYXRlUmVxRmFpbHVyZSsrPGJyLz5jb25uZWN0aW9uIFJTVCJdCiAgICBLIC0tPnx5ZXN8IE1bInRhcmdldCBwaWNrZWQgYnkgaGFzaCJdCiAgICBNIC0tPiBOeyJjb3VsZCB0aGUgY2hpbGQgYmUgYWRkZWQ8YnIvPnRvIHRoZSB0YXJnZXQgcXVldWU_In0KICAgIE4gLS0-fHllc3wgT1siVENQTWlncmF0ZVJlcVN1Y2Nlc3MrKyJdCiAgICBOIC0tPnxub3wgUFsiVENQTWlncmF0ZVJlcUZhaWx1cmUrKzxici8-Y29ubmVjdGlvbiBSU1QiXQ%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBBWyJjbG9zZSgpIG9yIHNodXRkb3duKCkiXSAtLT4gQlsidGNwX3NldF9zdGF0ZShzaywgVENQX0NMT1NFKSJdCiAgICBCIC0tPiBDWyJpbmV0X3VuaGFzaCB0aGVuIHJldXNlcG9ydF9zdG9wX2xpc3Rlbl9zb2NrKCkiXQogICAgQyAtLT4gRHsidGNwX21pZ3JhdGVfcmVxPTE8YnIvPm9yIGEgTUlHUkFURSBlQlBGIHByb2dyYW0gYXR0YWNoZWQ_In0KICAgIEQgLS0-fG5vfCBFWyJyZXVzZXBvcnRfZGV0YWNoX3NvY2soKTxici8-c2tfcmV1c2Vwb3J0X2NiID0gTlVMTCJdCiAgICBEIC0tPnx5ZXN8IEZbInNvY2tldCBtb3ZlcyB0byB0aGUgY2xvc2VkIHNlY3Rpb24sPGJyLz5ncm91cCBwb2ludGVyIHN1cnZpdmVzIl0KICAgIEUgLS0-IEdbImluZXRfY3NrX2xpc3Rlbl9zdG9wKCkiXQogICAgRiAtLT4gRwogICAgRyAtLT4gSFsiZm9yIGV2ZXJ5IGNoaWxkIGluIHRoZSBxdWV1ZTxici8-cmV1c2Vwb3J0X21pZ3JhdGVfc29jaygpIl0KICAgIEggLS0-IEl7ImRvZXMgdGhlIGdyb3VwIHBvaW50ZXIgc3Vydml2ZT8ifQogICAgSSAtLT58bm98IEpbInNpbGVudCBOVUxMPGJyLz5ubyBjb3VudGVyLCBjb25uZWN0aW9uIFJTVCJdCiAgICBJIC0tPnx5ZXN8IEt7ImFueSBsaXZlIGxpc3RlbmVyPyJ9CiAgICBLIC0tPnxub25lfCBMWyJUQ1BNaWdyYXRlUmVxRmFpbHVyZSsrPGJyLz5jb25uZWN0aW9uIFJTVCJdCiAgICBLIC0tPnx5ZXN8IE1bInRhcmdldCBwaWNrZWQgYnkgaGFzaCJdCiAgICBNIC0tPiBOeyJjb3VsZCB0aGUgY2hpbGQgYmUgYWRkZWQ8YnIvPnRvIHRoZSB0YXJnZXQgcXVldWU_In0KICAgIE4gLS0-fHllc3wgT1siVENQTWlncmF0ZVJlcVN1Y2Nlc3MrKyJdCiAgICBOIC0tPnxub3wgUFsiVENQTWlncmF0ZVJlcUZhaWx1cmUrKzxici8-Y29ubmVjdGlvbiBSU1QiXQ%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="826" height="2267"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;To confirm this I closed a listener with no other member left in the group: ten connections in the queue, the only listener closed, switch on.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;=== G: grupta BASKA dinleyici yokken kapatma (K=10) ===
  MIB delta: {'TCPMigrateReqSuccess': 0, 'TCPMigrateReqFailure': 10}
  istemci kaybi: 10/10
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where the counter speaks. Running the same experiment with the switch off also killed ten connections, but the counters stayed at zero. The practical conclusion is easy to read backwards: &lt;strong&gt;&lt;code&gt;TCPMigrateReqFailure&lt;/code&gt; is not a "how many connections did I lose" counter; it is a "I tried to migrate and found nowhere to go" counter.&lt;/strong&gt; Looking at it with the feature disabled and concluding "zero, so I have no problem" is like checking that a closed door's bell is not ringing and deciding nobody is home.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who picks the target: the docs say random, the code says hash
&lt;/h2&gt;

&lt;p&gt;The kernel documentation looks clear on this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Otherwise, the kernel will randomly pick an alive listener only if this option is enabled.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is no randomness in the source. &lt;code&gt;reuseport_migrate_sock&lt;/code&gt; takes the migrating socket's own hash (&lt;code&gt;hash = migrating_sk-&amp;gt;sk_hash&lt;/code&gt;) and picks the target with it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;j&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;reciprocal_scale&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;num_socks&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;reciprocal_scale&lt;/code&gt; is a multiply-and-shift that stands in for a modulo: &lt;code&gt;(u32)(((u64) val * ep_ro) &amp;gt;&amp;gt; 32)&lt;/code&gt;. Same input, same output. Measuring that turned out to be easy — I bound the clients to &lt;strong&gt;fixed source ports&lt;/strong&gt; and built the same 4-tuples twice, in the same namespace, against the same destination port:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tur 1: 41000-&amp;gt;B3, 41001-&amp;gt;B1, 41002-&amp;gt;B3, 41003-&amp;gt;B3, 41004-&amp;gt;B3, 41005-&amp;gt;B1, 41006-&amp;gt;B4, 41007-&amp;gt;B4
tur 1: gocen=8/8  dagilim=B1:2 B2:0 B3:4 B4:2
tur 2: 41000-&amp;gt;B3, 41001-&amp;gt;B1, 41002-&amp;gt;B3, 41003-&amp;gt;B3, 41004-&amp;gt;B3, 41005-&amp;gt;B1, 41006-&amp;gt;B4, 41007-&amp;gt;B4
tur 2: gocen=8/8  dagilim=B1:2 B2:0 B3:4 B4:2
  ortak kaynak port: 8, AYNI hedefe gidenler: 8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eight out of eight landed on the same listener in both rounds. You have to read "randomly" in the documentation as "unmanaged, without a policy"; if you read it as "statistically random", you will be wrong.&lt;/p&gt;

&lt;p&gt;The second half of that output pulled me in the wrong direction for a while: the distribution is &lt;strong&gt;B1:2 B2:0 B3:4 B4:2&lt;/strong&gt;, so one of the four listeners received none of the eight connections. My first reaction was "the hash does not spread evenly". Wrong reaction. For eight draws into four buckets that skew has a chi-square of 4.0, and the chance of seeing something that uneven or worse from a uniform hash is roughly 0.30 — indistinguishable from noise. Simulating the selection function itself (&lt;code&gt;reciprocal_scale&lt;/code&gt; over uniform 32-bit input) shows the opposite at scale: with 8,000 draws across four buckets, the max-to-min ratio over 2,000 trials has a median of 1.045 and a worst case of 1.15. There is no two-to-one load gap.&lt;/p&gt;

&lt;p&gt;The real issue is not the count but what the selection ignores. &lt;code&gt;reuseport_select_sock_by_hash&lt;/code&gt; does not look at queue depth or current load when it picks a target; if some member of the group uses &lt;code&gt;SO_INCOMING_CPU&lt;/code&gt; a CPU match enters the picture as well, and beyond that the only input is the connection's identity. So even when the &lt;strong&gt;number&lt;/strong&gt; of connections evens out over time, their &lt;strong&gt;cost&lt;/strong&gt; does not: a long-lived WebSocket and a 20-millisecond health check look identical to this selector.&lt;/p&gt;

&lt;p&gt;That determinism is not specific to migration, either. Normal distribution came out the same way under the same harness — four listeners up from the start, no migration, fixed source ports — and eight out of eight landed on the same listener in both rounds. But the mapping &lt;strong&gt;changed&lt;/strong&gt; between two separate containers, and that has a counterpart in the source as well: &lt;code&gt;inet_ehashfn&lt;/code&gt; computes the hash using &lt;code&gt;inet_ehash_secret + net_hash_mix(net)&lt;/code&gt;. The mapping therefore carries a per-boot random secret and a per-namespace mixer. It is reproducible on your machine, and different on mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Migration does not respect the target's backlog
&lt;/h2&gt;

&lt;p&gt;While reading the source I ran into something I was not looking for. The last step of a migration is &lt;code&gt;inet_csk_reqsk_queue_add&lt;/code&gt;, and that function never asks about the target's queue capacity:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="n"&gt;spin_lock&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;queue&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;rskq_lock&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;unlikely&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sk&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;sk_state&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;TCP_LISTEN&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;inet_child_forget&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sk&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;child&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;child&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="p"&gt;...&lt;/span&gt;
        &lt;span class="n"&gt;sk_acceptq_added&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sk&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The only check is "is the target still listening". There is no &lt;code&gt;sk_acceptq_is_full()&lt;/code&gt; call. On the normal path, accepting a new connection does make that check; on the migration path it does not.&lt;/p&gt;

&lt;p&gt;So it went to the lab. Twenty connections in A's queue, and B opened with &lt;code&gt;listen(1)&lt;/code&gt; — that is, having told the kernel "I want exactly one pending connection on this socket":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;B'nin ilan ettigi backlog: 1 | ss (A+B):
LISTEN 0      1      127.0.0.1:18085 0.0.0.0:*
LISTEN 20     128    127.0.0.1:18085 0.0.0.0:*

A kapandiktan SONRA, accept'ten ONCE ss:
LISTEN 20     1      127.0.0.1:18085 0.0.0.0:*
B kabul etti: 20/20  MIB={'TCPMigrateReqSuccess': 20, 'TCPMigrateReqFailure': 0}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Recv-Q 20&lt;/code&gt;, &lt;code&gt;Send-Q 1&lt;/code&gt;. A socket that asked for one has twenty waiting — twenty times the limit it declared. All of them were accepted cleanly afterwards, so this is not an accounting glitch; the queue really is that full.&lt;/p&gt;

&lt;p&gt;For most setups this is good news: migration does not give up halfway because the target's queue is tight. For a setup that deliberately keeps its queue short, the news is mixed — and here I had to correct an assumption of my own. I thought a short backlog was a fail-fast fuse: when the queue overflows the client fails quickly and the load balancer in front moves on to another server. Default Linux does not work that way. In the experiment in the next section, when the queue overflowed, not a single client got an error; the kernel dropped the SYNs silently and the clients retried. Setting &lt;code&gt;tcp_abort_on_overflow&lt;/code&gt; to 1 and repeating the experiment changed nothing, because that setting only engages when there is a request socket to reset, and here no request socket was ever created.&lt;/p&gt;

&lt;p&gt;What remains is a resource limit rather than a fuse: a low backlog bounds how many ready connections the application has to deal with at once, and migration does not honour that bound. If several workers close at the same time, the one still standing wakes up with a queue many times the limit it declared. I saw this in the source first, then measured it; I could not find it mentioned in the documentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which connections recover on their own
&lt;/h2&gt;

&lt;p&gt;To draw the boundary of the risk I ran one more experiment. I opened A with &lt;code&gt;listen(1)&lt;/code&gt; and threw twelve clients at it using non-blocking &lt;code&gt;connect()&lt;/code&gt;. The queue overflowed immediately, and &lt;code&gt;ss&lt;/code&gt; told me what the server side actually looked like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A'da accept kuyrugu (Recv-Q/Send-Q): LISTEN 2      1      127.0.0.1:19004 0.0.0.0:*
sunucu tarafi: SYN-RECV=0  ESTABLISHED=2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two connections finished their handshake and entered the queue. The SYNs of the remaining ten were never accepted — and rather than guess at that, I read it off the counters: &lt;code&gt;ListenOverflows&lt;/code&gt; and &lt;code&gt;ListenDrops&lt;/code&gt; both climb while &lt;code&gt;SYN-RECV&lt;/code&gt; stays at zero. The kernel did not even start building a request socket, because the queue was already full; it dropped the SYN.&lt;/p&gt;

&lt;p&gt;Then I added B (backlog 128) and closed A:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;tcp_migrate_req=0&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;tcp_migrate_req=1&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Accepted by B&lt;/td&gt;
&lt;td&gt;10 / 12&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;12 / 12&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arrived by migration (instantly)&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arrived by SYN retransmission&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;TCPMigrateReqSuccess&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What migration saved was exactly &lt;strong&gt;two&lt;/strong&gt; connections: the ones that had finished the handshake and had not been accepted. The counter says two as well. The other ten survived under both settings, but not thanks to migration.&lt;/p&gt;

&lt;p&gt;In the first round I waved at what saved them — "the client retries" — and that guess needed to become a measurement, because the ~0.86 seconds I saw read like a constant. Sampling &lt;code&gt;ListenOverflows&lt;/code&gt; once per second made the mechanism plain: the counter rises by exactly &lt;strong&gt;10&lt;/strong&gt; every second, for four seconds running. Each of the ten blocked clients retransmits its SYN once a second. Recovery then happens at the first retransmission after the new listener exists — so the delay depends on where the swap lands in that cycle. Shifting the swap moment gives 0.74, 0.86 and 0.58 seconds, and under 0.1 seconds when the swap coincides with a retransmission. Not a fixed delay budget: a wait that ranges from zero to one second.&lt;/p&gt;

&lt;p&gt;So the set &lt;code&gt;tcp_migrate_req&lt;/code&gt; protects is not "everything in flight at the moment of close". If the client is still trying, TCP does its own job. What cannot be recovered is the connection that the client considers established and is now waiting on, while the server has not yet picked it up. That also happens to be the most annoying failure mode on the user's side: request sent, no answer, connection reset. Whether it is safe to retry is something the client does not know.&lt;/p&gt;

&lt;p&gt;One more boundary, for honesty's sake: requests still mid-handshake (&lt;code&gt;TCP_NEW_SYN_RECV&lt;/code&gt;) have a separate migration path in the source, and that path runs not at &lt;code&gt;close()&lt;/code&gt; but when the SYN+ACK retransmission timer fires (inside &lt;code&gt;reqsk_timer_handler&lt;/code&gt;). I could not trigger that path in this experiment — the queue overflowed, so the requests were never created, and &lt;code&gt;SYN-RECV=0&lt;/code&gt;. Its behaviour here is therefore read from the source, not measured.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who is allowed into the group
&lt;/h2&gt;

&lt;p&gt;For migration to be on the table at all, the new worker has to have joined the &lt;strong&gt;same reuseport group&lt;/strong&gt;. That group's door has two locks; these outputs come from two separate runs, with the harness's own labels:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;=== SENARYO D (port 18086) gruba katilma kapilari ===
  SO_REUSEPORT'suz bind: OSError errno=98 (EADDRINUSE)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;=== F: gruba farkli UID ile katilma ===
  uid=65534 + SO_REUSEPORT bind: OSError errno=98 (EADDRINUSE)
  uid=0 (ayni kullanici) bind: BASARILI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first is expected. The second is the one that can catch you in production — and it should not really be a surprise, because &lt;code&gt;socket(7)&lt;/code&gt; states it outright: "To prevent port hijacking, all of the processes binding to the same address must have the same effective UID." Being written in the manual is not the same as being remembered on deployment day. A procedure that starts the new version under a different service account cannot bind to that port at all; the migration question never even comes up. And &lt;code&gt;EADDRINUSE&lt;/code&gt; will not tell you why — it says "port busy", not "different user".&lt;/p&gt;

&lt;p&gt;The ordering inside the group is not arbitrary either. &lt;code&gt;socket(7)&lt;/code&gt; defines the numbering: "Sockets are numbered in the order in which they are added to the group (that is, the order of &lt;code&gt;bind(2)&lt;/code&gt; calls for UDP sockets or the order of &lt;code&gt;listen(2)&lt;/code&gt; calls for TCP sockets)." That is the array the hash is scaled into. In other words, the order in which you start your workers is one of the inputs that decides which connection goes to which worker.&lt;/p&gt;

&lt;p&gt;One more run, because the documentation mentions &lt;code&gt;shutdown()&lt;/code&gt; alongside &lt;code&gt;close()&lt;/code&gt;. I called &lt;code&gt;shutdown(SHUT_RDWR)&lt;/code&gt; on the listening socket, leaving the file descriptor open:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A.shutdown(SHUT_RDWR) cagrildi, fd ACIK
B kabul etti: 10/10  MIB={'TCPMigrateReqSuccess': 10, 'TCPMigrateReqFailure': 0}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ten out of ten migrated with the switch on, zero out of ten with it off. If you are writing a graceful drain, that is useful: you can stop listening and hand off the queue without closing the descriptor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before you turn it on
&lt;/h2&gt;

&lt;p&gt;The part of the documentation that deserves the most attention is at the end, and it is a warning:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note that migration between listeners with different settings may crash applications. Let's say migration happens from listener A to B, and only B has TCP_SAVE_SYN enabled. B cannot read SYN data from the requests migrated from A.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Migration assumes the two listeners have identical socket options. If they differ, the target ends up reading a connection it did not set up using its own assumptions. During a version rollout, where the old and new worker's socket options can drift apart, that is a real risk — and that is precisely the moment you would most want migration. The kernel's suggested answer is to pick the target with a &lt;code&gt;BPF_SK_REUSEPORT_SELECT_OR_MIGRATE&lt;/code&gt; eBPF program and return &lt;code&gt;SK_DROP&lt;/code&gt; to cancel the migration when no suitable target exists. Flipping the bare sysctl means leaving that policy to "whatever the hash says".&lt;/p&gt;

&lt;p&gt;The decision framework I ended up with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If several processes listen on the same port and you close the old one during a rollout&lt;/strong&gt;, turn it on. nginx's &lt;code&gt;reuseport&lt;/code&gt; parameter does exactly that: per the documentation, "an individual listening socket for each worker process". A reload closes the old workers' sockets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you have a single listening socket&lt;/strong&gt;, nothing changes — but which kind of "single" you have shows up in the counters. If the socket was opened with &lt;code&gt;SO_REUSEPORT&lt;/code&gt; there is still a group: migration is attempted, finds nowhere to go, &lt;code&gt;Failure&lt;/code&gt; rises, and the connection dies anyway. If &lt;code&gt;SO_REUSEPORT&lt;/code&gt; is absent — which is the case for all 90 listeners on my own server — the &lt;code&gt;if (rcu_access_pointer(sk-&amp;gt;sk_reuseport_cb))&lt;/code&gt; gate inside &lt;code&gt;inet_unhash&lt;/code&gt; never opens, the migration code never runs, and both counters stay at zero. Either way the fix lives elsewhere: an approach like &lt;a href="https://mustafaerbay.com.tr/en/blog/tutorials/systemd-socket-activation-ile-zero-downtime-restart/" rel="noopener noreferrer"&gt;systemd socket activation&lt;/a&gt;, which leaves behind a socket that can be handed over.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you deliberately keep your queue short&lt;/strong&gt;, think twice. Migration does not honour that limit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If your workers' socket options are not identical&lt;/strong&gt;, avoid the bare sysctl: use an eBPF policy, or do not turn it on at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three commands are enough to check. First, do you even have a group:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ss &lt;span class="nt"&gt;-ltnH&lt;/span&gt; | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'{print $4}'&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; | &lt;span class="nb"&gt;uniq&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'$1&amp;gt;1'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the setting and the counters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /proc/sys/net/ipv4/tcp_migrate_req
nstat &lt;span class="nt"&gt;-az&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; migrate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both flags are needed, for different reasons: without &lt;code&gt;-a&lt;/code&gt;, &lt;code&gt;nstat&lt;/code&gt; prints the &lt;strong&gt;delta&lt;/strong&gt; against its own history file (which is kept per user, so running it under &lt;code&gt;sudo&lt;/code&gt; reads a different ledger), and without &lt;code&gt;-z&lt;/code&gt;, counters sitting at zero drop out of the listing entirely — so if no migration has been attempted yet, you will not see the line at all. And as the runs above show, while the feature is off these counters will not report your losses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the counter goes quiet
&lt;/h2&gt;

&lt;p&gt;This is a long story for a single byte. Looking back, what stayed with me is not the setting but the two gaps I ran into while chasing it.&lt;/p&gt;

&lt;p&gt;One is on the instrumentation side. With migration off, the kernel threw away twenty connections and not one counter moved — because that counter counts migrations that were attempted and failed, not the ones that were never attempted. A gauge reading zero does not mean the thing it measures is not happening; it does not even mean the gauge is connected. That is exactly why both counters sitting at zero on my own server tell me nothing.&lt;/p&gt;

&lt;p&gt;The other is whose point of view "zero downtime" is spoken from. The port was never idle, there was no gap in the process table, there is no error line in nginx's access log. Looked at from the server, there really was no downtime. The downtime happened on the side of twenty clients that had finished their handshake, and that side does not write into our ledger. If the client is still retrying, TCP repairs itself; if it has stopped, we are the ones creating the loss. It helps to think of this setting not as a performance knob but as where you draw the line between those two.&lt;/p&gt;

&lt;p&gt;A last note on currency. The feature arrived in 2021 with Linux 5.14 (&lt;code&gt;f9ac779f881c&lt;/code&gt;, "net: Introduce net.ipv4.tcp_migrate_req", Kuniyuki Iwashima) and it still sits in today's mainline in the same place with the same default. It has not been removed, renamed or marked discouraged; it has simply been off by default for five years and rarely discussed. The kernel's own test suite has a case called &lt;code&gt;migrate_reuseport.c&lt;/code&gt; — the right place to read if you want to see how they exercise the migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.kernel.org/doc/html/latest/networking/ip-sysctl.html" rel="noopener noreferrer"&gt;Documentation/networking/ip-sysctl.rst — tcp_migrate_req (kernel.org)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/blob/master/net/core/sock_reuseport.c" rel="noopener noreferrer"&gt;net/core/sock_reuseport.c — reuseport_migrate_sock and reuseport_stop_listen_sock (torvalds/linux)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/blob/master/net/ipv4/inet_connection_sock.c" rel="noopener noreferrer"&gt;net/ipv4/inet_connection_sock.c — inet_csk_listen_stop and inet_csk_reqsk_queue_add (torvalds/linux)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/blob/master/net/ipv4/proc.c" rel="noopener noreferrer"&gt;net/ipv4/proc.c — TCPMigrateReqSuccess and TCPMigrateReqFailure (torvalds/linux)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://git.kernel.org/pub/scm/docs/man-pages/man-pages.git/tree/man/man7/socket.7" rel="noopener noreferrer"&gt;socket(7) — SO_REUSEPORT and group numbering (kernel.org man-pages)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://nginx.org/en/docs/http/ngx_http_core_module.html" rel="noopener noreferrer"&gt;ngx_http_core_module — the reuseport and backlog parameters of the listen directive (nginx.org)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/blob/master/tools/testing/selftests/bpf/prog_tests/migrate_reuseport.c" rel="noopener noreferrer"&gt;tools/testing/selftests/bpf/prog_tests/migrate_reuseport.c — the kernel's migration tests (torvalds/linux)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>linux</category>
      <category>tcp</category>
      <category>network</category>
      <category>olcum</category>
    </item>
    <item>
      <title>I Enabled casefold, and the Directory Stayed Case-Sensitive in Turkish</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Sat, 03 Oct 2026 06:27:43 +0000</pubDate>
      <link>https://dev.to/merbayerp/i-enabled-casefold-and-the-directory-stayed-case-sensitive-in-turkish-2lj7</link>
      <guid>https://dev.to/merbayerp/i-enabled-casefold-and-the-directory-stayed-case-sensitive-in-turkish-2lj7</guid>
      <description>&lt;p&gt;Anyone who has moved a file server from Windows to Linux hits the same wall: &lt;code&gt;Rapor.docx&lt;/code&gt; and &lt;code&gt;rapor.docx&lt;/code&gt; are one file on Windows and two on ext4. The &lt;code&gt;casefold&lt;/code&gt; feature that arrived in ext4 with Linux 5.2 exists precisely to fix that — per-directory case insensitivity, implemented inside the filesystem itself.&lt;/p&gt;

&lt;p&gt;This week I sat down and turned it on. It worked: I found &lt;code&gt;Rapor.txt&lt;/code&gt; by asking for &lt;code&gt;RAPOR.TXT&lt;/code&gt;, and by asking for &lt;code&gt;RaPoR.TxT&lt;/code&gt;. Then I tried Turkish, and saw this — in the same directory, at the same time, &lt;code&gt;IŞIK.TXT&lt;/code&gt; and &lt;code&gt;işik.txt&lt;/code&gt; are &lt;strong&gt;one&lt;/strong&gt; file, while &lt;code&gt;IŞIK.TXT&lt;/code&gt; and &lt;code&gt;ışık.txt&lt;/code&gt; are &lt;strong&gt;two&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So my directory became case-insensitive in English and case-sensitive in Turkish. Not half-broken; perfectly consistent, just not with the rule I expected. The reason is a single line of kernel source, and I found it.&lt;/p&gt;

&lt;p&gt;The second finding bothers me more: &lt;code&gt;cp&lt;/code&gt;, &lt;code&gt;rsync&lt;/code&gt; and &lt;code&gt;tar&lt;/code&gt; — all three — copy two files into that directory, return zero, and leave one file behind. No error, no warning, clean exit code.&lt;/p&gt;

&lt;p&gt;A note on the terminal blocks below: they are the output of the runs as they happened, on a Turkish-locale shell, so some labels are Turkish (&lt;code&gt;BULDU&lt;/code&gt; = found, &lt;code&gt;YOK&lt;/code&gt; = missing, &lt;code&gt;ESLESTI&lt;/code&gt; = matched, &lt;code&gt;kalan_ad&lt;/code&gt; = surviving name, &lt;code&gt;icerik&lt;/code&gt; = content). I have left machine output exactly as it came out rather than retyping it in English.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lab
&lt;/h2&gt;

&lt;p&gt;Every measurement below comes from this environment, on fresh ext4 images on loop devices:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt; /etc/os-release&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PRETTY_NAME&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
Ubuntu 26.04 LTS
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;
7.0.0-34-generic
&lt;span class="nv"&gt;$ &lt;/span&gt;mke2fs &lt;span class="nt"&gt;-V&lt;/span&gt;
mke2fs 1.47.2 &lt;span class="o"&gt;(&lt;/span&gt;1-Jan-2025&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Setup is two commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;mkfs.ext4 &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="nt"&gt;-O&lt;/span&gt; casefold &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="nv"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;utf8 /var/tmp/cf.img
&lt;span class="nb"&gt;sudo &lt;/span&gt;mount &lt;span class="nt"&gt;-o&lt;/span&gt; loop /var/tmp/cf.img /mnt/cf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;casefold&lt;/code&gt; is a filesystem feature, but on its own it makes nothing insensitive. All it does is record in the superblock how text is encoded on this filesystem:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;tune2fs &lt;span class="nt"&gt;-l&lt;/span&gt; /var/tmp/cf.img | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"character encoding"&lt;/span&gt;
Character encoding:       utf8-12.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Insensitivity is per directory, switched on with the &lt;code&gt;+F&lt;/code&gt; inode flag. The kernel documentation says it in one sentence: "It is enabled by flipping the +F inode attribute of an empty directory." The empty-directory requirement is a real requirement, not advice:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; /mnt/cf/dolu &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;touch&lt;/span&gt; /mnt/cf/dolu/x
&lt;span class="nv"&gt;$ &lt;/span&gt;chattr +F /mnt/cf/dolu
chattr: Directory not empty &lt;span class="k"&gt;while &lt;/span&gt;setting flags on /mnt/cf/dolu
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On an empty directory it goes through, and once set, &lt;strong&gt;every directory you create underneath inherits the flag&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; /mnt/cf/ci &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; chattr +F /mnt/cf/ci
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; /mnt/cf/ci/alt &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; lsattr &lt;span class="nt"&gt;-d&lt;/span&gt; /mnt/cf/ci/alt
&lt;span class="nt"&gt;--------------e--F----&lt;/span&gt; /mnt/cf/ci/alt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In practice those two rules combine into this: you have to catch the &lt;strong&gt;top&lt;/strong&gt; of the tree you want insensitive while it is still empty. You cannot convert a populated share after the fact — you have to create a new tree and move things into it. The move itself is its own problem, and I will get to it.&lt;/p&gt;

&lt;p&gt;The mount point itself can never be insensitive, because the root of a fresh ext4 is never empty:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; /mnt/cf
&lt;span class="nb"&gt;.&lt;/span&gt;  ..  ci  dolu  lost+found
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;chattr +F /mnt/cf
chattr: Directory not empty &lt;span class="k"&gt;while &lt;/span&gt;setting flags on /mnt/cf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;lost+found&lt;/code&gt; closes that door from the start. Your insensitive tree will always be a subdirectory of the mount point.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, the portability trap: mke2fs does not consult your kernel
&lt;/h2&gt;

&lt;p&gt;I walked into this one myself. I created the &lt;code&gt;casefold&lt;/code&gt; image inside Docker Desktop's VM, &lt;code&gt;mkfs&lt;/code&gt; ran happily, and then it would not mount:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# uname -r
6.10.14-linuxkit
# mkfs.ext4 -q -O casefold -E encoding=utf8 /tmp/t.img &amp;amp;&amp;amp; echo BASARILI
BASARILI
# mount -o loop /tmp/t.img /mnt
mount: /mnt: wrong fs type, bad option, bad superblock on /dev/loop0,
       missing codepage or helper program, or other error.
# dmesg | tail -1
[309036.421209] EXT4-fs (loop0): Filesystem with casefold feature cannot be mounted without CONFIG_UNICODE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The cause is plain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# zcat /proc/config.gz | grep -i CONFIG_UNICODE&lt;/span&gt;
&lt;span class="c"&gt;# CONFIG_UNICODE is not set&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;mke2fs&lt;/code&gt; does not check whether your kernel supports the feature; it writes the flag into the superblock and calls it a day. So when you move a &lt;code&gt;casefold&lt;/code&gt; disk to a machine whose kernel was built without Unicode support, the disk is not corrupted — it simply &lt;strong&gt;never mounts&lt;/strong&gt;. Rescue environments, embedded devices, minimal VM images: all candidates.&lt;/p&gt;

&lt;p&gt;The picture across my own fleet looks like this. Docker Desktop's linuxkit kernel does not support it; VPS3 does, but nothing uses it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# VPS3&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"^CONFIG_UNICODE"&lt;/span&gt; /boot/config-&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;uname&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
6.8.0-142-generic
&lt;span class="nv"&gt;CONFIG_UNICODE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;y
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="k"&gt;for &lt;/span&gt;d &lt;span class="k"&gt;in &lt;/span&gt;sda1 sda16&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s2"&gt;"%s: "&lt;/span&gt; &lt;span class="nv"&gt;$d&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nb"&gt;sudo &lt;/span&gt;dumpe2fs &lt;span class="nt"&gt;-h&lt;/span&gt; /dev/&lt;span class="nv"&gt;$d&lt;/span&gt; 2&amp;gt;/dev/null | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; casefold &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"casefold YOK"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
&lt;/span&gt;sda1: casefold YOK
sda16: casefold YOK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Make &lt;code&gt;dumpe2fs -h | grep casefold&lt;/code&gt; a habit before you take a backup image. Ever since &lt;a href="https://mustafaerbay.com.tr/en/blog/tutorials/df-sifir-dedi-root-yazmaya-devam-etti/" rel="noopener noreferrer"&gt;the evening &lt;code&gt;df&lt;/code&gt; reported zero while root kept writing&lt;/a&gt;, I stopped doing filesystem work without looking at the feature flags.&lt;/p&gt;

&lt;h2&gt;
  
  
  Baseline behaviour: names are preserved, lookups flex
&lt;/h2&gt;

&lt;p&gt;Here is the part that works the way I expected — and it works well:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;touch&lt;/span&gt; /mnt/cf/ci/Rapor.txt
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="k"&gt;for &lt;/span&gt;n &lt;span class="k"&gt;in &lt;/span&gt;Rapor.txt rapor.txt RAPOR.TXT RaPoR.TxT&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
    &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s2"&gt;"/mnt/cf/ci/&lt;/span&gt;&lt;span class="nv"&gt;$n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"BULDU  &lt;/span&gt;&lt;span class="nv"&gt;$n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"YOK    &lt;/span&gt;&lt;span class="nv"&gt;$n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;done
&lt;/span&gt;BULDU  Rapor.txt
BULDU  rapor.txt
BULDU  RAPOR.TXT
BULDU  RaPoR.TxT
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; /mnt/cf/ci
Rapor.txt
alt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All four resolve to the same file, yet &lt;code&gt;ls&lt;/code&gt; still says &lt;code&gt;Rapor.txt&lt;/code&gt; (&lt;code&gt;alt&lt;/code&gt; is the inherited subdirectory from the previous step). In the kernel documentation's words, the behaviour is "name-preserving on the disk" — the name written to disk is a byte-per-byte match of what the user supplied. Only the &lt;strong&gt;comparison&lt;/strong&gt; is flexible, not the storage. Windows and NTFS do the same thing: keep the name as given, flex the comparison.&lt;/p&gt;

&lt;p&gt;This is where I saw the first silence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;touch&lt;/span&gt; /mnt/cf/ci/RAPOR.TXT&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"exit=&lt;/span&gt;&lt;span class="nv"&gt;$?&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;exit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; /mnt/cf/ci
Rapor.txt
alt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;touch&lt;/code&gt; did not create a new file; it opened the existing one and updated its &lt;code&gt;mtime&lt;/code&gt;. Exit code zero. If you want to state that your intent was "a new file", you need &lt;code&gt;O_EXCL&lt;/code&gt; — and then you learn the truth:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;/mnt/cf/ci/RAPOR.TXT&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_CREAT&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_EXCL&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_WRONLY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mo"&gt;0o644&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;FileExistsError&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Errno&lt;/span&gt; &lt;span class="mi"&gt;17&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="n"&gt;File&lt;/span&gt; &lt;span class="n"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;/mnt/cf/ci/RAPOR.TXT&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep that in mind: only the caller who asks gets to see &lt;code&gt;EEXIST&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turkish: one directory, two different rules
&lt;/h2&gt;

&lt;p&gt;Now the part that surprised me. I created a file called &lt;code&gt;IŞIK.TXT&lt;/code&gt; and looked for its variants:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Name on disk&lt;/th&gt;
&lt;th&gt;Lookup&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IŞIK.TXT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;IŞIK.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;match&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IŞIK.TXT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;işik.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;match&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IŞIK.TXT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ışık.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no match&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;IŞIK.TXT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Işık.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;no match&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;İŞLEM.TXT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;İŞLEM.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;match&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;İŞLEM.TXT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;işlem.txt&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no match&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;İŞLEM.TXT&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;i̇şlem.txt&lt;/code&gt; (&lt;code&gt;i&lt;/code&gt; + U+0307)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;match&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read that again. As far as this directory is concerned, the lowercase of &lt;code&gt;IŞIK.TXT&lt;/code&gt; is &lt;code&gt;işik&lt;/code&gt; — so the spelling that is &lt;strong&gt;wrong&lt;/strong&gt; in Turkish matches, and the correct one does not. The consequence shows up immediately: when I tried to create &lt;code&gt;ışık.txt&lt;/code&gt;, I got a second file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ışık.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_CREAT&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_EXCL&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_WRONLY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mo"&gt;0o644&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# created
&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;işik.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_CREAT&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_EXCL&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_WRONLY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mo"&gt;0o644&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# EEXIST
&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;IŞIK.TXT&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;İŞLEM.TXT&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ışık.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The directory now holds, to a Turkish reader, the same word twice in two cases — and the filesystem counts them as two different things.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why: the kernel drops the T rows
&lt;/h3&gt;

&lt;p&gt;This is not a bug; it is a deliberate choice. The kernel's Unicode tables are generated from the Unicode Character Database, and &lt;code&gt;fs/unicode/README.utf8data&lt;/code&gt; lists exactly which files feed them — &lt;a href="https://www.unicode.org/Public/12.1.0/ucd/CaseFolding.txt" rel="noopener noreferrer"&gt;&lt;code&gt;CaseFolding.txt&lt;/code&gt;&lt;/a&gt; among them. Every row in that file carries a status code: &lt;code&gt;C&lt;/code&gt; (common), &lt;code&gt;F&lt;/code&gt; (full), &lt;code&gt;S&lt;/code&gt; (simple), &lt;code&gt;T&lt;/code&gt; (&lt;strong&gt;Turkish/Azeri special case&lt;/strong&gt;).&lt;/p&gt;

&lt;p&gt;Inside &lt;code&gt;fs/unicode/mkutf8data.c&lt;/code&gt;, which generates the tables, a single line tells the whole story:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="cm"&gt;/* Use the C+F casefold. */&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sc"&gt;'C'&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sc"&gt;'F'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;T&lt;/code&gt; rows are discarded. Here are the &lt;code&gt;CaseFolding.txt&lt;/code&gt; entries for the letters involved — the annotations on the right are mine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;0049; C; 0069;       # LATIN CAPITAL LETTER I                 -&amp;gt; kept
0049; T; 0131;       # LATIN CAPITAL LETTER I                 -&amp;gt; DROPPED
0130; F; 0069 0307;  # LATIN CAPITAL LETTER I WITH DOT ABOVE  -&amp;gt; kept
0130; T; 0069;       # LATIN CAPITAL LETTER I WITH DOT ABOVE  -&amp;gt; DROPPED
015E; C; 015F;       # LATIN CAPITAL LETTER S WITH CEDILLA    -&amp;gt; kept
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The detail that closes the table is the one &lt;strong&gt;not&lt;/strong&gt; in that list: &lt;code&gt;CaseFolding.txt&lt;/code&gt; has no row at all for &lt;code&gt;ı&lt;/code&gt; (U+0131) or &lt;code&gt;ş&lt;/code&gt; (U+015F), so both fold to themselves. &lt;code&gt;UnicodeData.txt&lt;/code&gt; does record &lt;code&gt;I&lt;/code&gt; as the uppercase of &lt;code&gt;ı&lt;/code&gt;, but folding and uppercasing are different operations, and the filesystem uses folding.&lt;/p&gt;

&lt;p&gt;From there: &lt;code&gt;IŞIK&lt;/code&gt; folds to &lt;code&gt;işik&lt;/code&gt; (&lt;code&gt;I&lt;/code&gt; → &lt;code&gt;i&lt;/code&gt;, &lt;code&gt;Ş&lt;/code&gt; → &lt;code&gt;ş&lt;/code&gt;) while &lt;code&gt;ışık&lt;/code&gt; folds to &lt;code&gt;ışık&lt;/code&gt;. Different. &lt;code&gt;İŞLEM&lt;/code&gt; folds to &lt;code&gt;i̇şlem&lt;/code&gt;, that is &lt;code&gt;i&lt;/code&gt; plus a combining dot, whereas &lt;code&gt;işlem&lt;/code&gt; is a plain &lt;code&gt;i&lt;/code&gt;. Different again. All seven rows of the table come out of these entries.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;İ&lt;/code&gt; row has a subtlety. The canonical decomposition of U+0130 is already &lt;code&gt;0049 0307&lt;/code&gt;, that is capital &lt;code&gt;I&lt;/code&gt; plus the dot — and the &lt;code&gt;C&lt;/code&gt; row folds that &lt;code&gt;I&lt;/code&gt; to &lt;code&gt;i&lt;/code&gt;. So even without the &lt;code&gt;0130; F;&lt;/code&gt; row the result would land in the same place; two paths meet at one point.&lt;/p&gt;

&lt;p&gt;The reasoning behind the choice is defensible: a filesystem has no locale. When you mount the same disk from a Turkish machine and an English one, the answer to "which file is which" cannot change. Had they kept the &lt;code&gt;T&lt;/code&gt; rows, the file that &lt;code&gt;IŞIK.TXT&lt;/code&gt; resolves to would depend on the machine's &lt;code&gt;LANG&lt;/code&gt;. Between correctness and stability, they picked stability.&lt;/p&gt;

&lt;h3&gt;
  
  
  This is not an ext4 quirk
&lt;/h3&gt;

&lt;p&gt;I ran the identical test on my own Mac (macOS 26.6.2, APFS, the default insensitive volume):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;dosya 'IŞIK.TXT' varken 'ışık.txt' -&amp;gt; -
dosya 'IŞIK.TXT' varken 'işik.txt' -&amp;gt; ESLESTI
dosya 'İŞLEM.TXT' varken 'işlem.txt' -&amp;gt; -
dosya 'İŞLEM.TXT' varken 'i̇şlem.txt' -&amp;gt; ESLESTI
NFC 'rapor-é.txt' varken NFD arama -&amp;gt; ESLESTI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Line for line the same. So there is no story here about Linux not knowing Turkish; this is the shared stance of case-insensitive filesystems. It has worked this way on your Mac for years and you probably never noticed — because nobody wakes up wanting to put both &lt;code&gt;IŞIK&lt;/code&gt; and &lt;code&gt;ışık&lt;/code&gt; in the same folder. On a migrated file server with 200,000 files, it happens for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  casefold does normalization too, not just letters
&lt;/h2&gt;

&lt;p&gt;The sentence that caught my eye in the documentation: "The comparison algorithm is implemented by normalizing the strings to the Canonical decomposition form, as defined by Unicode, followed by a byte per byte comparison." So NFD normalization is part of the comparison. I tested it — wrote &lt;code&gt;é&lt;/code&gt; as a single code point (NFC, U+00E9) and looked it up as two (NFD, &lt;code&gt;e&lt;/code&gt; + U+0301):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;olusturuldu NFC 'rapor-é.txt' (12 bayt)
  NFD ile arama 'rapor-é.txt'  -&amp;gt; BULDU
  NFD buyuk harf 'RAPOR-É.TXT' -&amp;gt; BULDU
  NFC buyuk harf 'RAPOR-É.TXT' -&amp;gt; BULDU
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The control group is a directory on the same filesystem &lt;strong&gt;without&lt;/strong&gt; &lt;code&gt;+F&lt;/code&gt; — there, all four are separate files:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="c1"&gt;# /mnt/cf/plain (plain directory)
&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;RAPOR.TXT&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;          &lt;span class="c1"&gt;# with 'Rapor.txt' present
&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;rapor-é.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# with the NFC version present
&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;RAPOR.TXT&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Rapor.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;rapor-é.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;rapor-é.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at that last line: the directory holds &lt;strong&gt;two names that look identical&lt;/strong&gt;. One is NFC, the other NFD. &lt;code&gt;ls&lt;/code&gt; prints both as &lt;code&gt;rapor-é.txt&lt;/code&gt;. This bites you with archives coming from macOS, quite independently of case, and it drives people up the wall — this may well be the more valuable problem &lt;code&gt;casefold&lt;/code&gt; solves.&lt;/p&gt;

&lt;p&gt;The full lookup path works like this. The kernel runs the folding and the decomposition from one combined table; the split below is for readability:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVtOYW1lIHN1cHBsaWVkIGJ5IHVzZXJzcGFjZV0gLS0-IEJ7SXMgK0Ygc2V0IG9uIHRoZSBkaXJlY3Rvcnl9CiAgQiAtLSBObyAtLT4gWltCeXRlLWZvci1ieXRlIGNvbXBhcmU8YnIvPmV4YWN0IG1hdGNoIHJlcXVpcmVkXQogIEIgLS0gWWVzIC0tPiBDe0lzIHRoZSBuYW1lIHZhbGlkIFVURi04fQogIEMgLS0gTm8sIHN0cmljdCBvZmYgLS0-IFlbT3BhcXVlIGJ5dGUgc2VxdWVuY2U8YnIvPm5vIGNhc2Vmb2xkaW5nIGFwcGxpZWRdCiAgQyAtLSBObywgc3RyaWN0IG9uIC0tPiBYW0VJTlZBTDxici8-ZmlsZSBpcyBub3QgY3JlYXRlZF0KICBDIC0tIFllcyAtLT4gRFtGb2xkIHZpYSBDYXNlRm9sZGluZyBDK0Y8YnIvPlQgcm93cyBhcmUgZHJvcHBlZF0KICBEIC0tPiBFW05GRCBjYW5vbmljYWwgZGVjb21wb3NpdGlvbl0KICBFIC0tPiBGW0J5dGUtZm9yLWJ5dGUgY29tcGFyZV0KICBGIC0tPiBHW05hbWUgb24gZGlzayBzdGF5cyB1bmNoYW5nZWRd%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVtOYW1lIHN1cHBsaWVkIGJ5IHVzZXJzcGFjZV0gLS0-IEJ7SXMgK0Ygc2V0IG9uIHRoZSBkaXJlY3Rvcnl9CiAgQiAtLSBObyAtLT4gWltCeXRlLWZvci1ieXRlIGNvbXBhcmU8YnIvPmV4YWN0IG1hdGNoIHJlcXVpcmVkXQogIEIgLS0gWWVzIC0tPiBDe0lzIHRoZSBuYW1lIHZhbGlkIFVURi04fQogIEMgLS0gTm8sIHN0cmljdCBvZmYgLS0-IFlbT3BhcXVlIGJ5dGUgc2VxdWVuY2U8YnIvPm5vIGNhc2Vmb2xkaW5nIGFwcGxpZWRdCiAgQyAtLSBObywgc3RyaWN0IG9uIC0tPiBYW0VJTlZBTDxici8-ZmlsZSBpcyBub3QgY3JlYXRlZF0KICBDIC0tIFllcyAtLT4gRFtGb2xkIHZpYSBDYXNlRm9sZGluZyBDK0Y8YnIvPlQgcm93cyBhcmUgZHJvcHBlZF0KICBEIC0tPiBFW05GRCBjYW5vbmljYWwgZGVjb21wb3NpdGlvbl0KICBFIC0tPiBGW0J5dGUtZm9yLWJ5dGUgY29tcGFyZV0KICBGIC0tPiBHW05hbWUgb24gZGlzayBzdGF5cyB1bmNoYW5nZWRd%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="781" height="1185"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Three tools, zero exit codes, one file
&lt;/h2&gt;

&lt;p&gt;Everything so far was "interesting". This section is "be careful in production".&lt;/p&gt;

&lt;p&gt;On a plain ext4 I prepared two files whose names differ only in case, with different contents:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /var/tmp/kaynak &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;md5sum &lt;/span&gt;Rapor.txt RAPOR.TXT
RAPOR.TXT
Rapor.txt
a336ddc0133b9f2219b73d9fbdc43239  Rapor.txt
2b40728022b6f578b4f6a93288dce045  RAPOR.TXT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then I copied them into three separate, freshly created &lt;code&gt;casefold&lt;/code&gt; targets:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[cp]    exit=0  kalan_ad=Rapor.txt  icerik=YENI SURUM
[rsync] exit=0  kalan_ad=RAPOR.TXT  icerik=ESKI SURUM
[tar]   exit=0  kalan_ad=RAPOR.TXT  icerik=YENI SURUM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three returned zero. All three left &lt;strong&gt;one&lt;/strong&gt; file in the target. &lt;code&gt;rsync&lt;/code&gt; went further and printed &lt;code&gt;&amp;gt;f+++++++++&lt;/code&gt; on two separate lines — it claimed it had transferred two new files.&lt;/p&gt;

&lt;p&gt;Which file survives depends on the tool, and that is a trap of its own. &lt;code&gt;cp&lt;/code&gt; and &lt;code&gt;rsync&lt;/code&gt; keep the name they created first and overwrite the content from the second file; &lt;code&gt;tar&lt;/code&gt; unlinks and recreates on a collision, so it takes both the name and the content from the last entry. Run the same source into the same kind of target with three tools and you get three different results — and with &lt;code&gt;cp&lt;/code&gt; and &lt;code&gt;rsync&lt;/code&gt; you are left with a hybrid whose name came from one file and whose content came from the other.&lt;/p&gt;

&lt;p&gt;The one thing that does not change: one of your two files is gone, and no exit code said so.&lt;/p&gt;

&lt;h3&gt;
  
  
  Before migrating: scan, but scan correctly
&lt;/h3&gt;

&lt;p&gt;My first instinct was a &lt;code&gt;tolower&lt;/code&gt;-based &lt;code&gt;awk&lt;/code&gt; one-liner. I tested it against deliberate collisions and it came up short. There are six names in the source directory; three of them collide:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt; /var/tmp/tara
IŞIK.TXT
RAPOR.TXT
Rapor.txt
işik.txt
rapor-é.txt
rapor-é.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;awk&lt;/code&gt;/&lt;code&gt;tolower&lt;/code&gt; script saw only two of them — it &lt;strong&gt;missed&lt;/strong&gt; the NFC/NFD pair. What the kernel actually does came out of copying the same names into a &lt;code&gt;casefold&lt;/code&gt; directory: six names collapsed to &lt;strong&gt;three&lt;/strong&gt; in the target. So the pair it missed is real data loss.&lt;/p&gt;

&lt;p&gt;The right key is the one the kernel uses: casefold first, then NFD.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 - &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;PY&lt;/span&gt;&lt;span class="sh"&gt;'
import os, unicodedata, collections, sys
kok = sys.argv[1] if len(sys.argv) &amp;gt; 1 else '.'
g = collections.defaultdict(list)
for d, _, dosyalar in os.walk(kok):
    for f in dosyalar:
        g[(d, unicodedata.normalize('NFD', f.casefold()))].append(f)
for (d, _), v in sorted(g.items()):
    if len(v) &amp;gt; 1:
        print('CAKISMA:', d, sorted(v))
&lt;/span&gt;&lt;span class="no"&gt;PY
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the same directory it finds all three:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CAKISMA: ['IŞIK.TXT', 'işik.txt']
CAKISMA: ['rapor-é.txt', 'rapor-é.txt']
CAKISMA: ['RAPOR.TXT', 'Rapor.txt']
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not use this script without planting a deliberate collision first, or you will not be able to tell "found nothing" apart from "does not work". What it costs to over-trust a single measurement is laid out in &lt;a href="https://mustafaerbay.com.tr/en/blog/tutorials/copytruncate-cok-kucuk-bir-an-33879-satir-etti/" rel="noopener noreferrer"&gt;that "very small time slice" piece&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;One more caveat: Python's &lt;code&gt;casefold&lt;/code&gt; uses the Unicode tables of your installed interpreter, not 12.1, and it may not line up with the filesystem on &lt;code&gt;ı&lt;/code&gt;/&lt;code&gt;İ&lt;/code&gt;. The script gives you a candidate list, not a verdict — make the call by looking at the names.&lt;/p&gt;

&lt;h2&gt;
  
  
  strict: keeping invalid UTF-8 at the door
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;casefold&lt;/code&gt; has another quiet side. In the default configuration a filename that is not valid UTF-8 is accepted — but insensitivity &lt;strong&gt;does not work&lt;/strong&gt; for it. The documentation says so outright: when invalid strings are encountered, "it falls back to considering the entire string as an opaque byte sequence, which still allows the user to operate on that file, but the case-insensitive lookups won't work."&lt;/p&gt;

&lt;p&gt;Measured, exactly so:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;/mnt/cf/op/bozuk-&lt;/span&gt;&lt;span class="se"&gt;\xff\xfe&lt;/span&gt;&lt;span class="s"&gt;.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_CREAT&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;O_WRONLY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mo"&gt;0o644&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# created
&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;/mnt/cf/op/bozuk-&lt;/span&gt;&lt;span class="se"&gt;\xff\xfe&lt;/span&gt;&lt;span class="s"&gt;.txt&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="bp"&gt;True&lt;/span&gt;
&lt;span class="o"&gt;&amp;gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exists&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;b&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;/mnt/cf/op/BOZUK-&lt;/span&gt;&lt;span class="se"&gt;\xff\xfe&lt;/span&gt;&lt;span class="s"&gt;.TXT&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="bp"&gt;False&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result is a directory in which some names are insensitive and some are not. You cannot tell which is which without inspecting the bytes of the name. Closing that door is only possible at &lt;code&gt;mkfs&lt;/code&gt; time, with the &lt;code&gt;strict&lt;/code&gt; flag — and the &lt;code&gt;mke2fs&lt;/code&gt; man page notes the default is off:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;mkfs.ext4 &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="nt"&gt;-O&lt;/span&gt; casefold &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="nv"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;utf8,encoding_flags&lt;span class="o"&gt;=&lt;/span&gt;strict /var/tmp/st.img
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The outcome splits in two, and the split matters:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;strict +F dizin             : EINVAL (Invalid argument)
strict ama +F OLMAYAN dizin : olusturuldu (EVET)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;strict&lt;/code&gt; does not cover the whole filesystem; it only applies inside insensitive directories.&lt;/p&gt;

&lt;p&gt;And here a gap turned up that I had not expected: &lt;strong&gt;you cannot read back whether &lt;code&gt;strict&lt;/code&gt; is on.&lt;/strong&gt; I compared a strict and a non-strict image with three different tools, and all three said the same thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;tune2fs &lt;span class="nt"&gt;-l&lt;/span&gt;  /var/tmp/s2.img | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"encoding"&lt;/span&gt;
Character encoding:       utf8-12.1
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;dumpe2fs &lt;span class="nt"&gt;-h&lt;/span&gt; /var/tmp/s2.img | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"encoding"&lt;/span&gt;
Character encoding:       utf8-12.1
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;debugfs &lt;span class="nt"&gt;-R&lt;/span&gt; &lt;span class="s2"&gt;"features"&lt;/span&gt; /var/tmp/s2.img
Filesystem features: has_journal ext_attr ... casefold sparse_super ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No &lt;code&gt;Encoding flags&lt;/code&gt; line is ever printed. The only way to answer "is strict on for this disk" is to try writing an invalid UTF-8 name into a &lt;code&gt;+F&lt;/code&gt; directory and see whether you get &lt;code&gt;EINVAL&lt;/code&gt;. If you are building a new share, turn &lt;code&gt;encoding_flags=strict&lt;/code&gt; on from the start and write it down somewhere — you can neither add it later nor ask for it later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost: the first measurement was wrong
&lt;/h2&gt;

&lt;p&gt;I was curious about lookup cost. I created two directories of 20,000 files — one &lt;code&gt;+F&lt;/code&gt;, one plain — and timed 20,000 random &lt;code&gt;stat&lt;/code&gt; calls. The first result was this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;casefold (+F)  tam yazim     6.93 us/arama
duz dizin      tam yazim     1.65 us/arama
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A factor of 4.2. I was about to write "folding is expensive", except in the same run a lookup against the &lt;code&gt;casefold&lt;/code&gt; directory with a &lt;strong&gt;different spelling&lt;/strong&gt; came out at 1.86 µs. That would mean the case that requires folding is faster than the case that does not. Impossible; the measurement order had leaked in — whichever directory was measured first was paying for a cold dentry cache.&lt;/p&gt;

&lt;p&gt;I added warm-up passes, interleaved the scenarios, and took the median of three rounds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;casefold (+F)  tam yazim     medyan   1.25 us/arama  (1.36, 1.16, 1.25)  isabet=20000
casefold (+F)  farkli yazim  medyan   1.63 us/arama  (1.63, 1.54, 1.69)  isabet=20000
duz dizin      tam yazim     medyan   1.19 us/arama  (1.08, 1.19, 1.21)  isabet=20000
duz dizin      farkli yazim  medyan   0.87 us/arama  (0.87, 1.04, 0.86)  isabet=0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the real table, and even the comparison in its first two rows is weaker than I thought. For exact-spelling lookups, the &lt;code&gt;casefold&lt;/code&gt; directory's best round (1.16) beats the plain directory's worst (1.21); with three rounds I cannot claim to have measured that difference. The honest statement: &lt;strong&gt;there is no measurable difference for exact spellings.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When the name genuinely has to be folded, the gap does clear the rounds: 1.63 µs against 1.19, roughly 37%. The last row is not comparable, and let me underline that — all 20,000 lookups there are &lt;strong&gt;misses&lt;/strong&gt;, and a negative dentry is cheaper than a hit.&lt;/p&gt;

&lt;p&gt;Directory metadata grew negligibly too at 20,000 files: 708,608 bytes against 696,320, that is 1.8%. Do not go looking for a performance argument to reject this feature; the arguments live elsewhere.&lt;/p&gt;

&lt;p&gt;Nobody would have noticed if I had published the wrong measurement, so let me note it: a micro-benchmark taken in a single round and a single order tells you about your ordering, not about the thing you measured.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your backup does not carry &lt;code&gt;+F&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;This was nowhere in my plan, and I think it is the most expensive finding in the article. &lt;code&gt;+F&lt;/code&gt; is an inode flag, and archive formats have no place for it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;lsattr &lt;span class="nt"&gt;-d&lt;/span&gt; /mnt/y/canli
&lt;span class="nt"&gt;--------------e--F----&lt;/span&gt; /mnt/y/canli

&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;tar&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; /mnt/y &lt;span class="nt"&gt;-cf&lt;/span&gt; /var/tmp/yedek.tar canli
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;tar&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; /mnt/y/geri1 &lt;span class="nt"&gt;-xf&lt;/span&gt; /var/tmp/yedek.tar
&lt;span class="nv"&gt;$ &lt;/span&gt;lsattr &lt;span class="nt"&gt;-d&lt;/span&gt; /mnt/y/geri1/canli
&lt;span class="nt"&gt;--------------e-------&lt;/span&gt;

&lt;span class="nv"&gt;$ &lt;/span&gt;rsync &lt;span class="nt"&gt;-aAX&lt;/span&gt; /mnt/y/canli /mnt/y/geri2/
&lt;span class="nv"&gt;$ &lt;/span&gt;lsattr &lt;span class="nt"&gt;-d&lt;/span&gt; /mnt/y/geri2/canli
&lt;span class="nt"&gt;--------------e-------&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not even &lt;code&gt;rsync&lt;/code&gt;'s &lt;code&gt;-aAX&lt;/code&gt; carries it — those flags are for ACLs and extended attributes, not inode flags. So when you back up your insensitive share and restore it, what you get back is a &lt;strong&gt;sensitive&lt;/strong&gt; directory, and no command reports an error.&lt;/p&gt;

&lt;p&gt;What follows is worse: in the restored tree, &lt;code&gt;Rapor.txt&lt;/code&gt; and &lt;code&gt;RAPOR.TXT&lt;/code&gt; can sit side by side again. If the next restore target is &lt;code&gt;+F&lt;/code&gt;, the "three tools ate a file" trap from earlier in this article fires &lt;strong&gt;during recovery&lt;/strong&gt; — that is, while you are coming back from a disaster.&lt;/p&gt;

&lt;p&gt;The remedy is simple but has to be remembered: your restore procedure must create the empty &lt;code&gt;+F&lt;/code&gt; directories first and unpack into them. If you have not measured this in a recovery drill, you do not know it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security: anyone who can choose a name can choose a collision
&lt;/h2&gt;

&lt;p&gt;Insensitivity changes the answer to "which file is which". Anywhere an attacker controls a filename, that becomes a question about authority. Measured:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;chattr +F /mnt/y/guv
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'ssh-ed25519 AAAA...SAHIBI\n'&lt;/span&gt;   &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /mnt/y/guv/authorized_keys
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'ssh-ed25519 AAAA...BASKASI\n'&lt;/span&gt;  &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /mnt/y/guv/AUTHORIZED_KEYS
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt; /mnt/y/guv
authorized_keys
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /mnt/y/guv/authorized_keys
ssh-ed25519 AAAA...BASKASI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Someone with permission to write &lt;code&gt;AUTHORIZED_KEYS&lt;/code&gt; changed the contents of &lt;code&gt;authorized_keys&lt;/code&gt;. The filename stayed the same, &lt;code&gt;sshd&lt;/code&gt; reads the same path, and the content belongs to someone else. The same logic applies to every fixed-name file: &lt;code&gt;.htaccess&lt;/code&gt;, &lt;code&gt;.env&lt;/code&gt;, &lt;code&gt;index.php&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The rule that falls out: enable &lt;code&gt;casefold&lt;/code&gt; on shares that hold user data, not on trees that hold configuration and credentials. If they live in the same tree, give &lt;code&gt;+F&lt;/code&gt; to the data subtree rather than to a common ancestor.&lt;/p&gt;

&lt;h2&gt;
  
  
  Backing out: two separate doors
&lt;/h2&gt;

&lt;p&gt;If your plan is "I will turn it off if I don't like it", both doors are tighter than you think.&lt;/p&gt;

&lt;p&gt;The directory flag only comes off while the directory is empty — the same condition as putting it on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; /mnt/cf/bos &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; chattr +F /mnt/cf/bos &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; chattr &lt;span class="nt"&gt;-F&lt;/span&gt; /mnt/cf/bos
&lt;span class="nv"&gt;$ &lt;/span&gt;lsattr &lt;span class="nt"&gt;-d&lt;/span&gt; /mnt/cf/bos
&lt;span class="nt"&gt;--------------e-------&lt;/span&gt; /mnt/cf/bos          &lt;span class="c"&gt;# cleared&lt;/span&gt;

&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; /mnt/cf/dolu2 &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; chattr +F /mnt/cf/dolu2 &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;touch&lt;/span&gt; /mnt/cf/dolu2/a
&lt;span class="nv"&gt;$ &lt;/span&gt;chattr &lt;span class="nt"&gt;-F&lt;/span&gt; /mnt/cf/dolu2
chattr: Directory not empty &lt;span class="k"&gt;while &lt;/span&gt;setting flags on /mnt/cf/dolu2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The filesystem feature itself is more stubborn. On a fresh filesystem where no &lt;code&gt;+F&lt;/code&gt; directory was ever created, it clears without complaint. But once you have created one — &lt;strong&gt;even if you deleted it&lt;/strong&gt; — &lt;code&gt;tune2fs&lt;/code&gt; refuses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;mkdir&lt;/span&gt; /mnt/m/ci &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; chattr +F /mnt/m/ci &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;touch&lt;/span&gt; /mnt/m/ci/a
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;rm&lt;/span&gt; &lt;span class="nt"&gt;-rf&lt;/span&gt; /mnt/m/ci &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;sudo &lt;/span&gt;umount /mnt/m
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;tune2fs &lt;span class="nt"&gt;-O&lt;/span&gt; ^casefold /var/tmp/m.img
The casefold feature can&lt;span class="s1"&gt;'t be cleared when there are inodes with +F flag.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No directory, no file, and the flag still counts as present. The fix is &lt;code&gt;e2fsck&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;e2fsck &lt;span class="nt"&gt;-fy&lt;/span&gt; /var/tmp/m.img &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;tune2fs &lt;span class="nt"&gt;-O&lt;/span&gt; ^casefold /var/tmp/m.img
tune2fs 1.47.2 &lt;span class="o"&gt;(&lt;/span&gt;1-Jan-2025&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;tune2fs &lt;span class="nt"&gt;-l&lt;/span&gt; /var/tmp/m.img | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"^Filesystem features"&lt;/span&gt;
Filesystem features:      has_journal ext_attr resize_inode dir_index orphan_file
filetype extent 64bit flex_bg metadata_csum_seed sparse_super large_file huge_file
dir_nlink extra_isize metadata_csum
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;casefold&lt;/code&gt; is off the list. The interesting part is that &lt;code&gt;e2fsck&lt;/code&gt; does &lt;strong&gt;not&lt;/strong&gt; clear the flag. I read the deleted directory's inode before and after:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;-- after deletion --   Inode: 13   Type: directory   Flags: 0x40080000
-- after e2fsck --     Inode: 13   Type: directory   Flags: 0x40080000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;0x40000000&lt;/code&gt; is exactly &lt;code&gt;FS_CASEFOLD_FL&lt;/code&gt;. The flag byte sits in the inode table and stays there after &lt;code&gt;e2fsck&lt;/code&gt; too; what changes is that the inode no longer counts as in use, so &lt;code&gt;tune2fs&lt;/code&gt;'s scan skips it from then on. I did not read what &lt;code&gt;e2fsck&lt;/code&gt; thought it repaired; the behaviour is what the lab recorded.&lt;/p&gt;

&lt;p&gt;The error message is correct, just incomplete — it does not say "run &lt;code&gt;e2fsck -f&lt;/code&gt;". I hit that same wall on four separate attempts; I am writing it down so you do not hit it once.&lt;/p&gt;

&lt;h2&gt;
  
  
  A decision frame
&lt;/h2&gt;

&lt;p&gt;If I am building a new insensitive share:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;At &lt;code&gt;mkfs&lt;/code&gt; time, &lt;code&gt;-O casefold -E encoding=utf8,encoding_flags=strict&lt;/code&gt;. You can neither add &lt;code&gt;strict&lt;/code&gt; later nor read it back; write the decision down.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;chattr +F&lt;/code&gt; the &lt;strong&gt;root&lt;/strong&gt; of the tree while it is empty; subdirectories inherit it.&lt;/li&gt;
&lt;li&gt;The mount point is never empty because of &lt;code&gt;lost+found&lt;/code&gt;, so you will always work with a subdirectory.&lt;/li&gt;
&lt;li&gt;Verify &lt;code&gt;CONFIG_UNICODE=y&lt;/code&gt; on &lt;strong&gt;every&lt;/strong&gt; kernel that will mount this disk. Count the rescue environment.&lt;/li&gt;
&lt;li&gt;Keep configuration and credential files outside the insensitive tree.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If I am migrating an existing tree:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scan with a casefold + NFD key; plain &lt;code&gt;tolower&lt;/code&gt; misses normalization collisions.&lt;/li&gt;
&lt;li&gt;Prove the scanning script with a deliberate collision.&lt;/li&gt;
&lt;li&gt;Do not trust the exit codes of &lt;code&gt;cp&lt;/code&gt;/&lt;code&gt;rsync&lt;/code&gt;/&lt;code&gt;tar&lt;/code&gt;. Compare the file &lt;strong&gt;count&lt;/strong&gt; between source and target afterwards; that is the only check that works.&lt;/li&gt;
&lt;li&gt;Know that &lt;code&gt;I/ı&lt;/code&gt; and &lt;code&gt;İ/i&lt;/code&gt; pairs in Turkish names do &lt;strong&gt;not&lt;/strong&gt; count as collisions as far as the filesystem is concerned.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For backup and recovery:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The restore procedure has to create the &lt;code&gt;+F&lt;/code&gt; directories empty first. The archive does not carry the flag.&lt;/li&gt;
&lt;li&gt;In a drill, check &lt;code&gt;lsattr&lt;/code&gt; on the restored tree, and the file count too.&lt;/li&gt;
&lt;li&gt;Backing out with &lt;code&gt;^casefold&lt;/code&gt; needs &lt;code&gt;umount&lt;/code&gt; plus &lt;code&gt;e2fsck -f&lt;/code&gt;; in production that means planned downtime.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Places I would not touch at all:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Applications that compare filenames themselves. &lt;code&gt;readdir&lt;/code&gt; returns the name on disk; insensitivity stays in the kernel.&lt;/li&gt;
&lt;li&gt;Backup targets under continuous sync. Every sync from a sensitive source into an insensitive target carries a silent-merge risk.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;casefold&lt;/code&gt; is a good feature and it works. My mistake was reading the phrase "case-insensitive" with the meaning it has in my own language. The filesystem never promised me insensitivity; it promised an equivalence defined by &lt;strong&gt;Unicode&lt;/strong&gt;, independent of locale. The two coincide in English and diverge in Turkish, and the filesystem does not regard the second case as an exception — that is simply the rule.&lt;/p&gt;

&lt;p&gt;The lesson I take from this is bigger than filesystems: when a system says two things are "the same", do not build on top of it before asking whose "same" it means. &lt;code&gt;tolower&lt;/code&gt; has a locale; &lt;code&gt;casefold&lt;/code&gt; does not. &lt;code&gt;cp&lt;/code&gt;'s exit code says "I transferred", not "I transferred all of them". &lt;code&gt;rsync -aAX&lt;/code&gt; reads like "preserve everything" but does not cover inode flags. The gaps between those are silent — right up until one file is missing from a 200,000-file migration.&lt;/p&gt;

&lt;p&gt;Two current notes. &lt;code&gt;casefold&lt;/code&gt; is spreading: &lt;code&gt;tmpfs&lt;/code&gt; now supports &lt;code&gt;casefold=utf8-12.1.0&lt;/code&gt; and &lt;code&gt;strict_encoding&lt;/code&gt; as mount options, with the same &lt;code&gt;+F&lt;/code&gt; and inheritance rules. And the "your application does not know" gap above is closing — &lt;code&gt;FS_XFLAG_CASEFOLD&lt;/code&gt; and &lt;code&gt;FS_XFLAG_CASENONPRESERVING&lt;/code&gt; have landed in the kernel's userspace interface, so an application can now ask &lt;code&gt;FS_IOC_FSGETXATTR&lt;/code&gt; whether a directory is insensitive. The comment above the definition carries its own warning: do not assume the bit belongs to directories only.&lt;/p&gt;

&lt;p&gt;One more: the kernel's Unicode tables are still generated from 12.1.0, the release from 2019. For something like filename equivalence, a table frozen for seven years is, I suspect, a choice rather than an oversight.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.kernel.org/doc/html/latest/admin-guide/ext4.html" rel="noopener noreferrer"&gt;ext4 Data Structures and Algorithms — case-insensitive file name lookups (kernel.org)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/blob/master/fs/unicode/README.utf8data" rel="noopener noreferrer"&gt;fs/unicode/README.utf8data — Unicode 12.1.0 table sources (torvalds/linux)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/blob/master/fs/unicode/mkutf8data.c" rel="noopener noreferrer"&gt;fs/unicode/mkutf8data.c — "Use the C+F casefold" (torvalds/linux)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/blob/master/include/uapi/linux/fs.h" rel="noopener noreferrer"&gt;include/uapi/linux/fs.h — FS_CASEFOLD_FL and FS_XFLAG_CASEFOLD (torvalds/linux)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/tytso/e2fsprogs/blob/master/misc/mke2fs.8.in" rel="noopener noreferrer"&gt;mke2fs(8) — encoding and encoding_flags=strict (tytso/e2fsprogs)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/torvalds/linux/blob/master/Documentation/filesystems/tmpfs.rst" rel="noopener noreferrer"&gt;Documentation/filesystems/tmpfs.rst — casefold mount options (torvalds/linux)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ext4</category>
      <category>linux</category>
      <category>unicode</category>
      <category>dosyasistemi</category>
    </item>
    <item>
      <title>Every New Command I Type Deletes an Older One</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Fri, 02 Oct 2026 14:23:42 +0000</pubDate>
      <link>https://dev.to/merbayerp/every-new-command-i-type-deletes-an-older-one-cn2</link>
      <guid>https://dev.to/merbayerp/every-new-command-i-type-deletes-an-older-one-cn2</guid>
      <description>&lt;p&gt;It started with a simple question: how far back does my shell history go? How much of&lt;br&gt;
the years spent in a terminal is still sitting there? Answering it looked like a&lt;br&gt;
single-file job.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt; ~/.zsh_history
    1218 /Users/mustafaerbay/.zsh_history
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One thousand two hundred eighteen lines. But how many months is that? That is where I&lt;br&gt;
stalled, because the file contains not a single date — neither the oldest nor the newest&lt;br&gt;
line says when it was written. Looking elsewhere turned up &lt;code&gt;/etc/zshrc&lt;/code&gt; in the same&lt;br&gt;
neighbourhood, with these three lines inside:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;HISTFILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;ZDOTDIR&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;/.zsh_history
&lt;span class="nv"&gt;HISTSIZE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2000
&lt;span class="nv"&gt;SAVEHIST&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The file is owned by root, mode &lt;code&gt;r--r--r--&lt;/code&gt;, last modified on 13 August. This is not a&lt;br&gt;
setting of mine; it is a decision Apple shipped with the machine. And it says: at most&lt;br&gt;
&lt;strong&gt;one thousand&lt;/strong&gt; commands go into that file.&lt;/p&gt;

&lt;p&gt;When I finally counted the file properly, that was exactly the number: &lt;strong&gt;1,000&lt;/strong&gt;. Not&lt;br&gt;
one below the ceiling — precisely on it. Which means this file filled up months ago, and&lt;br&gt;
every command I have typed since then has silently deleted the oldest one.&lt;/p&gt;

&lt;p&gt;Every number below was measured on this machine on 2 October 2026: macOS 26.6.2&lt;br&gt;
(25G83), zsh 5.9, arm64. The file shifted underneath me mid-measurement — a terminal&lt;br&gt;
window I had forgotten about closed, and the size went from 52,982 to 53,283 bytes — so&lt;br&gt;
a copy was frozen at 16:47:35, and every figure comes from that single copy.&lt;/p&gt;
&lt;h2&gt;
  
  
  First I counted it wrong, and the wrong way of counting was the lesson
&lt;/h2&gt;

&lt;p&gt;My first attempt said 939. Not a rounding error but a thesis-inverting one: 939 means&lt;br&gt;
"sixty-one commands left before the ceiling"; 1,000 means "the ceiling is full and the&lt;br&gt;
losses started long ago."&lt;/p&gt;

&lt;p&gt;Here is the mistake. zsh records &lt;em&gt;events&lt;/em&gt;, not lines; a multi-line command is one event,&lt;br&gt;
with the line breaks marked by backslashes. The correct rule is a single sentence: &lt;strong&gt;an&lt;br&gt;
event ends at the first line that does not end in a backslash.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;total lines                       1218
lines ending in a backslash        218
=&amp;gt; event count                    1000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;My first script threw away blank lines as "not data." But none of the 61 blank lines in&lt;br&gt;
that file is an independent entry; every one of them sits in the middle of a multi-line&lt;br&gt;
command that happens to contain an empty line. Dropping them removed 61 lines &lt;em&gt;and&lt;/em&gt;&lt;br&gt;
stitched the surrounding commands together at the wrong seams.&lt;/p&gt;

&lt;p&gt;I verified it with zsh itself rather than with arithmetic alone. Under a temporary&lt;br&gt;
&lt;code&gt;ZDOTDIR&lt;/code&gt;, the shell read the file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;fc&lt;/span&gt; &lt;span class="nt"&gt;-R&lt;/span&gt; hist-snapshot.txt
&lt;span class="nv"&gt;$ &lt;/span&gt;print &lt;span class="k"&gt;${#&lt;/span&gt;&lt;span class="nv"&gt;history&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; print &lt;span class="nv"&gt;$HISTCMD&lt;/span&gt;
999
1000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;history&lt;/code&gt; array excludes the current event slot, so 999 + 1 = 1,000. The independent&lt;br&gt;
arithmetic agrees. The lesson is cheap to state and expensive to learn: when counting a&lt;br&gt;
file, use the rules of the program that writes it, not your own intuition.&lt;/p&gt;
&lt;h2&gt;
  
  
  What happens once the ceiling is full?
&lt;/h2&gt;

&lt;p&gt;This one went to the lab, reproduced without touching my real history file, inside a&lt;br&gt;
temporary &lt;code&gt;ZDOTDIR&lt;/code&gt;. A file of 995 fake entries, then 20 more commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;start:  995 entries, first=eski_0001  last=eski_0995
after: 1000 entries, first=eski_0016  last=yeni_20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The total should have been 1,015; the file stopped at 1,000 and the first fifteen&lt;br&gt;
entries were gone. Twenty more commands, and &lt;code&gt;eski_0016&lt;/code&gt; through &lt;code&gt;eski_0035&lt;/code&gt; went the&lt;br&gt;
same way. No error, no warning, no log line. My own file is in exactly this state.&lt;/p&gt;

&lt;p&gt;The single-session version is starker. A shell took 1,500 commands into memory, &lt;code&gt;fc -l&lt;/code&gt;&lt;br&gt;
counted all of them while the session ran, and then the shell closed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;in memory (fc -l): 1500
on disk:           1000
first line:        lab_komut_501
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Five hundred never reached the disk. In the zsh documentation &lt;code&gt;HISTSIZE&lt;/code&gt; is "the maximum&lt;br&gt;
number of events stored in the internal history list" and &lt;code&gt;SAVEHIST&lt;/code&gt; is "the maximum&lt;br&gt;
number of history events to save in the history file." Two numbers, two memories: one is&lt;br&gt;
what you remember during the session, the other is what you still have the next morning.&lt;/p&gt;
&lt;h2&gt;
  
  
  No entry has a time — but the shell will show you one
&lt;/h2&gt;

&lt;p&gt;How many of the thousand entries carry a timestamp? Zero. zsh writes timestamps in the&lt;br&gt;
&lt;code&gt;: &amp;lt;beginning time&amp;gt;:&amp;lt;elapsed seconds&amp;gt;;&amp;lt;command&amp;gt;&lt;/code&gt; form, and only when &lt;code&gt;EXTENDED_HISTORY&lt;/code&gt;&lt;br&gt;
is on. The documented default is off, and so is the state on this machine. There is not&lt;br&gt;
a single line starting with &lt;code&gt;HIST&lt;/code&gt; in my &lt;code&gt;~/.zshrc&lt;/code&gt;, so this too is an inherited default.&lt;/p&gt;

&lt;p&gt;So far this is annoying but honest: if the information is absent, it is absent. The sly&lt;br&gt;
part begins when you ask the shell. I listed a timestamp-free file from two separate&lt;br&gt;
shells four seconds apart, using &lt;code&gt;fc -l -t '%Y-%m-%d %H:%M:%S'&lt;/code&gt; so the seconds show:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read 1 — wall clock 16:40:46
    1  2026-10-02 16:40:46  eski_komut_bir
    2  2026-10-02 16:40:46  eski_komut_iki

read 2 — wall clock 16:40:50
    1  2026-10-02 16:40:50  eski_komut_bir
    2  2026-10-02 16:40:50  eski_komut_iki
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same two commands, two different times. The value shown is not when the command ran&lt;br&gt;
but when the file was &lt;em&gt;read&lt;/em&gt; — not even the file's modification time. I went looking for&lt;br&gt;
this in the documentation and could not find it: &lt;code&gt;fc&lt;/code&gt; only defines the format, not what&lt;br&gt;
gets printed for an undated file. So the sentence below rests on measurement, not on the&lt;br&gt;
manual: &lt;strong&gt;it fills an empty field and shows it to you.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The rule that falls out has now caught me twice on this blog: a tool showing you a number&lt;br&gt;
does not mean it knows that number. While&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/life/dizustum-yedi-gunde-76-kez-uyandi/" rel="noopener noreferrer"&gt;counting my laptop's sleep records&lt;/a&gt;&lt;br&gt;
the &lt;code&gt;uptime&lt;/code&gt; counter offered a reassuring figure, but what it measured was not "how long&lt;br&gt;
I stayed awake," it was "how long since the last reboot."&lt;/p&gt;
&lt;h2&gt;
  
  
  The order is not chronological either
&lt;/h2&gt;

&lt;p&gt;If there are no timestamps, is the ordering at least right? Two shells. A started first,&lt;br&gt;
wrote three commands, waited three seconds. B started half a second later, wrote three&lt;br&gt;
commands and exited immediately.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[t=0.5] B exited -&amp;gt;     [t=3.5] A exited too -&amp;gt;
kabukB_1                kabukB_1
kabukB_2                kabukB_2
kabukB_3                kabukB_3
                        kabukA_1
                        kabukA_2
                        kabukA_3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A's commands were typed &lt;em&gt;first&lt;/em&gt; and sit &lt;em&gt;last&lt;/em&gt; in the file. Because the definition of&lt;br&gt;
&lt;code&gt;HISTFILE&lt;/code&gt; is literally this: "the file to save the history in when an interactive shell&lt;br&gt;
exits." The moment of writing is not when you type the command, it is when you close the&lt;br&gt;
window. Thanks to the default &lt;code&gt;APPEND_HISTORY&lt;/code&gt; no session overwrites another; but the&lt;br&gt;
file's order is the order in which I closed terminal tabs.&lt;/p&gt;

&lt;p&gt;For anyone working across three screens that means the bottom line is not "what did I do&lt;br&gt;
last," it is "which window did I close last."&lt;/p&gt;
&lt;h2&gt;
  
  
  A crashed shell remembers nothing
&lt;/h2&gt;

&lt;p&gt;If the definition says "when an interactive shell exits," what does a shell that never&lt;br&gt;
exits do? Five commands, then &lt;code&gt;kill -9&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exit code: 137
ls: .../.zsh_history: No such file or directory
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The file was never created. Closing the same shell cleanly wrote three lines properly.&lt;br&gt;
So every time you force-restart the machine without closing your terminal, whatever you&lt;br&gt;
typed in that session is gone — and since the loss leaves no record, I cannot even count&lt;br&gt;
how often it has happened.&lt;/p&gt;

&lt;p&gt;The remedy is in the documentation: &lt;code&gt;INC_APPEND_HISTORY&lt;/code&gt; adds new lines to &lt;code&gt;$HISTFILE&lt;/code&gt;&lt;br&gt;
"incrementally (as soon as they are entered)." Default: off.&lt;/p&gt;
&lt;h2&gt;
  
  
  Commands the agent runs never arrive at all
&lt;/h2&gt;

&lt;p&gt;One number nagged at me. The &lt;code&gt;claude&lt;/code&gt; command appears in only 12 of the thousand&lt;br&gt;
entries, while most of the work on this machine happens &lt;em&gt;inside&lt;/em&gt; those twelve sessions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;zsh &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'echo etkilesimsiz_komut_calisti'&lt;/span&gt;
&lt;span class="go"&gt;etkilesimsiz_komut_calisti
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;zsh s.zsh
&lt;span class="go"&gt;betik_icinden
--- history file ---
0 lines
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero. The definition is explicit again: the file is written only when an &lt;strong&gt;interactive&lt;/strong&gt;&lt;br&gt;
shell exits. Nothing an agent, a script or a &lt;code&gt;cron&lt;/code&gt; job runs ever lands in this ledger.&lt;/p&gt;

&lt;p&gt;Twelve lines, twelve sessions, and no way to learn from this file how many commands ran&lt;br&gt;
inside them. Shell history is a record not of "what I did" but of "what I typed with my&lt;br&gt;
own hands." Those used to be the same thing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVtZb3UgdHlwZSBhIGNvbW1hbmRdIC0tPiBCe0ludGVyYWN0aXZlIHNoZWxsP30KICBCIC0tIE5vOiBzY3JpcHQgLyBhZ2VudCAvIGNyb24gLS0-IFhbV3JpdHRlbiBub3doZXJlXQogIEIgLS0gWWVzIC0tPiBDW0VudGVycyB0aGUgaW4tbWVtb3J5IGxpc3Q8YnIvPmNlaWxpbmcgSElTVFNJWkUgMjAwMF0KICBDIC0tPiBEe0hvdyBkaWQgdGhlIHNoZWxsIGVuZD99CiAgRCAtLSBraWxsIC05IC8gY3Jhc2ggLS0-IFlbVGhlIHdob2xlIHNlc3Npb24gaXMgbG9zdF0KICBEIC0tIGNsZWFuIGV4aXQgLS0-IEVbQXBwZW5kZWQgdG8gdGhlIGZpbGU8YnIvPm9yZGVyOiBleGl0IG9yZGVyXQogIEUgLS0-IEZ7SGFzIHRoZSBmaWxlIGZpbGxlZCBTQVZFSElTVCAxMDAwP30KICBGIC0tIFllcyAtLT4gR1tUaGUgb2xkZXN0IGVudHJ5IGlzIHNpbGVudGx5IGRyb3BwZWRdCiAgRiAtLSBObyAtLT4gSFtJdCBzdGF5cywgYnV0IHdpdGggbm8gdGltZXN0YW1wXQ%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVtZb3UgdHlwZSBhIGNvbW1hbmRdIC0tPiBCe0ludGVyYWN0aXZlIHNoZWxsP30KICBCIC0tIE5vOiBzY3JpcHQgLyBhZ2VudCAvIGNyb24gLS0-IFhbV3JpdHRlbiBub3doZXJlXQogIEIgLS0gWWVzIC0tPiBDW0VudGVycyB0aGUgaW4tbWVtb3J5IGxpc3Q8YnIvPmNlaWxpbmcgSElTVFNJWkUgMjAwMF0KICBDIC0tPiBEe0hvdyBkaWQgdGhlIHNoZWxsIGVuZD99CiAgRCAtLSBraWxsIC05IC8gY3Jhc2ggLS0-IFlbVGhlIHdob2xlIHNlc3Npb24gaXMgbG9zdF0KICBEIC0tIGNsZWFuIGV4aXQgLS0-IEVbQXBwZW5kZWQgdG8gdGhlIGZpbGU8YnIvPm9yZGVyOiBleGl0IG9yZGVyXQogIEUgLS0-IEZ7SGFzIHRoZSBmaWxlIGZpbGxlZCBTQVZFSElTVCAxMDAwP30KICBGIC0tIFllcyAtLT4gR1tUaGUgb2xkZXN0IGVudHJ5IGlzIHNpbGVudGx5IGRyb3BwZWRdCiAgRiAtLSBObyAtLT4gSFtJdCBzdGF5cywgYnV0IHdpdGggbm8gdGltZXN0YW1wXQ%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What do the surviving thousand commands say?
&lt;/h2&gt;

&lt;p&gt;After accounting for the losses, the survivors describe a narrower landscape than I&lt;br&gt;
expected:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Event count&lt;/td&gt;
&lt;td&gt;1,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-line commands&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Distinct command names&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Share of the top 10 commands&lt;/td&gt;
&lt;td&gt;70.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unique command texts&lt;/td&gt;
&lt;td&gt;373&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Entries eaten by repetition&lt;/td&gt;
&lt;td&gt;627 (62.7%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Command names used exactly once&lt;/td&gt;
&lt;td&gt;39&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The ranking: &lt;code&gt;ssh&lt;/code&gt; 158, &lt;code&gt;clear&lt;/code&gt; 126, &lt;code&gt;bstart&lt;/code&gt; 82, &lt;code&gt;cd&lt;/code&gt; 80, &lt;code&gt;fstart&lt;/code&gt; 71, &lt;code&gt;python&lt;/code&gt; 66,&lt;br&gt;
&lt;code&gt;yes&lt;/code&gt; 36, &lt;code&gt;git&lt;/code&gt; 31, &lt;code&gt;npm&lt;/code&gt; 30, &lt;code&gt;sudo&lt;/code&gt; 26, &lt;code&gt;seray&lt;/code&gt; 22, &lt;code&gt;ping&lt;/code&gt; 22, &lt;code&gt;ls&lt;/code&gt; 20, &lt;code&gt;scp&lt;/code&gt; 20.&lt;br&gt;
Ninety different tools, and a tenth of them cover roughly three quarters of the work.&lt;br&gt;
Thirty-nine commands appear exactly once. All 36 &lt;code&gt;yes&lt;/code&gt; lines are copies of a single text&lt;br&gt;
— one load test's invocation — so even the list flatters the variety.&lt;/p&gt;

&lt;p&gt;Three of them stand out: &lt;code&gt;bstart&lt;/code&gt; 82, &lt;code&gt;fstart&lt;/code&gt; 71, &lt;code&gt;seray&lt;/code&gt; 22, all shortcuts from my&lt;br&gt;
&lt;code&gt;.zshrc&lt;/code&gt;. &lt;code&gt;bstart&lt;/code&gt; enters a project directory, activates a virtual environment and starts&lt;br&gt;
a development server; &lt;code&gt;fstart&lt;/code&gt; only enters the front-end directory and calls&lt;br&gt;
&lt;code&gt;npm run dev&lt;/code&gt;; &lt;code&gt;seray&lt;/code&gt; starts nothing at all, it just enters the directory and activates&lt;br&gt;
the virtual environment. Together 175 entries, &lt;strong&gt;17.5%&lt;/strong&gt; of my history. Three words.&lt;/p&gt;

&lt;p&gt;Those 175 lines teach me nothing. Six months from now the history cannot tell me which&lt;br&gt;
command &lt;code&gt;bstart&lt;/code&gt; ran, because it is not written there; and if the definition changed, it&lt;br&gt;
will not recall the old one either. A shortcut gives back at reading time what it saved&lt;br&gt;
at typing time.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;clear&lt;/code&gt; appearing 126 times — exactly an eighth of the ledger — is its own joke: the&lt;br&gt;
command for clearing the screen was recorded as the very noise it was clearing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Repetition eats two thirds of the ceiling
&lt;/h2&gt;

&lt;p&gt;Of the thousand entries only 373 are distinct text; the remaining &lt;strong&gt;627 lines are&lt;br&gt;
repeats&lt;/strong&gt;. Think of it as a memory budget: the ceiling is fixed, and 62.7% of it goes to&lt;br&gt;
rewriting something already there. A rough derivation — not measured, but derived from&lt;br&gt;
two measured numbers: with the same distribution, dropping repeats would fit about&lt;br&gt;
&lt;strong&gt;2.7 times&lt;/strong&gt; today's variety into those thousand lines.&lt;/p&gt;

&lt;p&gt;zsh offers four switches for this, and all four ship off:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;HIST_IGNORE_DUPS&lt;/code&gt; — removes only an exact repeat of the previous event ("duplicates of
the previous event," in the manual's phrasing). The weakest one.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;HIST_IGNORE_ALL_DUPS&lt;/code&gt; — "if a new command line being added to the history list
duplicates an older one, the older command is removed from the list (even if it is not
the previous event)."&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;HIST_SAVE_NO_DUPS&lt;/code&gt; — "when writing out the history file, older commands that duplicate
newer ones are omitted." Memory stays as is; only the disk copy is tidied.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;HIST_EXPIRE_DUPS_FIRST&lt;/code&gt; — when room is needed, it causes "the oldest history event
that has a duplicate to be lost before losing a unique event from the list." The
documentation pairs this with keeping &lt;code&gt;HISTSIZE&lt;/code&gt; larger than &lt;code&gt;SAVEHIST&lt;/code&gt;; on macOS that
cushion already exists, 2,000 against 1,000. Apple left the pillow and never turned on
the switch that lies on it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The ledger next door has been silent for five months
&lt;/h2&gt;

&lt;p&gt;There is a second ledger in the same directory: &lt;code&gt;~/.bash_history&lt;/code&gt;. 7,207 bytes, 208&lt;br&gt;
lines, 207 entries, 109 of them distinct. Last modified 8 May 2026; not a single line&lt;br&gt;
added in five months.&lt;/p&gt;

&lt;p&gt;Its contents belong to a different era: &lt;code&gt;python&lt;/code&gt; 52, &lt;code&gt;clear&lt;/code&gt; 38, &lt;code&gt;cd&lt;/code&gt; 27, &lt;code&gt;rm&lt;/code&gt; 11. And 28&lt;br&gt;
lines begin directly with &lt;code&gt;#&lt;/code&gt; — instead of deleting a command I had put a hash in front&lt;br&gt;
of it and said "not right now."&lt;/p&gt;

&lt;p&gt;That file records nothing any more, because zsh has been the default shell for new&lt;br&gt;
accounts since macOS 10.15, as Apple's own documentation states. But it was never&lt;br&gt;
deleted. Two hundred and seven lines of a frozen period sit there, and unlike their&lt;br&gt;
neighbour they will never be trimmed. The condition for a complete memory, it seems, is&lt;br&gt;
that the memory is no longer in use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before enlarging the memory: what is in it?
&lt;/h2&gt;

&lt;p&gt;This file is plain text and it holds everything typed on the command line. Mine is&lt;br&gt;
&lt;code&gt;rw-------&lt;/code&gt;, readable only by me. Still, a crude scan was worth running: &lt;strong&gt;23&lt;/strong&gt; of the&lt;br&gt;
thousand entries contain one of the words &lt;code&gt;token&lt;/code&gt;, &lt;code&gt;secret&lt;/code&gt;, &lt;code&gt;password&lt;/code&gt; or &lt;code&gt;api key&lt;/code&gt;.&lt;br&gt;
Not all 23 carry a real secret — matching a word is enough to be counted — but until the&lt;br&gt;
scan I did not know that either, and asking the question correctly is the point. The good&lt;br&gt;
news: four of them use &lt;code&gt;read -rs&lt;/code&gt;, so the secret went into a silent prompt rather than the&lt;br&gt;
command line.&lt;/p&gt;

&lt;p&gt;Raising the ceiling also extends the half-life of that content. Three tools help:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;HIST_IGNORE_SPACE&lt;/code&gt; removes lines whose first character is a space — but the manual adds
an important caveat: "the command lingers in the internal history until the next command
is entered before it vanishes," so the disappearance is not instant.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;HISTORY_IGNORE&lt;/code&gt; takes a pattern and, "at the time history files are written… any
potential history entry that matches the pattern is skipped." Far stronger than the
leading-space trick for secret hygiene.&lt;/li&gt;
&lt;li&gt;And a one-line rule: shell history is not an audit trail. Who ran what, and when, is not
answered here; that needs
&lt;a href="https://mustafaerbay.com.tr/en/blog/tutorials/linux-auditd-ile-ayricalikli-komut-izleme-runbooku/" rel="noopener noreferrer"&gt;a command-auditing pipeline such as auditd&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Choosing the default on purpose
&lt;/h2&gt;

&lt;p&gt;There is a list of settings to change at the end of this, but the list is not the point.&lt;br&gt;
Until today shell history looked like a record to me; what I had was a &lt;em&gt;cache&lt;/em&gt;, and&lt;br&gt;
someone else had decided its capacity, its ordering and its expiry.&lt;/p&gt;

&lt;p&gt;Four questions for your own setup:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;What is the ceiling?&lt;/strong&gt; &lt;code&gt;echo $SAVEHIST&lt;/code&gt;. If you see 1,000, that is your system's
choice and not yours — roughly a month of memory for someone typing 30-40 commands a
day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Are timestamps kept?&lt;/strong&gt; Without &lt;code&gt;setopt EXTENDED_HISTORY&lt;/code&gt; there is no answer to "when
did I do this," and the shell will show you today's clock instead of saying so.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When does the write happen?&lt;/strong&gt; On exit, by default. &lt;code&gt;INC_APPEND_HISTORY&lt;/code&gt; writes as
soon as the command is entered; &lt;code&gt;INC_APPEND_HISTORY_TIME&lt;/code&gt; writes after it finishes and
records the duration; &lt;code&gt;SHARE_HISTORY&lt;/code&gt; additionally imports other sessions' lines. The
manual puts a clear warning here: "The three options should be considered mutually
exclusive." Pick one; do not turn on all three.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parallel shells?&lt;/strong&gt; &lt;code&gt;APPEND_HISTORY&lt;/code&gt; is on by default in zsh, so sessions do not
overwrite each other. If the file lives on a network share, &lt;code&gt;HIST_FCNTL_LOCK&lt;/code&gt; hands
locking to the system's &lt;code&gt;fcntl&lt;/code&gt; call — "where this method is available" — and the
manual says this avoids history corruption on NFS.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;There is a more radical route too: &lt;code&gt;atuin&lt;/code&gt;, &lt;code&gt;mcfly&lt;/code&gt; and &lt;code&gt;hstr&lt;/code&gt; keep history in a database&lt;br&gt;
instead of a flat file and remove all four complaints above — no timestamps, tight&lt;br&gt;
ceiling, wrong order, crashed sessions lost — in one move. The price is one more&lt;br&gt;
dependency in your shell.&lt;/p&gt;

&lt;p&gt;As this article goes out, the defaults are still running on this machine, because&lt;br&gt;
changing a setting is the job of a decision, not of a measurement. But I now know what&lt;br&gt;
the setting is and what it costs, which I did not before the measurement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Remembering Is Not a Default
&lt;/h2&gt;

&lt;p&gt;Back to the original question: how far back does this file go? I cannot say. There are no&lt;br&gt;
timestamps, so the age of the oldest line is unknown. The ceiling is full, so anything&lt;br&gt;
before it is already deleted. Crashed sessions were never written. Commands run by the&lt;br&gt;
agent and by scripts never entered. Four separate reasons, and none of them leaves a&lt;br&gt;
record.&lt;/p&gt;

&lt;p&gt;The real lesson is not about shell history. Most of the records in our lives work this&lt;br&gt;
way: there is a ceiling, there is a moment of writing, and there is a default handed to&lt;br&gt;
us without being asked. Until we open the ledger we assume it is boundless. Opening it&lt;br&gt;
shows that the ledger reflects not what it holds, but what it was &lt;em&gt;permitted&lt;/em&gt; to hold.&lt;/p&gt;

&lt;p&gt;Every command typed while writing this article took one of the oldest lines of my history&lt;br&gt;
with it. I will never know which ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/zsh-users/zsh/blob/master/Doc/Zsh/params.yo" rel="noopener noreferrer"&gt;zsh — Parameters Used By The Shell (HISTFILE, HISTSIZE, SAVEHIST, HISTORY_IGNORE)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/zsh-users/zsh/blob/master/Doc/Zsh/options.yo" rel="noopener noreferrer"&gt;zsh — History options (APPEND_HISTORY, EXTENDED_HISTORY, INC_APPEND_HISTORY, SHARE_HISTORY)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/zsh-users/zsh/blob/master/Src/hist.c" rel="noopener noreferrer"&gt;zsh — Src/hist.c, the savehistfile implementation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/zsh-users/zsh/tags" rel="noopener noreferrer"&gt;zsh release tags (5.9.2, 12 July 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/en-us/102360" rel="noopener noreferrer"&gt;Use zsh as the default shell on your Mac — Apple Support&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>zsh</category>
      <category>macos</category>
      <category>olcum</category>
      <category>aliskanlik</category>
    </item>
    <item>
      <title>I Left Sixty-Eight Backups; nginx Read Exactly One</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Fri, 02 Oct 2026 11:48:13 +0000</pubDate>
      <link>https://dev.to/merbayerp/i-left-sixty-eight-backups-nginx-read-exactly-one-1ape</link>
      <guid>https://dev.to/merbayerp/i-left-sixty-eight-backups-nginx-read-exactly-one-1ape</guid>
      <description>&lt;p&gt;Before I touch a live configuration file, my hands type the same line on their&lt;br&gt;
own: &lt;code&gt;cp -p file file.bak-$(date +%Y%m%d-%H%M%S)&lt;/code&gt;. An old reflex. If the change goes&lt;br&gt;
wrong, the copy is right there, two keystrokes away. On the evening of 28 September, while reworking the blog's nginx&lt;br&gt;
configuration, that line got typed again.&lt;/p&gt;

&lt;p&gt;Four days later the copy was still there. Not merely there — inside the&lt;br&gt;
directory nginx &lt;strong&gt;reads&lt;/strong&gt;. Over those four days nginx restarted once&lt;br&gt;
(30 September, 21:09) and reloaded its configuration seven times; every one of&lt;br&gt;
those times it parsed that backup too, took its &lt;code&gt;server&lt;/code&gt; block into memory, and&lt;br&gt;
told me so, in sixteen lines of warnings. That nobody read those warnings for four days is an inference&lt;br&gt;
from the backup still being there: once the "test is successful" line at the&lt;br&gt;
bottom of &lt;code&gt;nginx -t&lt;/code&gt; shows up, the job counts as done and the sixteen lines&lt;br&gt;
above it go unread.&lt;/p&gt;

&lt;p&gt;The backup did not take over the site. Nothing I did prevented it. This piece&lt;br&gt;
is about what that "nothing" actually was: the safety of the habit comes not&lt;br&gt;
from me but from two conventions I never chose.&lt;/p&gt;
&lt;h2&gt;
  
  
  First, the count: how widespread is the habit
&lt;/h2&gt;

&lt;p&gt;Wondering about one file is no reason to look at only one file. A scan of &lt;code&gt;/etc&lt;/code&gt;&lt;br&gt;
down to three levels, for every name that smells like a backup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;find /etc &lt;span class="nt"&gt;-maxdepth&lt;/span&gt; 3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="se"&gt;\(&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.bak"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.bak-*"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.old"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.disabled"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
     &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.orig"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.save"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*~"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.pre-*"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
     &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.before-*"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.next"&lt;/span&gt; &lt;span class="se"&gt;\)&lt;/span&gt; &lt;span class="nt"&gt;-type&lt;/span&gt; f | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The answer: &lt;strong&gt;68&lt;/strong&gt;. Broken down by directory:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Directory&lt;/th&gt;
&lt;th&gt;Copies&lt;/th&gt;
&lt;th&gt;How does the system read this directory?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/etc/nginx/sites-available&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;Never read (link targets only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/etc/systemd/system&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;Only recognised unit suffixes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;/etc&lt;/code&gt; (root)&lt;/td&gt;
&lt;td&gt;7&lt;/td&gt;
&lt;td&gt;Files are called by exact name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/etc/nginx/snippets&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Only via &lt;code&gt;include snippets/&amp;lt;name&amp;gt;&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/etc/nginx/conf.d&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;include conf.d/*.conf&lt;/code&gt; — suffix filtered&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/etc/nginx/sites-enabled&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;include sites-enabled/*&lt;/code&gt; — &lt;strong&gt;everything&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those six rows account for 40 files. The remaining 28 are scattered two and&lt;br&gt;
three at a time across &lt;code&gt;/etc&lt;/code&gt;: application config directories, &lt;code&gt;ssh&lt;/code&gt;, &lt;code&gt;apt&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;fail2ban&lt;/code&gt; — and five of them in the two directories already set aside for the&lt;br&gt;
purpose, &lt;code&gt;/etc/nginx/backups&lt;/code&gt; and &lt;code&gt;/etc/nginx/disabled-backups&lt;/code&gt;. The blog's deploy&lt;br&gt;
directory holds five more; those come later.&lt;/p&gt;

&lt;p&gt;Those three system directories got a separate check, because the claim in the&lt;br&gt;
title depends on them too: the line in &lt;code&gt;sshd_config&lt;/code&gt; reads&lt;br&gt;
&lt;code&gt;Include /etc/ssh/sshd_config.d/*.conf&lt;/code&gt;, so it is suffix filtered, and &lt;code&gt;apt&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;fail2ban&lt;/code&gt; go by extension as well. In all three, my backups are not read. The&lt;br&gt;
bare wildcard exists in exactly one place.&lt;/p&gt;

&lt;p&gt;The table was a relief, and immediately afterwards an irritation. The relief:&lt;br&gt;
67 of the 68 copies sit where the system does not care. The irritation comes&lt;br&gt;
from the same place — which ones the system ignores was not something I knew&lt;br&gt;
while dropping the copy. If sixty-seven of sixty-eight shots land, that is not&lt;br&gt;
marksmanship; that is the shape of the distribution.&lt;/p&gt;
&lt;h2&gt;
  
  
  Which directory gets read, and which does not
&lt;/h2&gt;

&lt;p&gt;What draws the line is not my attention. It is two lines inside &lt;code&gt;nginx.conf&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;include&lt;/span&gt; &lt;span class="n"&gt;/etc/nginx/conf.d/*.conf&lt;/span&gt;;
&lt;span class="k"&gt;include&lt;/span&gt; &lt;span class="n"&gt;/etc/nginx/sites-enabled/*&lt;/span&gt;;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first imposes a suffix. The file&lt;br&gt;
&lt;code&gt;mustafaerbay-cache.conf.bak-20260928-191601&lt;/code&gt; does not end in &lt;code&gt;.conf&lt;/code&gt;, so it&lt;br&gt;
never enters the &lt;code&gt;conf.d&lt;/code&gt; sweep — my backup sits there like an invisible stone.&lt;br&gt;
The second imposes nothing. A bare &lt;code&gt;*&lt;/code&gt; makes everything in that directory part&lt;br&gt;
of the configuration: &lt;code&gt;.bak&lt;/code&gt;, &lt;code&gt;.disabled&lt;/code&gt;, the &lt;code&gt;~&lt;/code&gt; file an editor left behind,&lt;br&gt;
all of it.&lt;/p&gt;

&lt;p&gt;Debian's &lt;code&gt;sites-available&lt;/code&gt; / &lt;code&gt;sites-enabled&lt;/code&gt; pair exists precisely for this. The&lt;br&gt;
real file lives in &lt;code&gt;sites-available&lt;/code&gt;, and &lt;code&gt;sites-enabled&lt;/code&gt; holds a symlink to it.&lt;br&gt;
In that arrangement, a copy dropped "beside" the file lands automatically on the&lt;br&gt;
&lt;code&gt;sites-available&lt;/code&gt; side — the side nginx does not look at. The convention itself&lt;br&gt;
is a guardrail.&lt;/p&gt;

&lt;p&gt;On the server, &lt;code&gt;sites-enabled&lt;/code&gt; holds 45 entries. Of those, &lt;strong&gt;41 are symlinks and&lt;br&gt;
4 are real files&lt;/strong&gt;. The blog's configuration is one of those four. Its backup is&lt;br&gt;
the second. So the copy went into one of the four places where the guardrail is&lt;br&gt;
switched off, and that is not a surprise either — the reason a real file lives&lt;br&gt;
there is also me. The live configuration now ships from the repository; the file&lt;br&gt;
on the server is byte-for-byte identical to the 222-line version under&lt;br&gt;
&lt;code&gt;deploy/nginx-live/&lt;/code&gt;. Automation wrote a plain file where the distribution&lt;br&gt;
expected a symlink. And then a copy went next to the plain file.&lt;/p&gt;

&lt;p&gt;nginx did not pass over this in silence. Its &lt;code&gt;nginx -t&lt;/code&gt; output carries sixteen&lt;br&gt;
lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;nginx: [warn] conflicting server name "mustafaerbay.com.tr" on 0.0.0.0:80, ignored
nginx: [warn] conflicting server name "www.mustafaerbay.com" on [::]:443, ignored
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four hostnames (&lt;code&gt;mustafaerbay.com.tr&lt;/code&gt;, &lt;code&gt;mustafaerbay.com&lt;/code&gt; and the &lt;code&gt;www&lt;/code&gt; form of&lt;br&gt;
each) times four listening sockets (IPv4/IPv6 over 80/443) equals sixteen&lt;br&gt;
warnings. Each says the same thing: you defined this name twice, and the second&lt;br&gt;
one is being ignored.&lt;/p&gt;
&lt;h2&gt;
  
  
  So which one does it ignore?
&lt;/h2&gt;

&lt;p&gt;The word "second" is the crux here, because the ordering is not mine. Rather&lt;br&gt;
than touch production, the question went to a lab: nginx 1.29.8 in Docker, two&lt;br&gt;
configuration files, the same &lt;code&gt;server_name&lt;/code&gt;, the same listening socket. One&lt;br&gt;
returns &lt;code&gt;GERCEK-KONFIG&lt;/code&gt;, the other &lt;code&gt;YEDEK-KONFIG&lt;/code&gt;. The files sit under &lt;code&gt;conf.d&lt;/code&gt;,&lt;br&gt;
but the include line was switched to the server's risky pattern —&lt;br&gt;
&lt;code&gt;include /etc/nginx/conf.d/*;&lt;/code&gt;, a bare wildcard. Otherwise &lt;code&gt;conf.d&lt;/code&gt;'s own suffix&lt;br&gt;
filter would have cancelled the whole experiment.&lt;/p&gt;

&lt;p&gt;In the first run the backup carries a &lt;strong&gt;suffix&lt;/strong&gt;, as on the server: &lt;code&gt;site&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;site.bak-20260928&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# nginx -T load order
/etc/nginx/conf.d/site
/etc/nginx/conf.d/site.bak-20260928

$ curl -H 'Host: lab.example.com' ...   → GERCEK-KONFIG
$ curl -H 'Host: baska.example.com' ... → GERCEK-KONFIG
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the second run a single thing changed — the backup was renamed to &lt;code&gt;bak-site&lt;/code&gt;.&lt;br&gt;
Same content, same line count, same permissions. Only the name now sorts ahead&lt;br&gt;
of the real one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# nginx -T load order
/etc/nginx/conf.d/bak-site
/etc/nginx/conf.d/site

$ curl -H 'Host: lab.example.com' ...   → YEDEK-KONFIG
$ curl -H 'Host: baska.example.com' ... → YEDEK-KONFIG
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The backup took over the site. It also answered the non-matching &lt;code&gt;Host&lt;/code&gt; header,&lt;br&gt;
because as the nginx documentation puts it, the default server is the first&lt;br&gt;
block listed for that port, and "the default server is a property of the listen&lt;br&gt;
port and not of the server name".&lt;/p&gt;

&lt;p&gt;Those lab files are synthetic; their only job was to report which one answers.&lt;br&gt;
The lab ran 1.29.8 while the server runs 1.30.0; nothing changed on the &lt;code&gt;include&lt;/code&gt;&lt;br&gt;
side — the &lt;code&gt;glob&lt;/code&gt; call in the nginx source still reads the same on today's&lt;br&gt;
&lt;code&gt;master&lt;/code&gt;.&lt;br&gt;
The backup on the server, by contrast, holds the 29 July version: 228 lines, 28&lt;br&gt;
of them different from the live file. Had the name sorted the other way, a&lt;br&gt;
two-month-old configuration would have been serving the site.&lt;/p&gt;

&lt;p&gt;The ordering comes from POSIX. nginx expands the pattern with &lt;code&gt;glob()&lt;/code&gt; and passes&lt;br&gt;
zero as the second argument:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;glob&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="kt"&gt;char&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;gl&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;pattern&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;gl&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;pglob&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Since &lt;code&gt;GLOB_NOSORT&lt;/code&gt; is not set, sorting stays on. The glibc&lt;br&gt;
&lt;a href="https://man7.org/linux/man-pages/man3/glob.3.html" rel="noopener noreferrer"&gt;manual&lt;/a&gt; on the server is&lt;br&gt;
terse: "By default, the returned pathnames are sorted."&lt;br&gt;
&lt;a href="https://pubs.opengroup.org/onlinepubs/9699919799/functions/glob.html" rel="noopener noreferrer"&gt;POSIX&lt;/a&gt;&lt;br&gt;
says what that order actually is: "Ordinarily, glob() sorts the matching&lt;br&gt;
pathnames according to the current setting of the LC_COLLATE category." So file&lt;br&gt;
names decide which configuration serves the site — though "alphabetical" is not&lt;br&gt;
quite the right word, since the server's locale sets the order.&lt;/p&gt;

&lt;p&gt;This is where it became clear that luck was not the operative force, and that is&lt;br&gt;
the part that nags: the &lt;code&gt;file.bak-timestamp&lt;/code&gt; shape is &lt;strong&gt;structurally&lt;/strong&gt; safe.&lt;br&gt;
Because it appends a suffix rather than a prefix, the copy always falls behind&lt;br&gt;
the original name; the string &lt;code&gt;site&lt;/code&gt; is a prefix of &lt;code&gt;site.bak-20260928&lt;/code&gt;, and the&lt;br&gt;
shorter one sorts first — a relation that holds whatever the collation setting&lt;br&gt;
is. Had the reflex settled on &lt;code&gt;bak-file&lt;/code&gt;, &lt;code&gt;00-file&lt;/code&gt; or&lt;br&gt;
&lt;code&gt;old-file&lt;/code&gt; instead, this would be an incident report. The&lt;br&gt;
safety of the habit comes from the direction of the naming, not from its purpose.&lt;/p&gt;
&lt;h2&gt;
  
  
  One habit, three other directories, three different verdicts
&lt;/h2&gt;

&lt;p&gt;The interesting part is that the same &lt;code&gt;cp&lt;/code&gt; reflex produces entirely different&lt;br&gt;
outcomes elsewhere.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVsiY3AgZmlsZSBmaWxlLmJhay1kYXRlIl0gLS0-IEJ7IldoaWNoIGRpcmVjdG9yeSBkaWQgdGhlIGNvcHkgbGFuZCBpbj8ifQogIEIgLS0-fCJzaXRlcy1lbmFibGVkIChpbmNsdWRlICopInwgQ1siQUxMIGZpbGVzIHJlYWQ8YnIvPk5hbWUgb3JkZXIgZGVjaWRlczxici8-RklSU1QgbG9hZGVkIHdpbnMiXQogIEIgLS0-fCJjb25mLmQgKGluY2x1ZGUgKi5jb25mKSJ8IERbIlN1ZmZpeCBmaWx0ZXI8YnIvPkNvcHkgaXMgbm90IHJlYWQiXQogIEIgLS0-fCIvZXRjL3N5c3RlbWQvc3lzdGVtInwgRVsiT25seSBrbm93biBzdWZmaXhlczxici8-VW5pdCBjb3VsZCBub3QgYmUgZm91bmQiXQogIEIgLS0-fCIqLnNlcnZpY2UuZCBkcm9wLWluInwgRlsiTGV4aWNvZ3JhcGhpYyBtZXJnZTxici8-TEFTVCBhcHBsaWVkIHdpbnMiXQogIEIgLS0-fCIvZXRjL2Nyb24uZCJ8IEdbIkRvdHRlZCBuYW1lIHNraXBwZWQ8YnIvPlNpbGVudGx5IG5ldmVyIHJ1bnMiXQogIEMgLS0-IEhbIlJpc2t5OiBsaXZlIGJlaGF2aW91ciBjaGFuZ2VzIl0KICBEIC0tPiBJWyJIYXJtbGVzcyBidXQgaW52aXNpYmxlIl0KICBFIC0tPiBJCiAgRiAtLT4gSAogIEcgLS0-IEpbIlJpc2t5IGluIHJldmVyc2U6IHRoZSBjb3B5IHlvdSBlZGl0IG5ldmVyIHJ1bnMiXQ%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVsiY3AgZmlsZSBmaWxlLmJhay1kYXRlIl0gLS0-IEJ7IldoaWNoIGRpcmVjdG9yeSBkaWQgdGhlIGNvcHkgbGFuZCBpbj8ifQogIEIgLS0-fCJzaXRlcy1lbmFibGVkIChpbmNsdWRlICopInwgQ1siQUxMIGZpbGVzIHJlYWQ8YnIvPk5hbWUgb3JkZXIgZGVjaWRlczxici8-RklSU1QgbG9hZGVkIHdpbnMiXQogIEIgLS0-fCJjb25mLmQgKGluY2x1ZGUgKi5jb25mKSJ8IERbIlN1ZmZpeCBmaWx0ZXI8YnIvPkNvcHkgaXMgbm90IHJlYWQiXQogIEIgLS0-fCIvZXRjL3N5c3RlbWQvc3lzdGVtInwgRVsiT25seSBrbm93biBzdWZmaXhlczxici8-VW5pdCBjb3VsZCBub3QgYmUgZm91bmQiXQogIEIgLS0-fCIqLnNlcnZpY2UuZCBkcm9wLWluInwgRlsiTGV4aWNvZ3JhcGhpYyBtZXJnZTxici8-TEFTVCBhcHBsaWVkIHdpbnMiXQogIEIgLS0-fCIvZXRjL2Nyb24uZCJ8IEdbIkRvdHRlZCBuYW1lIHNraXBwZWQ8YnIvPlNpbGVudGx5IG5ldmVyIHJ1bnMiXQogIEMgLS0-IEhbIlJpc2t5OiBsaXZlIGJlaGF2aW91ciBjaGFuZ2VzIl0KICBEIC0tPiBJWyJIYXJtbGVzcyBidXQgaW52aXNpYmxlIl0KICBFIC0tPiBJCiAgRiAtLT4gSAogIEcgLS0-IEpbIlJpc2t5IGluIHJldmVyc2U6IHRoZSBjb3B5IHlvdSBlZGl0IG5ldmVyIHJ1bnMiXQ%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="1275" height="702"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;All three verdicts got measured.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The systemd unit directory.&lt;/strong&gt; &lt;code&gt;/etc/systemd/system&lt;/code&gt; holds ten in-place&lt;br&gt;
backups: &lt;code&gt;alert-monitor.service.bak-20260703-131021&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;kopru-gateway.service.bak-20260730-115225&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;vpsman-agent.service.bak-20260911-1550&lt;/code&gt; and their siblings. systemd loads none&lt;br&gt;
of them, because none carries a recognised unit suffix. Asked directly, it is&lt;br&gt;
unambiguous:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ systemctl status alert-monitor.service.bak-20260703-131021
Unit alert-monitor.service.bak-20260703-131021.service could not be found.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It appended &lt;code&gt;.service&lt;/code&gt; to the name and found nothing. The list of loaded unit&lt;br&gt;
files matches zero of these patterns. All ten copies are entirely harmless — and&lt;br&gt;
entirely invisible. Anyone opening that directory has to read timestamps to work&lt;br&gt;
out which file is live.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Drop-in directories.&lt;/strong&gt; Here the rule inverts. In the words of the&lt;br&gt;
&lt;code&gt;systemd.unit(5)&lt;/code&gt; documentation, files with the &lt;code&gt;.conf&lt;/code&gt; suffix "will be merged in&lt;br&gt;
the alphanumeric order and parsed after the main unit file itself has been&lt;br&gt;
parsed", and multiple drop-ins with different names "are applied in lexicographic&lt;br&gt;
order". Whatever is applied later overrides the same setting from earlier. In&lt;br&gt;
nginx the first wins; here the last wins. Two directories, two opposite rules,&lt;br&gt;
one habit.&lt;/p&gt;

&lt;p&gt;Two units on the server have two drop-ins each. Under &lt;code&gt;kopru-web.service.d&lt;/code&gt; sit&lt;br&gt;
&lt;code&gt;20-live-readiness.conf&lt;/code&gt; and &lt;code&gt;order-key.conf&lt;/code&gt;, and &lt;code&gt;systemctl cat&lt;/code&gt; prints them in&lt;br&gt;
that order. They do not collide, because each defines a differently named&lt;br&gt;
credential and that setting is the kind that appends to a list. But that is&lt;br&gt;
luck: had both written the same single-valued setting, &lt;code&gt;order-key.conf&lt;/code&gt; would&lt;br&gt;
have won, and my assumption that the numbered file takes precedence because of&lt;br&gt;
its number would have survived untested.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;cron.d.&lt;/strong&gt; Here the trap runs in reverse. Debian's &lt;code&gt;cron(8)&lt;/code&gt; manual states that&lt;br&gt;
names in &lt;code&gt;/etc/cron.d&lt;/code&gt; "must consist solely of upper- and lower-case letters,&lt;br&gt;
digits, underscores, and hyphens"; a file containing dots is ignored. Copy a&lt;br&gt;
cron.d job to &lt;code&gt;job.bak&lt;/code&gt;, delete the original, and what remains is a file that&lt;br&gt;
never runs and never complains. My own &lt;code&gt;cron.d&lt;/code&gt; holds ten entries: nine ordinary&lt;br&gt;
jobs and a tenth called &lt;code&gt;.placeholder&lt;/code&gt; — a file silently ignored exactly as that&lt;br&gt;
rule intends.&lt;/p&gt;
&lt;h2&gt;
  
  
  The real cost of a copy is not being read
&lt;/h2&gt;

&lt;p&gt;Three of the seven copies at the root of &lt;code&gt;/etc&lt;/code&gt; are, as their names make plain,&lt;br&gt;
backups of files that carry credentials: one access-token file and two &lt;code&gt;env&lt;/code&gt;&lt;br&gt;
files. Their permissions checked out — all three are &lt;code&gt;-rw-------&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;root:root&lt;/code&gt;, same as the originals. There is no exposure here.&lt;/p&gt;

&lt;p&gt;The risk sits elsewhere: these copies freeze the old value of a token. Rotate the&lt;br&gt;
token and the copy does not rotate with it; what remains inside is either a&lt;br&gt;
string that no longer works or — worse — one that still does, because it was&lt;br&gt;
never revoked. Sitting in &lt;code&gt;/etc&lt;/code&gt; as a date-stamped file, in no inventory.&lt;/p&gt;

&lt;p&gt;On all three, the name and the timestamp disagree. Two of them carry &lt;code&gt;20260914&lt;/code&gt;&lt;br&gt;
in the name while their content was last written on 25 June — eighty-one days&lt;br&gt;
apart; on the third the gap is twenty-one days. The copies were made in a way&lt;br&gt;
that preserves timestamps, so the name records when the copy was taken and the&lt;br&gt;
timestamp records when the content was written.&lt;br&gt;
Anyone reading that directory without knowing the convention trusts the wrong&lt;br&gt;
file.&lt;/p&gt;
&lt;h2&gt;
  
  
  The alarm that rings about the wrong thing
&lt;/h2&gt;

&lt;p&gt;While counting backups, the blog's own working tree came under the same look,&lt;br&gt;
because the live site is served from it: the &lt;code&gt;blog-serve&lt;/code&gt; container mounts&lt;br&gt;
&lt;code&gt;/opt/mustafaerbay&lt;/code&gt; read-only as &lt;code&gt;/app&lt;/code&gt; and runs &lt;code&gt;dist/server/entry.mjs&lt;/code&gt; from&lt;br&gt;
inside it. In that directory &lt;code&gt;git status&lt;/code&gt; reports six modified tracked files and&lt;br&gt;
a HEAD stuck on 16 May. Today is 2 October; 139 days sit in between.&lt;/p&gt;

&lt;p&gt;On first read that looked alarming. A four-and-a-half-month-old HEAD on a live&lt;br&gt;
tree, six dirty files — as though edits nobody knows about had piled up on the&lt;br&gt;
server. Then each file went head to head with today's &lt;code&gt;origin/main&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File reported as "modified"&lt;/th&gt;
&lt;th&gt;Against today's &lt;code&gt;origin/main&lt;/code&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;deploy/backup.sh&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Identical (84 lines)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;scripts/notify-subscribers.mjs&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Identical (278 lines)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;package.json&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;package-lock.json&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;deploy/pipeline-health.sh&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;One line differs&lt;/strong&gt; (line 207)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;deploy/update.sh&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Not in the repository at all&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So four of the six files &lt;code&gt;git status&lt;/code&gt; calls "modified" have not changed at all;&lt;br&gt;
they merely look different against an old HEAD. The dirt is not drift, it is a&lt;br&gt;
stale reference. The alarm rings, but points at the wrong thing — and because it&lt;br&gt;
has pointed at the wrong thing for four and a half months, the habit of looking&lt;br&gt;
at it is long gone.&lt;/p&gt;

&lt;p&gt;The one-line difference turned out to be a familiar face. The health script on&lt;br&gt;
the server still prints the repository's old address in its alert emails:&lt;br&gt;
&lt;code&gt;github.com/mustafa_itwise/mustafaerbay&lt;/code&gt;. After the repository moved to the&lt;br&gt;
organisation on 28 September, three scripts still calling the old name got&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/career/depoyu-tasidim-uc-script-hala-eski-adi-cagiriyor/" rel="noopener noreferrer"&gt;cleaned up that day&lt;/a&gt;;&lt;br&gt;
this is the fourth. It was inside my own alerting machinery — in the very line&lt;br&gt;
meant to warn me when something breaks.&lt;/p&gt;
&lt;h2&gt;
  
  
  The debt that stays quiet
&lt;/h2&gt;

&lt;p&gt;The last row of that table is the real finding. &lt;code&gt;deploy/update.sh&lt;/code&gt; was deleted&lt;br&gt;
from the repository on 26 June, as part of the incremental-build work. The copy&lt;br&gt;
on the server was last modified on &lt;strong&gt;14 September&lt;/strong&gt;. A script that does not exist&lt;br&gt;
in the repository was touched 80 days after its deletion, and the copy beside it&lt;br&gt;
carries that same day's &lt;code&gt;migrate&lt;/code&gt; tag in its name. I am the only person who does&lt;br&gt;
that work on that server. It has three&lt;br&gt;
siblings: &lt;code&gt;update.sh.bak-1780461942&lt;/code&gt; and &lt;code&gt;update.sh.bak-20260914-migrate&lt;/code&gt; (both&lt;br&gt;
3 June, 07:45) and &lt;code&gt;update.sh.disabled&lt;/code&gt; (11 May).&lt;/p&gt;

&lt;p&gt;Worse, a unit file still calls it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Description=mustafaerbay.com deploy pull job
ExecStart=/bin/bash /opt/mustafaerbay/deploy/update.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The good news: the timer is off. For &lt;code&gt;mustafaerbay-deploy.timer&lt;/code&gt;, &lt;code&gt;is-enabled&lt;/code&gt;&lt;br&gt;
answers &lt;code&gt;disabled&lt;/code&gt;; the service is &lt;code&gt;static&lt;/code&gt; and &lt;code&gt;inactive&lt;/code&gt;, with not one journal&lt;br&gt;
line since 26 June. The bad news: that timer's content reads&lt;br&gt;
&lt;code&gt;OnUnitActiveSec=1min&lt;/code&gt;. A single &lt;code&gt;systemctl enable --now&lt;/code&gt; would start running,&lt;br&gt;
once a minute, a script no repository knows about. The gun is not loaded, but it&lt;br&gt;
sits next to the trigger with no label saying which repository, which version.&lt;/p&gt;

&lt;p&gt;The same directory also holds an 830 MB &lt;code&gt;dist-persistent&lt;/code&gt;: 13,445 files, the&lt;br&gt;
newest dated 29 June. Leftovers from a build experiment that never merged.&lt;br&gt;
Next to it, &lt;code&gt;dist-old&lt;/code&gt; was refreshed today — that one is not a backup but the&lt;br&gt;
rollback copy the deployment rotates on every run. The two sit side by side, one&lt;br&gt;
part of a live mechanism, one a ninety-five-day-old fossil. In a directory&lt;br&gt;
listing they are indistinguishable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the copy belongs
&lt;/h2&gt;

&lt;p&gt;After the measurements, the rule got sharper, and all of it hangs on one&lt;br&gt;
question: &lt;strong&gt;does anything read this directory, and by what criterion?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If the directory is read with a bare wildcard (&lt;code&gt;*&lt;/code&gt;), a copy goes there
&lt;strong&gt;never&lt;/strong&gt;. For the blog's configuration the right place is &lt;code&gt;sites-available&lt;/code&gt;
or a separate backup directory. The server already has &lt;code&gt;/etc/nginx/backups&lt;/code&gt;
and &lt;code&gt;/etc/nginx/disabled-backups&lt;/code&gt; — past me did the right thing, then forgot.&lt;/li&gt;
&lt;li&gt;If it is read with a suffix filter (&lt;code&gt;*.conf&lt;/code&gt;), the copy is safe but
&lt;strong&gt;invisible&lt;/strong&gt;. Invisibility is not free: three months later you work out which
file is live from its timestamp.&lt;/li&gt;
&lt;li&gt;Know the direction of the sort. In nginx a conflicting name goes to whichever
loads first — that one is not in the documentation; it is behaviour measured in
a lab and in the source. In systemd drop-ins it goes to whatever applies last,
and that one is documented. Naming a file &lt;code&gt;00-&lt;/code&gt; moves it to the front in one
system and to the back in the other.&lt;/li&gt;
&lt;li&gt;Use a suffix, not a prefix. &lt;code&gt;file.bak-timestamp&lt;/code&gt; falls behind the original
name; &lt;code&gt;bak-file&lt;/code&gt; jumps ahead of it. That single structural detail is what
makes the habit safe.&lt;/li&gt;
&lt;li&gt;If configuration ships from a repository, break the symlink convention
deliberately, not by accident. A deploy script that writes a plain file also
switches off the protection the distribution handed you.&lt;/li&gt;
&lt;li&gt;Build the detection too, because &lt;code&gt;nginx -t&lt;/code&gt; returns &lt;strong&gt;zero&lt;/strong&gt; alongside those
sixteen warnings — measured in the lab. No deployment gate will ever see them.
Three checks are enough: treat warnings as fatal with
&lt;code&gt;nginx -t 2&amp;gt;&amp;amp;1 | grep -q "\[warn\]"&lt;/code&gt;, read the real load order with
&lt;code&gt;nginx -T | grep "# configuration file"&lt;/code&gt;, and count the non-symlinks with
&lt;code&gt;find /etc/nginx/sites-enabled -maxdepth 1 -type f&lt;/code&gt; (mine returned four).&lt;/li&gt;
&lt;li&gt;If you leave a backup of a file that carries credentials, add that backup to
the rotation list too. Even a copy with correct permissions can keep an
unrevoked old secret alive indefinitely.&lt;/li&gt;
&lt;li&gt;For a git working copy sitting on a live tree, either keep HEAD current or
stop believing &lt;code&gt;git status&lt;/code&gt; in that directory. There is no middle ground: a
stale HEAD manufactures fake debt daily, and hides the real debt inside that
noise.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Of sixty-eight copies only one mattered, and what kept that one from doing harm&lt;br&gt;
were two conventions I had never thought about: a naming habit that grows by&lt;br&gt;
suffix, and Debian's symlink farm. Neither is my design. One arrived through&lt;br&gt;
muscle memory, the other comes bundled with the distribution.&lt;/p&gt;

&lt;p&gt;The finding that stings most is the warning lines. For four days nginx said, in&lt;br&gt;
sixteen lines, "you defined this name twice"; those lines got skipped, and the&lt;br&gt;
one at the bottom — "test is successful" — was taken as the end of the job. The &lt;code&gt;git status&lt;/code&gt; output that shows six dirty files has been lying&lt;br&gt;
every day for four and a half months, so I stopped looking at it. An alarm that&lt;br&gt;
rings about the wrong thing is more dangerous than one that stays silent,&lt;br&gt;
because it trains you against your own instrument. The day half-finished work&lt;br&gt;
got counted four ways and&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/career/yetmis-dal-biriktirdim-altisi-gercekten-yarim-kaldi/" rel="noopener noreferrer"&gt;produced four different answers&lt;/a&gt;&lt;br&gt;
taught the same lesson; apparently learning it once is not enough.&lt;/p&gt;

&lt;p&gt;The copy left inside that directory was not a safety net. It was part of the&lt;br&gt;
configuration. What separated the two was the alphabetical order of its name.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://nginx.org/en/docs/http/request_processing.html" rel="noopener noreferrer"&gt;nginx — How nginx processes a request&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://nginx.org/en/docs/ngx_core_module.html" rel="noopener noreferrer"&gt;nginx — ngx_core_module: include&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://nginx.org/en/docs/http/server_names.html" rel="noopener noreferrer"&gt;nginx — Server names&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/nginx/nginx/blob/master/src/os/unix/ngx_files.c" rel="noopener noreferrer"&gt;nginx source — ngx_open_glob (src/os/unix/ngx_files.c)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/systemd/systemd/blob/main/man/systemd.unit.xml" rel="noopener noreferrer"&gt;systemd — systemd.unit(5) manual source&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://manpages.debian.org/trixie/cron/cron.8.en.html" rel="noopener noreferrer"&gt;Debian — cron(8)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>sistemyonetimi</category>
      <category>nginx</category>
      <category>systemd</category>
      <category>olcum</category>
    </item>
    <item>
      <title>I Tested ispmanager in Three Rounds: One Leak Fixed, Six Findings Retracted</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:39:18 +0000</pubDate>
      <link>https://dev.to/merbayerp/i-tested-ispmanager-in-three-rounds-one-leak-fixed-six-findings-retracted-afp</link>
      <guid>https://dev.to/merbayerp/i-tested-ispmanager-in-three-rounds-one-leak-fixed-six-findings-retracted-afp</guid>
      <description>&lt;p&gt;&lt;em&gt;Disclosure: ispmanager provided a free test licence for this review; I was not paid. In our correspondence they also mentioned a free month of use and a "longer-term partnership" that could include free ongoing use of the panel; I deferred all of that until after testing and accepted none of it. The vendor has no control over the test scope, the findings, the severity ratings or this text. The text was sent to them for a technical accuracy check only: on 18 September, and in its post-second-round form on 21 September. The ispmanager links in this article exist so readers can reach the documentation; nothing was received in return for them.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On 10 September an email arrived. Aleksandra from ispmanager had been reading the blog; they wanted to grow in the Turkish market and were looking for an honest review from the perspective of someone who normally wouldn't need a panel. My relationship with control panels has never been good. On my own servers I prefer the terminal, because when something breaks I know exactly whom to ask: myself.&lt;/p&gt;

&lt;p&gt;That is also what made the offer interesting. The question was never "is the UI easy?" It was what the panel takes on, what it keeps out of sight, and whether it tells the operator the truth when something breaks. In my first reply I set a single condition: independence. I would test in my own environment and reach my own conclusions, and the final word on both positive and negative findings would be mine. They agreed.&lt;/p&gt;

&lt;p&gt;Three weeks later I have 23 findings, three rounds of testing and an outcome I did not expect. The vendor first rejected the most serious finding, accepted it after the evidence from my second round, and closed it in version 6.153; on 1 October I verified that fix in my own lab. I, in turn, retracted the substance of six of my own findings, because when I measured again, I turned out to be wrong. This article covers both. In my view, the value of a review is not in its number of findings but in how those findings change when new evidence arrives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three rounds, one environment
&lt;/h2&gt;

&lt;p&gt;The test environment was not a production server, and I should say so up front: Docker Desktop on a Mac, running a privileged AlmaLinux 9.8 container booted with &lt;code&gt;systemd&lt;/code&gt;. Because ispmanager's AlmaLinux repository ships only x86_64 packages, it ran under x86_64 emulation on Apple Silicon.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Round&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Version&lt;/th&gt;
&lt;th&gt;Focus&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;14 September&lt;/td&gt;
&lt;td&gt;Lite 6.150.4 (core 5.445.2)&lt;/td&gt;
&lt;td&gt;13 areas, 23 findings, 15-page report&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;21 September&lt;/td&gt;
&lt;td&gt;Lite 6.150.4, clean install&lt;/td&gt;
&lt;td&gt;Measuring the vendor's objections&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;1 October&lt;/td&gt;
&lt;td&gt;Lite 6.153.0 (core 5.448.0), clean install; stable 6.152.2 for comparison&lt;/td&gt;
&lt;td&gt;The SEC-02 fix + re-measuring the findings&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first round used the test licence the vendor sent. In the second and third rounds that licence's activation key was no longer accepted, so I used the trial licence the installer issues automatically to this IP. Round two installed the same packages as round one; round three installed 6.153.0 with the current installer script from the vendor's server.&lt;/p&gt;

&lt;p&gt;The method was the same every time. I exercised each feature through the panel's API first, then verified it on the operating system: the filesystem, the process table, &lt;code&gt;iptables&lt;/code&gt;, &lt;code&gt;systemd&lt;/code&gt;, the generated configuration files. On top of that I added hostile scenarios: symlink escapes, cross-tenant read attempts, killed services, a broken nginx configuration, changes made behind the panel's back.&lt;/p&gt;

&lt;p&gt;Running in a container produced three side effects, which I separated out from the start: IP management expects an &lt;code&gt;ifcfg-eth0&lt;/code&gt; file, panel login depends on the &lt;code&gt;sshd&lt;/code&gt; PAM profile, and MariaDB fails on first start. This round I finally found the cause of the last one. I'll come back to it in a minute, because it is small but instructive.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the panel takes ownership of
&lt;/h2&gt;

&lt;p&gt;This was the most important question for me, and the answer was more mature than I expected. The installed stack is a classic split: nginx terminates TLS and serves static content, and PHP runs either under MPM-ITK, one of Apache's third-party &lt;a href="https://httpd.apache.org/docs/current/mpm.html" rel="noopener noreferrer"&gt;MPM&lt;/a&gt; modules, with &lt;code&gt;mod_php&lt;/code&gt; under a separate UID per site, or in per-site PHP-FPM pools over a Unix socket. The panel itself lives in a separate HTTP server, &lt;code&gt;ihttpd&lt;/code&gt;, on port 1500.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IExSCiAgICBDWyJDbGllbnQ8YnIvPkhUVFAvSFRUUFMiXSAtLT4gTlsibmdpbnggMS4zMDxici8-VExTICsgc3RhdGljICsgc3ltbGluayBwcm90ZWN0aW9uIl0KICAgIE4gLS0-fCJtb2RfcGhwIG1vZGUifCBBWyJBcGFjaGUgTVBNLUlUSzxici8-cGVyLXNpdGUgVUlEIl0KICAgIE4gLS0-fCJGYXN0Q0dJInwgRlsiUEhQLUZQTTxici8-cGVyLXNpdGUgcG9vbCArIHNvY2tldCJdCiAgICBBIC0tPiBNWyJNYXJpYURCIDEwLjUiXQogICAgRiAtLT4gTQogICAgUFsiaWh0dHBkIDoxNTAwPGJyLz50aGUgcGFuZWwgaXRzZWxmIl0gLS4tPnwibWFuYWdlcyJ8IE4KICAgIFAgLS4tPnwibWFuYWdlcyJ8IEEKICAgIFAgLS4tPnwibWFuYWdlcyJ8IEY%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IExSCiAgICBDWyJDbGllbnQ8YnIvPkhUVFAvSFRUUFMiXSAtLT4gTlsibmdpbnggMS4zMDxici8-VExTICsgc3RhdGljICsgc3ltbGluayBwcm90ZWN0aW9uIl0KICAgIE4gLS0-fCJtb2RfcGhwIG1vZGUifCBBWyJBcGFjaGUgTVBNLUlUSzxici8-cGVyLXNpdGUgVUlEIl0KICAgIE4gLS0-fCJGYXN0Q0dJInwgRlsiUEhQLUZQTTxici8-cGVyLXNpdGUgcG9vbCArIHNvY2tldCJdCiAgICBBIC0tPiBNWyJNYXJpYURCIDEwLjUiXQogICAgRiAtLT4gTQogICAgUFsiaWh0dHBkIDoxNTAwPGJyLz50aGUgcGFuZWwgaXRzZWxmIl0gLS4tPnwibWFuYWdlcyJ8IE4KICAgIFAgLS4tPnwibWFuYWdlcyJ8IEEKICAgIFAgLS4tPnwibWFuYWdlcyJ8IEY%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="1137" height="319"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Configuration is generated from templates, and the distribution's own &lt;code&gt;nginx.conf&lt;/code&gt; only gets &lt;code&gt;include&lt;/code&gt; lines added; the file is not overwritten. Every site and user has &lt;code&gt;vhosts-resources/&amp;lt;site&amp;gt;/&lt;/code&gt; and &lt;code&gt;users-resources/&amp;lt;user&amp;gt;/&lt;/code&gt; directories; that is the way to extend things without breaking the panel. In the first round I added a line by hand to a vhost file the panel had generated, and after the panel rewrote the file the line was still there. A good sign for anyone who lives in fear of "don't touch it, the panel will overwrite it".&lt;/p&gt;

&lt;h2&gt;
  
  
  What held up
&lt;/h2&gt;

&lt;p&gt;This list is long, and all of it was measured. nginx returned 403 for a web-reachable symlink pointing at &lt;code&gt;/etc&lt;/code&gt;; the panel writes &lt;a href="https://nginx.org/en/docs/http/ngx_http_core_module.html#disable_symlinks" rel="noopener noreferrer"&gt;&lt;code&gt;disable_symlinks if_not_owner&lt;/code&gt;&lt;/a&gt; into the vhosts. Apache MPM-ITK really does run each site under its owner's UID; the test script under &lt;code&gt;mod_php&lt;/code&gt; ran as &lt;code&gt;uid=1010&lt;/code&gt; (&lt;code&gt;usera&lt;/code&gt;). A database user created through the panel could create and write tables in its own database and got &lt;code&gt;ERROR 1044&lt;/code&gt; on its neighbour's; I repeated that on 6.153 on 1 October as well.&lt;/p&gt;

&lt;p&gt;An archive taken with &lt;code&gt;backup2&lt;/code&gt; contained files with their ownership preserved, a valid &lt;code&gt;mysqldump&lt;/code&gt; and metadata. I deleted the site's files and restored them from the archive, and the site returned 200 again. Self-signed certificate generation, assignment to a site and the HTTP-to-HTTPS 301 redirect all worked. I injected a broken nginx configuration: the reload was refused and the running workers were left alone. The &lt;code&gt;history_users&lt;/code&gt; table records who did what, from which IP, and which fields changed, with a timestamp. Firewall rules added through the panel show up as commented, persistent &lt;code&gt;iptables&lt;/code&gt; entries and are removed cleanly. During testing my own IP got caught by the brute-force lockout, so that works too.&lt;/p&gt;

&lt;p&gt;Round three added two items to this list, and both came from places where I had been wrong: the panel's own metadata does go into the full backup, and backups can be managed through the API on Lite. I'll come back to both below.&lt;/p&gt;

&lt;h2&gt;
  
  
  SEC-02: three rounds of one password
&lt;/h2&gt;

&lt;p&gt;This turned out to be the longest-lived finding of the review, and in my view the most instructive.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Round one.&lt;/strong&gt; You register the database server in the panel with the root user and a password. Then you call the &lt;code&gt;func=db.server&lt;/code&gt; list: the response carries the password you saved, in plain text, in the &lt;code&gt;password=&lt;/code&gt; field. I rated it High in the report and wrote "including reseller accounts" under impact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The vendor's objection (18 September, in writing).&lt;/strong&gt; The engineering team assessed it as "by design". The reasoning: &lt;code&gt;db.server.edit&lt;/code&gt; is available only to the superadmin, the superadmin is equivalent to root on the server, so seeing the password gives them no new capability. The last sentence of their reply left the door open: if I could show a scenario in which someone other than the superadmin obtains the credential, they would reassess.&lt;/p&gt;

&lt;p&gt;That reply changed my initial assessment. I withdrew the reseller claim, because I had not measured it. But nobody had measured the "superadmin only" claim either.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Round two (21 September).&lt;/strong&gt; I created an ordinary administrator with &lt;code&gt;admin.edit&lt;/code&gt;, &lt;code&gt;level=29&lt;/code&gt;. The superadmin is &lt;code&gt;level=30&lt;/code&gt;. This account stayed within the limits the panel documents: the file manager, the shell, creating administrators and &lt;code&gt;db.server.edit&lt;/code&gt; were all denied. But it could call the &lt;code&gt;db.server&lt;/code&gt; list, and the list carried the password:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;authinfo=opadmin:...&amp;amp;func=db.server
id=1 name=MySQL type=mysql host=127.0.0.1 username=root password=MgrRoot#2026 savedver=10.5.29

authinfo=opadmin:...&amp;amp;func=file
ERROR access(function): You have insufficient permissions to execute the function 'file'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(The password is a lab-only test value.) The table in the vendor's &lt;a href="https://www.ispmanager.com/docs/ispmanager/user-levels" rel="noopener noreferrer"&gt;account levels&lt;/a&gt; documentation shows "DBMS management", "File manager", "Command execution" and "Shell client" as unavailable to the administrator. Everything that was denied matched the documentation exactly; the password in the &lt;code&gt;db.server&lt;/code&gt; list, however, leaked out beneath the documented boundary. And it was not a placeholder value: what the panel had stored was the real password of the MariaDB root account, in other words the key to every database on the server. I sent this to the vendor on 21 September, together with the reproduction steps.&lt;/p&gt;

&lt;p&gt;The whole weight of the threat-model debate rested on a single detail. The vendor had looked at the &lt;code&gt;db.server.edit&lt;/code&gt; form; the password was being returned by the &lt;code&gt;db.server&lt;/code&gt; list. Four letters separate the two function names, and one role separates their privilege levels.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The vendor's reply (23 September).&lt;/strong&gt; It was clear and direct: "You are correct: our previous response referred to the wrong API call." They reproduced the scenario in their own environment. They didn't hide their own perspective either: in their model the administrator is a trusted, high-privilege role, and as a rule they don't hide from an administrator what an administrator can see. But they agreed that a stored root password surfacing for a role that is denied the file manager, the shell and administrator management sits below that line, and wrote that they would remove the password from the &lt;code&gt;db.server&lt;/code&gt; response, targeting version 6.153, due on 29 September.&lt;/p&gt;

&lt;p&gt;I don't see this often. The vendor put forward a claim in their first reply, I measured it, the measurement refuted it, and the vendor revised its initial assessment in writing. To me that is more valuable than an "I found a vulnerability" story: a disagreement was settled by measurement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Round three (1 October).&lt;/strong&gt; According to ispmanager's &lt;a href="https://www.ispmanager.com/changelog" rel="noopener noreferrer"&gt;changelog&lt;/a&gt;, 6.153.0 shipped a day late, on 30 September, in the &lt;strong&gt;beta&lt;/strong&gt; channel; the stable channel moved to 6.152.2 the same day and is still there as of 2 October. I passed &lt;code&gt;--release 6.153&lt;/code&gt; to the installer, installed a clean 6.153.0 and repeated the second-round scenario exactly. Same administrator role, same call, three output formats:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# opadmin (level 29), 6.153.0
out=text      id=1 name=MySQL type=mysql remote_access=off host=localhost hostandport=127.0.0.1:3306 savedver=10.5.29
out=JSONdata  [{"id":"1","name":"MySQL","type":"mysql","remote_access":"off","host":"localhost","hostandport":"127.0.0.1:3306","savedver":"10.5.29"}]
out=xml       &amp;lt;elem&amp;gt;&amp;lt;id&amp;gt;1&amp;lt;/id&amp;gt;&amp;lt;name&amp;gt;MySQL&amp;lt;/name&amp;gt;...&amp;lt;savedver&amp;gt;10.5.29&amp;lt;/savedver&amp;gt;&amp;lt;/elem&amp;gt;

# root (level 30), 6.153.0
out=text      id=1 name=MySQL ... username=root password=MgrRoot#2026 savedver=10.5.29
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The administrator's response no longer contains a &lt;code&gt;password&lt;/code&gt; field, nor a &lt;code&gt;username&lt;/code&gt; field, in any of the three formats: text, XML and JSONdata. The superadmin's list still carries the password. That is consistent with the threat model the vendor defended from the start: the superadmin is root anyway. If it were up to me, I would mask the password on read paths there too, but that is a defence-in-depth preference, not a boundary violation.&lt;/p&gt;

&lt;p&gt;I also checked that the fix did not break anything else. The administrator is still denied &lt;code&gt;db.server.edit&lt;/code&gt;, the file manager, the shell, creating administrators and the scheduler. A user-level account cannot call &lt;code&gt;db.server&lt;/code&gt; at all. The administrator can create a database on behalf of a user, and the list keeps working without exposing the stored credentials. The test password does not appear anywhere in the panel's log files under &lt;code&gt;/usr/local/mgr5/var&lt;/code&gt;. One detail remains: although the documentation table shows "DBMS management" as unavailable to the administrator, the administrator can still call the &lt;code&gt;db.server&lt;/code&gt; list, now without the password. Documentation and behaviour still don't fully match, but no secret leaks any more.&lt;/p&gt;

&lt;p&gt;The detail that matters most for readers: for now the fix exists only in the beta channel. I ran the same scenario in a second container on 6.152.2, which is what the installer fetches with &lt;code&gt;--release stable&lt;/code&gt;. The administrator account still receives the password in plain text:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# opadmin (level 29), 6.152.2 (stable)
out=text  id=1 name=MySQL ... username=root password=MgrRoot#2026 savedver=10.5.29
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The vendor says it updates the beta channel every two weeks and the stable channel once a month; it hasn't yet said when 6.153 will reach stable. Until then, anyone on stable has two options: move to 6.153 in the beta channel, or create no administrator accounts other than the superadmin. If level-29 administrator accounts could reach the database server list before 6.153, the password may have been exposed to them; change the MariaDB root password after upgrading as well. The upgrade itself doesn't require this; the reason is that exposure window. How to do it is in the day-one list below, because the wrong method breaks the panel.&lt;/p&gt;

&lt;p&gt;One more note. The 6.153.0 changelog entry has three "fixed a security issue" lines; two are rated low and one medium, and none of them is named. I don't know which one is SEC-02, and I won't guess. What I do know is that the behaviour I measured has changed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBBWyIxNCBTZXAgLSBSb3VuZCAxPGJyLz5kYi5zZXJ2ZXIgcmV0dXJucyB0aGUgcGFzc3dvcmQgaW4gcGxhaW4gdGV4dCJdIC0tPiBCWyIxOCBTZXAgLSBWZW5kb3I8YnIvPnN1cGVyYWRtaW4gb25seSwgYnkgZGVzaWduIl0KICAgIEIgLS0-IENbIjIxIFNlcCAtIFJvdW5kIDI8YnIvPmEgbGV2ZWwgMjkgYWRtaW5pc3RyYXRvciBhbHNvIGdldHMgaXQiXQogICAgQyAtLT4gRFsiMjMgU2VwIC0gVmVuZG9yPGJyLz53ZSBsb29rZWQgYXQgdGhlIHdyb25nIEFQSSBjYWxsLCByZW1vdmFsIGluIDYuMTUzIl0KICAgIEQgLS0-IEVbIjMwIFNlcCAtIDYuMTUzLjAgaW4gdGhlIGJldGEgY2hhbm5lbCJdCiAgICBFIC0tPiBGWyIxIE9jdCAtIFJvdW5kIDM8YnIvPm5vIHBhc3N3b3JkIGluIHRoZSBhZG1pbiByZXNwb25zZSwgbm8gcmVncmVzc2lvbiJd%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBBWyIxNCBTZXAgLSBSb3VuZCAxPGJyLz5kYi5zZXJ2ZXIgcmV0dXJucyB0aGUgcGFzc3dvcmQgaW4gcGxhaW4gdGV4dCJdIC0tPiBCWyIxOCBTZXAgLSBWZW5kb3I8YnIvPnN1cGVyYWRtaW4gb25seSwgYnkgZGVzaWduIl0KICAgIEIgLS0-IENbIjIxIFNlcCAtIFJvdW5kIDI8YnIvPmEgbGV2ZWwgMjkgYWRtaW5pc3RyYXRvciBhbHNvIGdldHMgaXQiXQogICAgQyAtLT4gRFsiMjMgU2VwIC0gVmVuZG9yPGJyLz53ZSBsb29rZWQgYXQgdGhlIHdyb25nIEFQSSBjYWxsLCByZW1vdmFsIGluIDYuMTUzIl0KICAgIEQgLS0-IEVbIjMwIFNlcCAtIDYuMTUzLjAgaW4gdGhlIGJldGEgY2hhbm5lbCJdCiAgICBFIC0tPiBGWyIxIE9jdCAtIFJvdW5kIDM8YnIvPm5vIHBhc3N3b3JkIGluIHRoZSBhZG1pbiByZXNwb25zZSwgbm8gcmVncmVzc2lvbiJd%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="276" height="830"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  SEC-01: open_basedir still does not carry over to FPM
&lt;/h2&gt;

&lt;p&gt;When you create a site with the default settings (&lt;code&gt;mod_php&lt;/code&gt;), &lt;a href="https://www.php.net/manual/en/ini.core.php#ini.open-basedir" rel="noopener noreferrer"&gt;&lt;code&gt;open_basedir&lt;/code&gt;&lt;/a&gt; is set correctly: limited to &lt;code&gt;/var/www/&amp;lt;user&amp;gt;/data:.&lt;/code&gt;, and a test script cannot read &lt;code&gt;/etc/passwd&lt;/code&gt;. Switch the same site to the PHP-FPM mode the panel supports, call the same script, and the restriction is gone. The panel keeps showing &lt;code&gt;basedir=on&lt;/code&gt;; the generated pool file has no &lt;code&gt;open_basedir&lt;/code&gt;, the site-level include file is zero bytes, and the user-level one only carries &lt;code&gt;upload_tmp_dir&lt;/code&gt; and &lt;code&gt;session.save_path&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The vendor reproduced this in their first reply, agreed that PHP modes should behave consistently, and put a fix for a future release on their backlog. But they disputed the severity: in the multi-tenant model the real boundary is at the operating-system level, each pool runs under its own system user, so one tenant's script cannot read another tenant's files even without &lt;code&gt;open_basedir&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That objection was measurable, and I had not measured it in the first report. &lt;code&gt;/etc/passwd&lt;/code&gt; is world-readable anyway. In round two I tried a direct cross-tenant read, and on 1 October I repeated it on 6.153. User A's FPM site tries to read a file on user B's site:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# 6.153.0, a.local (PHP-FPM, isp-php84)
SAPI=fpm-fcgi user=usera uid=1010 open_basedir=[]
/etc/passwd                                 =&amp;gt; OK 1796 bytes
/var/www/userb/data/www/b.local/secret.txt  =&amp;gt; FAIL (Permission denied)
/usr/local/mgr5/etc/ispmgr.conf             =&amp;gt; FAIL (Permission denied)
/etc/shadow                                 =&amp;gt; FAIL (Permission denied)

# 6.153.0, a2.local (mod_php), same user
SAPI=apache2handler user=usera uid=1010 open_basedir=[/var/www/usera/data:.]
/etc/passwd                                 =&amp;gt; FAIL (Operation not permitted)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result goes the vendor's way. The mechanism is worth learning, too. B's &lt;code&gt;data&lt;/code&gt; directory is &lt;code&gt;drwx---r-x userb:mgrsecure&lt;/code&gt;: "others" can read and traverse it, which is what nginx and Apache need to serve static files. But the &lt;code&gt;mgrsecure&lt;/code&gt; group has no rights at all, and the panel adds every tenant to that group. On Linux, &lt;a href="https://git.kernel.org/pub/scm/docs/man-pages/man-pages.git/tree/man/man7/path_resolution.7" rel="noopener noreferrer"&gt;permission checking&lt;/a&gt; picks the first matching class out of owner, group and others and applies only that class's bits; once the group matches, "others" is never consulted. Tenants cannot enter each other's directories; service users can. A clever layout; I had missed it in the first report.&lt;/p&gt;

&lt;p&gt;That is why I lowered SEC-01 from &lt;strong&gt;High to Medium&lt;/strong&gt; in round two, and that is where it stays. A defence layer that the panel shows as enabled silently disappearing on a supported mode switch is a real defect; world-readable system files are exposed. But the tenant boundary held in every measurement. The 6.153 changelog has no line about this, and the behaviour has not changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I retracted six of my findings
&lt;/h2&gt;

&lt;p&gt;In round three I also re-measured the findings other than SEC-02, or opened and read their documentation. I did not re-measure LIF-01 (uninstall) or LIF-02 (resources outside the panel); for API-02 and PLT-02 I only checked the documentation.&lt;/p&gt;

&lt;p&gt;I also need to say one thing plainly. On 23 September the vendor wrote that they had sent feedback on the whole file, and according to my working notes that feedback included design explanations for ROB-04, FUN-04, FUN-05 and API-02. I verified everything below by measuring it myself or opening the documentation, not by trusting those explanations. But for some of these retractions, it was the vendor who pointed me in the right direction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ROB-03 (monitoring).&lt;/strong&gt; In the first report I wrote "the panel monitors services but does not bring a failed one back". In round two I stopped nginx and waited: the panel's &lt;code&gt;srvmon&lt;/code&gt; monitoring job brought it back with &lt;code&gt;systemctl restart nginx.service&lt;/code&gt; after 9 minutes 47 seconds. I had waited only 20 seconds before passing judgement. On 1 October I repeated it on 6.153: 11 minutes 58 seconds. In 6.153's crontab the &lt;code&gt;srvmon&lt;/code&gt; job runs every 15 minutes (&lt;code&gt;*/15&lt;/code&gt;), and in both measurements the restart landed on the hour. So on this install a failed service can stay down for close to a quarter of an hour. In round two the panel's service list showed nginx as &lt;code&gt;started&lt;/code&gt; throughout; on 6.153 it correctly showed &lt;code&gt;stopped&lt;/code&gt; from start to finish, which I note as an improvement. One more confession: in this round's first measurement nginx came back after one minute. The cause was not monitoring; another test I was running at the same time had made the panel rewrite a site configuration. I threw that measurement away and repeated the test without touching the machine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ROB-04 (panel metadata not in the backup).&lt;/strong&gt; Wrong. On 1 October I ran &lt;code&gt;backup2 -m full&lt;/code&gt; and opened the &lt;code&gt;root&lt;/code&gt; account's archive. It contains &lt;code&gt;/usr/local/mgr5/etc/ispmgr.db&lt;/code&gt;, &lt;code&gt;core_core.db&lt;/code&gt;, their &lt;code&gt;-wal&lt;/code&gt; and &lt;code&gt;-shm&lt;/code&gt; files and &lt;code&gt;ispmgr.conf&lt;/code&gt;; 2,203 files in total. In the first round I had looked at the user archives and never opened root's.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FUN-02 (backups cannot be managed via the API).&lt;/strong&gt; Wrong, but in an interesting way. In the first round I called the &lt;code&gt;backup.plan&lt;/code&gt; and &lt;code&gt;backup.storages&lt;/code&gt; functions, got &lt;code&gt;module is missing&lt;/code&gt;, and wrote "on Lite, backup management is effectively CLI-only". In round three a banner inside one of the responses pointed at a &lt;code&gt;backup2.schedule&lt;/code&gt; function. I tried it: &lt;code&gt;backup2.settings&lt;/code&gt;, &lt;code&gt;backup2.schedule&lt;/code&gt; and &lt;code&gt;backup2.list&lt;/code&gt; work. The &lt;a href="https://www.ispmanager.com/docs/ispmanager/ispmanager-api" rel="noopener noreferrer"&gt;API reference&lt;/a&gt; lists both families; the &lt;code&gt;backup.plan&lt;/code&gt; family is absent on Lite, the &lt;code&gt;backup2&lt;/code&gt; family is present. What remains is a documentation clarity note: both families are documented side by side, and the docs don't say which works in which edition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FUN-03 (no WAF on Lite).&lt;/strong&gt; Wrong. The WAF is an optional component installed from the &lt;a href="https://www.ispmanager.com/docs/ispmanager/waf" rel="noopener noreferrer"&gt;Software configuration&lt;/a&gt; screen; on 6.153 the form shows a &lt;code&gt;package_nginx_modsecurity=off&lt;/code&gt; option. I had seen that it was not installed and concluded that it did not exist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FUN-04 (no supervision for Node.js apps).&lt;/strong&gt; Wrong. I had looked for a systemd unit, didn't find one, and wrote "no mechanism that restarts the app after a crash was visible". The panel manages Node.js and Python applications with PM2; &lt;code&gt;/usr/lib/ispnodejs/bin/pm2&lt;/code&gt; is installed, the &lt;a href="https://www.ispmanager.com/docs/ispmanager/managing-a-nodejs-project" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; gives that same path for diagnostics, and the 6.152 changelog contains a fix related to the PM2 interpreter. I had been looking for the wrong tool. I did not separately measure a restart after a crash; this retraction rests on the installed tool and the documentation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FUN-05 (the system user cannot log in over FTP).&lt;/strong&gt; Wrong. The &lt;a href="https://www.ispmanager.com/docs/ispmanager/ftp-server-configuration" rel="noopener noreferrer"&gt;FTP server documentation&lt;/a&gt; describes virtual FTP users as the design itself: FTP access without creating real system accounts. I had also noted "the docs say &lt;code&gt;password&lt;/code&gt;, the API wants &lt;code&gt;passwd&lt;/code&gt;"; the API reference says &lt;code&gt;passwd&lt;/code&gt; too, so that note was wrong as well.&lt;/p&gt;

&lt;p&gt;In three places I corrected a finding without retracting it. In &lt;strong&gt;PLT-01&lt;/strong&gt; (no ARM packages) the behaviour is unchanged; the aarch64 path in the 6.153 repository also returns 404. But the &lt;a href="https://www.ispmanager.com/docs/ispmanager/system-requirements" rel="noopener noreferrer"&gt;system requirements&lt;/a&gt; page says "use the server version of an operating system with the x64 architecture". I had written up a documented limitation as a Medium defect; I downgraded it to an informational note. In &lt;strong&gt;FUN-01&lt;/strong&gt; (Python does not work) I had given "Phusion Passenger is not installed" as the cause; in ispmanager 6, Python &lt;a href="https://www.ispmanager.com/docs/ispmanager/working-with-python" rel="noopener noreferrer"&gt;runs under PM2&lt;/a&gt;, and the Python component is disabled in a default install. The remaining defect lies elsewhere, below. In &lt;strong&gt;SEC-03&lt;/strong&gt; I had written that AlmaLinux 9's default password scheme is "yescrypt"; the container's &lt;code&gt;/etc/login.defs&lt;/code&gt; says &lt;code&gt;ENCRYPT_METHOD SHA512&lt;/code&gt;. The finding stands; the comparison was wrong.&lt;/p&gt;

&lt;p&gt;There is no point hiding any of this. If I value the vendor revising its initial assessment of SEC-02, I have to apply the same standard to myself. The common denominator is clear, too: everywhere I was wrong, I had either not waited long enough or drawn a conclusion without opening the documentation. In the first report my numbers were confident; my evidence was not as strong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is still open on 6.153
&lt;/h2&gt;

&lt;p&gt;I reproduced every remaining medium-severity finding on 6.153.0 on 1 October. None of them has changed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEC-03.&lt;/strong&gt; Users created through the panel are written to &lt;code&gt;/etc/shadow&lt;/code&gt; with the &lt;code&gt;$1$&lt;/code&gt; prefix, that is, md5crypt. The &lt;a href="https://github.com/besser82/libxcrypt/blob/develop/doc/crypt.5" rel="noopener noreferrer"&gt;libxcrypt documentation&lt;/a&gt; is explicit about this scheme: MD5 is so cheap on modern hardware that it should not be used for new hashes, and its processing cost is not adjustable. On the same machine &lt;code&gt;login.defs&lt;/code&gt; says SHA512; the panel bypasses the system default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEC-04.&lt;/strong&gt; MariaDB listens on &lt;code&gt;*:3306&lt;/code&gt;, and the single multiport rule in the default firewall opens &lt;code&gt;20:22,25,80,443,110,143,465,587,993,995,53,3306,5432,1500&lt;/code&gt; to every address. The two anonymous accounts the install leaves behind (&lt;code&gt;''@localhost&lt;/code&gt; and &lt;code&gt;''@&amp;lt;hostname&amp;gt;&lt;/code&gt;) are still in the table. Since every named account is &lt;code&gt;@localhost&lt;/code&gt;, remote login should not work at the moment. I did not try it. But the day the first &lt;code&gt;@%&lt;/code&gt; account is created, nothing is left at the door.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEC-05.&lt;/strong&gt; The panel's root login calls PAM with the &lt;a href="https://github.com/linux-pam/linux-pam/blob/master/doc/man/pam_start.3.xml" rel="noopener noreferrer"&gt;&lt;code&gt;sshd&lt;/code&gt; service name&lt;/a&gt;. On 1 October I temporarily removed &lt;code&gt;/etc/pam.d/sshd&lt;/code&gt;, and an API call with the correct password returned &lt;code&gt;ERROR auth(badpassword)&lt;/code&gt;; once the file was back, the same call worked. Every hardening change to sshd's PAM stack, an MFA module or a deny rule for instance, also affects whether the panel can be reached.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SEC-06.&lt;/strong&gt; This round I measured it with a control group. You send &lt;code&gt;secure=on&lt;/code&gt; for a site owned by a user whose SSL limit is off, the API returns &lt;code&gt;OK&lt;/code&gt;, and when you read the site back it says &lt;code&gt;secure=off&lt;/code&gt; and nginx has no 443 line. Turn the limit on for the same user and send the request: because I passed the certificate name in the wrong format, I first got an explicit validation error (&lt;code&gt;ERROR value(ssl_cert)&lt;/code&gt;), so with the limit on, the panel does tell you what's wrong. With the correct name: &lt;code&gt;secure=on&lt;/code&gt;, port 443 is listening, and HTTPS returns 200. I turned the limit off and sent the same correct parameters again: &lt;code&gt;OK&lt;/code&gt; again, &lt;code&gt;secure=off&lt;/code&gt; again. The operator believes HTTPS is on.&lt;/p&gt;

&lt;p&gt;What is left of &lt;strong&gt;FUN-01&lt;/strong&gt; belongs to the same family. With the Python component not installed, sending &lt;code&gt;python=on&lt;/code&gt; for a site still returns &lt;code&gt;OK&lt;/code&gt;, and the site stays without Python. I lowered it to Low, because this time it is not a security setting; but the pattern is the same.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ROB-01.&lt;/strong&gt; The panel connects to MariaDB through &lt;code&gt;root@localhost&lt;/code&gt; for its own database operations. After a routine hardening step, &lt;code&gt;ALTER USER 'root'@'localhost' IDENTIFIED BY '...'&lt;/code&gt;, I tried to create a database: &lt;code&gt;ERROR db(connect): ... Access denied for user 'root'@'localhost'&lt;/code&gt;. There is no health indicator and no startup check. Reverting fixes it. There is nothing wrong with &lt;a href="https://mariadb.com/docs/server/reference/plugins/authentication-plugins/authentication-plugin-unix-socket" rel="noopener noreferrer"&gt;unix_socket&lt;/a&gt; authentication itself; the problem is that the panel never tells the operator about this dependency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ROB-02.&lt;/strong&gt; Here is the small story I promised. In all three rounds MariaDB failed during installation. A &lt;code&gt;.cache&lt;/code&gt; subdirectory appears in &lt;code&gt;/var/lib/mysql&lt;/code&gt;, and &lt;code&gt;mariadb-prepare-db-dir&lt;/code&gt; refuses to initialise a non-empty directory. This round I looked inside: &lt;code&gt;.cache/rosetta&lt;/code&gt;. The &lt;code&gt;mysql&lt;/code&gt; user's home directory is &lt;code&gt;/var/lib/mysql&lt;/code&gt;. Under emulation, running any process as a user creates &lt;code&gt;.cache/rosetta&lt;/code&gt; in that user's home directory; I reproduced this with a separate test user by running &lt;code&gt;runuser -u rtest -- /bin/true&lt;/code&gt;. So the trigger is my environment; you will not see this on a real x86 server.&lt;/p&gt;

&lt;p&gt;The product finding is not the trigger but what happens next. The 6.153 install log has &lt;code&gt;Error in POSTIN scriptlet in rpm package coremanager-pkg-mysql&lt;/code&gt;, three &lt;code&gt;Job for mariadb.service failed&lt;/code&gt; lines, and then seven &lt;code&gt;The server wasn't added. Retrying...&lt;/code&gt; lines, backing off from 1 to 64 seconds. The installation still ends with &lt;code&gt;EXIT=0&lt;/code&gt; and the message "Your newly installed ispmanager panel is available". What's left is a panel with no database server registered, and the operator finds out when they try to create their first database and get &lt;code&gt;notconfigured(nodbserver)&lt;/code&gt;. The retry loop in the &lt;code&gt;coremanager-pkg-mysql&lt;/code&gt; package's post-install script and a "pending registration" marker file show that the vendor has been looking at this class of problem; the 6.152.0 changelog also has a fix for the default database server not appearing in the list after installation. The same thing happened in the stable 6.152.2 container. Failure is still reported as success.&lt;/p&gt;

&lt;h2&gt;
  
  
  The full table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;ID&lt;/th&gt;
&lt;th&gt;Initial rating&lt;/th&gt;
&lt;th&gt;As of 1 October&lt;/th&gt;
&lt;th&gt;Note&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SEC-01&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Medium, open&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No &lt;code&gt;open_basedir&lt;/code&gt; under FPM; tenant boundary holds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SEC-02&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Fixed (6.153.0)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No password in the admin response; independently verified&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SEC-03&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium, open&lt;/td&gt;
&lt;td&gt;md5crypt (&lt;code&gt;$1$&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SEC-04&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium, open&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;*:3306&lt;/code&gt;, world-open rule, anonymous accounts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SEC-05&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium, open&lt;/td&gt;
&lt;td&gt;Panel login tied to the &lt;code&gt;sshd&lt;/code&gt; PAM profile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SEC-06&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium, open&lt;/td&gt;
&lt;td&gt;SSL request swallowed with &lt;code&gt;OK&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ROB-01&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium, open&lt;/td&gt;
&lt;td&gt;Panel silently breaks when the root credential changes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ROB-02&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium, open&lt;/td&gt;
&lt;td&gt;Install "succeeds" without a working DB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FUN-01&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;python=on&lt;/code&gt; → &lt;code&gt;OK&lt;/code&gt; without the Python component&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FUN-02&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;backup2&lt;/code&gt; API exists; docs mix two families&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PLT-01&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;x64 is a documented requirement&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ROB-03&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;Recovery exists: &lt;code&gt;srvmon&lt;/code&gt;, &lt;code&gt;*/15&lt;/code&gt; on 6.153 (measured 11 min 58 s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ROB-04&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Retracted&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The root backup includes the panel databases&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ROB-05&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Low, open&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;restore2&lt;/code&gt; date and file options&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FUN-03&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Retracted&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The WAF is an optional component&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FUN-04&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Retracted&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;PM2 supervision&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FUN-05&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Retracted&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Virtual FTP users are by design&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API-01&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;Docs say &lt;code&gt;elid&lt;/code&gt;; &lt;code&gt;name=&lt;/code&gt; is ignored without an error&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LIF-01&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Low (not re-measured)&lt;/td&gt;
&lt;td&gt;Leftovers after uninstall&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API-02&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;Password in the URL; a session-key alternative is documented&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LIF-02&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;Info (not re-measured)&lt;/td&gt;
&lt;td&gt;Resources created outside the panel don't appear in its lists&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PLT-02&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;The installer requires SELinux to be disabled (documented)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OBS-01&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;Info&lt;/td&gt;
&lt;td&gt;13 MB &lt;code&gt;ispmgr.log&lt;/code&gt; in ~16 min, including install and two test runs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In numbers: as of 1 October, 6.153.0 has 7 Medium, 3 Low and 8 informational items open. One finding has been fixed and four have been fully retracted. When I say I retracted the substance of six findings, I mean those four plus ROB-03 and FUN-02, whose claims collapsed and which were reduced to informational notes. PLT-01 also became an informational note, but I didn't retract it: the behaviour I measured was real, it was just documented. On 6.153 there is no open High-severity finding left; on stable 6.152.2, SEC-02 is still open.&lt;/p&gt;

&lt;p&gt;Two rows need more explanation. For API-01 I sent &lt;code&gt;name=nginx&lt;/code&gt; to &lt;code&gt;services.restart&lt;/code&gt;; the panel returned no error and did not touch nginx. It worked with &lt;code&gt;elid=nginx&lt;/code&gt;. The &lt;a href="https://www.ispmanager.com/docs/ispmanager/ispmanager-api" rel="noopener noreferrer"&gt;API reference&lt;/a&gt; already says &lt;code&gt;elid&lt;/code&gt;, so this is less a defect than a note about an API that silently swallows parameters it doesn't recognise. For API-02, carrying credentials in the URL via &lt;code&gt;authinfo=&lt;/code&gt; is protected in transit by TLS. The panel masks the value as &lt;code&gt;authinfo=*&lt;/code&gt; in its own access log, but the same is not guaranteed for the log of a proxy in between; however, the &lt;a href="https://www.ispmanager.com/docs/ispmanager/guide-to-ispmanager-software-api" rel="noopener noreferrer"&gt;API guide&lt;/a&gt; also describes a session ID valid for one hour, one-time key login and an IP whitelist for &lt;code&gt;authinfo&lt;/code&gt;. The choice is the operator's.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two themes
&lt;/h2&gt;

&lt;p&gt;Put the remaining findings side by side and most of them cluster in two places.&lt;/p&gt;

&lt;p&gt;The first is silent failure semantics. A limit rejection is swallowed (SEC-06, FUN-01). The installer "succeeds" with a broken database (ROB-02). The panel silently breaks when the root credential changes (ROB-01). A configured defence layer disappears on a supported mode switch while the UI still shows it as enabled (SEC-01). The product tells the truth, but not loudly enough. On a security-related setting, "OK, but not applied" is the worst answer there is.&lt;/p&gt;

&lt;p&gt;The second is a tight coupling to the host's identity infrastructure: the &lt;code&gt;sshd&lt;/code&gt; profile in PAM, the root credential in the database, and SELinux, which the installer &lt;a href="https://www.ispmanager.com/docs/ispmanager/ispmanager-installation-guide" rel="noopener noreferrer"&gt;requires to be disabled&lt;/a&gt;. The panel does not carry its own identity; it leans on the host's. Neither requires an architectural rewrite; both call for a clearer definition of the ownership boundary and of the failure contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you are installing ispmanager: a day-one list
&lt;/h2&gt;

&lt;p&gt;I am writing these based on the behaviour I measured on 6.153.0. They are worth re-checking as versions change.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Version.&lt;/strong&gt; If you are going to give administrators panel accounts, install 6.153.0 or later, or hold off on administrator accounts until you can. As of 2 October, 6.153 is in the beta channel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database port.&lt;/strong&gt; Check the listening address with &lt;code&gt;ss -ltnp | grep 3306&lt;/code&gt;. If you don't need remote access, remove 3306 and 5432 from the firewall's multiport rule or narrow the source, and pull the daemon back to localhost with &lt;a href="https://mariadb.com/docs/server/server-management/variables-and-modes/server-system-variables#bind_address" rel="noopener noreferrer"&gt;bind-address&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anonymous accounts.&lt;/strong&gt; Delete the rows with an empty user name from &lt;code&gt;SELECT user,host FROM mysql.user&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FPM sites.&lt;/strong&gt; On every site you switch to PHP-FPM, check &lt;code&gt;ini_get('open_basedir')&lt;/code&gt; with a small script. If it is empty, there is no restriction, whatever the panel shows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whether the install really finished.&lt;/strong&gt; After the installer finishes, verify that &lt;code&gt;mgrctl -m ispmgr db.server&lt;/code&gt; is not empty. The "success" message does not guarantee it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credential changes.&lt;/strong&gt; Before you change the MariaDB root credential or touch sshd's PAM stack, write down in your runbook that the panel depends on them, and try a panel operation after the change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Panel port.&lt;/strong&gt; The same multiport rule also opens 1500, the panel itself, to every address. If the panel should only be reachable from a management network, narrow the source.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The password after moving to 6.153.&lt;/strong&gt; If level-29 administrator accounts could reach the database server list before 6.153, they may have read the stored root password; change the MariaDB root password. The upgrade itself doesn't require this. Keep &lt;code&gt;unix_socket&lt;/code&gt; while doing it (&lt;code&gt;ALTER USER 'root'@'localhost' IDENTIFIED VIA unix_socket OR mysql_native_password USING PASSWORD('...')&lt;/code&gt;), then save the new password in the panel with &lt;code&gt;db.server.edit&lt;/code&gt;. I tried this sequence on 6.153 and the panel kept working; a plain &lt;code&gt;IDENTIFIED BY&lt;/code&gt; triggers ROB-01.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backups.&lt;/strong&gt; Open the &lt;code&gt;root&lt;/code&gt; archive of a full backup yourself once and see that it contains the panel databases. I wrote it up wrong without looking.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What the vendor did
&lt;/h2&gt;

&lt;p&gt;I sent the report on the night of 14 September. On 15 September the reply came: "we've shared it with the team, we'd love a full article". On 18 September the engineering team's written assessment arrived: they reproduced SEC-01 and put it on the backlog, and objected to SEC-02 on threat-model grounds. For the remaining findings the summary was: some make good sense and will be acted on, in a smaller number the intended design wasn't fully accounted for in the report, nothing is urgent, everything goes into the backlog. On 23 September they revised their assessment of SEC-02 and gave a fix date. On 1 October Aleksandra asked whether I had run the test, when the article would be out, and, with a smiley, whether there would be "at least one link".&lt;/p&gt;

&lt;p&gt;In round three I saw how right that "intended design wasn't fully accounted for" sentence was: FTP, the WAF, PM2 and the backup API fell squarely into that category. It made me think again that a good vendor response is not saying "you're right" to every finding. A good response reproduces, concedes where it agrees, and gives a testable justification where it doesn't. ispmanager did that. When I tested the testable justification, it turned out to be wrong in one place, and they accepted that. The answer to the link question is in this article too: I linked their documentation generously, because readers should be able to check my claims on the vendor's own pages.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who it is for, and what I would do
&lt;/h2&gt;

&lt;p&gt;ispmanager's natural territory is web hosting, reseller setups, and servers that are run by several operators or handed over to customers. The panel makes sense wherever someone who doesn't know every corner of Linux has to do routine work from a safe surface. Tenant isolation, template discipline and backup integrity are solid enough for that. For anyone on ARM infrastructure the door is closed, and documented as such; for those who manage configuration heavily outside the panel, LIF-02 is a nuisance. For anyone thinking of building their own PaaS layer instead of a classic hosting panel, &lt;a href="https://mustafaerbay.com.tr/en/blog/tutorials/coolify-ile-self-host-paas-kurulumunda-guncel-yaklasim/" rel="noopener noreferrer"&gt;what I wrote about Coolify&lt;/a&gt; describes a different trade-off.&lt;/p&gt;

&lt;p&gt;My default is the terminal, and that hasn't changed. But this test changed my view of where a panel creates value. In a setup where I am not the only operator, and want to hand part of the work to someone who doesn't go deep into the system, I would take ispmanager seriously. That is not the answer I would have given before I started. My conditions are clear too: 6.153 or later, the day-one list above, and checking &lt;code&gt;open_basedir&lt;/code&gt; on FPM sites myself until SEC-01 is closed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;I tested ispmanager Lite at its edges, not on the happy path. Template discipline, OS-level isolation, backup integrity, the TLS lifecycle, the audit trail and atomic reloads genuinely held up. Of the two High findings, one was lowered to Medium by measurement and the other was closed in 6.153, which I verified independently; the fix has not reached the stable channel yet. Seven medium-severity findings remain open on 6.153, and most of them are symptoms of the same habit: failing silently.&lt;/p&gt;

&lt;p&gt;Both sides changed their minds during this process. The vendor revised its initial threat-model assessment of SEC-02 after reproducing the privilege-boundary case, and decided to fix the behaviour. I lowered SEC-01's severity, deleted ROB-03's "no recovery" sentence and retracted the substance of six of my findings. Two days ago I wrote about how the patch that closed twenty CodeQL findings &lt;a href="https://mustafaerbay.com.tr/en/blog/career/yirmi-bulguyu-kapattim-yirmi-birincisi-benim-yamamdi/" rel="noopener noreferrer"&gt;turned out to be the source of the twenty-first&lt;/a&gt;. This week I learned the same lesson from a different direction. A control panel should make your work easier on a good day and tell you plainly what happened on a bad one. The same standard applies to a review.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Environment: ispmanager Lite 6.150.4 (core 5.445.2) and 6.153.0 (core 5.448.0), AlmaLinux 9.8, a &lt;code&gt;--privileged&lt;/code&gt; systemd container on Docker Desktop, x86_64 emulation on Apple Silicon. Round 1: 14 September, vendor test licence. Round 2 (21 September) and Round 3 (1 October): IP-bound trial licence, clean install; in Round 3, 6.152.2 (core 5.447.2) in a separate container for the stable-channel comparison. The three environment-induced conditions (ifcfg, the sshd PAM profile, MariaDB &lt;code&gt;.cache&lt;/code&gt;) were not counted in the severity totals. Not tested: Let's Encrypt end to end (there was no real domain), disk quotas (container filesystem), and behaviour on a real virtual machine or physical server. The raw command logs are kept.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://nginx.org/en/docs/http/ngx_http_core_module.html#disable_symlinks" rel="noopener noreferrer"&gt;nginx — disable_symlinks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://httpd.apache.org/docs/current/mpm.html" rel="noopener noreferrer"&gt;Apache HTTP Server — Multi-Processing Modules&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/php/doc-en/blob/master/appendices/ini.core.xml" rel="noopener noreferrer"&gt;PHP documentation — open_basedir (ini.core)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/php/doc-en/blob/master/install/fpm/configuration.xml" rel="noopener noreferrer"&gt;PHP documentation — PHP-FPM configuration, php_admin_value&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/linux-pam/linux-pam/blob/master/doc/man/pam_start.3.xml" rel="noopener noreferrer"&gt;Linux-PAM — pam_start(3) and the service name&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/besser82/libxcrypt/blob/develop/doc/crypt.5" rel="noopener noreferrer"&gt;libxcrypt — crypt(5), md5crypt and yescrypt&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://git.kernel.org/pub/scm/docs/man-pages/man-pages.git/tree/man/man7/path_resolution.7" rel="noopener noreferrer"&gt;Linux man-pages — path_resolution(7), permission checks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/MariaDB/server/blob/main/plugin/auth_socket/auth_socket.c" rel="noopener noreferrer"&gt;MariaDB — unix_socket authentication plugin (source)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ispmanager</category>
      <category>kontrolpaneli</category>
      <category>security</category>
      <category>phpfpm</category>
    </item>
    <item>
      <title>My Quarantine Ledger Has 1,945 Rows and Not One Address</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Fri, 02 Oct 2026 07:25:46 +0000</pubDate>
      <link>https://dev.to/merbayerp/my-quarantine-ledger-has-1945-rows-and-not-one-address-1a72</link>
      <guid>https://dev.to/merbayerp/my-quarantine-ledger-has-1945-rows-and-not-one-address-1a72</guid>
      <description>&lt;p&gt;The question was innocent enough. I wanted to know where a file sitting in my Downloads folder had come from. I couldn't remember its name, my browser history was long, and asking the file itself looked like the shorter path.&lt;/p&gt;

&lt;p&gt;macOS has a place for exactly this. Downloaded files get an extended attribute called &lt;code&gt;com.apple.quarantine&lt;/code&gt;, and LaunchServices keeps a SQLite ledger alongside it. I opened both. The ledger held &lt;strong&gt;1,945 records&lt;/strong&gt;, a tidy run from June 2025 to late last night. Then I looked at the address column: all 1,945 rows were &lt;strong&gt;empty&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That empty column pulled me into a two-hour measurement. Here is what came out of it: the quarantine mark is not what we think it is. It isn't a record saying "this file came from here"; it's a flag saying "this file needs the user's consent". And whether that flag survives into the next file is decided not by the operating system but by the code of whatever tool touches the file — I tried nine paths, five of them inherited the mark and four lost it silently.&lt;/p&gt;

&lt;p&gt;Every measurement below comes from one machine at one moment: macOS 26.6.2 (25G83), arm64, counted on the morning of 2 October 2026. Both the ledger and the folder are live, so tomorrow's numbers will differ — what matters isn't the numbers but the relationships between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the Mark: Four Fields, Zero Documentation
&lt;/h2&gt;

&lt;p&gt;Checking whether a file carries the mark takes one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;xattr &lt;span class="nt"&gt;-p&lt;/span&gt; com.apple.quarantine ~/Downloads/&amp;lt;file&amp;gt;
&lt;span class="c"&gt;# 0281;6a8435e5;Chrome;B71407DA-5626-4ACB-8201-7A62855420FA&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four semicolon-separated fields: a flag number, a hexadecimal timestamp, the name of the app that applied the mark, and a UUID. Apple documents this string's format nowhere — I looked, and I didn't find it. But the LaunchServices quarantine properties are documented and map onto the fields one by one. For &lt;code&gt;kLSQuarantineAgentNameKey&lt;/code&gt; the documentation says it is the app name of the quarantining agent, and adds that when the key is absent from the caller's dictionary, the agent name is set automatically to the current process name. That is why the third field reads &lt;code&gt;Chrome&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Decoding the timestamp takes a single line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"import datetime;print(datetime.datetime.fromtimestamp(int('6abf42eb',16)))"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So far, familiar territory. What's interesting is whose job the mark actually is. I assumed quarantining a download happened &lt;em&gt;automatically&lt;/em&gt;. It doesn't. The documentation for the &lt;code&gt;LSFileQuarantineEnabled&lt;/code&gt; key puts it in one line: it is a Boolean saying whether files the app creates get quarantined by default. The mark is something the downloading app volunteers to apply.&lt;/p&gt;

&lt;p&gt;Easy enough to test. I pulled the same file with &lt;code&gt;curl&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; astro-logo.svg https://astro.build/favicon.svg
xattr &lt;span class="nt"&gt;-l&lt;/span&gt; astro-logo.svg
&lt;span class="c"&gt;# com.apple.provenance:&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No quarantine. Only &lt;code&gt;com.apple.provenance&lt;/code&gt;, which I'll come back to. Chrome would have left a mark; &lt;code&gt;curl&lt;/code&gt; doesn't, because &lt;code&gt;curl&lt;/code&gt; isn't a member of that club. That was the first lesson: &lt;strong&gt;the absence of a mark doesn't tell you the file is safe, it tells you the tool that fetched it doesn't apply marks.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Volunteering isn't the only door, though. The system carries an exception list of its own that can override an app's Info.plist:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;plutil &lt;span class="nt"&gt;-p&lt;/span&gt; /System/Library/CoreServices/CoreTypes.bundle/Contents/Resources/Exceptions.plist &lt;span class="se"&gt;\&lt;/span&gt;
  | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; LSFileQuarantineEnabled
&lt;span class="c"&gt;# 18&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On this machine, eighteen bundle identifiers are forced to apply marks by the system even without carrying the key themselves — Firefox, Thunderbird, Transmission, Opera, Azureus and a few browsers that no longer exist. The file sits inside root-owned, signed system storage; a user can't edit the list. So you can't tell whether an app applies marks by reading its Info.plist alone. That half the list is decade-old dead software is its own detail: quarantine exceptions have piled up like a museum.&lt;/p&gt;

&lt;p&gt;I scanned my own Downloads folder. Including subdirectories, 217 of 306 files carry a mark and &lt;strong&gt;89 do not&lt;/strong&gt;. The distribution of agent names shows who volunteers: Chrome 77, WhatsApp 45, Claude 42, Outlook 22, Preview 7, Excel 6, Word 2 — plus 16 files where the agent field is empty.&lt;/p&gt;

&lt;h2&gt;
  
  
  A 1,945-Row Ledger With Zero Addresses
&lt;/h2&gt;

&lt;p&gt;The mark sits on the file, but there's a central ledger too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;sqlite3 ~/Library/Preferences/com.apple.LaunchServices.QuarantineEventsV2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"SELECT count(*) FROM LSQuarantineEvent;"&lt;/span&gt;
&lt;span class="c"&gt;# 1945&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This 425 KB file has stored every quarantine event from 2025-06-26 11:50:06 to 2026-10-01 22:51:21. Broken down by agent, it's a more honest summary than browser history: Chrome 1031, Cyberduck 873, ChatGPT 21, Opera 9, Homebrew Cask 9, Atlas 2.&lt;/p&gt;

&lt;p&gt;The schema cheered me up, because the columns I wanted were right there: &lt;code&gt;LSQuarantineDataURLString&lt;/code&gt; and &lt;code&gt;LSQuarantineOriginURLString&lt;/code&gt;. Their documentation is clear too — &lt;code&gt;kLSQuarantineDataURLKey&lt;/code&gt; is the actual URL of the quarantined item, and &lt;code&gt;kLSQuarantineOriginURLKey&lt;/code&gt; is the URL of the resource originally hosting it; for web downloads the page on which the user started the download, for attachments the email message or calendar event the item was attached to.&lt;/p&gt;

&lt;p&gt;Then I counted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;sqlite3 ~/Library/Preferences/com.apple.LaunchServices.QuarantineEventsV2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"SELECT sum(LSQuarantineDataURLString IS NOT NULL),
          sum(LSQuarantineOriginURLString IS NOT NULL),
          count(*) FROM LSQuarantineEvent;"&lt;/span&gt;
&lt;span class="c"&gt;# 0|0|1945&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Zero. Zero. One thousand nine hundred forty-five. The columns exist, the schema is unchanged, the documentation still talks about URLs — and on this machine not one of fifteen months' worth of records stores an address. The file side has a similar gap: of the 277 top-level files, only 86 carry the &lt;code&gt;com.apple.metadata:kMDItemWhereFroms&lt;/code&gt; attribute, so 69% of them have no "where from" information at all. (The mark ratio was counted over 306 files including subdirectories; I kept these two attribute scans to the top level, which is why the denominators differ.)&lt;/p&gt;

&lt;p&gt;Putting the ledger next to the files opened a second crack. The agents marking files on disk include WhatsApp, Claude, Outlook, Preview, Excel and Word; &lt;strong&gt;none of those six appears in the ledger at all&lt;/strong&gt;. The reverse holds too: Cyberduck wrote 873 events into the ledger and has not a single marked file in Downloads. Chrome is the only name on both lists. So writing an xattr onto a file and writing a row into the ledger are not the same job; they happen on separate paths, done by separate implementations. Seeing a mark and concluding "then there must be a ledger row for it" would be wrong for six applications on this machine.&lt;/p&gt;

&lt;p&gt;I don't know why, and I'm not going to invent a reason. What I do know is what it means in practice: &lt;strong&gt;this ledger is not a forensic record.&lt;/strong&gt; It tells you which app applied how many marks and when; it does not tell you that a given file came from a given address. If you're investigating a file's origin and leaning on this table, count the column first and confirm it isn't empty.&lt;/p&gt;

&lt;p&gt;One note on opening it: don't write to the live database, copy it first and work on the copy. The timestamp is kept in the Core Data epoch, so making it readable needs &lt;code&gt;datetime(LSQuarantineTimeStamp + 978307200, 'unixepoch', 'localtime')&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  I Tried to Decode the Flag Field, Then Stopped
&lt;/h2&gt;

&lt;p&gt;That first number is very tempting. I had a ready-made sample, so I counted: &lt;code&gt;0281&lt;/code&gt; on 71 files, &lt;code&gt;0087&lt;/code&gt; on 42, &lt;code&gt;0083&lt;/code&gt; on 39, &lt;code&gt;0081&lt;/code&gt; on 30, &lt;code&gt;0283&lt;/code&gt; on 18, &lt;code&gt;0082&lt;/code&gt; on 13, &lt;code&gt;0287&lt;/code&gt; on 2, &lt;code&gt;03c1&lt;/code&gt; on one, &lt;code&gt;0086&lt;/code&gt; on one.&lt;/p&gt;

&lt;p&gt;In my own tests I saw a pattern. When I hand-wrote &lt;code&gt;0083&lt;/code&gt; onto an archive and extracted it, the files that came out read &lt;code&gt;0283&lt;/code&gt; every single time — the &lt;code&gt;0x0200&lt;/code&gt; bit added, the agent field cleared, the UUID preserved. (The timestamp sometimes matched the source's and sometimes carried the current moment; I couldn't pin down which is chosen when, so I'm not writing a rule for it.) "Fine," I said, "&lt;code&gt;0x0200&lt;/code&gt; marks an inherited mark." It was a nice hypothesis. Then I tested it in the field: of the 92 files carrying the &lt;code&gt;0x0200&lt;/code&gt; bit, &lt;strong&gt;76 have a non-empty agent name&lt;/strong&gt;. If an inherited mark is supposed to clear the agent, those 76 files refute the idea. The other direction held perfectly, though: all 16 files with an empty agent field carry &lt;code&gt;0x0200&lt;/code&gt;. A one-way relationship isn't enough to write a rule on, but it's worth noting.&lt;/p&gt;

&lt;p&gt;So I stopped there. Apple doesn't document these bits, and I'm not going to write up a pattern derived from three measurements as if it were a rule. The practical takeaway still holds: &lt;strong&gt;don't build policy on the flag field.&lt;/strong&gt; Whether a file is marked is reliable information; trying to read the mark's subtext means reverse-engineering an undocumented binary format that can change quietly in the next release.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who Carries the Mark, and Who Drops It
&lt;/h2&gt;

&lt;p&gt;This was my real question. When you extract a marked archive, do the files inside inherit the mark?&lt;/p&gt;

&lt;p&gt;I set the experiment up cleanly: packed a small two-file directory as zip, tar.gz and uncompressed tar, applied the same mark to all three by hand, then opened or copied them nine different ways.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;Q&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"0083;&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%x'&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; +%s&lt;span class="si"&gt;))&lt;/span&gt;&lt;span class="s2"&gt;;SafariTest;&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;uuidgen&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
xattr &lt;span class="nt"&gt;-w&lt;/span&gt; com.apple.quarantine &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$Q&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; package.zip
unzip &lt;span class="nt"&gt;-q&lt;/span&gt; package.zip &lt;span class="nt"&gt;-d&lt;/span&gt; A
xattr &lt;span class="nt"&gt;-p&lt;/span&gt; com.apple.quarantine A/package/a.txt
&lt;span class="c"&gt;# 0283;6abf42eb;;DCC5A817-FC07-440B-AF95-0BDEEB1405C7&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The results came out sharper than I expected:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;th&gt;Mark inherited?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;unzip -q package.zip&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Yes (&lt;code&gt;0283&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ditto -x -k package.zip B&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Yes (&lt;code&gt;0283&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;tar -xf package.tar&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Yes (&lt;code&gt;0283&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;tar -xf package.zip&lt;/code&gt; (bsdtar reads zip)&lt;/td&gt;
&lt;td&gt;Yes (&lt;code&gt;0283&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cp package.zip copy.zip&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Yes (&lt;code&gt;0283&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;`gzip -dc package.tgz \&lt;/td&gt;
&lt;td&gt;tar -xf -`&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cat package.zip &amp;gt; new.zip&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;dd if=package.zip of=new.bin&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;rsync package.zip new.bin&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I tried Python too, in three different orders: open the input before creating the output, do the reverse, create a second file in the same process. The mark survived none of them. I repeated the plain read-then-write case separately with Apple's &lt;code&gt;/usr/bin/python3&lt;/code&gt; and with the python.org 3.13 build — both produced an unmarked file.&lt;/p&gt;

&lt;p&gt;That table settles one thing: &lt;strong&gt;inheritance is not an automatic behavior of the file system or the kernel.&lt;/strong&gt; If it were, &lt;code&gt;cat&lt;/code&gt; and &lt;code&gt;cp&lt;/code&gt; would behave the same way — both read the same file given by name, and one carries the mark while the other doesn't. What separates them is the tools' own code: &lt;code&gt;unzip&lt;/code&gt;, &lt;code&gt;ditto&lt;/code&gt;, &lt;code&gt;bsdtar&lt;/code&gt; and &lt;code&gt;cp&lt;/code&gt; implement inheritance, while &lt;code&gt;cat&lt;/code&gt;, &lt;code&gt;dd&lt;/code&gt;, &lt;code&gt;rsync&lt;/code&gt; and Python just move bytes.&lt;/p&gt;

&lt;p&gt;Two conditions have to hold at once, and &lt;code&gt;bsdtar&lt;/code&gt; demonstrates that by itself: the same tool appears on both sides of the table. &lt;code&gt;tar -xf package.tar&lt;/code&gt; carries the mark and &lt;code&gt;gzip -dc package.tgz | tar -xf -&lt;/code&gt; doesn't — because in the second case the archive arrives as a stream, so there is no source file and therefore no attribute to read. The tool has to know about inheritance &lt;strong&gt;and&lt;/strong&gt; have a named source to read the mark from. (The &lt;code&gt;rsync&lt;/code&gt; row was measured with openrsync, this machine's default, with no extra flags; results may differ with flags that carry xattrs.)&lt;/p&gt;

&lt;p&gt;One detail stands out: &lt;code&gt;cp&lt;/code&gt; does not copy the mark &lt;em&gt;verbatim&lt;/em&gt;. The source reads &lt;code&gt;0083;…;SafariTest;UUID&lt;/code&gt; while the copy reads &lt;code&gt;0283;…;;UUID&lt;/code&gt;. So the attribute isn't being duplicated, it's being rewritten by the system. The same thing happens to files coming out of archives.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVtDaHJvbWUgwrcgV2hhdHNBcHAgwrcgT3V0bG9vazxici8-TFNGaWxlUXVhcmFudGluZUVuYWJsZWQgb25dIC0tPnxhcHBsaWVzIG1hcmt8IEJbZG93bmxvYWRlZCBmaWxlPGJyLz5jb20uYXBwbGUucXVhcmFudGluZV0KICBDW2N1cmwgwrcgd2dldCDCtyBzY3BdIC0tPnxubyBtYXJrfCBEW3VubWFya2VkIGZpbGVdCiAgQiAtLT4gRXt3aGljaCB0b29sIHRvdWNoZWQgaXR9CiAgRSAtLT58dW56aXAgwrcgZGl0dG8gwrcgYnNkdGFyIMK3IGNwfCBGW21hcmsgaW5oZXJpdGVkPGJyLz5mbGFncyByZXdyaXR0ZW5dCiAgRSAtLT58Y2F0IMK3IGRkIMK3IHJzeW5jIMK3IHB5dGhvbjxici8-Z3ppcCBwaXBlIHRhcnwgR1ttYXJrIGRyb3BwZWRdCiAgRiAtLT4gSHtpcyBpdCBhIE1hY2gtTyBiaW5hcnl9CiAgRyAtLT4gSVtubyBjaGVjazxici8-ZXhlYyBhbGxvd2VkXQogIEggLS0-fHllcywgdW5hcHByb3ZlZHwgSltTSUdLSUxMIGF0IGV4ZWNdCiAgSCAtLT58c2hlbGwgc2NyaXB0fCBJ%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVtDaHJvbWUgwrcgV2hhdHNBcHAgwrcgT3V0bG9vazxici8-TFNGaWxlUXVhcmFudGluZUVuYWJsZWQgb25dIC0tPnxhcHBsaWVzIG1hcmt8IEJbZG93bmxvYWRlZCBmaWxlPGJyLz5jb20uYXBwbGUucXVhcmFudGluZV0KICBDW2N1cmwgwrcgd2dldCDCtyBzY3BdIC0tPnxubyBtYXJrfCBEW3VubWFya2VkIGZpbGVdCiAgQiAtLT4gRXt3aGljaCB0b29sIHRvdWNoZWQgaXR9CiAgRSAtLT58dW56aXAgwrcgZGl0dG8gwrcgYnNkdGFyIMK3IGNwfCBGW21hcmsgaW5oZXJpdGVkPGJyLz5mbGFncyByZXdyaXR0ZW5dCiAgRSAtLT58Y2F0IMK3IGRkIMK3IHJzeW5jIMK3IHB5dGhvbjxici8-Z3ppcCBwaXBlIHRhcnwgR1ttYXJrIGRyb3BwZWRdCiAgRiAtLT4gSHtpcyBpdCBhIE1hY2gtTyBiaW5hcnl9CiAgRyAtLT4gSVtubyBjaGVjazxici8-ZXhlYyBhbGxvd2VkXQogIEggLS0-fHllcywgdW5hcHByb3ZlZHwgSltTSUdLSUxMIGF0IGV4ZWNdCiAgSCAtLT58c2hlbGwgc2NyaXB0fCBJ%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="628" height="1098"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Apple's 2026 Patch List Includes These Four Tools
&lt;/h2&gt;

&lt;p&gt;After building the table I went to Apple's security notes, mostly to check whether my measurement was still current. The answer both supported the table and showed its limits.&lt;/p&gt;

&lt;p&gt;The macOS Tahoe 26.5 notes (11 May 2026) carry three entries: CVE-2026-28849 in the BOM component, CVE-2026-28900 in libarchive, CVE-2026-28914 in &lt;code&gt;zip&lt;/code&gt;. All three describe the same impact: a maliciously crafted ZIP archive may bypass Gatekeeper checks. The descriptions differ, though — the line about a file quarantine bypass being addressed with additional checks appears in only one of the three, the libarchive entry. The same release closes a similar bypass via disk images in the Kernel (CVE-2026-28954), and that entry carries the same description.&lt;/p&gt;

&lt;p&gt;The 26.7 notes (14 September 2026) add CVE-2026-65399 in the &lt;code&gt;copyfile&lt;/code&gt; component: an archive may be able to bypass Gatekeeper. The same release also closes Gatekeeper or sandbox bypasses in CoreServices, the Kernel, and a component named &lt;code&gt;quarantine&lt;/code&gt; outright.&lt;/p&gt;

&lt;p&gt;Now lay the two lists on top of each other. The four tools that carried the mark in my measurement: &lt;code&gt;ditto&lt;/code&gt; (BOM), &lt;code&gt;bsdtar&lt;/code&gt; (libarchive), &lt;code&gt;unzip&lt;/code&gt; (zip), &lt;code&gt;cp&lt;/code&gt; (copyfile). All four are among the components Apple patched during 2026. That's no coincidence — because inheritance is implemented by the tools, the bugs live in the tools too; every archive format is a separate implementation and a separate bypass surface.&lt;/p&gt;

&lt;p&gt;But Apple's list is longer than my table, and that's information in itself: the same two releases patched the Kernel, CoreServices and &lt;code&gt;quarantine&lt;/code&gt; as well, and one bypass came through disk images. The quarantine bypass surface doesn't end with archive tools. The table above covers only its most hand-testable part.&lt;/p&gt;

&lt;p&gt;What this means in practice: this chain working correctly depends on your OS version. This machine runs 26.6.2, so the &lt;code&gt;copyfile&lt;/code&gt; fix isn't in this kernel — which also means the &lt;code&gt;cp&lt;/code&gt; row was measured against an unpatched &lt;code&gt;copyfile&lt;/code&gt;, and I haven't confirmed it behaves the same on a patched 26.7. If you have a workflow that relies on mark inheritance, staying current isn't a comfort, it's a precondition of the mechanism.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Mark Actually Blocks
&lt;/h2&gt;

&lt;p&gt;So far I've looked at how the mark travels. But what does being marked actually do?&lt;/p&gt;

&lt;p&gt;The Gatekeeper documentation says it plainly: approval is requested the first time downloaded software is opened, so that nobody is tricked into running executable code they took for a plain data file. The approval path is spelled out too: System Settings, Privacy &amp;amp; Security, "Open Anyway".&lt;/p&gt;

&lt;p&gt;I pictured a click-through dialog. Measuring it from the terminal, I met something much harsher. I compiled a three-line C program, left one copy clean and marked the other:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;clang &lt;span class="nt"&gt;-o&lt;/span&gt; t1 y.c &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cp &lt;/span&gt;t1 t2
./t1                      &lt;span class="c"&gt;# BINARY RAN, exit=0&lt;/span&gt;
xattr &lt;span class="nt"&gt;-w&lt;/span&gt; com.apple.quarantine &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$Q&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; t2
./t2                      &lt;span class="c"&gt;# no output, exit=137&lt;/span&gt;
xattr &lt;span class="nt"&gt;-d&lt;/span&gt; com.apple.quarantine t2
./t2                      &lt;span class="c"&gt;# BINARY RAN, exit=0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;137, that is 128+9: SIGKILL. The same binary, the same file system, one attribute of difference. With the mark the process is killed before it starts running; remove the mark and it comes back. No dialog, no warning, no error message — I went looking with &lt;code&gt;log show&lt;/code&gt; and found nothing: a three-minute window, &lt;code&gt;--info --debug&lt;/code&gt; on, predicates covering &lt;code&gt;process == "kernel"&lt;/code&gt;, the AMFI sender image and the &lt;code&gt;syspolicy&lt;/code&gt; subsystem. There may be lines a non-root &lt;code&gt;log show&lt;/code&gt; can't see; if someone can surface them, the method is right here and easy to refute. If anyone has spent hours wondering why a marked binary dies silently when called from a script, the answer is here.&lt;/p&gt;

&lt;p&gt;Two more tests, both instructive.&lt;/p&gt;

&lt;p&gt;I marked a copy of Apple-signed &lt;code&gt;/bin/echo&lt;/code&gt;: it died with &lt;code&gt;exit=137&lt;/code&gt; as well. So the rule isn't "unsigned binaries"; once something stops being a system binary, any unapproved executable is refused while the mark is on it.&lt;/p&gt;

&lt;p&gt;Then I put the same mark on a shell script:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'#!/bin/sh\necho "SCRIPT RAN"\n'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; x.sh &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;chmod&lt;/span&gt; +x x.sh
xattr &lt;span class="nt"&gt;-w&lt;/span&gt; com.apple.quarantine &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$Q&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; x.sh
./x.sh      &lt;span class="c"&gt;# SCRIPT RAN&lt;/span&gt;
sh x.sh     &lt;span class="c"&gt;# SCRIPT RAN&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both ran. A marked script meets no resistance at all, because the thing being &lt;code&gt;exec&lt;/code&gt;'d isn't the script, it's &lt;code&gt;/bin/sh&lt;/code&gt;. For a plain binary launched straight from the terminal, the gate stands in front of the Mach-O; interpreted code never walks through it. That isn't the same as saying the mark only concerns binaries — app bundles, pkg installers and disk images are separate paths, and the Kernel entry in 26.5 closes the disk image one precisely.&lt;/p&gt;

&lt;p&gt;I owe a warning here: the fix for this section is not "delete the mark". For a binary you compiled yourself, &lt;code&gt;xattr -d&lt;/code&gt; is reasonable; for one you distribute to other people, the right path is signing it, notarizing it, and attaching the ticket with &lt;code&gt;xcrun stapler staple&lt;/code&gt;. Apple's notarization documentation says it plainly: when the user first installs or runs the software, the presence of a ticket tells Gatekeeper that Apple notarized it. Make mark-deletion your distribution strategy and you've asked your users to switch Gatekeeper off.&lt;/p&gt;

&lt;p&gt;One more thing worth measuring: what does &lt;code&gt;spctl&lt;/code&gt; say? I assessed two copies of a notarized app, one marked and one clean. Both returned &lt;code&gt;accepted source=Notarized Developer ID&lt;/code&gt;, and &lt;code&gt;--context kLSDownloadedFileContext&lt;/code&gt; didn't change the answer either. So &lt;code&gt;spctl -a&lt;/code&gt; reports the state of the signature and the notarization, not the presence of the mark. You have to ask the two questions separately: &lt;code&gt;spctl&lt;/code&gt;/&lt;code&gt;codesign&lt;/code&gt; for the signature, &lt;code&gt;xattr&lt;/code&gt; for the mark.&lt;/p&gt;

&lt;p&gt;And the most open spot in the chain: &lt;code&gt;curl … | sh&lt;/code&gt; never creates a file. There is no object to mark, so there is no consent to ask for. Gatekeeper has nothing to say on that path — which is not a flaw, it's the scope of the design.&lt;/p&gt;

&lt;h2&gt;
  
  
  What You're Deleting When You Clean Up
&lt;/h2&gt;

&lt;p&gt;The internet is full of "if the mark is in your way, delete it" advice. Usually with &lt;code&gt;xattr -cr&lt;/code&gt;, meaning all attributes at once. So I counted what goes, running &lt;code&gt;xattr -c&lt;/code&gt; on a file with three attributes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;before: com.apple.metadata:kMDItemWhereFroms com.apple.provenance com.apple.quarantine
after : com.apple.provenance
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The quarantine went, and the "where from" information went with it. So in removing the mark you also remove the only origin record you had — the very record that's already missing on 69% of files. If your goal is to clear the mark, &lt;code&gt;xattr -d com.apple.quarantine&lt;/code&gt; takes exactly one attribute; &lt;code&gt;-c&lt;/code&gt; sweeps the floor.&lt;/p&gt;

&lt;p&gt;The interesting part is that &lt;code&gt;com.apple.provenance&lt;/code&gt; isn't removed. It was on the file I pulled with &lt;code&gt;curl&lt;/code&gt;, and it's on 193 of the 277 files in Downloads. I couldn't find any Apple documentation for this attribute, so I won't make claims about what it holds; all I'm noting is that it's more widespread than quarantine and that &lt;code&gt;xattr -c&lt;/code&gt; can't remove it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Half-Hour Check on Your Own Machine
&lt;/h2&gt;

&lt;p&gt;If you want to run the same measurements yourself, here's the order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Count how many files in your Downloads folder are marked. If the ratio isn't 100%, you have tools that don't apply marks either — the agent field tells you which.&lt;/li&gt;
&lt;li&gt;Count the rows in the ledger and check whether the address columns are populated. If you plan to investigate origins, know in advance whether you're leaning on an empty column.&lt;/li&gt;
&lt;li&gt;Check your version with &lt;code&gt;sw_vers&lt;/code&gt;. The quarantine bypasses in the 26.5 and 26.7 notes concern archive paths; if you rely on mark inheritance, being current is part of the mechanism.&lt;/li&gt;
&lt;li&gt;If your build or deployment pipeline extracts archives, choose the tool deliberately. A &lt;code&gt;tar&lt;/code&gt; fed through a pipe drops the mark — sometimes that's exactly what you want, but don't let it happen by accident.&lt;/li&gt;
&lt;li&gt;If you need to clean a mark, use &lt;code&gt;-d&lt;/code&gt;, not &lt;code&gt;-c&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;If the machine belongs to a fleet, some of these decisions may not be yours: the Gatekeeper documentation notes that users can override the policy and open any software unless a device management service restricts that. Don't assume the behavior you measure on your own machine will hold on a managed one.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And one habit to avoid: treating the absence of a mark as a trust signal. Eighty-nine of my files are unmarked, and not one of them earned that by passing a check; the tool that fetched them simply didn't apply one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Mark Is a Question, Not an Address
&lt;/h2&gt;

&lt;p&gt;I started this measurement wondering where a file came from, and I never got an answer to that question. What I got was more useful: I learned not what the quarantine mechanism is, but what it &lt;em&gt;isn't&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Quarantine is not an origin record. It's a note an app leaves saying "I brought this in, let the user approve it once". Leaving the note is voluntary, carrying it is the tools' business, the ledger keeps no addresses, the only thing it blocks is executable binaries, and when it's removed no trace remains.&lt;/p&gt;

&lt;p&gt;What does knowing that change? One general habit around security mechanisms: don't treat the presence of a flag as proof that a check was performed. The mark is a question mark, not a certificate. Black boxes spend their worst nights without telling anyone; this particular box at least confesses what it doesn't do, as long as you count.&lt;/p&gt;

&lt;p&gt;Incidentally, there's another example of this machine's own ledgers misleading me: &lt;a href="https://mustafaerbay.com.tr/en/blog/technology/launchd-21881-kez-denedi-ben-bir-kez-bakmadim/" rel="noopener noreferrer"&gt;launchd tried 21,881 times and I never looked once&lt;/a&gt;. The counters are right there, and nobody reads them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/guide/security/gatekeeper-and-runtime-protection-sec5599b66df/web" rel="noopener noreferrer"&gt;Apple Platform Security — Gatekeeper and runtime protection in macOS&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/en-us/102445" rel="noopener noreferrer"&gt;Apple Support — Safely open apps on your Mac&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.apple.com/documentation/bundleresources/information-property-list/lsfilequarantineenabled" rel="noopener noreferrer"&gt;Apple Developer — LSFileQuarantineEnabled&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.apple.com/documentation/coreservices/klsquarantineagentnamekey" rel="noopener noreferrer"&gt;Apple Developer — kLSQuarantineAgentNameKey&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.apple.com/documentation/coreservices/klsquarantineoriginurlkey" rel="noopener noreferrer"&gt;Apple Developer — kLSQuarantineOriginURLKey&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.apple.com/documentation/coreservices/klsquarantinedataurlkey" rel="noopener noreferrer"&gt;Apple Developer — kLSQuarantineDataURLKey&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developer.apple.com/documentation/security/notarizing-macos-software-before-distribution" rel="noopener noreferrer"&gt;Apple Developer — Notarizing macOS software before distribution&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/en-us/127115" rel="noopener noreferrer"&gt;Apple Support — About the security content of macOS Tahoe 26.5&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/en-us/149042" rel="noopener noreferrer"&gt;Apple Support — About the security content of macOS Tahoe 26.7&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>macos</category>
      <category>security</category>
      <category>gatekeeper</category>
      <category>karantina</category>
    </item>
    <item>
      <title>I Have 91 Backups. None of Them Are This Machine's.</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:21:41 +0000</pubDate>
      <link>https://dev.to/merbayerp/i-have-91-backups-none-of-them-are-this-machines-3hj0</link>
      <guid>https://dev.to/merbayerp/i-have-91-backups-none-of-them-are-this-machines-3hj0</guid>
      <description>&lt;p&gt;I was actually looking at something else. I opened &lt;code&gt;~/Backups&lt;/code&gt; to check whether a side&lt;br&gt;
project's nightly dumps were still landing, saw the listing, and relaxed — 91 files, dates&lt;br&gt;
marching along neatly. Then, in the same terminal, I typed a command I had not planned to&lt;br&gt;
type: &lt;code&gt;tmutil destinationinfo&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The answer was one line.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tmutil: No destinations configured.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What came after made me write this article twice. In the first version I said "I have no&lt;br&gt;
backup," and I believed I had proved it. Then I opened the places I had looked at too&lt;br&gt;
quickly: I did have a backup. It covered exactly my most irreplaceable files, in fact —&lt;br&gt;
and I had forgotten about it.&lt;/p&gt;

&lt;p&gt;So the thesis here is not "I neglected my backups." It's the more uncomfortable one:&lt;br&gt;
&lt;strong&gt;I didn't know what my backup covered.&lt;/strong&gt; I had a working copy, a dead copy, and a pile of&lt;br&gt;
things with no copy at all, and the map in my head got all three wrong. Every number below&lt;br&gt;
was measured on this machine on 1 October 2026.&lt;/p&gt;
&lt;h2&gt;
  
  
  The backup that works: 91 dumps, 760 megabytes, zero surprises
&lt;/h2&gt;

&lt;p&gt;Good news first. Inside &lt;code&gt;~/Library/LaunchAgents&lt;/code&gt; there is a job called&lt;br&gt;
&lt;code&gt;com.burncpu.backup-pull&lt;/code&gt; that wakes up every day at 10:00. What it does is modest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;rsync &lt;span class="nt"&gt;-az&lt;/span&gt; &lt;span class="nt"&gt;--timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;60 vps3:/opt/burncpu/backups/ &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$DEST&lt;/span&gt;&lt;span class="s2"&gt;/files/"&lt;/span&gt;
find &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$DEST&lt;/span&gt;&lt;span class="s2"&gt;/files"&lt;/span&gt; &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"burncpu-*.sql.gz"&lt;/span&gt; &lt;span class="nt"&gt;-mtime&lt;/span&gt; +90 &lt;span class="nt"&gt;-delete&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It mirrors the nightly dumps from the server to this Mac; it never deletes to match the&lt;br&gt;
source, it only prunes anything older than 90 days locally. The result: &lt;strong&gt;91 &lt;code&gt;.sql.gz&lt;/code&gt;&lt;br&gt;
files&lt;/strong&gt;, &lt;strong&gt;760 MB&lt;/strong&gt; in total. The oldest is from 3 July, the newest is&lt;br&gt;
&lt;code&gt;burncpu-2026-10-01T00-00-02Z.sql.gz&lt;/code&gt;. Its log is dry and clear:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[2026-09-30 10:00:05] pull basliyor
[2026-09-30 10:00:22] ok — 91 dosya, 767M
[2026-10-01 10:00:05] pull basliyor
[2026-10-01 10:00:15] ok — 91 dosya, 761M
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I misread this job's health at first glance. The &lt;code&gt;launchd.err.log&lt;/code&gt; next to it was dated 11&lt;br&gt;
June, and I thought I had found a backup that had been stale for four months. But the file&lt;br&gt;
is 0 bytes: it was born in June when the job was installed and has found nothing to write&lt;br&gt;
since. The staleness of the error log was proof of health, not of failure.&lt;/p&gt;

&lt;p&gt;That was a preview of the rest of this article: the same silence carries both good and bad&lt;br&gt;
news, and I read it wrong both times on the first pass.&lt;/p&gt;

&lt;p&gt;One detail: those 91 dumps are not a &lt;strong&gt;backup&lt;/strong&gt; for this machine. Their original lives on&lt;br&gt;
VPS3; the copy here is the second instance. If the Mac dies, that data lives on.&lt;/p&gt;
&lt;h2&gt;
  
  
  The machine's own backup: nothing at the system level
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;tmutil destinationinfo&lt;/code&gt; can't find a configured destination. &lt;code&gt;tmutil listlocalsnapshots /&lt;/code&gt;&lt;br&gt;
prints an empty heading and falls silent. &lt;code&gt;diskutil list&lt;/code&gt; shows a single physical disk:&lt;br&gt;
&lt;code&gt;/dev/disk0&lt;/code&gt;, carrying a 460 GiB (494.4 GB) main APFS container; under &lt;code&gt;/Volumes&lt;/code&gt; there is&lt;br&gt;
nothing but the system's own boot volume. The data volume is 90% full, 386 GiB in use.&lt;/p&gt;

&lt;p&gt;The tooling side says the same: &lt;code&gt;restic&lt;/code&gt;, &lt;code&gt;borg&lt;/code&gt;, &lt;code&gt;borgmatic&lt;/code&gt;, &lt;code&gt;kopia&lt;/code&gt;, &lt;code&gt;duplicacy&lt;/code&gt;, &lt;code&gt;arq&lt;/code&gt;&lt;br&gt;
— none installed, no backup application under &lt;code&gt;/Applications&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I need a parenthesis here, because I fell straight into a hole I dug myself. In my first&lt;br&gt;
pass I ran &lt;code&gt;du&lt;/code&gt; on the iCloud folder, saw &lt;strong&gt;60 KB&lt;/strong&gt;, and wrote "sync is off, the data is&lt;br&gt;
here." Wrong. The folder holds 211 files and every one of them has &lt;code&gt;blocks&lt;/code&gt; set to 0 —&lt;br&gt;
they're uploaded to iCloud and evicted locally. &lt;code&gt;du&lt;/code&gt; counts blocks, so it reports 60 KB.&lt;br&gt;
The &lt;code&gt;Desktop&lt;/code&gt; and &lt;code&gt;Documents&lt;/code&gt; entries inside it aren't "empty shells" either; they're&lt;br&gt;
symbolic links to the real folders, and &lt;code&gt;du&lt;/code&gt; doesn't follow links, so they show as 0 bytes.&lt;br&gt;
What I took as proof that sync was off was proof that it was working.&lt;/p&gt;

&lt;p&gt;A week ago I wrote&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/life/diskimde-ne-kadar-yer-var-bes-cevap-13-gib-fark/" rel="noopener noreferrer"&gt;a separate article&lt;/a&gt;&lt;br&gt;
about how the "free space" number on this very disk changes depending on who you ask. I&lt;br&gt;
forgot that article's lesson in my own measurement one week later. This seems likely to&lt;br&gt;
keep happening, so I'm writing it down here.&lt;/p&gt;
&lt;h2&gt;
  
  
  I didn't forget to turn Time Machine on — I have no disk to turn it on with
&lt;/h2&gt;

&lt;p&gt;Apple's documentation sets two conditions for Time Machine. Capacity: a device with "at&lt;br&gt;
least twice the storage capacity of your Mac." Dedication: you use that disk only for Time&lt;br&gt;
Machine. Twice a 460 GiB disk means, roughly, a 1 TB external drive — and it means that&lt;br&gt;
drive exists solely for this purpose. I don't own one. So what's missing isn't a setting,&lt;br&gt;
it's &lt;strong&gt;hardware&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Local snapshots don't close the gap either. As far as Apple's page says, they're kept on&lt;br&gt;
the Mac's own startup disk, taken roughly every hour, and held for 24 hours. A copy that&lt;br&gt;
lives on the same disk is no help when the disk itself is what goes.&lt;/p&gt;

&lt;p&gt;My list is empty. I don't know exactly why, and I won't present my guess as fact: the&lt;br&gt;
destination may simply never have been configured, or Apple's other condition on the same&lt;br&gt;
page may not hold — it says Time Machine "stores snapshots only on disks that have plenty&lt;br&gt;
of free space," and my data volume is 90% full. Two candidate causes; my data isn't enough&lt;br&gt;
to choose between them.&lt;/p&gt;
&lt;h2&gt;
  
  
  What would I lose: 37 repos, 209 commits
&lt;/h2&gt;

&lt;p&gt;"I have no backup" is dramatic on its own, and empty. What matters is how much&lt;br&gt;
&lt;em&gt;irreplaceable&lt;/em&gt; material sits on that single disk. &lt;code&gt;~/IdeaProjects&lt;/code&gt; is 35 GB. I counted.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Git repositories (&lt;code&gt;.git&lt;/code&gt; directory, two levels deep)&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repositories with no remote (&lt;code&gt;git remote&lt;/code&gt;) at all&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Of those, ones with real history&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Commits in those 8 repositories&lt;/td&gt;
&lt;td&gt;209&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tracked files in those 8 repositories&lt;/td&gt;
&lt;td&gt;1,332&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repositories &lt;strong&gt;ahead&lt;/strong&gt; of their remote (unpushed)&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unpushed commits&lt;/td&gt;
&lt;td&gt;143&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dirty repositories&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Modified files&lt;/td&gt;
&lt;td&gt;275&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Untracked entries / actual files&lt;/td&gt;
&lt;td&gt;230 / 1,518&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I'm writing the counting method down so it can be checked: I searched for &lt;code&gt;.git&lt;/code&gt;&lt;br&gt;
&lt;strong&gt;directories&lt;/strong&gt;, which leaves four git worktrees out of the list — their commits already&lt;br&gt;
live in the main repository.&lt;/p&gt;

&lt;p&gt;I also have to correct that "230," because my first pass had a units error in it: 230 is&lt;br&gt;
the number of entries git &lt;strong&gt;collapses&lt;/strong&gt;, so most of them are directories, not individual&lt;br&gt;
files. The real untracked file count is &lt;strong&gt;1,518&lt;/strong&gt;. And that number shouldn't be taken at&lt;br&gt;
face value either: one empty scratch repository alone accounts for 1,118 of them, another&lt;br&gt;
for 196. So most of the 1,518 is noise. The real work is in the places I sampled: a new&lt;br&gt;
cost-rollup module, a barcode handler, six new test files, and database migrations numbered&lt;br&gt;
from the tenth through the twenty-second. Migrations are the kind of file that takes the&lt;br&gt;
schema with it when it disappears.&lt;/p&gt;

&lt;p&gt;I'm not writing the repository names down; most of them are client work. But I do need to&lt;br&gt;
mention one: a 916 MB portal, 19 commits, its last two commits made on 29 September, 44&lt;br&gt;
minutes apart. So it's live work from days ago. Remotes configured: zero. One of those&lt;br&gt;
commit messages contains this: "restore the quote package and pricing work that was lost in&lt;br&gt;
production." The code that rescues work lost in production is not itself anywhere it could&lt;br&gt;
be rescued from.&lt;/p&gt;

&lt;p&gt;Let me say exactly what the check covers when I say "no remote": &lt;code&gt;git remote&lt;/code&gt; returns&lt;br&gt;
empty, meaning &lt;strong&gt;there is no configured path to push to&lt;/strong&gt;. That doesn't prove the code was&lt;br&gt;
never copied anywhere — as we're about to see, some of it was. But on a recovery night, "I&lt;br&gt;
think I copied it somewhere" doesn't count as a plan.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where git doesn't look — and where it looks by accident
&lt;/h2&gt;

&lt;p&gt;I counted the &lt;code&gt;.env&lt;/code&gt; style files across the projects: &lt;strong&gt;18&lt;/strong&gt;. Then I asked each one, "is git&lt;br&gt;
tracking you?" On my first attempt the answer was "none," and I was pleased with that. My&lt;br&gt;
method was broken: I asked about each file by its bare name instead of its path relative to&lt;br&gt;
the repository root, so git naturally found nothing.&lt;/p&gt;

&lt;p&gt;Asked properly: &lt;strong&gt;16 are untracked, 2 are tracked.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The repositories holding those two have an &lt;code&gt;.env&lt;/code&gt; rule in their &lt;code&gt;.gitignore&lt;/code&gt;. So the rule&lt;br&gt;
was written and the file is inside anyway — added before the rule, or force-added. Their&lt;br&gt;
contents are two lines of domain and API address, not a secret leak; but that isn't the&lt;br&gt;
point.&lt;/p&gt;

&lt;p&gt;The point is this: what git protects is what git looks at, and I was wrong about both&lt;br&gt;
sides of that boundary. 16 files sit outside the scope, by my own choice, despite being the&lt;br&gt;
hardest class to reproduce. 2 files sit inside it while I believed they were out. I can&lt;br&gt;
rewrite a lost source file; I cannot rewrite a lost key — I'd have to go to the provider,&lt;br&gt;
generate a new one, and redeploy it everywhere it was used.&lt;/p&gt;
&lt;h2&gt;
  
  
  The destination I named "backup" is dead — the one I didn't name works
&lt;/h2&gt;

&lt;p&gt;Hoping for better, I looked at the &lt;code&gt;rclone&lt;/code&gt; side. Two remotes are configured:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;gdrive:
gdrive-yedek:
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second one is named, literally, "yedek" — Turkish for backup. I tried listing it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CRITICAL: Failed to create file system for "gdrive-yedek:":
couldn't find root directory ID: ... couldn't fetch token:
invalid_grant: maybe token expired? - try refreshing it
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In the first version of this article I stopped right here. I wrote "the destination I named&lt;br&gt;
backup cannot authenticate," closed the section, and that was drama built on selective&lt;br&gt;
evidence. I saw two remotes and tested only the broken one.&lt;/p&gt;

&lt;p&gt;When I opened the other, seven folders came back. Three of them matter:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A "Projeler" folder dated 22–30 April&lt;/strong&gt;: ~100,000 objects, 5.6 GiB. Two branches
underneath, &lt;code&gt;lokal&lt;/code&gt; and &lt;code&gt;harici&lt;/code&gt;, with six project folders under &lt;code&gt;lokal&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A folder dated 18 September&lt;/strong&gt;: Android release signing material and a restore note. A
signing key is even less replaceable than an API key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A key escrow dated 11 July&lt;/strong&gt;: GPG-encrypted, 1,788 bytes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So an offsite copy &lt;strong&gt;exists&lt;/strong&gt;, encryption &lt;strong&gt;exists&lt;/strong&gt;, a restore note &lt;strong&gt;exists&lt;/strong&gt;. My&lt;br&gt;
sentence "I have no copy anywhere" was wrong, and I had to delete it.&lt;/p&gt;

&lt;p&gt;But the real finding starts here. I compared the names of the six project folders in that&lt;br&gt;
April copy against the list of nine repositories with no remote: &lt;strong&gt;none of them match.&lt;/strong&gt;&lt;br&gt;
Worse, the earliest first commit among those nine is 7 July. So the copy sitting in the&lt;br&gt;
cloud was taken &lt;strong&gt;ten weeks before&lt;/strong&gt; all the work I had flagged as at risk. However&lt;br&gt;
well-intentioned, that folder does not protect me — what it would need to protect hadn't&lt;br&gt;
been born yet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IExSCiAgVlBTW1ZQUzMgZGF0YWJhc2VdIC0tPnxkYWlseSAxMDowMHwgTUFDWyhUaGlzIE1hYyldCiAgTUFDIC0tPnwyMi0zMCBBcHJpbCwgNS42IEdpQnwgR0RbZ2RyaXZlOiB3b3Jrczxici8-c2l4IHByb2plY3QgZm9sZGVyc10KICBNQUMgLS0-fDE4IFNlcCArIDExIEp1bHwgR0QyW2dkcml2ZTogc2lnbmluZyArIGVzY3Jvd10KICBNQUMgLS4tPnxubyBkZXN0aW5hdGlvbnwgVE1bVGltZSBNYWNoaW5lXQogIE1BQyAtLi0-fGludmFsaWRfZ3JhbnR8IEdEWVtnZHJpdmUteWVkZWs6IGRlYWRdCiAgUkVQT1s5IHJlcG9zIHdpdGggbm8gcmVtb3RlPGJyLz5maXJzdCBjb21taXQgNyBKdWx5XSAtLi0-fE5PVCBpbiB0aGUgQXByaWwgY29weXwgR0QKICBzdHlsZSBUTSBzdHJva2UtZGFzaGFycmF5OiA1IDUKICBzdHlsZSBHRFkgc3Ryb2tlLWRhc2hhcnJheTogNSA1%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IExSCiAgVlBTW1ZQUzMgZGF0YWJhc2VdIC0tPnxkYWlseSAxMDowMHwgTUFDWyhUaGlzIE1hYyldCiAgTUFDIC0tPnwyMi0zMCBBcHJpbCwgNS42IEdpQnwgR0RbZ2RyaXZlOiB3b3Jrczxici8-c2l4IHByb2plY3QgZm9sZGVyc10KICBNQUMgLS0-fDE4IFNlcCArIDExIEp1bHwgR0QyW2dkcml2ZTogc2lnbmluZyArIGVzY3Jvd10KICBNQUMgLS4tPnxubyBkZXN0aW5hdGlvbnwgVE1bVGltZSBNYWNoaW5lXQogIE1BQyAtLi0-fGludmFsaWRfZ3JhbnR8IEdEWVtnZHJpdmUteWVkZWs6IGRlYWRdCiAgUkVQT1s5IHJlcG9zIHdpdGggbm8gcmVtb3RlPGJyLz5maXJzdCBjb21taXQgNyBKdWx5XSAtLi0-fE5PVCBpbiB0aGUgQXByaWwgY29weXwgR0QKICBzdHlsZSBUTSBzdHJva2UtZGFzaGFycmF5OiA1IDUKICBzdHlsZSBHRFkgc3Ryb2tlLWRhc2hhcnJheTogNSA1%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram" width="970" height="432"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the best evidence for the "silence is the default" thesis. Nobody told me anything&lt;br&gt;
false; I never asked anyone the right question. The destination with "backup" in its name&lt;br&gt;
misled me precisely because it had a name; the one without a name escaped me precisely&lt;br&gt;
because it didn't. An inventory should be kept by measurement, not by label.&lt;/p&gt;

&lt;p&gt;I also don't know why the token died. Google's documentation lists several causes for&lt;br&gt;
&lt;code&gt;invalid_grant&lt;/code&gt;: access being revoked, a refresh token going unused for six months, a&lt;br&gt;
password change while Gmail scopes are granted, exceeding the limit of 100 tokens per&lt;br&gt;
account — in which case, in the documentation's own words, the oldest are invalidated&lt;br&gt;
&lt;em&gt;without warning&lt;/em&gt; — and, for projects whose publishing status is "Testing," expiry in 7&lt;br&gt;
days. I can't say which. What every item shares is that none of them is obliged to tell me.&lt;/p&gt;

&lt;p&gt;I've caught this pattern before. A launchd job&lt;br&gt;
&lt;a href="https://mustafaerbay.com.tr/en/blog/technology/launchd-21881-kez-denedi-ben-bir-kez-bakmadim/" rel="noopener noreferrer"&gt;tried 21,881 times and failed every time&lt;/a&gt;,&lt;br&gt;
and I never looked once.&lt;/p&gt;
&lt;h2&gt;
  
  
  Today, for the first time, I tested a dump
&lt;/h2&gt;

&lt;p&gt;Last question: can those 91 dumps I called good actually be restored?&lt;/p&gt;

&lt;p&gt;I searched the pulling script for &lt;code&gt;restore&lt;/code&gt;, &lt;code&gt;verify&lt;/code&gt;, &lt;code&gt;gunzip -t&lt;/code&gt;, &lt;code&gt;zcat&lt;/code&gt;, &lt;code&gt;psql&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;pg_restore&lt;/code&gt;. &lt;strong&gt;Zero hits.&lt;/strong&gt; The script copies and prunes; it does not test. For ninety-one&lt;br&gt;
days files arrived, and not one was ever opened.&lt;/p&gt;

&lt;p&gt;Today I tested the newest. &lt;code&gt;gzip -t&lt;/code&gt; passed, and &lt;code&gt;zcat&lt;/code&gt; gave up its first lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;--
-- PostgreSQL database dump
--
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That proves one thing: the dump's stream was &lt;strong&gt;not truncated&lt;/strong&gt;. A &lt;code&gt;pg_dump | gzip&lt;/code&gt; that&lt;br&gt;
died halfway would fail its checksum. That's all I know, and I won't claim more. &lt;code&gt;gzip -t&lt;/code&gt;&lt;br&gt;
verifies the integrity of the compression container, not the consistency of the SQL inside&lt;br&gt;
it; a missing role, a missing extension, a broken ordering would all be invisible.&lt;br&gt;
PostgreSQL's documentation, describing a plain-text dump, says what to do with one: to&lt;br&gt;
restore it, you feed it to &lt;code&gt;psql&lt;/code&gt;. A dump becomes a backup only when a database eats it.&lt;/p&gt;

&lt;p&gt;CISA's backup guidance asks for three things: 3-2-1 (three copies, two different media, one&lt;br&gt;
offsite), protecting the backups themselves (physical security, &lt;strong&gt;encryption&lt;/strong&gt;, offline&lt;br&gt;
copies), and &lt;strong&gt;testing&lt;/strong&gt; the procedures. Let me draw my own scorecard honestly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Copies&lt;/th&gt;
&lt;th&gt;Offsite&lt;/th&gt;
&lt;th&gt;Encrypted&lt;/th&gt;
&lt;th&gt;Tested&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Side project database&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;never&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;April project copy&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;never&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Key escrow&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;yes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;never&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Work since 7 July&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;no&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Here's what's interesting: the best job I did is in the smallest file. A 1,788-byte GPG&lt;br&gt;
escrow covers two of the three legs at once. The rest was built out of that day's urgency,&lt;br&gt;
not out of habit.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm going to do, ordered by risk
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Add remotes to the 9 repositories that have none.&lt;/strong&gt; 209 commits stop being
single-copy in exchange for one evening's work. Client code goes to private
repositories. Cheapest action, biggest gain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Push the 143 unpushed commits, decide about the 275 modified files.&lt;/strong&gt; Migrations
first. Most of the 1,518 untracked files are junk, but I can't know that without sorting
them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the inventory by measurement, not by label.&lt;/strong&gt; This is the real lesson of this
article. I didn't know what the April copy covered; unless I write down what today's
copy covers, I won't know that either in three months.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Either refresh the dead token or delete it.&lt;/strong&gt; A non-working path with "backup" in its
name is worse than no path at all — judging by the fact that it made me write this
article twice, it did the same to me.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Buy a 1 TB disk and turn Time Machine on.&lt;/strong&gt; And remember that a backup disk left
permanently connected is also exposed to ransomware; CISA's "offline copies" item exists
exactly for that.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;An honest note: I did not do item 1 today. If I don't do it tomorrow, this paragraph will&lt;br&gt;
be sitting here to embarrass me.&lt;/p&gt;

&lt;h2&gt;
  
  
  Backup isn't a tool question, it's an inventory question
&lt;/h2&gt;

&lt;p&gt;What I take from this measurement isn't "I neglected my backups" — that's what I thought&lt;br&gt;
when I first wrote it, and I was wrong. I did make backups. There is a daily, automatic, 91-file one; a 5.6 GiB copy of my projects went to the cloud; a key escrow got encrypted with GPG. I couldn't remember what any of them covered.&lt;/p&gt;

&lt;p&gt;Because the thing that's easy to automate is not the thing that's expensive to lose. The&lt;br&gt;
database on the server dumps with one command and already exists in two places. A&lt;br&gt;
half-finished migration, an uncommitted test, a key shown exactly once: none of them dumps&lt;br&gt;
with one command. Backup effort ran downhill, like water, along the path of least&lt;br&gt;
resistance — and nobody reminded me where that path ended up.&lt;/p&gt;

&lt;p&gt;Back in June I published an article on this blog arguing that backups are everyone's&lt;br&gt;
responsibility. While writing it, I had not run &lt;code&gt;tmutil destinationinfo&lt;/code&gt; on my own machine.&lt;br&gt;
The distance between those two things bothers me, so I'm putting it in writing: giving&lt;br&gt;
advice is easier than measuring.&lt;/p&gt;

&lt;p&gt;The default sound a system makes is silence. Time Machine didn't report having no&lt;br&gt;
destination, the token didn't announce its death, the April copy didn't mention going&lt;br&gt;
stale. And today I read a 0-byte error log first as a fault and then as proof of health,&lt;br&gt;
read iCloud's evicted files as "sync is off," and wrote the sentence "I have no copy&lt;br&gt;
anywhere" only to delete it. I misread the same silence three times in one day.&lt;/p&gt;

&lt;p&gt;So don't ask your own setup "do I have a backup?" You'll answer that one with a clear&lt;br&gt;
conscience. Ask this instead: what exactly does my backup cover, up to what date — and how&lt;br&gt;
do I know? I asked today, and I didn't know half the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/en-us/104984" rel="noopener noreferrer"&gt;Back up your Mac with Time Machine — Apple Support&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://support.apple.com/en-us/102154" rel="noopener noreferrer"&gt;About Time Machine local snapshots — Apple Support&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://developers.google.com/identity/protocols/oauth2" rel="noopener noreferrer"&gt;Using OAuth 2.0 to Access Google APIs — Google for Developers&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.cisa.gov/audiences/small-and-medium-businesses/secure-your-business/back-up-business-data" rel="noopener noreferrer"&gt;Back Up Business Data — CISA&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.postgresql.org/docs/current/app-pgdump.html" rel="noopener noreferrer"&gt;pg_dump — PostgreSQL Documentation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>backup</category>
      <category>dataloss</category>
      <category>macos</category>
      <category>silentfailure</category>
    </item>
    <item>
      <title>I Piled Up Seventy Branches; Six Were Really Unfinished</title>
      <dc:creator>Mustafa ERBAY</dc:creator>
      <pubDate>Thu, 01 Oct 2026 11:53:07 +0000</pubDate>
      <link>https://dev.to/merbayerp/i-piled-up-seventy-branches-six-were-really-unfinished-3bf8</link>
      <guid>https://dev.to/merbayerp/i-piled-up-seventy-branches-six-were-really-unfinished-3bf8</guid>
      <description>&lt;p&gt;I typed &lt;code&gt;git branch&lt;/code&gt; in the blog repo and the screen started scrolling. I counted:&lt;br&gt;
71 lines. One is &lt;code&gt;main&lt;/code&gt;; the other 70 are places where I once started something and&lt;br&gt;
walked away.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;71 local branches (including main) · 56 remote branches
tip commit by month:  2026-05: 4   06: 15   07: 21   08: 7   09: 23   10: 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every number in this article belongs to a single moment: the morning of 1 October 2026,&lt;br&gt;
with &lt;code&gt;origin/main&lt;/code&gt; at &lt;code&gt;4702bc4e&lt;/code&gt;. A branch list is a living thing; by the time you read&lt;br&gt;
this it will have changed on my side too.&lt;/p&gt;

&lt;p&gt;Four of them date back to May. They have been sitting there for four and a half months&lt;br&gt;
and I cannot remember what any of them are.&lt;/p&gt;

&lt;p&gt;What I felt looking at that list was not a technical feeling. It was the "I have this&lt;br&gt;
much unfinished work" feeling — a to-do list nobody wrote, growing on its own. I wanted,&lt;br&gt;
for once, to turn that background hum into a number. This article is the story of that&lt;br&gt;
count, and the count ended differently than I expected: I asked about unfinished work in&lt;br&gt;
four separate ways, got four separate answers, and the real debt turned out to be exactly&lt;br&gt;
where I had not been looking.&lt;/p&gt;
&lt;h2&gt;
  
  
  Seventy-one lines, zero information
&lt;/h2&gt;

&lt;p&gt;Let's admit something first: the length of &lt;code&gt;git branch&lt;/code&gt; output measures nothing. Creating&lt;br&gt;
a branch is free; it amounts to writing a forty-character commit id somewhere. And I&lt;br&gt;
created one in every agent session, every experiment, every "let me just poke at this"&lt;br&gt;
moment all year. Of the 70 side branches, 23 carry the &lt;code&gt;claude/&lt;/code&gt; prefix — working areas&lt;br&gt;
that were opened automatically.&lt;/p&gt;

&lt;p&gt;Still, every time I open that list it says the same thing to me: &lt;em&gt;you left a lot of work&lt;br&gt;
half-done.&lt;/em&gt; Here is the interesting part — git does not manufacture my guilt, I do; git&lt;br&gt;
merely stores it, and I pay interest on it with every &lt;code&gt;branch&lt;/code&gt; call. I was also reluctant&lt;br&gt;
to delete the branches, because of the fear that "there might be something in there." I&lt;br&gt;
had never once checked whether that fear was justified. Not in four and a half months.&lt;/p&gt;

&lt;p&gt;To get from a number to information, I had to ask a question. I did not know how many&lt;br&gt;
questions it would take.&lt;/p&gt;
&lt;h2&gt;
  
  
  First question: ancestor or not
&lt;/h2&gt;

&lt;p&gt;The most familiar question, and the easiest. &lt;code&gt;git branch --merged&lt;/code&gt; and &lt;code&gt;--no-merged&lt;/code&gt;. The&lt;br&gt;
documentation describes what they do without any embellishment:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"With &lt;code&gt;--merged&lt;/code&gt;, only branches merged into the named commit (i.e. the branches whose&lt;br&gt;
tip commits are reachable from the named commit) will be listed."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So the test is one single thing: can the branch's &lt;em&gt;tip commit&lt;/em&gt; be found by walking&lt;br&gt;
backwards from the named commit? Reachability. A genealogy query.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;git branch --merged origin/main    → 46
git branch --no-merged origin/main → 24
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Forty-six branches are closed, provably. That leaves 24, and my first reflex was to read&lt;br&gt;
that as "24 unfinished jobs." So this was the numerical equivalent of four and a half&lt;br&gt;
months of humming.&lt;/p&gt;

&lt;p&gt;It was not. Because this command does not know how I work.&lt;/p&gt;
&lt;h2&gt;
  
  
  Second question: is the content the same
&lt;/h2&gt;

&lt;p&gt;Most of the work I do by hand in this repo lands through pull requests, squashed; the rest&lt;br&gt;
of the commit traffic is the bot's content generation. GitHub's own documentation states&lt;br&gt;
the outcome of a squash plainly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Squashing turns all commits in the pull request into one commit on the base branch."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;GitLab describes the same behaviour on its side as "Squash and merge combines multiple&lt;br&gt;
small commits into a single meaningful commit." So this is not a GitHub quirk but a common&lt;br&gt;
way of merging.&lt;/p&gt;

&lt;p&gt;The result: a new commit, a new identity. My branch's tip commit will &lt;em&gt;never&lt;/em&gt; appear in&lt;br&gt;
&lt;code&gt;main&lt;/code&gt;'s history, because what entered &lt;code&gt;main&lt;/code&gt; is not a copy of it but a summary. The&lt;br&gt;
content landed; the genealogy did not. &lt;code&gt;--no-merged&lt;/code&gt; cannot distinguish this case, and it&lt;br&gt;
is not supposed to — I asked it a different question.&lt;/p&gt;

&lt;p&gt;The command that asks the right question is &lt;code&gt;git cherry&lt;/code&gt;. It looks at the diff, not the&lt;br&gt;
genealogy:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The equivalence test is based on the diff, after removing whitespace and line numbers."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A commit marked &lt;code&gt;-&lt;/code&gt; has an equivalent upstream; one marked &lt;code&gt;+&lt;/code&gt; does not. The idea&lt;br&gt;
underneath is &lt;code&gt;git patch-id&lt;/code&gt;; its documentation defines a patch id as "a sum of SHA-1 of&lt;br&gt;
the file diffs associated with a patch, with line numbers ignored," and it ties the two&lt;br&gt;
commands together explicitly: "git-cherry shows what commits from a branch have patch ID&lt;br&gt;
equivalent commits in some upstream branch." So if a patch does the same work, it yields&lt;br&gt;
the same fingerprint regardless of where it was applied or which commit id it carries.&lt;/p&gt;

&lt;p&gt;I ran all 24 branches through this question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;commits on the 24 branches that main lacks : 34
of those, with an equivalent upstream      : 20
without an equivalent                      : 14

branches with no unique commit at all      : 15
branches with unique commits               :  9
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now &lt;em&gt;this&lt;/em&gt; number says something. Fifteen of the 24 "unfinished" branches are &lt;strong&gt;completely&lt;br&gt;
empty&lt;/strong&gt; — the work every commit inside them does is already in &lt;code&gt;main&lt;/code&gt;, just under a&lt;br&gt;
different identity. &lt;code&gt;fix/pace-date&lt;/code&gt;, &lt;code&gt;fix/pace-runner&lt;/code&gt;, &lt;code&gt;fix/pace-url-encoding&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;fix/devto-mermaid-restore&lt;/code&gt;... The output is unambiguous, too; for &lt;code&gt;fix/pace-runner&lt;/code&gt; the&lt;br&gt;
command returns a single line, and the mark at the start of that line says everything:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git cherry &lt;span class="nt"&gt;-v&lt;/span&gt; origin/main fix/pace-runner
- c36fff1b539c6066e23c02e3d03b2fe99a171f9f fix&lt;span class="o"&gt;(&lt;/span&gt;ci&lt;span class="o"&gt;)&lt;/span&gt;: pace-check hosted runner yerine self-hosted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Minus. It has an equivalent. The work inside a branch I had been reluctant to delete for&lt;br&gt;
sixteen days had already landed in &lt;code&gt;main&lt;/code&gt;. Fifteen of the 24 suspects dissolved at this&lt;br&gt;
step.&lt;/p&gt;

&lt;p&gt;I could have stopped here and congratulated myself. Nine instead of 24, how nice. But&lt;br&gt;
there is something &lt;code&gt;git cherry&lt;/code&gt; cannot tell you either, and this time it errs in the&lt;br&gt;
opposite direction: it compares diffs, not intentions. If I &lt;em&gt;rewrote&lt;/em&gt; the same work, the&lt;br&gt;
two patches look different and the commit gets marked &lt;code&gt;+&lt;/code&gt;. So 14 does not mean "open&lt;br&gt;
work"; it means "possibly open work."&lt;/p&gt;

&lt;p&gt;That leaves exactly one method, and it cannot be automated: open the nine branches one by&lt;br&gt;
one and read them.&lt;/p&gt;
&lt;h2&gt;
  
  
  Third question: was this work really not done
&lt;/h2&gt;

&lt;p&gt;I opened the nine branches in order, looked at which files they touched, then searched for&lt;br&gt;
those files in &lt;code&gt;origin/main&lt;/code&gt;. For three of them the answer was clear — the work landed, the&lt;br&gt;
branch is redundant:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;claude/calendar-atomic-write&lt;/code&gt;&lt;/strong&gt; — the commit message reads "wip: local working session&lt;br&gt;
sync." Inside it are &lt;code&gt;AnimatedPostCover.astro&lt;/code&gt; and &lt;code&gt;copy-blog-assets.mjs&lt;/code&gt;. Both are sitting&lt;br&gt;
in &lt;code&gt;main&lt;/code&gt;. The patch looks different because this commit froze half a session as-is; the&lt;br&gt;
work itself landed through a proper PR.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;fix/topic-research-quota-fallback&lt;/code&gt;&lt;/strong&gt; — two commits, 894 added lines; the job of decoupling&lt;br&gt;
topic verification from the Gemini quota. I searched &lt;code&gt;main&lt;/code&gt; for the &lt;code&gt;SEARCH_STOPWORDS&lt;/code&gt; and&lt;br&gt;
&lt;code&gt;TR_FOLD&lt;/code&gt; definitions; both are there. It landed, but then evolved, which is why the diff&lt;br&gt;
no longer matches.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;fix/self-host-workflows&lt;/code&gt;&lt;/strong&gt; — on 21 July I opened this to move the keep-alive and lighthouse&lt;br&gt;
jobs onto my own runner. In &lt;code&gt;main&lt;/code&gt; today, line 27 of &lt;code&gt;keep-alive.yml&lt;/code&gt; and line 17 of&lt;br&gt;
&lt;code&gt;lighthouse-audit.yml&lt;/code&gt; say the same thing: &lt;code&gt;runs-on: [self-hosted, itwise-mac]&lt;/code&gt;. The goal&lt;br&gt;
was reached — two months later, by a different commit, with a note dated 26 September. The&lt;br&gt;
branch's intention came true; the branch itself became redundant.&lt;/p&gt;

&lt;p&gt;Three branches, four commits, zero debt. The remaining six are another story.&lt;/p&gt;
&lt;h2&gt;
  
  
  Six that are real
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVs3MCBzaWRlIGJyYW5jaGVzXSAtLT58Z2l0IGJyYW5jaCAtLW1lcmdlZHwgQls0NiBjbG9zZWRdCiAgQSAtLT58LS1uby1tZXJnZWR8IENbMjQgbG9vayBvcGVuXQogIEMgLS0-fGdpdCBjaGVycnk6IGVxdWl2YWxlbnQgZm91bmR8IERbMTUgYnJhbmNoZXMgZW1wdHldCiAgQyAtLT58Z2l0IGNoZXJyeTogbm8gZXF1aXZhbGVudHwgRVs5IGJyYW5jaGVzIC8gMTQgY29tbWl0c10KICBFIC0tPnxyZWFkIGJ5IGhhbmQ6IHdvcmsgaXMgaW4gbWFpbnwgRlszIGxhbmRlZCBhbm90aGVyIHdheV0KICBFIC0tPnxyZWFkIGJ5IGhhbmQ6IHdvcmsgTk9UIGluIG1haW58IEdbNiBnZW51aW5lbHkgb3Blbl0%3Ftype%3Dpng%26bgColor%3Dwhite" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgQVs3MCBzaWRlIGJyYW5jaGVzXSAtLT58Z2l0IGJyYW5jaCAtLW1lcmdlZHwgQls0NiBjbG9zZWRdCiAgQSAtLT58LS1uby1tZXJnZWR8IENbMjQgbG9vayBvcGVuXQogIEMgLS0-fGdpdCBjaGVycnk6IGVxdWl2YWxlbnQgZm91bmR8IERbMTUgYnJhbmNoZXMgZW1wdHldCiAgQyAtLT58Z2l0IGNoZXJyeTogbm8gZXF1aXZhbGVudHwgRVs5IGJyYW5jaGVzIC8gMTQgY29tbWl0c10KICBFIC0tPnxyZWFkIGJ5IGhhbmQ6IHdvcmsgaXMgaW4gbWFpbnwgRlszIGxhbmRlZCBhbm90aGVyIHdheV0KICBFIC0tPnxyZWFkIGJ5IGhhbmQ6IHdvcmsgTk9UIGluIG1haW58IEdbNiBnZW51aW5lbHkgb3Blbl0%3Ftype%3Dpng%26bgColor%3Dwhite" alt="Diagram"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For each of the six I looked for the evidence in &lt;code&gt;main&lt;/code&gt; and failed to find it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Branch&lt;/th&gt;
&lt;th&gt;Work&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude/elegant-bassi-fa298f&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;never-404 hardening, 5 commits, 9 June&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;500.astro&lt;/code&gt;, &lt;code&gt;en/500.astro&lt;/code&gt;, &lt;code&gt;maintenance.html&lt;/code&gt; absent; zero trace of 500 in the middleware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fix/callout-structure-gate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Callout structure gate, 27 July&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;mdx-structure.ts&lt;/code&gt; absent; &lt;code&gt;validateMdxStructure&lt;/code&gt; is imported by neither of the two validators&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fix/accuracy-source-coverage&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;widening the source allowlist, 21 August&lt;/td&gt;
&lt;td&gt;the allowlist lacks &lt;code&gt;caddyserver.com&lt;/code&gt;, &lt;code&gt;libvirt.org&lt;/code&gt;, &lt;code&gt;linux-kvm.org&lt;/code&gt;, &lt;code&gt;isc.org&lt;/code&gt;, &lt;code&gt;letsencrypt.org&lt;/code&gt;; &lt;code&gt;source-policy.mjs&lt;/code&gt; is 242 lines in &lt;code&gt;main&lt;/code&gt;, 292 on the branch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;fix/unconfirmed-rejection-not-permanent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;stopping an unverified rejection from eating the queue, 2 September&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;topic-research-chain.mjs&lt;/code&gt; in &lt;code&gt;main&lt;/code&gt; is 283 lines, zero relevant trace&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude/dreamy-maxwell-17b977&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;caption timeout + outage breaker, 18 September&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;provider-outage.mjs&lt;/code&gt; absent; &lt;code&gt;main&lt;/code&gt; still uses 5- and 15-second timeouts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude/beautiful-noether-e1c690&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the ispmanager review article, 21 September&lt;/td&gt;
&lt;td&gt;file absent from &lt;code&gt;main&lt;/code&gt;; zero hits across the 1300 records of the live feed; URL returns 404&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Seventy came down to six. So 91 percent of the humming was not real.&lt;/p&gt;
&lt;h2&gt;
  
  
  Where age says nothing
&lt;/h2&gt;

&lt;p&gt;Once the count was done I was left holding an intuition, and I wanted to test that too:&lt;br&gt;
an old branch is a dead branch, a new branch is live work. It sounds reasonable. I sorted&lt;br&gt;
both lists by date.&lt;/p&gt;

&lt;p&gt;Of the 15 empty branches, the oldest is 16 May (138 days) and the newest is 23 September.&lt;br&gt;
Of the 6 genuinely open ones, the oldest is 9 June (114 days) and the newest is&lt;br&gt;
21 September. The two ranges almost entirely overlap.&lt;/p&gt;

&lt;p&gt;The intuition was wrong — but exactly how it was wrong is the interesting part. One side&lt;br&gt;
of it holds: of the four branches left over from May, one had already been merged and the&lt;br&gt;
remaining three are still on the list, and all three are empty; so the fear that "they&lt;br&gt;
have been sitting there for four and a half months" had precisely zero substance behind&lt;br&gt;
it. But the 114-day-old &lt;code&gt;elegant-bassi&lt;/code&gt; is sitting right there too, carrying five commits&lt;br&gt;
of real work. Age indicates neither innocence nor debt. An aged branch has either found&lt;br&gt;
its own way in or is waiting somewhere nobody looks, and you cannot read which from its&lt;br&gt;
date.&lt;/p&gt;

&lt;p&gt;That is why there is no shortcut around the third step. Whatever filter I tried, the only&lt;br&gt;
thing that answered the question was looking at the file.&lt;/p&gt;
&lt;h2&gt;
  
  
  What the minus sign does not say
&lt;/h2&gt;

&lt;p&gt;The method's most important blind spot was sitting inside my own data.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;claude/beautiful-noether-e1c690&lt;/code&gt; carries two commits. One &lt;code&gt;+&lt;/code&gt;, one &lt;code&gt;-&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;git cherry &lt;span class="nt"&gt;-v&lt;/span&gt; origin/main claude/beautiful-noether-e1c690   &lt;span class="c"&gt;# SHAs abbreviated&lt;/span&gt;
+ fd33e470  feat: yeni makale — ispmanager Lite&lt;span class="s1"&gt;'ın Altına Baktım: 23 Bulgu
- ea466a4b  fix(ispmanager): final denetim — yayın 25 Eyl, disclosure, kanıt düzeltmeleri
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The equivalent of the minus-marked &lt;code&gt;ea466a4b&lt;/code&gt; really does exist in &lt;code&gt;main&lt;/code&gt;'s history:&lt;br&gt;
&lt;code&gt;6003feb2&lt;/code&gt;. But there is one more line right after it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;fd17e0ff Revert "fix(ispmanager): final denetim — yayın 25 Eyl, ..."
         This reverts commit 6003feb2381fc9f5877c6ef687ed1f58dac3e77b.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the minus in &lt;code&gt;git cherry&lt;/code&gt; does not mean "this work is live"; it means "an equivalent of&lt;br&gt;
this entered upstream at some point." If it entered and was later reverted, the mark does&lt;br&gt;
not change. I caught it only because this branch also carries a &lt;code&gt;+&lt;/code&gt; — had the branch held&lt;br&gt;
that &lt;code&gt;-&lt;/code&gt; commit alone, my four-step method would have called it "empty, safe to delete."&lt;br&gt;
And the work is not live.&lt;/p&gt;

&lt;p&gt;Catching your own method with your own data is a strange feeling. The thing you are&lt;br&gt;
measuring turns out to be measuring your instrument too.&lt;/p&gt;
&lt;h2&gt;
  
  
  Two correct notes, two jobs never done
&lt;/h2&gt;

&lt;p&gt;Two of the six are where I actually stopped — but not for the reason I expected.&lt;/p&gt;

&lt;p&gt;For &lt;code&gt;claude/dreamy-maxwell-17b977&lt;/code&gt; I wrote this on 18 September: the social posting&lt;br&gt;
pipeline hung for 40 minutes, the cause was a fetch without a timeout, the fix is this&lt;br&gt;
module. The note's last line reads: &lt;em&gt;"Awaiting verification: check this line in the first&lt;br&gt;
social-post log after the merge."&lt;/em&gt; I asked &lt;code&gt;gh&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;gh &lt;span class="nb"&gt;pr &lt;/span&gt;view 13 &lt;span class="nt"&gt;--json&lt;/span&gt; number,state,mergedAt,headRefName
&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"headRefName"&lt;/span&gt;:&lt;span class="s2"&gt;"claude/dreamy-maxwell-17b977"&lt;/span&gt;,&lt;span class="s2"&gt;"mergedAt"&lt;/span&gt;:null,&lt;span class="s2"&gt;"number"&lt;/span&gt;:13,&lt;span class="s2"&gt;"state"&lt;/span&gt;:&lt;span class="s2"&gt;"OPEN"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Still open. Open for thirteen days. The note was not wrong — on the contrary, it said&lt;br&gt;
plainly that it was &lt;em&gt;waiting&lt;/em&gt; for a merge. That merge simply never happened, and nobody&lt;br&gt;
went back to check. The code in &lt;code&gt;main&lt;/code&gt; today still runs with the old timeouts; that&lt;br&gt;
40-minute hang could happen again tomorrow.&lt;/p&gt;

&lt;p&gt;The second one is starker. For the ispmanager review, my note records in three separate&lt;br&gt;
places that the article never went live, including the verification of the live 404. The&lt;br&gt;
note even contains the command to run on 25 September: revert the two revert commits, push,&lt;br&gt;
trigger the deploy run. The command was written correctly. It was never run. Today the&lt;br&gt;
article is absent from the 1300 records of the live feed and the URL still returns 404.&lt;/p&gt;

&lt;p&gt;The story I expected was "my notes misled me." The real story is more uncomfortable:&lt;br&gt;
neither git nor my notes were wrong. Both were correct, both were right there, and nobody&lt;br&gt;
went back to either. I had nothing at all that distinguishes a decision written down from&lt;br&gt;
a job carried out — and that gap was sitting inside the branch list the whole time,&lt;br&gt;
invisible because a list of names cannot show it.&lt;/p&gt;

&lt;p&gt;The feeling was familiar. Five days ago I wrote about &lt;a href="https://mustafaerbay.com.tr/en/blog/career/dogru-hipotezi-iki-ay-erken-kurdum/" rel="noopener noreferrer"&gt;forming the right hypothesis two&lt;br&gt;
months early and losing it&lt;/a&gt;; the day&lt;br&gt;
after that, about &lt;a href="https://mustafaerbay.com.tr/en/blog/career/doksan-saniyede-geri-aldim-bes-saat-bozuk-kaldi/" rel="noopener noreferrer"&gt;rolling back in ninety seconds and leaving things broken for five&lt;br&gt;
hours&lt;/a&gt;. All three share&lt;br&gt;
one thing: the record was there, and nobody read it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I count now
&lt;/h2&gt;

&lt;p&gt;Four steps, in order. If you want to try it on your own setup:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Do not mistake a number for information.&lt;/strong&gt; &lt;code&gt;git branch | wc -l&lt;/code&gt; produces a feeling,
not data. The only job of this step is to accept that counting is not enough.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask the genealogy&lt;/strong&gt; (&lt;code&gt;--merged&lt;/code&gt; / &lt;code&gt;--no-merged&lt;/code&gt;). Which branches closed is proven.
The ones that look open are &lt;em&gt;suspects&lt;/em&gt;, not convicts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask the diff&lt;/strong&gt; (&lt;code&gt;git cherry origin/main &amp;lt;branch&amp;gt;&lt;/code&gt;). This is where you catch work that
landed via squash. For me, 15 of the 24 suspects dissolved at this step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the survivors by hand.&lt;/strong&gt; Rewritten work does not resemble the diff; no automation
that skips this step can answer "was this work done."&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Three traps caught me while applying the fourth step, and I am writing all of them down:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A &lt;code&gt;-&lt;/code&gt; does not mean "live"&lt;/strong&gt; (the section above). Work whose equivalent landed and was
then reverted also shows as &lt;code&gt;-&lt;/code&gt;. If in doubt, look at what came after the equivalent
commit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;git branch -d&lt;/code&gt; refuses these branches.&lt;/strong&gt; The documentation's condition is explicit:
the branch must be fully merged in its upstream branch or in &lt;code&gt;HEAD&lt;/code&gt;. Those 15 branches
are &lt;code&gt;--no-merged&lt;/code&gt; by definition, so you have to use &lt;code&gt;-D&lt;/code&gt;. The evidence you gather does
not &lt;em&gt;complement&lt;/em&gt; the safety of &lt;code&gt;-d&lt;/code&gt;; it replaces it — so do not type &lt;code&gt;-D&lt;/code&gt; without
gathering it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A branch attached to a worktree cannot be deleted.&lt;/strong&gt; In this repo four branches are
currently held by a worktree, one of them from the "genuinely open" six. You have to
remove the worktree first.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And one area the recipe does not cover at all: all of this is local. I have 56 remote&lt;br&gt;
branches against 70 local ones, and deleting a local branch does not delete&lt;br&gt;
&lt;code&gt;origin/&amp;lt;branch&amp;gt;&lt;/code&gt;. A branch that exists only on the remote never appears in &lt;code&gt;git branch&lt;/code&gt;&lt;br&gt;
output at all — so this count is itself incomplete.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took away
&lt;/h2&gt;

&lt;p&gt;The count did not show me my unfinished work. It showed me how I audit myself.&lt;/p&gt;

&lt;p&gt;For four and a half months I looked at a seventy-line list and felt vaguely bad, because&lt;br&gt;
staying vague was free. Reading nine things one by one was uncomfortable — and it was&lt;br&gt;
precisely that discomfort that found two jobs never done. Vague guilt turns out to be a&lt;br&gt;
defense mechanism: as long as I know nothing for certain, I never have to admit that&lt;br&gt;
something clearly did not get done either.&lt;/p&gt;

&lt;p&gt;A number is a feeling, a diff is a fact, and only reading tells you which one is right.&lt;br&gt;
But I saved the real lesson for last: in this count, no record lied. Git was right, my&lt;br&gt;
notes were right, the PR's status was in plain view. The only thing missing was someone&lt;br&gt;
going back after writing. Black boxes spend their worst nights without telling anyone;&lt;br&gt;
mine had recorded everything properly — nobody opened the lid.&lt;/p&gt;

&lt;h2&gt;
  
  
  Official Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/git/git/blob/master/Documentation/git-branch.adoc" rel="noopener noreferrer"&gt;git-branch documentation (&lt;code&gt;--merged&lt;/code&gt;, &lt;code&gt;--no-merged&lt;/code&gt;, the &lt;code&gt;-d&lt;/code&gt; condition)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/git/git/blob/master/Documentation/git-cherry.adoc" rel="noopener noreferrer"&gt;git-cherry documentation (diff-based equivalence test)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/git/git/blob/master/Documentation/git-patch-id.adoc" rel="noopener noreferrer"&gt;git-patch-id documentation (patch id and its link to git-cherry)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/incorporating-changes-from-a-pull-request/about-pull-request-merges" rel="noopener noreferrer"&gt;GitHub: pull request merge methods and squash behaviour&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.gitlab.com/user/project/merge_requests/squash_and_merge/" rel="noopener noreferrer"&gt;GitLab: squash and merge behaviour&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>git</category>
      <category>measurement</category>
      <category>technicaldebt</category>
      <category>worklife</category>
    </item>
  </channel>
</rss>
