This blog's publishing pipeline wakes up three times a day: 08:30, 13:30, 16:30.
Each time it picks a topic, researches it, writes it in Turkish and English,
renders a cover, pushes it through the gates and puts it live. Six months of
that. Where it keeps breaking is something I have already
laid out in numbers.
In the middle of September the pipeline did not miss a single day. Pace on
target, gates green, alarms quiet. And over that same stretch, the 44 posts it
produced all sat in the same category.
Once the count was in, the number was not the uncomfortable part. The
uncomfortable part was that nothing had broken. Had it broken, it would have
told me. It did not break; it just kept writing into the same shelf.
The queue ran out on 3 September
The pipeline has a topic queue: scripts/content-calendar.json. There are 514
entries in it, 500 of them published. Each entry carries a source note —
recording where a topic came from is a habit that started embarrassingly late,
but from 3 September onward the notes are explicit:
"source": "nobetci-uretimi (kuyruk tukendi, taze konu)"
A second record agrees on the date. When the old generation workflow was retired
on 9 September, the note left at the top of the file reads: "Bu cron ise 3 Eyl'den
beri SIFIR makale uretti: takvim kuyrugu tukendi" — this cron has produced zero
articles since 3 September, the calendar queue is empty. Two separate places,
the same day.
So no queued topic was left and the pipeline started finding its own. That part
was no secret to me; I simply did not read it as a problem. I did not want
production to stop when the queue emptied, and the fallback behaviour looked
sound: find a mechanism the blog has never covered, verify it against the primary
source, measure it in a lab, write it up.
The behaviour was sound; the outcome was not.
The fingerprint of a reflex
The window from 3 to 22 September — the day the queue emptied through the day the
warning arrived — is 20 days. Counting from the frontmatter gives 65 posts:
| Category | Posts | Share |
|---|---|---|
| technology | 59 | 90.8% |
| tutorials | 5 | 7.7% |
| life | 1 | 1.5% |
| career | 0 | 0% |
On 15 of those 20 days every single post was technology. The longest unbroken
run is 8–21 September: 14 days, 44 posts, one category, no interruption. The
last career post before the gap went out on 1 September and the next one on 23
September — 22 days apart. For life the gap is 20 days, for tutorials 17.
technology, meanwhile, never went quiet in that window at all: it published on
all 20 of the 20 days.
A table does not quite convey a run like that; a list does. The nine posts
published between 13 and 15 September, in order: HSTS preload, Authenticated
Origin Pulls, fs.protected_regular, kTLS, tcp_ecn, TCP keepalive, dm-verity,
KSM, ssh -J. Nine distinct mechanisms, nine distinct measurements — and all
nine on the same shelf. For three days, a reader of this blog was offered exactly
one kind of writing: a lab note that flips a kernel knob and shows the result
differing from expectation. A good kind on its own. Three days running, it stops
being a genre and becomes a tic.
The tags point the same way. Of those 65 posts, 47 (72.3%) carry the linux tag
and 28 carry kernel. Of those 59 technology posts, 43 carry the linux tag
and 27 carry kernel or sysctl. None of them were bad posts; they
were measured, sourced, tested in my own lab. They were simply all being filed
in one place.
A control group: what came out while the queue was full
A number like that is misleading on its own; it needs a comparison. The window
immediately before, 14 August to 2 September, while the queue was still full:
19 days, 59 posts.
| Window | Posts/day | technology | tutorials | career | life |
linux tag |
|---|---|---|---|---|---|---|
| 14 Aug – 2 Sep (queue full) | 3.11 | 39.0% | 44.1% | 10.2% | 6.8% | 3.4% |
| 3 – 22 Sep (queue empty) | 3.25 | 90.8% | 7.7% | 0% | 1.5% | 72.3% |
Exactly one column refuses to move: posts per day. It went from 3.11 to 3.25.
Throughput held; composition collapsed.
While the queue was full, the variety was not a virtue of mine. The balance was
built months earlier, when the topics were written into the list by hand, and the
list carried it. The list ran out, the balance went with it, and what was left was the reflex.
The volume indicator never turned red
The pipeline has gates and they work: the source policy (at least three primary
sources across two distinct hosts), the duplicate scan, word-count and
reading-time consistency, MDX parsing, 93 policy tests, a publication-stalled
alarm, an intraday pace check. Some of them do look at the individual post: the duplicate scan asks "has this
topic run before?", the source policy asks "are the sources in order?" So the
post-scale version of "what did it publish?" is covered. What is never asked is
the scale above it: the ratio between posts. No gate puts two posts side by
side and asks whether they resemble each other.
One of them deserves a closer look, because it is the watchman that counts the
published output every hour: daily-pace-check. The comment at the top of the
file states its job in one line — it checks the maximum-four-posts-per-day rule
once an hour. The step itself does exactly that: it counts the commits on main
since the start of the day whose message begins with feat: yeni makale, and
errors out only if the count exceeds four. GitHub's documentation defines the
schedule event in a single sentence of its own: "The schedule event allows you
to trigger a workflow at a scheduled time." What fires it is not a judgement about
content but a clock — and the same page noting that the trigger can be delayed
under high load, with queued jobs sometimes dropped, is a reminder that it is a
timer rather than an opinion. For two weeks that clock kept finding a number that
did not exceed four — on 10 and 17 September exactly four posts went out, and
neither day broke the rule. It was entirely right to stay silent; that was the
question I had given it.
Google's SRE book names two questions for a monitoring system: "what's broken,
and why?" My indicators answered both correctly — nothing was broken. But the
SLO chapter of the SRE workbook has a sharper sentence: if a disruption is not
captured by any indicator, that is "a strong sign that your SLO lacks coverage".
The gap in my coverage had a name, and the name was composition.
The Prometheus alerting guide lands in the same place: alert on symptoms rather
than causes, and aim for "as few alerts as possible". Sound advice. But defining
the symptom is my job, and I had defined it as "publishing stopped". Publishing
did not stop. The blog narrowed to one subject — a symptom that definition
simply did not contain.
The warning came from a reader
The best evidence of the pipeline's blindness to composition sits in the
calendar's own record. The source note on the topic pool added on 23 September
begins like this:
"KONU HAVUZU (23 Eyl 2026, kullanici talebi: makaleler hep technology
cikiyordu, 7 Eyl'den beri kategori cesitliligi yok)"
Translated: topic pool, 23 September 2026, user request — the articles kept
coming out as technology, no category variety since 7 September. The note is
unambiguous about the source of the signal. What flagged the broken balance was
not an indicator, a test or a report; it was somebody reading the blog. From 7 to
23 September is 16 days. That is the detection latency: sixteen days and zero
automatic signals.
There is no soft way to phrase this. The person who has spent six months
counting and writing up every place this pipeline breaks failed to catch, by his
own measurement, the crudest deviation in its most visible output. And what the
side that did catch it said was not a subtle metric: the articles keep coming out
as technology.
One correction I owe here: "since 7 September" is the note's own phrasing, not my
measurement. One of the four posts published on 7 September is tutorials; the
unbroken run begins the next day. What the record supports is this: 15 days
counting from the start of the run, or 13 days counting from the day the detector
I am about to build would have fired. Either way, the side that produced the
signal does not change.
Back-testing the detector I did not have
"Which indicator would have seen this?" is a question I did not want to leave
hanging, because questions like that usually hang there and quietly die as good
intentions. So here is a simple rule: slide a seven-day window, and if it holds at
least seven posts, compute the category shares; if the dominant category crosses a
threshold, warn. Then apply the rule backwards to every day from 1 April to 3
October — a range covering 1,276 of the archive's 1,308 posts; the remaining 32
carry older dates.
At a 70% threshold: 22 warning days in six months, in three clusters — 4–7 August
(dominant tutorials, 19 of 27 posts), 26–27 August (12 of 17) and 9–24
September (16 days; at first trigger, 19 of 24 posts technology).
At 80% a single cluster survives: 10–24 September. No other day fires at all.
So a threshold exists that drives false alarms to zero, and that threshold would
have caught the deviation on 10 September. The reader spoke on 23 September.
Thirteen days apart, and 41 more posts went out in that gap.
The weakness of this test belongs in the article too: the threshold was found in
data where the deviation was already known. 80% fits this six-month archive well;
a week whose topic distribution narrows by design — a series going deep on one
product, a cluster of posts after an incident — would trip the same threshold
unfairly. So the right shape is not "warn and halt" but "warn and ask": a
question about whether the narrowing is a deliberate choice or a reflex. Still,
to claim the threshold sits roughly in the right place, there is six months of
real trigger history behind it, which beats having none.
There is also a more honest route for anyone picking a threshold without the
benefit of hindsight: instead of writing a fixed number, let your own history set
it. Apply the same sliding window to the past, derive the distribution of the
dominant category's share, and make an upper percentile — p95, say — the
threshold; your normal gets defined by you, rather than borrowing my 80%. Window
length follows the same logic: a window needs at least a dozen units in it, or a
two-item day reads as "monoculture" by chance. For a pipeline producing three
posts a day, seven days clears that bar; for somebody closing two tickets a week,
the same rule needs a thirty-day window or it never fires at all.
The count itself is about as hard as a shell loop. Category and date are already
in every post's frontmatter:
for f in src/content/blog/*/*.mdx; do
case "$f" in *.en.mdx) continue;; esac
awk -F': *' '/^publishDate:/{d=$2} /^category:/{gsub(/"/,"",$2); c=$2} END{print d, c}' "$f"
done | sort | awk '$1>="2026-09-03" && $1<="2026-09-22" {n[$2]++; t++}
END{for (k in n) printf "%-11s %3d %%%.1f\n", k, n[k], 100*n[k]/t}'
technology 59 %90.8
tutorials 5 %7.7
life 1 %1.5
That a query this small went unrun for sixteen days is the most expensive detail
in this piece. Measuring was not hard; thinking of measuring was hard — because
every counter was green, and a green counter does not provoke questions.
Why the reflex kept going to the same place
This is the part that interests me most, because this part is mine, not the
pipeline's.
The topic-finding method goes like this: grep the archive, find a mechanism never
covered, verify it against the primary source, measure it. That method behaves
like a cost function — and kernel settings are by far its cheapest candidates.
You flip a sysctl in one session, measure, compare, and put the evidence on
screen. Cheap to verify, fresh evidence, low risk of being wrong.
Career and life pieces demand the same evidentiary bar at a far higher price.
The evidence for a kernel setting is on my own machine, now, in the output of one
command. The evidence for a career piece is scattered across six months of git
history, CI logs, the actual state of servers, or 1,300 published posts. It has
to be searched, extracted, counted, and the count verified twice — the windows,
the control group and the back-test in this article are exactly that expensive
part. The method was picking the most easily verifiable topic; I had been reading
that as the best topic.
A distinction is worth drawing here: the reflex did not produce bad work. I still
stand behind the technical content of those 44 posts; they were measured and
their sources held. What the reflex broke was not individual posts but the ratio
between them. And a ratio is a property no single post contains — it exists only
when you look from above.
The chapter on automation in the SRE book fits precisely here: automation is a
"force multiplier, not a panacea", and much of its value comes from consistency —
"very few of us will ever be as consistent as a machine". The book scopes that
consistency narrowly: the execution of well-scoped, known procedures. What my
pipeline was executing was not a procedure but a choice. The catch is that
consistency never asks whether the decision was right. The pipeline applied my
preference 44 times in a row with flawless fidelity. A human would have grown
tired, grown bored, switched subjects. The machine did not get bored.
I also do not think this failure is specific to automation. Looking at my own
weeks, the same curve shows up: what I take up to learn is usually dictated not
by curiosity but by how provable the thing is. I go where I can measure. Part of
what I call expertise is the sediment of a reflex like that.
Writing the rule into the system: compute, do not choose
Two things happened on 23 September. A new topic pool went into the calendar,
with category balance taken into account. And a step that runs before topic
selection was added to the production recipe: compute the target category from
the last-written date, take the one that has gone longest without a post, and
produce only there on that run.
The detail that matters is that the rule is code rather than a sentence. Writing
"mind the category variety" into the recipe would have achieved nothing, because
minding runs at the same moment as the reflex that finds the cheapest topic, and
it loses that race. The rule is now a step that computes the target category from
the calendar and imposes it. Not a statement of intent; a gate.
The price of that distinction has come due before: there was a line in the deploy
workflow I took for a guard, and months later it turned out to be
blocking nothing at all.
A rule being written down does not mean it is enforced. That post closed on
"being able to count the places the gates do not look"; the difference here is
that this time the counting actually happened instead of the sentence being
repeated.
Eleven days later — and the limits of this measurement
From 23 September to 3 October: 11 days, 34 posts.
| Category | Posts | Share |
|---|---|---|
| technology | 10 | 29.4% |
| tutorials | 8 | 23.5% |
| career | 8 | 23.5% |
| life | 8 | 23.5% |
This table does not count the post you are reading; with it the window holds 35
posts, career rises to 9 (25.7%) and posts per day to 3.18.
On none of those 11 days did a day stay in a single category; in the previous
window that ratio was 15 of 20. The average number of distinct categories per day
went from 1.25 to 2.82, and the share of posts tagged linux fell from 72.3% to
44.1%.
One more look at throughput, this time with the same denominator across all three
windows — per calendar day, counting days with no publication too: 2.95 · 3.25 ·
3.18. A number moving inside a ten-percent band. Composition collapsed and then
recovered while volume did not budge; the balance was not bought by cutting
output.
Now, the way this table should not be read: "the rule worked, here is the proof."
Two things changed on the same day; a topic pool built with category balance also
went in. The pool alone could have produced this table. A measurement separating
the two is not something I have, and not something I set up. Without a
discriminating test there is no telling which intervention is doing the work —
the same trap I fell into
two months earlier with a LinkedIn hypothesis. And 11 days is a short window next to 20.
The honest statement: composition recovered, and the cause cannot be assigned to
a single intervention.
And one more correction. "We will see when the pool runs dry again" is what I was
about to write; the calendar says it has run dry already. Of 514 entries, 500
are generated and the remaining 14 are rejected — zero topics waiting. The 23
September pool lasted eleven days. So the rule is already running on its own, and
the post you are reading is the product of exactly such a run: rotation said
"career", the career queue was empty, and the pipeline found the topic itself.
This time what it found was not a kernel setting but its own publishing record.
The real gain from this piece is not even the rotation rule
— it is that a query which measures composition now exists, and noticing a
fourteen-day run no longer requires waiting for a reader.
The reflex moved up a level
Running the same eye over the draft of this post produced a finding that stung,
because it was right. All nine career posts published since 23 September carry a
counted number in the title: 67 repairs, seven servers, two months, ninety
seconds, three scripts, twenty findings, seventy branches, sixty-eight backups —
and the forty-four posts you are reading about. The category field recovered; the
title shape settled into a single mould.
The rotation rule looks at the category field. It does not look at the shape of
the title, the structure of the piece, or the kind of evidence used. The reflex
did not disappear; it moved outside the field being measured. Put the constraint
anywhere and repetition accumulates one notch above it — and my new composition
query is blind to this by construction, because it too counts only categories.
Starting to measure one monoculture does not mean you have started counting the
dimensions you are not measuring.
This paragraph is not here as a gesture of modesty. If the thesis holds, its first
casualty should be the post itself: your coverage reaches exactly as far as the
dimension you measure, and the next blind spot is waiting one level above wherever
you put the gate.
Applying it to your own work
For anyone reading this without a blog pipeline, the same questions in short
form:
- What is the composition of your last 20 units? The last 20 tickets you closed, the last 20 reports you wrote, the last 20 proposals you sent. Split them into categories and count. Not volume — distribution.
- Which of your indicators would have seen this drift? If none would, "everything is green" means no more than "the three things I measure are fine".
- What is carrying your balance? A list, a plan, a client portfolio — or your appetite on the day? When the list ends, the balance ends with it.
- Is cost making your choices? Easily verified work is not the same as the right work. Your reflex goes to the cheapest evidence, and the shape of your expertise comes out of that.
- Did you write a rule or a gate? "I will be careful" is not a rule. A rule is computed, imposed, and visible when skipped.
- Who is your detector? If the answer is a person — a reader, a client, a teammate — that person is your monitoring system, and it does not scale.
Conclusion
A system looking healthy does not mean it is doing the right thing; it means it
answers the question I gave it well. An indicator that measures volume cannot, by
definition, see composition collapse: with 44 posts out of one category, the
counter still writes the same three.
Variety is not a by-product of productivity. What arrives as a by-product is
repetition — because repetition is cheaper, faster, and its evidence is already
at hand. If balance is wanted, it has to be written somewhere as a constraint,
into the system or into the day. Otherwise the reflex does the choosing, and a
reflex always reaches for the nearest shelf.
Official Sources
- SRE Book — Monitoring Distributed Systems: the two questions monitoring must answer
- SRE Workbook — Implementing SLOs: a disruption no indicator captures is a coverage gap
- SRE Book — Automation at Google: the value of consistency and how automation spreads mistakes at scale
- Prometheus — Alerting: alert on symptoms and keep the number of alerts small
- GitHub Actions — the
scheduleevent: triggers a workflow at a scheduled time, and can be delayed under load
Top comments (0)