This blog's duplicate-post problem has been solved twice. By hand in June: 38
posts deleted, a 301 put in place for every deleted address. By mechanism at the
end of July: a similarity gate that tests a candidate title against the archive
at generation time. The second one still holds — 222 posts have gone out since
it landed, this one aside, and there is not a single new pair among them.
This morning that same rule went the other way, onto the archive itself. There
were 1,313 posts in it at that moment: 861,328 comparisons, 19.3 seconds. The
result: 36 pairs, 27 clusters, 61 posts. All live, all returning 200.
So the gate was not built wrong. It was simply never pointed at the archive.
This reads like a blog story, but the shape is familiar: anyone who adds a lint
rule, puts NOT NULL on a column or wires a new check into CI draws the same
line. The rule protects everything after that commit; the lines written before
it, the rows that went into the schema earlier, the files that passed through
before the check — those stay exactly where they are. And adding the rule does a
very good impression of being finished.
In June I deleted thirty-eight posts
On 15 June at 12:41 a consolidation went in: 38 posts that told the same narrow
story in different words were deleted, one canonical was picked per cluster, and
the deleted ones were 301'd through src/lib/redirects.ts. The note at the top
of that file sums up the logic of the day:
// AI pipeline ayni dar konuyu Nisan-Haziran boyunca tekrar uretti; her tekrar
// kumesi icin tek kanonik secildi, digerleri buraya 301'lendi (Google
// cannibalization + thin/duplicate content temizligi). Benzersiz bolumler
// silinmeden once kanoniklere graftlandi -> icerik kaybi yok.
38 entries, 7 canonicals. The two biggest clusters tied at eight posts each:
PostgreSQL WAL bloat and ERP supply-chain data flow. Behind them, at seven
apiece, came database index selection and ERP bill-of-materials denormalisation.
Before deleting anything, the unique sections of each cluster went into its
canonical, because the posts were not bad; they just answered the same question
eight times.
It felt done. I wrote the map, wired it into the middleware, confirmed the
single-hop 301, closed the tab.
The next day the map turned out to be wrong
On 16 June at 15:34, 34 more entries went into the same file and the map grew
from 38 to 72. But the real fix was not in the count. That commit's own message
spells out day one's mistake:
Mevcut harita (38) yalnizca kategori-li URL'leri yakaliyordu — ama eski
URL'ler KATEGORISIZ (/blog/<slug>/) oldugu icin resolveBlogRedirect
eslestiremiyordu (kullanicinin gordugu 5-adimda 404'unun gercek sebebi buydu).
The existing 38-entry map only matched category-prefixed URLs, it says, but the
old URLs had no category segment, so the resolver never matched them at all.
Everything I verified was correct; the one thing I never verified was the shape
the old links arrived in. The map expected /blog/<category>/<slug>/; what
arrived was /blog/<slug>/. Day one's work was not incomplete, part of it was
wrong.
And what surfaced it was not a measurement of mine. The first line of the commit
message names the source: the dead-link list in Analytics, Google's index and
backlink 404s, and a visitor hitting a 404 on the 5-adimda address. There was
a signal. It was not mine, and neither was the cost.
The fix had two parts: the redirect resolver was made category-flexible, and the
34 missing posts went into the map — the five that belonged to a cluster to its
canonical, the 29 one-off deletions to their category page. Of today's 72
entries, 29 go to a category page and 43 to nine canonical posts.
That is the character of hand-done cleanup: it ends when it ends, it never runs
itself again, and somebody else tells you where it went wrong.
At the end of July the gate went in
On 28 July at 11:27, scripts/lib/topic-dedup.ts went live. During generation it
tests the candidate title against every published title and rejects the topic if
any of four conditions holds:
const duplicate =
sameNormalized ||
(intersection >= 3 && jaccard >= 0.54) ||
(intersection >= 3 && containment >= 0.78) ||
trigramDice >= 0.84;
Three similarity measures. Jaccard looks at the share of common words.
Containment catches a short title dissolving inside a longer one — "Switch
Hardening: Is It Always a Necessary Step?" against "Switch Hardening: Does Every
Device Need the Same Detail?" is exactly that. Trigram Dice is a character-level
measure that ignores word boundaries. On top of that sits a 54-word stop list
("steps", "guide", "how", "setup") and light Turkish stemming; without it every
"Setting Up X" title would collide with every "Setting Up Y".
Describing those three cost me a sentence I had already written. The first draft
said all three catch a different failure; then came a count of which rule
actually decides. In 26 of the 36 pairs more than one rule fires together, in 7
only containment, in 3 only Jaccard. The number of pairs where Trigram Dice
decides on its own: zero. More than that, the typical bridge this measure
exists for — "API Versiyonlama" against "API Versioning" — scores 0.54 against a
threshold of 0.84. The rule has never once fired in this archive. Rather like
the 'icin' entry appearing twice in the stop list: a detail nobody notices
without reading the code.
The gate's record is good. 222 posts have been published since 28 July and among
those 222 there is not one pair that its own rule would call a duplicate. The
last pair to get through went live eight hours before the gate was merged: "5
Ways for Junior Developers to Stand Out in the AI Era" on the tutorials shelf,
dated 17 July, and "4 Steps for Junior Developers to Stand Out in the AI Era" on
the career shelf, timestamped 28 July at 03:17. Eight hours. I could not ask for
a cleaner record of when a mechanism started working.
This morning the rule went the other way
The gate is a module with an exported function. Pointing the same rule at the
archive therefore needs no new similarity metric — importing the module and
comparing every post with every other post is enough:
const topics = collectPublishedTopics('src/content/blog');
const hits = [];
for (let i = 0; i < topics.length; i++) {
for (let k = i + 1; k < topics.length; k++) {
if (findClosestTopic(topics[i].title, [topics[k]]).duplicate) {
hits.push([topics[i].slug, topics[k].slug]);
}
}
}
console.log(hits.length);
861,328 comparisons, 19.3 seconds, 36 pairs. No pre-filter, no sampling, no
threshold of my own — whatever title the gate rejects today is what gets counted.
Split by the later member's date, the pairs fall out like this:
| Window | Pairs | What it means |
|---|---|---|
| Before 16 June | 28 | The hand cleanup walked past these |
| 16 June – 28 July | 7 | The 42 days between cleanup and gate |
| 28 July onward | 1 | Only the Junior pair |
The single pair on the third row is the boundary case above: its date lands on
the day the gate merged, its timestamp eight hours earlier. So no pair has been
published since the gate started working.
The row that bothers me is the first one. The cleanup took 38 posts; among the
835 posts dated on or before that same date and still standing today, 28 pairs
remain. The hand-done work was not a sweep, it was a sequence of judgement calls
— and nothing recorded where those calls stopped.
The duplication is in the intent, not the text
The bodies of those 36 pairs went through a comparison too: code blocks
stripped, text overlap measured in five-word windows. Median 0.21%, highest
0.81%. Not one pair clears 1%.
The number looks small, and that is the actual finding. No text-similarity tool
would flag any of these 36 pairs. No copy-paste, no repeated paragraphs, no
rewording. Each post was written from scratch, with different sentences and
different examples. What overlaps is not the text: it is the question being
answered.
The biggest cluster shows it bare. Between 21 May and 6 June — 17 days, three
categories — six posts went out, and all six ask the same question:
| Date | Category | Title |
|---|---|---|
| 21 May | life | API Versioning: URI or Header? A Pragmatic Choice |
| 23 May | technology | API Versioning Strategy: URI or Header? A Pragmatic Choice |
| 27 May | tutorials | The Cost of API Versioning: URI or Header? |
| 28 May | tutorials | API Versioning: URI vs Header – Which Is More Practical? |
| 1 Jun | tutorials | API Versioning Strategies: Pragmatic Approaches |
| 6 Jun | life | API Versioning Strategy: Simple Approach or Future-Proof Solution? |
All six addresses returned 200 this morning, and each one points rel="canonical"
at itself. Google's own documentation is explicit about what happens next: if you
do not specify a canonical, "Google will identify which version of the URL is
objectively the best version to show to users in Search." Which of the six shows
up in search results is therefore not up to me. Not making that call is also a
call; I just do not get to see its outcome.
The method of the June cleanup was right: pick one canonical per cluster, move
the unique sections into it, 301 the rest. A 301 is the tool the standard itself
defines for this — RFC 9110 says the target resource "has been assigned a new
permanent URI and any future references to this resource ought to use one of the
enclosed URIs", and Google writes that its indexing pipeline reads the redirect
as a canonicalisation signal. The right tool was in my hands; it just never reached
any of these 27 clusters.
Three more clusters, three different lessons
There is no point walking through the remaining 26, but three of them say
different things.
ERP multi-tenant architecture — three posts, three shelves. 23 May life, 29
May career, 3 June technology. Eleven days, the same architectural question
answered from three categories. The fault is not in the topic: having three
shelves was enough to ask the same question three times.
The swap fire — three posts, one event. 9 May, technology: "VPS Swap Fire: A
Nightmare That Started With a Kernel CVE Patch". 14 May, life: "Swap Fire on My
VPS: A Nightmare That Started With a Kernel CVE Patch". The difference between
those two titles is a possessive. A third slots in between them on 12 May: "Swap
Fire on My 7.6GB VPS: A Nightmare That Started With a Kernel Patch". What
separates the three is not an editorial angle but cosmetics: a number added here,
"CVE" dropped there. Three more clusters shipped both halves on the same day — 3
April, ERP DMZ "pattern" and "design"; 13 May, two Docker network monitoring
guides; 16 May, "3 Practical Strategies" and "3 Practical Approaches". Not even
same-day comparison was happening.
Offline-first sync — three posts, one shelf. 16, 29 and 30 May, all three in
tutorials. This cluster takes away the "different categories, that's why I
missed it" defence: all three sit on the same shelf, two weeks apart, side by
side in the same directory.
The reader's side: forty-seven pairs, zero links
Every number so far is from my side. From the reader's side it looks like this:
across the 27 clusters there are 47 post pairs that are each other's neighbours,
and in none of those 47 does either post link to the other. Not one internal
link — all 27 clusters are closed. Someone looking for an answer on API
versioning lands on one of six pages and has no way of learning that the other
five exist.
The 61 posts in these clusters come to 102,751 readable words. At 200 words a
minute that is eight and a half hours of reading, for 27 questions. Not because
any of them is badly written — none of them is; they just answer the same
question without knowing about each other.
The missing internal links look like an oversight, but I think the real
information is this: had those posts been linked to each other, I would have seen
the clusters without having to count them. An internal link is also a note to
yourself saying "I have written about this before".
Why the gate cannot see the archive
The reason is not the threshold, it is the direction. The gate's signature is
findClosestTopic(candidate, existingTopics) — one candidate, one archive.
collectPublishedTopics walks the archive and indexes it; the index is rebuilt
on every run, so freshness is not the issue. The issue is that nothing ever
calls that function for two members of the archive. The comparison is one-way —
the new looks at the old, the old looks at nothing.
The tests do not cover it either. There are 16 test files under scripts/tests/
and one of them does check for duplicates; but that test looks for repeated JSON
keys in the calendar file and for the same slug appearing in two entries — exact
matches. No test compares published posts pairwise.
Had I run this scan once on the day the gate went in, the table would have been
on my screen that day. It costs 19 seconds. I did not run it, because installing
a gate feels like the end of the job. In fact every new gate draws a debt line on
its install date: everything after it is protected, everything before it stays
put — and that line is not written down in any file.
The debt starts in the plan
The same scan then went over the topic queue — the topic strings of the 519
entries in scripts/content-calendar.json, compared against each other. Five
pairs. So duplication does not begin at generation time; in some cases the plan
was already written twice.
One of them is an exact match: the entry "The Untold Side of Working Remotely
Abroad" sits in the calendar twice — 4 July and 8 July, same category, same
targetWords, character-for-character the same title. Both are marked
generated: true. In the repository that post has exactly one creation commit: 8
July at 19:27. The calendar claims two productions; git shows one.
A gate that would catch these entries is in fact sitting in the repository, and
it is well written: scripts/verify-topic-calendar.ts applies the same
similarity module to the calendar, and by pushing every accepted entry onto its
comparison list it tests entries against each other too. The check I was looking
for has already been written.
The reason it stays quiet is this single line:
const pending = calendar.entries.filter((e) => !e.generated && !e.rejected);
The gate only looks at pending entries. When this scan ran, the calendar
held 519 entries: 505 generated, 14 rejected, zero pending. Its output when run just now:
"0 bekleyen konu, 1314 yayımlanmış konu ile karşılaştırıldı" — zero pending
topics, compared against 1,314 published ones. Both duplicate "Working Remotely
Abroad" entries are marked generated: true, so the gate does not look at them
by definition. And this validator is not called from any workflow; it waits for
someone to type npm run verify:topics.
That turns out to be the sharpest example of my own argument: a correctly
written gate that audits zero of 519 entries and never runs in CI.
The same blindness is not limited to one gate
The duplicate gate is not alone in this story. The same pattern shows up on
another gate in the archive: the source policy that requires at least three
primary sources per post.
Of 1,313 posts, 1,090 (83%) either have no "Official Sources" section at all
or fewer than three links in it. Broken down by month, the gate's install date is
readable straight from the data:
| Month | Posts | Without three sources |
|---|---|---|
| 2026-05 | 274 | 273 |
| 2026-06 | 232 | 232 |
| 2026-07 | 200 | 182 |
| 2026-08 | 97 | 2 |
| 2026-09 | 96 | 1 |
From July to August the rate drops from 91% to 2%. The gate went in somewhere in
that window and has held ever since. All six of the API versioning posts above
carry zero official sources — a technical question answered six times, not once
tied to a primary source.
Same shape, same outcome: the gate faces forward, the archive stays put. The
point is not tuning a threshold; it is not treating gate installation as the
finish line.
The gate looks in one direction and in one language
A third boundary surfaced while the archive scan was running:
collectPublishedTopics skips isEnglishPost files while building its index.
So all 36 pairs above are between Turkish posts. Every one of them has an
English twin, and those twins carry exactly the same clusters; the English answer
to the API versioning question also sits at six separate addresses, all six
returning 200, all six pointing rel="canonical" at themselves.
On the redirect side this is not a problem: a single entry in the
src/lib/redirects.ts map serves both /blog/<slug>/ and /en/blog/<slug>/,
because the English address is the Turkish one with an /en prefix. That is also
why the 301s are resolved by hand in src/middleware.ts rather than through
Astro's own redirects config — one map for two languages, and a single hop. On
the measurement side there is no such symmetry: the gate counts Turkish and never
sees English.
The false positive deserves saying out loud
All 36 pairs went through a read, one by one. There is one pair where the gate is
plainly wrong: "There Is No Such Thing as a Perfect Product: The Naked Truth of
20 Years" on the career shelf, dated 2024, against "There Is No Such Thing as
Perfect Architecture" on technology, dated 2026. The gate matched them at
containment 0.80 — the second title is short, four of its five meaningful words
sit inside the first, and the one word left outside is the word carrying the
subject: "architecture". Its Jaccard is 0.40, so the primary measure objects; the
containment rule is deciding on its own.
Then there are the borderline ones — the two Ollama posts, for instance. One
walks through setting up a local LLM, the other covers using a local model for
code completion to cut cloud dependency. Same tool, different question — or two
faces of the same question. In clusters like that, the answer to "same or not"
lives in editorial judgement, not in a metric. Which is exactly why the scan's
output has to be a list, not a verdict.
At generation time this false positive costs nothing: the agent picks another
topic and nobody notices. In the archive it is not free — a job that deletes
automatically on that rule would needlessly 301 one of two genuinely different
posts. A criterion can be cheap in one direction and expensive in reverse. The
scan is becoming a standing test, but without delete authority: the output will
be a list, a human makes the call, and known false positives live in an
allowlist.
The category shelf was noise too
22 of the 36 pairs span two different categories. The same question answered
separately under technology, tutorials, life and career; the API versioning
cluster alone spreads across three shelves. In that period the category was not a
property of the topic, it was a coin flipped at generation time.
When the category label started carrying meaning is its own story: 44 posts in a
row landing on the same shelf
after the queue ran dry. That post already named this gap — "no gate puts two
posts side by side and asks whether they resemble each other", it says. Naming
turns out not to be measuring: eighteen more days passed between naming the gap
and measuring it, and the measurement took 19 seconds. What this one adds is the
second part: a gate's install date is a boundary of its own, and that boundary
shows up on no indicator.
The shape of a fix
What I take from this is not about duplicate content. It is about the shape
of a fix: the same fix is either carried out as a migration or installed as an
invariant, and those two do not have the same half-life.
These are the questions I now ask before fixing anything in my own pipeline:
- Does this fix run once, or on every run? Run once, and it corrects today's photograph, not tomorrow's flow.
- From what date does the gate protect? Has the data from before that date been scanned once with the same rule? If the scan takes 19 seconds, I have no excuse.
- Is the criterion safe in reverse? A false positive that is cheap in production can turn into data loss in the archive.
- Who will report the part left undone? In June that job fell to the dead-link list in Analytics and a visitor who hit a 404. The owner of the signal should not be the reader.
The price of asking is the same price the pipeline's 67 repairs
taught me: every mechanism
brings its own blind spot along with it. The gate protects the candidate against
the old; it protects the old against nothing.
One thing has not been measured, and saying so is part of the report: what these
61 posts actually cost. Which of the six API pages Google indexes, whether these
clusters draw search traffic at all — none of that has been checked. Most of this
blog's traffic has already turned out to be bots, so my guess is that the damage
is small. The reason for fixing it is not search ranking but the archive itself:
so a reader walking through one of six doors is not left unaware of the other
five.
The work in front of me is clear and small. The 27 clusters break down like this:
one of six posts, three of three, the remaining 23 of two. So most of the job is
binary calls; the big clusters are the labour-intensive minority.
The criterion for picking a canonical will differ from June's, because of one
data point I did not have that day: none of these posts has a sources section. So
"keep the one with the best sources" is off the table. What is left, applied in
order:
- The fullest body stays (in the API cluster that is the 27 May post, 2,314 readable words; the thinnest is 28 May at 1,011).
- The category that fits the topic stays; on a tie, category rotation decides the right shelf for that subject.
- The unique sections of the rest move into the canonical, then get 301'd — no content lost, the same graft method as June.
- Once merged, a sources section goes in; the merged post becomes able to pass today's gate as well.
The only new thing is that this time the scan runs on every run — output a list,
no authority, false positives in an allowlist. Because an archive is not a place
you clean once; it is a place that grows quietly as long as nobody looks at it.
Top comments (0)