DEV Community

Nerav Doshi
Nerav Doshi

Posted on Originally published at pipelineandprompts.com

Grew the Collection from 5 Entries to 741 Real Chunks

Every retrieval test since Entry 05 has run against the same handful of entries — three documents in Entries 05-06, twenty after Entry 07's chunking. Realistic enough to prove the mechanics work, not realistic enough to trust the actual rankings. Growing this into something closer to the real published catalog was overdue, and getting there cleanly took two wrong turns first.

First attempt pointed the bulk-embed script at what looked like the right content folder and got back three files that were in Repo scaffolding, not articles — the real posts weren't there at all. Hugo commonly organizes content as page bundles, each article living in its own subfolder as <slug>/index.md, and the script's flat glob("*.md") only looked one level deep. Fixed by pointing at the actual posts directory directly rather than guessing at folder depth.

Second run picked up all 45 real files — the 15 "Today I Ran" entries plus 30 other published articles — and embedded 741 chunks. Looked like a clean win. It wasn't, quite: the first run's three junk files were still sitting in the collection, uncleaned, and worse, the 15 "Today I Ran" entries that were already in the store from Entries 05-06 got re-embedded under a different ID scheme than before. The original pipeline stored single-chunk files under their plain filename (02-oc-cli-mentor-system-prompt.md); this script's chunking always appends _chunk0, _chunk1, even for short files. Same content, two ID formats, both present at once — a self-inflicted version of the exact cross-document ranking bias Entry 07 found by accident with a real multi-chunk document.

Rather than hand-picking which of ~766 entries to delete, wiped the whole collection and rebuilt from nothing:

import chromadb
client = chromadb.PersistentClient(path="./chroma_db")
client.delete_collection("today_i_ran_notes")
Enter fullscreen mode Exit fullscreen mode

Then re-ran the bulk script fresh. Clean this time:

Files found:      45
Files processed:  45
Files failed:     0
Chunks embedded:  741
Collection size:  0 -> 741
Enter fullscreen mode Exit fullscreen mode

No duplicates, no scaffolding files, consistent chunking and ID scheme across every document. From 5 test entries to 741 real chunks drawn from the actual catalog, in one clean pass.

One thing to glag before next test: this collection now includes the "Today I Ran" series' own back catalog — Entries 01 through 15 — alongside everything else. The index can now retrieve passages from its own past entries about building the index. Querying this store about chunking bias will likely surface Entry 07, which is the entry that discovered chunking bias, embedded into the very store that bias affects. Recursive in a way that's more interesting than useful, but worth noting before the next retrieval test runs against this and something from the series' own history shows up as the top result.

The real payoff of this entry isn't the number 741 — it's that every retrieval test from here forward is running against something that resembles the actual size and shape of a real knowledge base, instead of a handful of hand-picked documents chosen to make a specific point work. Entry 07's chunking-bias fix (the per-document result cap flagged back then, still not built) matters a lot more now than it did against 20 entries.

Top comments (0)