I assumed that if search found the right document, that was good enough for a conversational assistant answering questions about my work.
My portfolio had grown beyond a collection of project pages. Between independent projects, technical writing, career history, and engineering documentation across repositories, there was increasingly more material to navigate. I was building retrieval behind a site chat assistant so visitors could ask about that work without hunting through sources themselves.
The first indexing pass treated each whole document as one searchable item, enough to prove ingest, stable identifiers, and scripted checks before wiring chat.
Some documents covered several distinct topics: a case overview, a block on how attribution was measured, runbook notes in the same file. A narrow question needs one of those sections, not an average of the whole document. With one search result per document, similarity ranking cannot prefer the attribution section over the overview when both share the same parent.
I needed retrieval to return individual passages, not whole documents, without changing how the underlying content was authored. The pipeline normalized mixed sources into Open Knowledge Format (OKF) documents, split documents that covered several topics into independently searchable passages, and indexed them for similarity search:
Portfolio, repos, and writing
→ OKF structured documents
→ Passage splitting
one document, several topics
├─ Overview passage
├─ Attribution passage
└─ Runbook passage
each independently searchable; retains source reference
→ OpenAI embeddings
→ Postgres + pgvector
→ Similarity search
→ Top matching passages
Splits follow Markdown headings (#, ##, and so on), not fixed character counts, so each passage lines up with how the source was already organized. Headings inside fenced code blocks are ignored so examples do not become false section breaks. When a section was still too long, the splitter used paragraph boundaries rather than cutting through the middle of a paragraph.
What the retrieval tests showed
Across 35 structured documents from the selected sources, splitting produced 61 independently searchable passages. I built a small set of automated tests using questions where I knew what information should be returned. Each passage kept a stable id and a pointer to its parent document so checks could score document match and section match separately.
One test asked about attribution in growth experiments. The relevant write-up covered several topics, so finding the document alone wasn't enough. The test checked whether search returned the specific section about attribution measurement.
It did. The attribution passage ranked first.
I also tested what happened when the information wasn't there.
For a deliberately unrelated question about nuclear reactors, search still returned passages from my career history. They were the closest available matches, even though none could answer the question.
That meant the chat assistant would need both the retrieved passages and their distance scores to help determine whether there was enough evidence to answer.
Expanding the portfolio corpus
I expanded the corpus to include more of the material already available across my portfolio: project case studies, published articles, and documentation about the tools and workflows I'd built. The retrieval architecture stayed the same. What changed was the information available for search.
| Stage | Structured documents | Searchable passages |
|---|---|---|
| Initial retrieval experiment | 35 | 61 |
| Portfolio corpus expansion | 89 | 209 |
After expanding the portfolio corpus, a question about agent memory still exposed a gap. The answer lived in architecture notes for Savepoints, a side project where I capture durable learnings from agent sessions, and those notes were not in the index yet.
Expanding beyond the portfolio
I added selected engineering documentation from my other repositories, including the Savepoints architecture notes, alongside reviewed career material. I kept the index limited to material suitable for public use.
After indexing the additional material, search could finally find the architecture notes about agent memory.
For that question, related articles about agent portability occupied the first nine results. The Savepoints notes appeared at rank 10.
The evaluation passed because it searched the top 10 results, while the retrieval CLI returned only five by default. With that setting, the relevant Savepoints notes would have been missed.
That left a decision for when I wired up the chat assistant: how many passages should each search return? Five wasn't necessarily wrong, and ten wasn't necessarily right. The answer depended on what was in the index, what visitors asked, and which passages ranked highest.
Takeaway: Retrieval quality depends on both the content being indexed and how search is configured. Passing retrieval tests is a useful starting point, but the settings still need to be evaluated against the questions the assistant is expected to answer.
Top comments (4)
Long time no see! This post is incredibly timely. I’m working on agent testing in my new role at a new company, and the ideas you laid out here helped me expand my test scenarios. This is fantastic!👏
Long time no see! Great to hear from you again, and congrats on the new role! Really glad the post was useful.
I'm curious what additional test scenarios it inspired, especially since I'm still working through the evaluation side. Unlike my other projects, the quality of the retrieval results can shift as the content in the vector database grows and changes over time.
Thanks!
One key test scenario I’m building targets ranking edge cases — relevant documents get pushed far down the result list even though they hold the correct answer.
We’re building a three‑tier system: agent platform, individual agents, and an agent OS. The knowledge base component is still being refined, and we plan to submit it for testing next week. I’ll reach out to discuss this once we have real-world scenario tests up and running.😝
The two results near the end meet in a single comparison if the distance check runs on each passage. The reactor question showed that every query gets its nearest passages back, so an answer has to be gated on distance, and the Savepoints question showed a passage that does answer sitting at rank 10 behind nine on a neighbouring topic. If the distance of that rank 10 passage is no better than the best distance the reactor question got, no per-passage cutoff keeps one and refuses the other. If it is clearly better, a cutoff with k at 10 admits it and still turns the reactor question away. Both searches have already run, so the comparison is free. If the check looks only at each question's best passage instead, the two come apart, and k alone decides whether the Savepoints notes arrive.
The tests also hide the movement that caused this, because they score a pass at a cutoff. Recording the rank of the expected passage for each question on each index build would have shown the Savepoints notes arriving at rank 10 instead of a pass, and if a later batch pushes the attribution passage from first to fourth, it would show that too, well before the passage leaves the top five. With a rank per question, choosing between five and ten becomes reading the largest rank among the questions the assistant has to answer, and the cost of that choice is the number of passages that ride along on every other question.