Loose end from Entry 14: switching mcp_search_server.py from stdio to SSE broke the Entry 08 Claude Code connection, which was registered expecting a spawned process, not a network server. Fixing it turned out to be the easy part:
claude mcp remove today-i-ran-notes
claude mcp add --transport sse today-i-ran-notes http://localhost:8090/sse
claude mcp list
✔ Connected, server still running from Entry 14. Clean.
Re-ran Entry 08's original test — ask it to use the tool, search for the oc pod-status question. It worked, and gave a solid narrative answer identifying Entry 02's wrong command. But I wanted the actual distance score this time, not just the summary, so I asked for the raw tool output directly. That's where this got more interesting than "confirmed the reconnect works."
Two calls came back, for two slightly different phrasings of basically the same question:
| Query | Distance to 02-oc-cli-mentor-system-prompt.md
|
|---|---|
| "how to check pod status with oc" | 430.1 |
| "oc get pods command output troubleshooting" | 375.8 |
Same target document, same intent, a 54-point swing just from rewording. For comparison, the gap between the best and worst result within a single query was 442.6 → 461.1 — about 18 points. The variance from paraphrasing the question was three times larger than the variance between a genuinely-relevant result and a marginal one in the same result set.
That's a sharper version of the open question from Entry 13, and it resolves cleanly once you see both numbers side by side. Entry 13 guessed the embedding model might be rewarding lexical/syntactic overlap over pure meaning — a broken, command-shaped query scoring tighter than a clean one. This data points at something more basic underneath that: distance isn't calibrated for "is this relevant" at all, full stop. It's only meaningful for ranking results within one specific query's wording. Compare distances across different phrasings of the same question, and the number tells you more about word choice than about relevance.
There's also a reproducibility result worth stating plainly, since it cuts the other way: Entry 14's n8n test and this one both queried the same exact string — "how do I check pod status with oc" — through different clients, and both landed on 437.7. That's not in tension with today's finding. It's the other half of it: identical query text reliably gives identical distance, regardless of client or transport. Reword the query, even slightly, and the number moves more than you'd expect from meaning alone.
One more thing worth being honest about, since it affects how much to trust the narrative answer from the first run: Claude Code's accurate quote of the wrong oc command — the one from Entry 02 — didn't actually come from the MCP tool. The tool's results are truncated to ~200 characters per snippet, and that quote wasn't in any of them. The session separately grepped and read the full file directly off disk. So the earlier clean-looking answer wasn't a pure test of the RAG pipeline; it was the RAG pipeline plus Claude Code's own filesystem access filling the gap. Worth knowing before holding that result up as proof the retrieval system alone is doing the work — some of it was something else entirely.
Collection size matters here too, and it's worth naming rather than letting the clean distance numbers imply otherwise: every query in this entry returned the exact same five results, just reordered. The index doesn't have much in it yet. These rankings are real, but they're ranking a very short list.
Top comments (0)