<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hardcore Engineer</title>
    <description>The latest articles on DEV Community by Hardcore Engineer (@kaziava).</description>
    <link>https://dev.to/kaziava</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4120551%2Fc4b02a5c-476e-4df2-a30d-a6a1c02d6bd6.png</url>
      <title>DEV Community: Hardcore Engineer</title>
      <link>https://dev.to/kaziava</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kaziava"/>
    <language>en</language>
    <item>
      <title>The Negative Test That Passed for the Wrong Reason</title>
      <dc:creator>Hardcore Engineer</dc:creator>
      <pubDate>Sat, 19 Sep 2026 15:49:30 +0000</pubDate>
      <link>https://dev.to/kaziava/the-negative-test-that-passed-for-the-wrong-reason-3koh</link>
      <guid>https://dev.to/kaziava/the-negative-test-that-passed-for-the-wrong-reason-3koh</guid>
      <description>&lt;p&gt;Last february we swapped our embedding model from text-embedding-3-small to bge-large, thinking "better embedder, better RAG." our golden set stayed green. two negative tests that had been failing for months suddenly passed. looked like a win.&lt;/p&gt;

&lt;p&gt;When the retriever fails to pull the relevant chunks, the LLM never sees the trap. if it's a polite model, it refuses to answer, and your test passes. but it passed for the wrong reason: the retrieval failed, not the model. the test looked green while the system was&lt;br&gt;
broken.&lt;/p&gt;

&lt;p&gt;This is the same shape as Serhiy Kucherenko's finding that "the corpus decides more than the model does," just quieter. his refusal traps worked because the prompt boundary held. ours failed silently because the retrieval boundary broke.&lt;/p&gt;
&lt;h2&gt;
  
  
  The fix: trap_chunk_id and embedder tags
&lt;/h2&gt;

&lt;p&gt;Every negative test in our golden set now carries two stamps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;trap_chunk_id&lt;/strong&gt;: the id of the chunk containing the trap, stamped at authoring time. our CI step fails the run if that chunk isn't in the retrieved set, regardless of what the LLM says. "test did not run" becomes a third verdict next to pass and fail.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;embedder tag&lt;/strong&gt;: the name of the embedding model under which the test was last validated. after any swap, lines with the old tag are stale by default — they stay in the file, but the CI summary flags them as "unverified under current embedder" until someone re-runs and re-dates them.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The mechanics are dumb simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# golden_set.json
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;neg_17&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;question&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;what was the q3 2019 revenue?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expected_answer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;not in documents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trap_chunk_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk_847&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;embedder_tag&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text-embedding-3-small-2024-02-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;last_validated&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2024-02-15&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# ci_eval.py
&lt;/span&gt;&lt;span class="n"&gt;retrieved_chunk_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;run_retrieval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;trap_chunk_id&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;retrieved_chunk_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;verdict&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;test did not run&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;llm_answer&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;expected_answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;verdict&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pass&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;verdict&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When we swapped to bge-large in february, the two negative tests that "passed" immediately turned red: trap_chunk_id wasn't in the retrieved set. the new embedder was worse at semantic similarity for those specific traps. we rolled back the swap, fixed the chunking strategy, and only then re-validated the tests under the new embedder.&lt;/p&gt;

&lt;h2&gt;
  
  
  The maintenance cost nobody talks about
&lt;/h2&gt;

&lt;p&gt;The catch is real: trap_chunk_id moves whenever you re-chunk or re-parse documents. every pipeline change triggers a re-stamping step:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;re-run the full document ingestion&lt;/li&gt;
&lt;li&gt;for each negative test, find the new chunk id containing the trap&lt;/li&gt;
&lt;li&gt;update the golden set with the new trap_chunk_id&lt;/li&gt;
&lt;li&gt;re-run eval under the current embedder&lt;/li&gt;
&lt;li&gt;update the embedder_tag and last_validated timestamp&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It's maybe 20 minutes of work per pipeline change, but it's non-negotiable. skip it and you're trusting verdicts that describe a retrieval that no longer exists.&lt;/p&gt;

&lt;p&gt;Same goes for the embedder tag: after any swap, every line with the old tag is stale. our CI summary now has a section called "unverified under current embedder" that lists all the tests waiting for re-validation. cheap bookkeeping, but it stopped us from shipping broken retrieval twice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why cosine similarity can't save you
&lt;/h2&gt;

&lt;p&gt;Amritpal singh wrote about this exact problem: his refusal threshold had no effect because the similarity distributions for answerable and unanswerable questions overlapped. cosine similarity can't tell "about the topic" from "answers the question." both get high scores.&lt;/p&gt;

&lt;p&gt;Our trap_chunk_id approach sidesteps the overlap problem entirely. instead of trying to find a threshold that separates the distributions (which doesn't exist), we assert that the specific chunk containing the trap must be retrieved. if it's not, the test didn't run.&lt;/p&gt;

&lt;p&gt;This is the same insight as hybrid search: single-vector similarity will always fail on negatives because high similarity ≠ correct answer. BM25 catches exact matches, vectors catch semantic similarity, metadata filters catch structural context. but for negative tests, none of that matters if the trap chunk isn't retrieved at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  The checklist
&lt;/h2&gt;

&lt;p&gt;If you're running negative tests in your RAG eval:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt; every negative test carries a trap_chunk_id&lt;/li&gt;
&lt;li&gt; CI fails if that chunk isn't retrieved (verdict: "test did not run")&lt;/li&gt;
&lt;li&gt; every test carries an embedder_tag from when it was last validated&lt;/li&gt;
&lt;li&gt; after any embedder swap, tests with the old tag are flagged as stale&lt;/li&gt;
&lt;li&gt; re-chunking triggers a re-stamping of all trap_chunk_ids&lt;/li&gt;
&lt;li&gt; the CI summary shows "unverified under current embedder" tests prominently&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your eval doesn't have these, you're probably shipping retrieval bugs that look like model bugs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The broader lesson
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://dev.to/izgorodin/how-can-i-prevent-my-ai-coding-assistant-from-repeating-fixed-mistakes-across-sessions-2kf7"&gt;Edward Izgorodin's&lt;/a&gt; framing is the sharpest I've seen: "writing the correction down is not the fix, because the memory that carried the mistake is still in the store and still eligible." But there's a subtler failure mode: &lt;strong&gt;evaluation fails without observability into whether the test actually ran&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A negative test that passes because the LLM refused is not evidence that the retrieval worked. it's evidence that the LLM is polite. the only way to know if the retrieval worked is to check whether the trap chunk was actually retrieved.&lt;/p&gt;

&lt;p&gt;This is the same shape as observability in production: you don't trust that a service is healthy because it returns 200 OK. you check that it actually did the work. same logic applies to eval.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where to go deeper
&lt;/h2&gt;

&lt;p&gt;This post is the mechanics. if you want the full audit checklist for production RAG (parsing, chunking, metadata, hybrid search, reranking, eval), i packaged it as a PDF: &lt;strong&gt;"RAG checklist: 10 checks to stop LLM hallucinations on tables and numbers"&lt;/strong&gt; plus a &lt;strong&gt;golden set template with 20 questions&lt;/strong&gt; covering digits, contracts, dates, comparisons, negatives, and paraphrases.&lt;/p&gt;

&lt;p&gt;Both are free in my telegram channel: &lt;strong&gt;&lt;a href="https://t.me/llmops_engineering" rel="noopener noreferrer"&gt;@llmops_engineering&lt;/a&gt;&lt;/strong&gt; — comment "+" on the pinned post and i'll send you both PDFs. no email signup, no spam, just the files.&lt;/p&gt;

&lt;p&gt;The channel is where i post production RAG war stories, vector DB benchmarks, and the boring LLMOps stuff that actually ships. if this post saved you a week of debugging, the channel will save you a month.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What's your eval setup for negative tests? have you hit the "test passed but retrieval was broken" failure mode? drop your approach in the comments — curious how other teams handle this.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>debugging</category>
      <category>llm</category>
      <category>rag</category>
      <category>testing</category>
    </item>
    <item>
      <title>Why your RAG hallucinates on tables (and a minimal local GraphRAG starter to fix it)</title>
      <dc:creator>Hardcore Engineer</dc:creator>
      <pubDate>Fri, 11 Sep 2026 09:42:03 +0000</pubDate>
      <link>https://dev.to/kaziava/why-your-rag-hallucinates-on-tables-and-a-minimal-local-graphrag-starter-to-fix-it-2e24</link>
      <guid>https://dev.to/kaziava/why-your-rag-hallucinates-on-tables-and-a-minimal-local-graphrag-starter-to-fix-it-2e24</guid>
      <description>&lt;p&gt;If you've ever built a RAG system for corporate documents, you've probably hit this wall: you ask "What was the revenue in Q3?", and the LLM confidently hallucinates a number with three extra zeros.&lt;/p&gt;

&lt;p&gt;This isn't a model bug. It's an architectural blind spot. As recent research from Microsoft points out, naive 500-token chunking destroys table structures. Cosine similarity over embeddings just doesn't understand rows and columns.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix: GraphRAG
&lt;/h2&gt;

&lt;p&gt;Instead of flat vector search, we extract entities and relationships into a Knowledge Graph. When a user asks about numbers, the LLM translates the question into a Cypher query, and the graph returns exact data. Zero hallucinated digits.&lt;/p&gt;

&lt;h2&gt;
  
  
  I open-sourced a minimal local starter
&lt;/h2&gt;

&lt;p&gt;To prove this works without sending sensitive corporate data to cloud APIs, I packaged my local GraphRAG stack into a minimal, production-oriented starter repo:&lt;br&gt;
👉 &lt;strong&gt;&lt;a href="https://github.com/kaziava/local-graphrag-starter" rel="noopener noreferrer"&gt;github.com/kaziava/local-graphrag-starter&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Stack:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;PyMuPDF&lt;/strong&gt; for local PDF parsing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LangChain&lt;/strong&gt; (&lt;code&gt;LLMGraphTransformer&lt;/code&gt;) to extract the graph&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Neo4j&lt;/strong&gt; (via docker-compose) as the graph DB&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ollama&lt;/strong&gt; (running &lt;code&gt;llama3.1:8b&lt;/code&gt; or &lt;code&gt;qwen2.5:3b&lt;/code&gt;) for local inference&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How it works:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Start Neo4j&lt;/span&gt;
docker compose up &lt;span class="nt"&gt;-d&lt;/span&gt;

&lt;span class="c"&gt;# 2. Build the graph from your PDF&lt;/span&gt;
python main.py ingest data/report.pdf

&lt;span class="c"&gt;# 3. Ask a question (NL -&amp;gt; Cypher -&amp;gt; Exact Answer)&lt;/span&gt;
python main.py ask &lt;span class="s2"&gt;"What was Apple's revenue in Q3 2024?"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Benchmarks (MacBook M2, 16GB RAM)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Parse 50-page PDF: ~45s&lt;/li&gt;
&lt;li&gt;Load graph: ~10s&lt;/li&gt;
&lt;li&gt;Answer a question: 3–5s&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost per query: $0&lt;/strong&gt; (compared to ~$15–20 for the same volume via GPT-4 API)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Honest Limitations &amp;amp; Early Stage Status
&lt;/h2&gt;

&lt;p&gt;🚧 This is an early-stage reference architecture. Local 8B models are weaker than frontier LLMs on complex multi-hop reasoning, and complex tables might still need parser tuning.&lt;/p&gt;

&lt;p&gt;If you try running it on your machine and hit any OS-specific bugs, &lt;strong&gt;please open an Issue&lt;/strong&gt; on GitHub! I'm actively maintaining it and will fix things fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Want to dive deeper into the code?
&lt;/h2&gt;

&lt;p&gt;I regularly share raw benchmark scripts, Docker configs, and architectural diagrams from my production LLMOps experience. I document this primarily in Russian on my Telegram channel (&lt;strong&gt;&lt;a href="https://t.me/llmops_engineering" rel="noopener noreferrer"&gt;@llmops_engineering&lt;/a&gt;&lt;/strong&gt;), but the code snippets, schematics, and engineering discussions are universal. Feel free to join the engineer chat there or reach out.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What graph DB are you using for your RAG setups? Let me know in the comments!&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>machinelearning</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
