<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Igor Eduardo</title>
    <description>The latest articles on DEV Community by Igor Eduardo (@nomad-link-id).</description>
    <link>https://dev.to/nomad-link-id</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3929927%2F8c3a1bbb-7afd-475e-baad-ed984b036b51.png</url>
      <title>DEV Community: Igor Eduardo</title>
      <link>https://dev.to/nomad-link-id</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nomad-link-id"/>
    <language>en</language>
    <item>
      <title>Done Is Testimony. Terminal State Is the Grade.</title>
      <dc:creator>Igor Eduardo</dc:creator>
      <pubDate>Mon, 05 Oct 2026 15:54:00 +0000</pubDate>
      <link>https://dev.to/nomad-link-id/done-is-testimony-terminal-state-is-the-grade-og8</link>
      <guid>https://dev.to/nomad-link-id/done-is-testimony-terminal-state-is-the-grade-og8</guid>
      <description>&lt;p&gt;Two things landed in the same week and they are the same argument.&lt;/p&gt;

&lt;p&gt;Microsoft and Hugging Face published &lt;a href="https://huggingface.co/blog/microsoft/thinkingbox" rel="noopener noreferrer"&gt;ThinkingBox&lt;/a&gt;, a benchmark that grades agents on the backend state and side effects they leave behind instead of the sentences they generate, and then asks whether they can do it twenty times in a row. Their opening example is an agent that makes nine clean tool calls, closes a ticket as resolved, and is wrong: the carrier exception was still open, so the required end state was "on hold." A grader reading tool calls sees nine well-formed calls. The database disagrees.&lt;/p&gt;

&lt;p&gt;On Dev.to, &lt;a href="https://dev.to/james_anderson_h/the-witness-was-the-suspect-why-ai-audit-logs-cant-be-trusted-2190"&gt;The Witness Was the Suspect&lt;/a&gt; made the same point from the incident side: when the agent writes its own success log, the record of the incident was authored by the cause of the incident.&lt;/p&gt;

&lt;p&gt;Here is what I would require before I trusted an agent's "done," and what I would not claim yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stake
&lt;/h2&gt;

&lt;p&gt;Teams will ship the wrong default.&lt;/p&gt;

&lt;p&gt;Most agent dashboards still answer "did it say it succeeded?" Status strings, tool-call traces, and a tidy final message are all produced by the same run that took the action. They are useful for debugging. They are not a grade. When the eval and the agent share a witness, a confident wrong action scores the same as a correct one, and the failure shows up later as a customer, a ledger, or an auditor.&lt;/p&gt;

&lt;p&gt;AI failures rarely fail loud. They fail plausible. That is exactly why the grade has to come from somewhere the agent cannot write.&lt;/p&gt;

&lt;h2&gt;
  
  
  Opinion (one sentence)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;An agent's "done" is testimony, not evidence. The eval contract should grade the terminal state through a read path the agent does not control, and it should grade it more than once.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is a preference with a limiter, not a benchmark result. I have not run ThinkingBox, and its numbers (507 stateful workflows, 20 runs each) are theirs. What I am formalizing is the shape: separate the actor from the grader, and treat "said done while the state is red" as a named failure, not as noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  What separating the witness looks like (menu, not recipe)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Question it answers&lt;/th&gt;
&lt;th&gt;Failure if skipped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Required end state&lt;/td&gt;
&lt;td&gt;What should the record look like when the task is truly finished?&lt;/td&gt;
&lt;td&gt;"Resolved" accepted when "on hold" was correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Independent read path&lt;/td&gt;
&lt;td&gt;Who reads the final state, and can the agent influence that read?&lt;/td&gt;
&lt;td&gt;The agent's own log grades the agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Side-effect ledger&lt;/td&gt;
&lt;td&gt;What else changed (tickets, refunds, emails, rows)?&lt;/td&gt;
&lt;td&gt;Correct primary record, wrong collateral writes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Done-while-red&lt;/td&gt;
&lt;td&gt;Did the run claim success while the state check failed?&lt;/td&gt;
&lt;td&gt;Confident wrong actions blend into the pass rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeat consistency&lt;/td&gt;
&lt;td&gt;Does the same task reach the same end state across N runs?&lt;/td&gt;
&lt;td&gt;One lucky green run treated as a capability&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of that needs a private harness. It needs someone to write down the required end state before the run, and a grader that reads the system of record after the run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why repeats belong in the contract
&lt;/h2&gt;

&lt;p&gt;ThinkingBox's "twenty times in a row" framing is the part I would steal first.&lt;/p&gt;

&lt;p&gt;A single green run tells you the agent &lt;em&gt;can&lt;/em&gt; reach the right state. Production needs to know whether it &lt;em&gt;will&lt;/em&gt;. If the same ticket ends "resolved" on some runs and "on hold" on others, you do not have a capability with a small error rate. You have a coin with a nice transcript. I would report consistency next to the pass rate, not bury it in an appendix.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this connects to retrieval
&lt;/h2&gt;

&lt;p&gt;My public work is mostly retrieve-first systems, and the same split shows up there. A RAG answer can cite cleanly and still be wrong about the source, the same way an agent can log cleanly and still leave the wrong record. In both cases the fix is the same move: grade against something the generator did not produce. For retrieval that is a gold source set; for agents it is the terminal state.&lt;/p&gt;

&lt;p&gt;The hybrid retrieval pipeline I keep public (&lt;a href="https://github.com/nomad-link-id/hybrid-rag-pipeline" rel="noopener noreferrer"&gt;hybrid-rag-pipeline&lt;/a&gt;) is built around that habit: measure the retrieval half on its own before letting a fluent answer vouch for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Steal this check
&lt;/h2&gt;

&lt;p&gt;Pick one agent workflow you already run and add three things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Write the required end state first.&lt;/strong&gt; Not "the agent responds," but "the ticket is on hold, no refund issued, customer told why."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Grade from the system of record.&lt;/strong&gt; Read the final state with a query the agent never touches, and ignore the agent's own status message for scoring.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count done-while-red separately.&lt;/strong&gt; Every run where the agent claimed success but the state check failed goes in its own column. That column is the one that becomes an incident.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then run it more than once and report how often it lands in the same place.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I am not claiming
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;I am not reporting ThinkingBox results or reproducing their benchmark.&lt;/li&gt;
&lt;li&gt;I am not saying traces and logs are useless. They are how you debug. They are just the wrong thing to score.&lt;/li&gt;
&lt;li&gt;I am not publishing thresholds. What counts as acceptable consistency depends on what the action costs when it is wrong.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The point
&lt;/h2&gt;

&lt;p&gt;The process that acted should not be the only witness that it worked. Grade the record, not the report.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Igor Eduardo builds retrieve-first and evaluation-as-contract systems. More at &lt;a href="https://igoreduardo.com" rel="noopener noreferrer"&gt;igoreduardo.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>evaluation</category>
      <category>testing</category>
    </item>
    <item>
      <title>Green Scores Without a Pinned Harness Are Screenshots</title>
      <dc:creator>Igor Eduardo</dc:creator>
      <pubDate>Mon, 28 Sep 2026 16:09:46 +0000</pubDate>
      <link>https://dev.to/nomad-link-id/green-scores-without-a-pinned-harness-are-screenshots-3666</link>
      <guid>https://dev.to/nomad-link-id/green-scores-without-a-pinned-harness-are-screenshots-3666</guid>
      <description>&lt;p&gt;Everyone is talking about Evaluation Cards and reproducible eval benches — the EvalEval × UK AISI write-up on Hugging Face, plus a week of discourse where green numbers still hide the wrong harness (and DeepEval-style “which half is broken?” chatter keeps landing). That heat is a live market object, not a niche footnote.&lt;/p&gt;

&lt;p&gt;Here is what I would require before I trusted a green eval score for high-stakes retrieval or generation — and what I would not claim yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stake
&lt;/h2&gt;

&lt;p&gt;Teams will ship the wrong default.&lt;/p&gt;

&lt;p&gt;A dashboard can show a comforting pass rate while the report never pins &lt;em&gt;how&lt;/em&gt; the score was earned: which protocol, which compute budget, which dataset slice, and whether the number measured retrieval, generation, or a blended soup. Celebrate that number and you optimize a screenshot. The next engineer cannot reproduce the gate, and the next incident cannot tell which half failed.&lt;/p&gt;

&lt;p&gt;If your domain is identifier-heavy, multi-hop, or regulated-adjacent, a fluent answer after an unpinned harness is not a win. It is a silent miss with a pretty badge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Opinion (one sentence)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A green score without a pinned harness — protocol, compute, dataset slice, and which half of the pipeline it measures — is not an eval contract; it is a screenshot.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is a preference with a limiter, not a SOTA claim. I have not run EvalEval in production. AISI card fields and anyone else’s leaderboard numbers stay theirs. What I am formalizing is the &lt;em&gt;harness pin&lt;/em&gt; as a first-class part of eval-as-contract, from the same reliable-systems lane I already publish on.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a pinned harness looks like (menu, not recipe)
&lt;/h2&gt;

&lt;p&gt;Evaluation Cards earned attention because they make the boring fields visible: protocol, compute, data cut, scoring rules. You do not need their exact schema to steal the &lt;em&gt;shape&lt;/em&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Question it answers&lt;/th&gt;
&lt;th&gt;Failure if skipped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Protocol&lt;/td&gt;
&lt;td&gt;What procedure produced this number?&lt;/td&gt;
&lt;td&gt;“We eval’d it” with no replay path&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compute&lt;/td&gt;
&lt;td&gt;What budget / hardware / call limits?&lt;/td&gt;
&lt;td&gt;Incomparable scores across teams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dataset slice&lt;/td&gt;
&lt;td&gt;Which queries / docs / holdouts?&lt;/td&gt;
&lt;td&gt;Student-subset vibes framed as prod&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pipeline half&lt;/td&gt;
&lt;td&gt;Retrieval, generation, or joint?&lt;/td&gt;
&lt;td&gt;Green faithfulness while recall is broken (or the reverse)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The DeepEval-style lesson that keeps circulating is the same shape: a single RAG score can hide &lt;em&gt;which half&lt;/em&gt; failed. Pin the half. Prefer two honest columns over one blended crown.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check you can run without my internals
&lt;/h2&gt;

&lt;p&gt;You do not need my private stacks (and I will not publish them). Steal this menu-only check:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pin the harness block before the score.&lt;/strong&gt; Same report header every time: protocol name/version, compute class, dataset slice id, and pipeline half (retrieval / generation / joint). If any field is missing, the number is not ready to celebrate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split halves when you report.&lt;/strong&gt; For the same query set, publish retrieval quality and generation/faithfulness as separate columns. One green blended score is not enough — “which half is broken?” should be answerable from the table alone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Freeze the evaluator before you chase labels.&lt;/strong&gt; If the judge, metric, or prompt changed mid-board, say so. Moving the goalposts mid-week and claiming progress is theater.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a missing-evidence arm.&lt;/strong&gt; Same task with the gold document removed (or retrieval forced empty). If generation still narrates a confident answer, the eval contract is incomplete — regardless of how green the happy-path column looks.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No thresholds copied from anyone’s blog as “our production numbers.” No paste recipe. Calibrate cutoffs on &lt;em&gt;your&lt;/em&gt; corpus; where missing evidence is costly, treat empty retrieval as a failure in the harness, not as a soft success.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why query-class honesty transfers
&lt;/h2&gt;

&lt;p&gt;On Portuguese clinical text, BM25 and dense retrieval solved different query classes; fusing them beat either alone on a public 500-query study. Exact terms and conceptual phrasing fail in different ways. That finding is checkable: open code at &lt;a href="https://github.com/nomad-link-id/hybrid-rag-pipeline" rel="noopener noreferrer"&gt;&lt;code&gt;nomad-link-id/hybrid-rag-pipeline&lt;/code&gt;&lt;/a&gt;, companion write-up on Dev.to.&lt;/p&gt;

&lt;p&gt;The lesson that transfers to Evaluation Cards is not “clone our fusion.” It is &lt;strong&gt;eval honesty on query classes and pipeline halves&lt;/strong&gt;. If your harness cannot say &lt;em&gt;which class&lt;/em&gt; and &lt;em&gt;which half&lt;/em&gt; produced the green cell, you are not ready to ship the number into a decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I will not claim
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;That EvalEval, UK AISI cards, or any one vendor harness is universally correct for every workload. Workload-honest eval first.&lt;/li&gt;
&lt;li&gt;That we “ran EvalEval inside Cortexa/DocMinds.” We did not. Trend-jacking with fake usage is spam.&lt;/li&gt;
&lt;li&gt;That a pinned harness alone proves end-to-end answer quality. Downstream faithfulness and missing-doc policy remain separate gates.&lt;/li&gt;
&lt;li&gt;Internal thresholds, prompts, or clone guides for private products. Menu only: pin the fields, split the halves — not the spice blend.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why this formalizes who we are
&lt;/h2&gt;

&lt;p&gt;The market is amplifying Evaluation Cards and green eval dashboards. The thesis I want indexed next to that heat:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Eval as contract&lt;/strong&gt; — protocol, compute, slice, and pipeline half are gates, not vibes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliable systems&lt;/strong&gt; — production AI fails on harness mismatch and wrong-half celebration — not only on model choice.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieve-first&lt;/strong&gt; — when the retrieval half is unpinned, generation fluency is not evidence.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If a sharp reader walks away believing &lt;em&gt;this engineer will not celebrate a green score without a pinned harness&lt;/em&gt;, the post did its job. Follows that come from that are the point — not volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  Soft pointers (contribution first)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Public hybrid pipeline (BM25 + dense + RRF): &lt;a href="https://github.com/nomad-link-id/hybrid-rag-pipeline" rel="noopener noreferrer"&gt;https://github.com/nomad-link-id/hybrid-rag-pipeline&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Empirical companion (500 clinical queries, open methodology): &lt;a href="https://dev.to/nomad-link-id/two-retrieval-methods-are-better-than-one-evidence-from-500-clinical-queries-4g41"&gt;Dev.to — Two Retrieval Methods Are Better Than One&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Site / builder trail: &lt;a href="https://igoreduardo.com" rel="noopener noreferrer"&gt;https://igoreduardo.com&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Market object that sparked this note: &lt;a href="https://huggingface.co/blog/evaleval-aisi" rel="noopener noreferrer"&gt;EvalEval × UK AISI Evaluation Cards (HF)&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Stop here. No recipe. No “we proved SOTA.” Preference + limiter + public trail.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;By Igor Eduardo · Austin, TX · Engineer of reliable search and AI systems for high-stakes science · &lt;a href="https://igoreduardo.com" rel="noopener noreferrer"&gt;https://igoreduardo.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>evaluation</category>
      <category>rag</category>
      <category>ai</category>
      <category>llm</category>
    </item>
    <item>
      <title>Ranking Isn't Judging: What the Jev Wave Still Owes Retrieval</title>
      <dc:creator>Igor Eduardo</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:58:50 +0000</pubDate>
      <link>https://dev.to/nomad-link-id/ranking-isnt-judging-what-the-jev-wave-still-owes-retrieval-301m</link>
      <guid>https://dev.to/nomad-link-id/ranking-isnt-judging-what-the-jev-wave-still-owes-retrieval-301m</guid>
      <description>&lt;h1&gt;
  
  
  Ranking Isn't Judging: What the Jev Wave Still Owes Retrieval
&lt;/h1&gt;

&lt;p&gt;Everyone is talking about typed judgment models for RAG — Jev-class scorers that keep or drop passages, not only reorder them. Hugging Face just got a clean write-up of &lt;code&gt;jev-reranker&lt;/code&gt;: hybrid pool in, relevance filter out, fewer documents reaching the generator. That is a live market object, not a niche blog.&lt;/p&gt;

&lt;p&gt;Here is what I would require before I trusted “we upgraded ranking to judging” for high-stakes retrieval — and what I would not claim yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stake
&lt;/h2&gt;

&lt;p&gt;Teams will ship the wrong default.&lt;/p&gt;

&lt;p&gt;A hybrid lexical + dense pool can look fine on ranking metrics while still shipping distracting neighbors into the prompt. Rerank asks &lt;em&gt;in what order?&lt;/em&gt; Relevance filtering asks &lt;em&gt;should this context reach the generator at all?&lt;/em&gt; Those are different contracts. Collapsing them into one score hides the broken half.&lt;/p&gt;

&lt;p&gt;If your domain is identifier-heavy, multi-hop, or regulated-adjacent, a fluent answer after a weak retained set is not a win. It is a silent miss.&lt;/p&gt;

&lt;h2&gt;
  
  
  Opinion (one sentence)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Ranking answers closeness; judging needs an explicit leave-out policy plus a faithfulness gate that is scored separately — never a single blended “quality” number.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is a preference with a limiter, not a SOTA claim. I have not run Jev in production. Their NanoHotpotQA numbers stay theirs. What I am formalizing is the &lt;em&gt;judgment type&lt;/em&gt; split, from the same retrieve-first / eval-as-contract lane I already publish on.&lt;/p&gt;

&lt;h2&gt;
  
  
  What ranking gets right — and where it stops
&lt;/h2&gt;

&lt;p&gt;Hybrid retrieval earned its keep for a reason. On Portuguese clinical text, BM25 and dense retrieval solved different query classes; fusing them beat either alone on a public 500-query study. Exact terms (scores, drug names, identifiers) and conceptual phrasing fail in different ways. That finding is checkable: open code at &lt;a href="https://github.com/nomad-link-id/hybrid-rag-pipeline" rel="noopener noreferrer"&gt;&lt;code&gt;nomad-link-id/hybrid-rag-pipeline&lt;/code&gt;&lt;/a&gt;, companion write-up on Dev.to, Zenodo preprint under CC BY.&lt;/p&gt;

&lt;p&gt;Reranking on top of that pool is a natural next step. Cross-encoders and decision models both try to push the useful passages up. Fine — as a &lt;em&gt;ranking&lt;/em&gt; layer.&lt;/p&gt;

&lt;p&gt;Judging is different work:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Judgment&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Failure if skipped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Rank&lt;/td&gt;
&lt;td&gt;Which of these hits is closer?&lt;/td&gt;
&lt;td&gt;Wrong order; still often has &lt;em&gt;something&lt;/em&gt; in context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Filter / leave-out&lt;/td&gt;
&lt;td&gt;Should this hit reach the generator?&lt;/td&gt;
&lt;td&gt;Distractors in the prompt; token waste&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Faithfulness&lt;/td&gt;
&lt;td&gt;Does the final answer’s claim live in the cited source?&lt;/td&gt;
&lt;td&gt;Fluent invention with a real cite list&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A leave-out policy is not “sort harder.” When the filter clears the deck, the system needs a first-class &lt;strong&gt;missing-evidence&lt;/strong&gt; outcome — not a silent fallback into guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check you can run without my internals
&lt;/h2&gt;

&lt;p&gt;You do not need my private stacks (and I will not publish them). Steal this menu-only check:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Split the report.&lt;/strong&gt; For the same query set, publish (a) ranking quality on the candidate pool and (b) retained-set size / leave-out rate after the filter. One number is not enough — the HF &lt;code&gt;jev-reranker&lt;/code&gt; table that pairs nDCG with documents retained is the right &lt;em&gt;shape&lt;/em&gt; of honesty, independent of their thresholds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pin a missing-evidence arm.&lt;/strong&gt; Same task with the filter forced to retain nothing (or with the gold document removed from the pool). If the generator still narrates a confident answer, the eval contract is incomplete.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep faithfulness separate.&lt;/strong&gt; For every quoted span, verify membership against the &lt;em&gt;cited&lt;/em&gt; source id. Fail closed on miss. Ranking/relevance can still look fine while the quote-span check fails — I have seen that failure mode in public discourse this week, and it matches retrieve-first practice.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No thresholds copied from anyone’s blog as “our production numbers.” No paste recipe. Calibrate cutoffs on &lt;em&gt;your&lt;/em&gt; corpus; where missing evidence is costly, treat empty retained sets as failures in the harness, not as soft successes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hybrid pools make the split sharper
&lt;/h2&gt;

&lt;p&gt;Once you run BM25 and dense in parallel, the candidate set is &lt;em&gt;intentionally&lt;/em&gt; broad. That is a feature: complementary methods catch different query classes. It is also why a leave-out layer matters more, not less.&lt;/p&gt;

&lt;p&gt;A wide hybrid pool without a filter ships more distractors. A filter without a missing-evidence policy ships fluent guesses when nothing survived. Rerank alone does not fix either failure — it only reorders the same set.&lt;/p&gt;

&lt;p&gt;So when the market says “just add a judgment model,” I hear two separate jobs: upgrade the &lt;em&gt;order&lt;/em&gt; when you still want generation, and upgrade the &lt;em&gt;admit/deny&lt;/em&gt; policy when generation should be allowed to refuse. Publish both outcomes in the eval report. Prefer workload-honest tables over a universal crown for any one store, model, or scorer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I will not claim
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;That Jev (or any typed decision model) is universally better than a cross-encoder or a lexical filter. Workload-honest eval first.&lt;/li&gt;
&lt;li&gt;That filter quality alone proves end-to-end answer quality. Downstream faithfulness and missing-doc policy are separate gates.&lt;/li&gt;
&lt;li&gt;That we “shipped Jev inside Cortexa/DocMinds.” We did not. Trend-jacking with fake usage is spam.&lt;/li&gt;
&lt;li&gt;Internal thresholds, prompts, or clone guides for private products. Menu only: what we serve is the judgment split and the eval shape — not the spice blend.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why this formalizes who we are
&lt;/h2&gt;

&lt;p&gt;The market is amplifying judgment models. The thesis I want indexed next to that heat:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Retrieve-first&lt;/strong&gt; — evidence path before orchestration theater.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hybrid / complementary search&lt;/strong&gt; — lexical and dense cover different query classes; fusion needs honest eval, not a universal crown.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Eval as contract&lt;/strong&gt; — exact-match, faithfulness, context precision/recall, paired leave-out tests — gates, not vibes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If a sharp reader walks away believing &lt;em&gt;this engineer will not collapse ranking into judging, and will fail closed when evidence is missing&lt;/em&gt;, the post did its job. Follows that come from that are the point — not volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  Soft pointers (contribution first)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Public hybrid pipeline (BM25 + dense + RRF): &lt;a href="https://github.com/nomad-link-id/hybrid-rag-pipeline" rel="noopener noreferrer"&gt;https://github.com/nomad-link-id/hybrid-rag-pipeline&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Empirical companion (500 clinical queries, open methodology): &lt;a href="https://dev.to/nomad-link-id/two-retrieval-methods-are-better-than-one-evidence-from-500-clinical-queries-4g41"&gt;Dev.to — Two Retrieval Methods Are Better Than One&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Preprint (CC BY): &lt;a href="https://doi.org/10.5281/zenodo.19686739" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.19686739&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Market object that sparked this note: &lt;a href="https://huggingface.co/blog/hotchpotch/introducing-jev-reranker" rel="noopener noreferrer"&gt;Introducing jev-reranker (HF)&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Stop here. No recipe. No “we proved SOTA.” Preference + limiter + public trail.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;By Igor Eduardo · Austin, TX · Engineer of reliable search and AI systems for high-stakes science · &lt;a href="https://igoreduardo.com" rel="noopener noreferrer"&gt;https://igoreduardo.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>rag</category>
      <category>evaluation</category>
      <category>ai</category>
    </item>
    <item>
      <title>Two Retrieval Methods Are Better Than One: Evidence from 500 Clinical Queries</title>
      <dc:creator>Igor Eduardo</dc:creator>
      <pubDate>Wed, 13 May 2026 19:14:30 +0000</pubDate>
      <link>https://dev.to/nomad-link-id/two-retrieval-methods-are-better-than-one-evidence-from-500-clinical-queries-4g41</link>
      <guid>https://dev.to/nomad-link-id/two-retrieval-methods-are-better-than-one-evidence-from-500-clinical-queries-4g41</guid>
      <description>&lt;p&gt;When I set out to evaluate retrieval configurations for Portuguese clinical text, I expected one method to dominate. Instead, I found something more interesting: BM25 and dense retrieval solve &lt;em&gt;different&lt;/em&gt; questions. Neither is a substitute for the other.&lt;/p&gt;

&lt;p&gt;This post summarizes the methodology and results from a 500-query empirical study of hybrid retrieval for clinical question answering. All code is open source: &lt;a href="https://github.com/nomad-link-id/hybrid-rag-pipeline" rel="noopener noreferrer"&gt;https://github.com/nomad-link-id/hybrid-rag-pipeline&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup
&lt;/h2&gt;

&lt;p&gt;500 clinical queries across 6 medical specialties (cardiology, endocrinology, infectology, nephrology, neurology, oncology). Each query has a single reference answer grounded in a specific passage from clinical documentation.&lt;/p&gt;

&lt;p&gt;Four retrieval configurations were evaluated:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;BM25-only&lt;/td&gt;
&lt;td&gt;BM25 with Portuguese stopword removal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dense-only&lt;/td&gt;
&lt;td&gt;BioBERTpt embeddings, cosine similarity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hybrid-RRF&lt;/td&gt;
&lt;td&gt;BM25 + dense via Reciprocal Rank Fusion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hybrid-Rerank&lt;/td&gt;
&lt;td&gt;RRF candidates re-ranked with cross-encoder&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What Is Reciprocal Rank Fusion?
&lt;/h2&gt;

&lt;p&gt;RRF combines ranked lists from multiple retrievers without requiring score normalization:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;rrf_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rankings&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]],&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;ranking&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rankings&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;doc_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ranking&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;sorted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;scores&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;reverse&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;Recall@5&lt;/th&gt;
&lt;th&gt;MRR&lt;/th&gt;
&lt;th&gt;Citation F1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;BM25-only&lt;/td&gt;
&lt;td&gt;0.71&lt;/td&gt;
&lt;td&gt;0.64&lt;/td&gt;
&lt;td&gt;0.82&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dense-only&lt;/td&gt;
&lt;td&gt;0.68&lt;/td&gt;
&lt;td&gt;0.61&lt;/td&gt;
&lt;td&gt;0.78&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hybrid-RRF&lt;/td&gt;
&lt;td&gt;0.84&lt;/td&gt;
&lt;td&gt;0.77&lt;/td&gt;
&lt;td&gt;0.91&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hybrid-Rerank&lt;/td&gt;
&lt;td&gt;0.86&lt;/td&gt;
&lt;td&gt;0.79&lt;/td&gt;
&lt;td&gt;0.93&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The Complementarity Finding
&lt;/h2&gt;

&lt;p&gt;McNemar's test on BM25-only versus dense-only:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;BM25 correct, dense incorrect: 89 queries&lt;/li&gt;
&lt;li&gt;Dense correct, BM25 incorrect: 57 queries&lt;/li&gt;
&lt;li&gt;McNemar chi2 = 39.55, p &amp;lt; 0.001&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The asymmetry is statistically significant. Dense-only missed 22.2% of queries that BM25 solved. You need both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Citation Verification
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Deterministic approach&lt;/strong&gt; (BM25 score threshold + exact n-gram overlap): 461/500 citations verified.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt-based LLM approach&lt;/strong&gt; (same passages, ask LLM "does this support the answer?"): 1/500.&lt;/p&gt;

&lt;p&gt;The difference is task design, not model quality. A deterministic check measures actual textual overlap; a prompt check measures the model's opinion of the overlap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inter-Annotator Agreement
&lt;/h2&gt;

&lt;p&gt;100 query-response pairs independently annotated by two reviewers. Cohen's kappa = 0.954 — near-perfect agreement on what constitutes correct retrieval for clinical text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Takeaway
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Run both BM25 and dense retrieval&lt;/li&gt;
&lt;li&gt;Use RRF to merge results&lt;/li&gt;
&lt;li&gt;Implement deterministic citation verification&lt;/li&gt;
&lt;li&gt;Measure complementarity with McNemar's test on your domain&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Code and Data
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/nomad-link-id/hybrid-rag-pipeline" rel="noopener noreferrer"&gt;https://github.com/nomad-link-id/hybrid-rag-pipeline&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/nomad-link-id/citation-guard" rel="noopener noreferrer"&gt;https://github.com/nomad-link-id/citation-guard&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Preprint: &lt;a href="https://doi.org/10.5281/zenodo.19686739" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.19686739&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Igor Eduardo | igoreduardo.com | ORCID: 0009-0005-6288-1135&lt;/p&gt;

</description>
      <category>python</category>
      <category>rag</category>
      <category>ai</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
