<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Josh Hodgetts</title>
    <description>The latest articles on DEV Community by Josh Hodgetts (@josh_hodgetts_f9f91ff3a23).</description>
    <link>https://dev.to/josh_hodgetts_f9f91ff3a23</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4160733%2Fb715ada9-6680-4d67-82e4-6fbd1a6d218a.png</url>
      <title>DEV Community: Josh Hodgetts</title>
      <link>https://dev.to/josh_hodgetts_f9f91ff3a23</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/josh_hodgetts_f9f91ff3a23"/>
    <language>en</language>
    <item>
      <title>We removed 98.77% of an LLM’s input. Accuracy went up.</title>
      <dc:creator>Josh Hodgetts</dc:creator>
      <pubDate>Mon, 05 Oct 2026 23:50:37 +0000</pubDate>
      <link>https://dev.to/josh_hodgetts_f9f91ff3a23/we-removed-9877-of-an-llms-input-accuracy-went-up-1dof</link>
      <guid>https://dev.to/josh_hodgetts_f9f91ff3a23/we-removed-9877-of-an-llms-input-accuracy-went-up-1dof</guid>
      <description>&lt;p&gt;The obvious fix for an AI app missing information is to give it more context.&lt;/p&gt;

&lt;p&gt;But what happens when the context contains three versions of the truth?&lt;/p&gt;

&lt;p&gt;An old price. A replacement price. A correction entered today that applies to last week.&lt;/p&gt;

&lt;p&gt;All three can be relevant to the question. Only some belong in the answer.&lt;/p&gt;

&lt;p&gt;I’m the founder of Jylus. We ran a frozen 528-question benchmark comparing three ways of supplying evidence to the same model.&lt;/p&gt;

&lt;p&gt;Here’s what we observed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evidence supplied&lt;/th&gt;
&lt;th&gt;Average input tokens&lt;/th&gt;
&lt;th&gt;Strict accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Full context&lt;/td&gt;
&lt;td&gt;205,129&lt;/td&gt;
&lt;td&gt;78.79%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BM25-only RAG&lt;/td&gt;
&lt;td&gt;8,128&lt;/td&gt;
&lt;td&gt;76.14%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Jylus Context Pack&lt;/td&gt;
&lt;td&gt;2,532&lt;/td&gt;
&lt;td&gt;100% observed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same Gemini 3.1 Flash Lite settings. Same questions. Same deterministic scorer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The BM25 row is the interesting part.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It used much less input than full context, but accuracy fell. Reducing tokens alone didn’t solve the problem.&lt;/p&gt;

&lt;p&gt;The Jylus path used 98.77% less input than full context and scored higher on this workload. That does not isolate which part of evidence preparation caused the improvement, but it gives us something concrete to investigate.&lt;/p&gt;

&lt;p&gt;Consider this small example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Record A
Event time: Wednesday, 09:00
Recorded: Wednesday, 09:05
Observation: Temperature warning triggered.

Record B
Event time: Wednesday, 09:00
Recorded: Friday, 10:00
Correction: Wednesday's warning was caused by a faulty sensor.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now ask:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What caused Wednesday’s warning, using everything we know now?&lt;/li&gt;
&lt;li&gt;What did we know about the warning on Wednesday?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first answer can use Friday’s correction.&lt;/p&gt;

&lt;p&gt;The second must preserve the uncertainty that existed on Wednesday. It cannot quietly borrow knowledge from Friday.&lt;/p&gt;

&lt;p&gt;Returning both records is useful, but it leaves the model with another job: deciding which evidence is admissible for the question.&lt;/p&gt;

&lt;p&gt;That distinction matters in support histories, changing subscriptions, incident investigations and any application where a later update changes how an earlier event should be understood.&lt;/p&gt;

&lt;p&gt;Jylus sits between your data and your model. It retrieves evidence, resolves state and relationships, and compiles a bounded Context Pack with source references, retaining relevant conflicts and gaps.&lt;/p&gt;

&lt;p&gt;The model then reasons over that prepared evidence.&lt;/p&gt;

&lt;p&gt;For developers, the output should let you inspect questions such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which source supports this fact?&lt;/li&gt;
&lt;li&gt;Does this fact apply to the time the user asked about?&lt;/li&gt;
&lt;li&gt;Has another record superseded it?&lt;/li&gt;
&lt;li&gt;Do the sources disagree?&lt;/li&gt;
&lt;li&gt;Is there enough evidence to answer at all?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A token budget should constrain the amount of evidence supplied without hiding the uncertainty needed to interpret it.&lt;/p&gt;

&lt;p&gt;There are limits to what our result establishes. This was &lt;strong&gt;our own frozen adversarial benchmark across four data domains&lt;/strong&gt;, and it has not been independently reproduced. The 100% figure means 528/528 under this test’s strict scoring—not universal accuracy. BM25-only retrieval also does not represent every RAG architecture.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://jylus.ai/benchmark-methodology" rel="noopener noreferrer"&gt;methodology is public&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you want to challenge the evidence preparation yourself, the &lt;a href="https://jylus.ai/try" rel="noopener noreferrer"&gt;Jylus playground&lt;/a&gt; accepts synthetic records without an account and lets you inspect the Context Pack. That is a way to examine its behaviour, separate from reproducing the benchmark.&lt;/p&gt;

&lt;p&gt;Bring an awkward case: a backdated correction, two conflicting records or a question whose answer is genuinely absent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What’s the smallest set of records that makes your AI app give a confidently wrong answer?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>rag</category>
    </item>
    <item>
      <title>I built Jylus: 40.05 NDCG@10 on TEMPO, with a playground to test your own evidence</title>
      <dc:creator>Josh Hodgetts</dc:creator>
      <pubDate>Sun, 04 Oct 2026 02:37:38 +0000</pubDate>
      <link>https://dev.to/josh_hodgetts_f9f91ff3a23/i-built-jylus-4005-ndcg10-on-tempo-with-a-playground-to-test-your-own-evidence-1m41</link>
      <guid>https://dev.to/josh_hodgetts_f9f91ff3a23/i-built-jylus-4005-ndcg10-on-tempo-with-a-playground-to-test-your-own-evidence-1m41</guid>
      <description>&lt;p&gt;An AI app can retrieve relevant records and still leave the model to untangle which facts apply.&lt;/p&gt;

&lt;p&gt;I built Jylus to handle that work before the model reasons.&lt;/p&gt;

&lt;p&gt;Jylus sits between your data and your model. It resolves state and relationships, identifies conflicting or missing evidence, and returns a compact Context Pack with source references.&lt;/p&gt;

&lt;p&gt;For developers, the goal is less custom evidence-handling code between your data and your AI application.&lt;/p&gt;

&lt;p&gt;A measured result&lt;/p&gt;

&lt;p&gt;On our complete 1,730-query TEMPO run through the public API, Jylus reported:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;40.048 NDCG@10&lt;/li&gt;
&lt;li&gt;39.745 Recall@10&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are retrieval measurements from our evaluation, not answer-accuracy percentages or an officially accepted leaderboard position.&lt;/p&gt;

&lt;p&gt;"Evaluation methodology" (&lt;a href="https://jylus.ai/benchmark-methodology" rel="noopener noreferrer"&gt;https://jylus.ai/benchmark-methodology&lt;/a&gt;)&lt;/p&gt;

&lt;p&gt;See what your model would receive&lt;/p&gt;

&lt;p&gt;The "Jylus playground" (&lt;a href="https://jylus.ai/try" rel="noopener noreferrer"&gt;https://jylus.ai/try&lt;/a&gt;) works without an account:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Select “Conflicting records” or “Late correction.”&lt;/li&gt;
&lt;li&gt;Inspect the supplied records and question.&lt;/li&gt;
&lt;li&gt;Run the trial.&lt;/li&gt;
&lt;li&gt;Examine the facts, source proofs, conflicts and missing evidence in the returned pack.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then try a small example from your own application. The playground returns evidence; your model remains responsible for generating the answer.&lt;/p&gt;

&lt;p&gt;For an application integration, "/api/v1/analyze" exposes the context preparation step through the API.&lt;/p&gt;

&lt;p&gt;I’m the founder. I’d particularly like feedback from developers currently writing their own logic to reconcile changing records before passing context to a model.&lt;/p&gt;

&lt;p&gt;What would you test first?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
