The obvious fix for an AI app missing information is to give it more context.
But what happens when the context contains three versions of the truth?
An old price. A replacement price. A correction entered today that applies to last week.
All three can be relevant to the question. Only some belong in the answer.
I’m the founder of Jylus. We ran a frozen 528-question benchmark comparing three ways of supplying evidence to the same model.
Here’s what we observed:
| Evidence supplied | Average input tokens | Strict accuracy |
|---|---|---|
| Full context | 205,129 | 78.79% |
| BM25-only RAG | 8,128 | 76.14% |
| Jylus Context Pack | 2,532 | 100% observed |
Same Gemini 3.1 Flash Lite settings. Same questions. Same deterministic scorer.
The BM25 row is the interesting part.
It used much less input than full context, but accuracy fell. Reducing tokens alone didn’t solve the problem.
The Jylus path used 98.77% less input than full context and scored higher on this workload. That does not isolate which part of evidence preparation caused the improvement, but it gives us something concrete to investigate.
Consider this small example:
Record A
Event time: Wednesday, 09:00
Recorded: Wednesday, 09:05
Observation: Temperature warning triggered.
Record B
Event time: Wednesday, 09:00
Recorded: Friday, 10:00
Correction: Wednesday's warning was caused by a faulty sensor.
Now ask:
- What caused Wednesday’s warning, using everything we know now?
- What did we know about the warning on Wednesday?
The first answer can use Friday’s correction.
The second must preserve the uncertainty that existed on Wednesday. It cannot quietly borrow knowledge from Friday.
Returning both records is useful, but it leaves the model with another job: deciding which evidence is admissible for the question.
That distinction matters in support histories, changing subscriptions, incident investigations and any application where a later update changes how an earlier event should be understood.
Jylus sits between your data and your model. It retrieves evidence, resolves state and relationships, and compiles a bounded Context Pack with source references, retaining relevant conflicts and gaps.
The model then reasons over that prepared evidence.
For developers, the output should let you inspect questions such as:
- Which source supports this fact?
- Does this fact apply to the time the user asked about?
- Has another record superseded it?
- Do the sources disagree?
- Is there enough evidence to answer at all?
A token budget should constrain the amount of evidence supplied without hiding the uncertainty needed to interpret it.
There are limits to what our result establishes. This was our own frozen adversarial benchmark across four data domains, and it has not been independently reproduced. The 100% figure means 528/528 under this test’s strict scoring—not universal accuracy. BM25-only retrieval also does not represent every RAG architecture.
If you want to challenge the evidence preparation yourself, the Jylus playground accepts synthetic records without an account and lets you inspect the Context Pack. That is a way to examine its behaviour, separate from reproducing the benchmark.
Bring an awkward case: a backdated correction, two conflicting records or a question whose answer is genuinely absent.
What’s the smallest set of records that makes your AI app give a confidently wrong answer?
Top comments (0)