DEV Community

Cover image for I Realized I Couldn't Trust AI With My Research
Temiloluwa Valentine
Temiloluwa Valentine Subscriber

Posted on

I Realized I Couldn't Trust AI With My Research

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

Imagine remembering clearly that you left your keys on the kitchen counter, only to find that they aren't there. Someone tells you that you never left them there, but you remember doing it.

Who do you trust: your memory, or the evidence in front of you?

For people living with dementia and other forms of memory impairment, losing confidence in their own memory can become part of everyday life.

This is one of the reasons I'm building AI memory. I'm exploring systems that can remember what happens around a person over time: where objects were last seen, what changed, who was there, and what happened.

The goal is that when someone asks, “Where did I leave my keys?”, the system doesn't generate the most likely answer. It retrieves what actually happened.

But while building this, I ran into a different version of the same problem.

I didn't know which memory architecture I could trust.

Should continuous video become captions? Persistent entities? Episodic events? Knowledge graphs? And when that memory grows over days or months, how should it be retrieved?

I turned to research papers for answers.

The problem was that almost every approach had results showing that it worked, but those results weren't always directly comparable. Different papers used different benchmarks, models, subsets, metrics, and experimental conditions.

A higher number in one paper doesn't automatically mean its architecture is better than another architecture tested somewhere else.

And when I used AI to help reason across the literature, I found another problem: it could produce a convincing comparison where the papers were real, the citations were real, and the numbers were real, while the conclusion itself wasn't actually supported by a comparable experiment.

That was the problem I couldn't ignore.

I was building something to help people trust their memory. I couldn't build it on evidence I couldn't trust.

So I built Research Council.

Research Council doesn't treat a research question as a prompt that needs an answer. It treats it as an investigation.

Relevant evidence is retrieved from my research knowledge base first. Four agents examine that evidence in parallel:

Researcher A builds the strongest evidence-supported case.

Researcher B independently investigates the question.

Contradiction Hunter actively searches for conflicting evidence, incompatible experimental setups, and conclusions that shouldn't be compared directly.

Evidence Auditor examines whether claims are actually supported by the retrieved material.

Their claims then pass through deterministic verification before reaching the Judge, which reasons over the evidence that survived verification.

The result isn't required to be an answer I want.

It can be:

SUPPORTED. CONTRADICTED. INSUFFICIENT EVIDENCE.

That last outcome is important to me.

If two architectures have never actually been compared under compatible conditions, Research Council shouldn't manufacture a winner just because both papers contain numbers.

Sometimes the most useful research result is knowing:

We don't have enough evidence yet.

Research Council started as something I needed for my own AI-memory research. I wanted a way to move from “this paper says this works” to “what does the evidence actually allow me to conclude?”

And that principle connects directly back to why I'm interested in memory systems in the first place:

If I'm going to build AI that people may eventually depend on to remember, I need to be just as careful about what I teach that system to believe.

Demo

Code

GitHub logo Valentinetemi / research-council

An evidence-led AI research council that queries a Sanity Knowledge Base, investigates with four parallel agents, verifies citations, and delivers evidence-backed verdicts.

Research Council

An evidence-led research tool that helps you investigate technical questions without forcing a conclusion the sources cannot support.

Research Council retrieves evidence from a Sanity Knowledge Base, runs four research agents in parallel, verifies cited quotes and numbers in deterministic code, and asks a Judge to evaluate the verified evidence.

The verdict can be:

  • SUPPORTED
  • CONTRADICTED
  • MIXED
  • INSUFFICIENT EVIDENCE

Sometimes the most useful result is knowing that the available evidence cannot settle the question.

Why I Built It

I started Research Council while researching architectures for an external memory system for people living with dementia.

Should continuous video become captions, persistent entities, episodic events, or knowledge graphs? How should those memories be retrieved over time?

Research papers offered promising results, but they often used different benchmarks, models, subsets, metrics, and experimental conditions. A higher score in one paper did not necessarily mean its architecture was better than another.

…

How I Used Sanity

Every memory needs a place to live. For Research Council, that place is a Sanity Knowledge Base.

What I pointed Sanity Context at

I gave Sanity Context 13 research papers on AI memory, from agents that build long-term memory from video to benchmarks that test whether a system can remember where an object was last seen. Sanity turned them into cited entries, and every entry stays linked to the paper it came from.

Then I wrote instructions for how the Knowledge Base should remember:

  • Keep the benchmark, subset, model and metric behind every number
  • Never merge or average numbers across papers
  • Keep a paper's claims about its own method separate from what other papers say about it

That last rule matters most. In memory research, a result can look very different when another team tests it in a different setting. Without that separation, the Knowledge Base would blur two different measurements into one, the same way a failing memory blends two different moments into a single false one.

The Sanity Context tools I used

My app connects to a Sanity Context MCP endpoint and uses three tools:

  • initial_context gives the Knowledge Base outline, fetched once and placed in the system prompt
  • knowledge_base_search runs keyword search across the entries
  • knowledge_base_read reads the full text of the entries that matter

What the agents do with the content

The model never decides what to look up. My code runs the searches, reads the most relevant entries, and hands that evidence to all four agents at once. That way, no agent can quietly skip a source that disagrees with its argument.

After the agents make their claims, the verifier goes back to Sanity. It calls knowledge_base_read again for every cited entry and checks each quote against the real text, word for word. It never trusts what an agent says it read. It checks the source itself.

This only works because the content is structured. If the papers were just one big pile of text, there would be nothing to check a claim against. Because every entry is cited and linked, every claim can be traced back to the exact place it came from, or rejected.

One full investigation takes just 5 model calls: four agents in parallel, then the Judge. That means it runs on Gemini's free tier.

Sanity Project Details

  • Sanity project ID: 873zrpb9
  • Knowledge Base: Evidence Lab: Physical Memory research (id: kbeZdQSQKQXx), built from 13 research papers (46 source documents) with custom instructions for how results should be recorded
  • MCP endpoint: "Research Council," serving the Knowledge Base through initial_context, knowledge_base_search and knowledge_base_read

Agent Session

Claude Code section - https://dev.to/agent_sessions/claude-code-session-anrjak
Codex section - [https://dev.to/agent_sessions/codex-session-z6msn1](https://dev.to/agent_sessions/codex-session-z6msn1

Top comments (0)