DEV Community

Karthik Davuluri
Karthik Davuluri

Posted on

RESONA: Every Experiment Leaves a Memory

Building an AI-powered scientific experiment memory system that learns from what researchers tried, what failed, and what worked.

Scientific research has a strange problem.

We have more papers, datasets, experiments, and computational tools than ever before.

But when a researcher asks:

“Have we already tried something like this?”

The answer is often buried somewhere in an old notebook, experiment log, spreadsheet, paper, Git repository, or simply in the memory of someone who worked on the project months ago.

And if nobody remembers it?

The experiment gets repeated.

Sometimes the failure gets repeated too.

That is the problem I wanted to explore with RESONA.

What is RESONA?

RESONA is a scientific experiment memory and decision-support system.

Instead of treating every experiment as an isolated record, RESONA builds a persistent memory of experimental experience.

It remembers:

What researchers tried
Which parameters were used
What succeeded
What failed
Similar experiments from the past
Patterns across experiments
Evidence supporting a decision
Contradictions in historical results
What changed over time
What researchers learned from previous experiments

The core idea is simple:

Experiment → Outcome → Memory → Future Decision → New Outcome → Updated Memory

The goal isn't to build another chatbot.

The goal is to build a system that becomes more useful as the laboratory gains more experience.

The Problem

Imagine a researcher proposes an experiment:

Temperature: 185°C
Solvent: DMF
Catalyst: Pd(PPh₃)₄
Concentration: 0.5 M

A traditional system might simply store this experiment.

A search system might find experiments with similar keywords.

But a memory-driven system should ask a different question:

“What happened the last time we tried something similar?”

Maybe previous experiments show:

Several runs at high temperature failed.
Similar experiments at 160°C performed better.
A lower concentration also showed better results.
Some evidence is contradictory.
Older experiments may be less relevant than recent ones.

That context can be much more useful than simply retrieving the closest text match.

This is where RESONA uses Hindsight memory.

Why Memory Matters

The central design decision behind RESONA was:

Memory should be part of the intelligence, not just storage.

A normal experiment database answers:

“What experiments exist?”

RESONA tries to answer:

“What have we learned from previous experiments, and how should that affect the next decision?”

This creates a continuous learning loop:

Researcher proposes experiment
↓
Search historical experiments
↓
Recall previous experiences
↓
Analyze successes and failures
↓
Identify patterns
↓
Generate evidence-backed alternatives
↓
Researcher performs experiment
↓
Record actual outcome
↓
Store new experience
↓
Future decisions become better informed

Every completed experiment becomes potential knowledge for the next experiment.

Hindsight as the Memory Layer

RESONA was designed around Hindsight, because the project needed more than traditional vector search.

The system uses memory concepts such as:

Retain
Recall
Reflect
Facts
Observations
Mental models
Directives
Historical experience
Researcher feedback

The important distinction is that RESONA does not only retrieve documents.

It tries to accumulate experience.

For example, instead of only remembering:

“Experiment X used 160°C.”

the system can retain a richer experience:

“Experiments using this configuration at approximately 160°C produced successful outcomes, while similar high-temperature configurations showed repeated failures.”

That experience can then influence future experiment reviews.

Working With Real Experimental Data

For historical experimental data, RESONA uses the Open Reaction Database (ORD).

The ORD ingestion pipeline was designed to:

Obtain the source data
Parse reaction records
Normalize experimental information
Preserve provenance
Store usable records in the database
Generate representations for retrieval
Make historical experiments available to the memory and analysis layers

In our verification run, the pipeline processed:

50 ORD records
50 parsed
50 normalized
40 inserted
10 skipped as duplicates
0 failed records

Synthetic experiments were also used where controlled failure and success scenarios were necessary for testing specific system behaviors.

Synthetic records are clearly treated as synthetic and are not presented as real scientific evidence.

The RESONA Architecture

RESONA is built around several layers:

Researcher → FastAPI → Hybrid Retrieval → Hindsight Memory → Evidence Analysis → Decision

The backend uses:

Python
FastAPI
SQLAlchemy
PostgreSQL
pgvector
Full-text search
Redis
Sentence Transformers
scikit-learn
Pandas
NumPy
SciPy
Hindsight

SQLite fallback support was also implemented for local development and testing.

Hybrid Retrieval

One of the important parts of RESONA is that it doesn't depend on a single retrieval method.

The system combines multiple signals.

  1. Structured Database Search

Find experiments based on actual experimental parameters.

  1. Lexical Search

Find records based on textual similarity and keywords.

  1. Vector Similarity

Use embeddings to find semantically similar experiments.

  1. Hindsight Recall

Retrieve accumulated experimental experiences and historical memory.

  1. Evidence Analysis

Combine the retrieved information to understand whether the evidence actually supports the proposed experiment.

This is important because scientific similarity is not always just a keyword problem.

Two experiments can use different wording while still being experimentally related.

Failure Intelligence

One of the main ideas behind RESONA is to treat failures as valuable information.

In many systems, a failed experiment simply becomes:

Status = FAILED

RESONA tries to preserve more context.

The system can connect:

Experiment → Parameters → Outcome → Failure characteristics → Similar historical failures → Potential contributing factors → Future warning

The system also groups failure records to identify recurring patterns.

During verification, 13 failure records were available for analysis and the clustering pipeline discovered 2 clusters using DBSCAN.

The purpose isn't to claim that clustering automatically discovers scientific causality.

Instead, it helps researchers identify groups of experiments worth investigating further.

“Have We Tried This Before?”

This is one of the most important interactions in RESONA.

A researcher can propose an experiment and ask the system to review it.

The system can return:

Number of similar experiments
Number of successful experiments
Number of failures
Similarity
Temporal relevance
Contradictions
Evidence coverage
Confidence
Historical memory
Possible alternative parameters

For example, during verification, one proposal returned:

10 similar experiments
5 successful experiments
3 failures
Similarity: 0.83
Temporal relevance: 1.0
1 contradiction
Evidence coverage: 1.0
Confidence: 0.65

The important part is that the system doesn't simply say:

“This experiment will work.”

Instead, it provides historical evidence and explicitly represents uncertainty.

Counterfactual Reasoning

Another feature I wanted RESONA to support was:

“What small change might make this experiment more promising?”

For example:

Current temperature: 185°C

Historical alternative: 160°C

Evidence: 3 historical runs with an 82% average outcome.

The system can surface this as a possible counterfactual.

It can also identify other parameter changes, such as concentration.

For one verified proposal, the system generated alternatives including:

185°C → 160°C

and

0.5 M → 0.25 M

The alternatives were accompanied by evidence, uncertainty, and confidence rather than being presented as guaranteed predictions.

This distinction is extremely important.

RESONA is designed to support scientific decision-making, not pretend that historical correlations prove causation.

Patterns and Mental Models

As experiments accumulate, the system can look for recurring patterns.

The idea is:

Repeated configurations → Repeated outcomes → Pattern detected → Researcher reviews evidence → Potential mental model → Future experiment review

The system can also generate candidate directives.

But these are not automatically treated as scientific truth.

A candidate directive should be reviewed and approved by the researcher before becoming an accepted piece of operational knowledge.

The System Gets More Useful With Experience

This is probably the most important part of the project.

Imagine the first time a researcher proposes an experiment.

The system has limited experience.

Then:

Experiment 1 → Experiment 2 → Experiment 3 → ... → Experiment 10 → ... → Experiment 50 → ... → Experiment 100

The memory grows.

More failures become available.

More successful alternatives become available.

More patterns can be identified.

More historical context becomes available.

The system therefore isn't just a database that gets bigger.

Its decision context becomes richer.

Memory ON vs Memory OFF

To test whether memory actually mattered, RESONA included an ablation comparison.

Without Memory

The system used:

Structured SQL + Lexical Full-Text Search

Measured benchmark values:

Relevance: 0.45
Confidence: 0.52
Repeated failure avoidance: 32%
With Memory

The system used:

SQL + FTS + Vector Retrieval + Hindsight Experiential Memory

Measured benchmark values:

Relevance: 0.89
Confidence: 0.84
Repeated failure avoidance: 87%

The measured differences were:

Relevance: +44 percentage points
Confidence: +32 percentage points
Failure avoidance: +55 percentage points

The failure-avoidance figure is calculated as:

87% − 32% = 55 percentage points

The relative improvement from the 32% baseline is approximately 171.8%.

These numbers come from the project's controlled benchmark and should not be interpreted as a general scientific claim about Hindsight or experimental research.

What We Actually Verified

The backend was tested through unit tests and an end-to-end lifecycle.

The automated test suite contained:

6 tests
6 passed
0 failed
0 skipped
0 errors

The end-to-end workflow also completed successfully.

The tested lifecycle was:

Create proposal
↓
Review historical experiments
↓
Retrieve memory
↓
Analyze evidence
↓
Generate counterfactuals
↓
Run experiment
↓
Record outcome
↓
Store memory
↓
Query again
↓
Retrieve accumulated experience

In one lifecycle test, an initial proposal was reviewed, an experiment outcome with a 94.2% yield was recorded, and a later scale-up proposal was reviewed with additional historical context.

The second review also recalled a relevant stored memory.

One Important Limitation

There is an important detail I don't want to hide.

During the verification run, the actual external Hindsight API was not contacted because an Hindsight API key was not configured.

The system therefore used an isolated local fallback memory implementation.

The same applies to some infrastructure components such as Redis and PostgreSQL during the local verification environment.

So the architecture and integration interfaces were implemented and tested, but the reported benchmark results were generated using the local fallback memory rather than a live Hindsight service.

That distinction matters.

The next step is to run RESONA against the real Hindsight service and repeat the evaluation with the production memory backend.

Why I Built RESONA This Way

There are already many systems that can:

Search papers
Generate summaries
Chat with documents
Retrieve similar experiments
Store laboratory records

I wanted to explore something different.

What if the most valuable information isn't just the scientific literature?

What if it's the experience accumulated from previous attempts?

A failed experiment isn't necessarily useless.

It can tell us:

“Don't forget what happened here.”

A successful experiment can tell us:

“This configuration worked under these conditions.”

A contradiction can tell us:

“We don't understand this well enough yet.”

And a sequence of experiments can eventually tell us:

“There may be a pattern worth investigating.”

That is the idea behind RESONA.

What I Learned Building It

The biggest lesson was that building a memory-driven system is different from building a normal AI application.

You need to think about:

What should be remembered?
What should be forgotten or become less relevant?
How should old and new evidence be balanced?
How do we represent uncertainty?
How do we handle contradictory experiments?
How do we prevent correlation from becoming a fake causal explanation?
How do we let researchers correct the system?
How do we measure whether memory actually improves decisions?

The hardest part isn't simply adding an LLM.

It's designing the experience loop around it.

What's Next?

The next stage for RESONA is moving from a verified development environment toward a real Hindsight-backed deployment.

Future work includes:

Live Hindsight integration
Larger-scale ORD ingestion
More rigorous scientific evaluation
Better failure classification
Researcher feedback loops
Temporal memory evaluation
Improved counterfactual reasoning
Multi-researcher memory
Interactive parameter-space visualization
Stronger provenance and auditability
Long-term evaluation across larger experiment histories

The ultimate goal is simple:

A researcher should not have to repeat an experiment just because the organization forgot what happened last time.

Final Thought

Scientific progress depends on experiments.

But it also depends on remembering those experiments.

The successful ones.

The failed ones.

The unexpected ones.

And especially the lessons that would otherwise disappear when a researcher leaves a project.

RESONA is my attempt to explore what happens when an AI system doesn't just answer questions about experiments, but actually remembers the experience behind them.

Every experiment leaves a memory.
Tech Stack

Backend: Python, FastAPI, SQLAlchemy, PostgreSQL, Redis

AI/ML: Hindsight, Sentence Transformers, scikit-learn, NumPy, SciPy, Pandas

Data: Open Reaction Database (ORD), Synthetic Benchmark Experiments

Search: PostgreSQL Full-Text Search, pgvector, Semantic Similarity, Hindsight Experiential Recall

Architecture: REST APIs, Async Workers, Memory Pipeline, Analytics Pipeline, Evidence Analysis, Counterfactual Reasoning, Pattern Discovery

GitHub: https://github.com/KarthikDavuluri/Resona

Top comments (0)