The second time I ran a classification experiment on the same churn dataset, the AI planner already knew something about it. Not because I told it. Not because I hard-coded anything. It knew because a previous experiment had left a lesson behind — stored in a cloud memory bank — and the agent recalled it, read it, and used it before making a single decision.
That moment changed how I think about ML experiment pipelines.
The Problem: ML Experiments Have No Memory
Every ML experiment produces artifacts: a trained model, a metric report, a configuration file. What they almost never produce is transferable judgment — the kind of reasoning a human engineer accumulates over time. "That dataset was too small to trust a perfect ROC-AUC score." "Logistic regression worked surprisingly well here." "Don't use cross-validation kwargs in the model constructor."
Standard pipelines don't capture this. They log numbers. They don't distill lessons. So the next time you run an experiment — even on the same data — you start from zero. The pipeline has no recollection of what worked, what failed, or what the data actually looked like.
I built ML Memory to change that. The core idea is simple: after every experiment, the agent writes a structured lesson into a persistent memory bank. Before the next experiment, it reads from that bank. The planning step is informed by actual prior experience, not just the current data.
What ML Memory Does
ML Memory is a full-stack application for AI-driven binary classification experiments. You upload a CSV, specify a target column, and the system takes over: profiling your dataset, recalling relevant past lessons, planning a model, training it, evaluating it, reflecting on the results, and storing a new lesson for the future.
The stack is deliberately pragmatic:
- FastAPI — REST backend, experiment lifecycle management, SQLite persistence
- LangGraph — orchestrates the multi-step agent workflow as a compiled state graph
-
Groq — LLM inference (planning and reflection nodes), using
openai/gpt-oss-120b -
scikit-learn — the actual ML engine:
LogisticRegression,DecisionTreeClassifier,RandomForestClassifierwith full preprocessing pipelines -
Hindsight — the persistent memory layer, handling both storage (
retain) and semantic retrieval (recall) - Next.js — frontend dashboard with experiment history, detail views, and a Memory Center
The application stores experiment records in SQLite for the UI and metrics history. The lessons — the synthesized knowledge from each run — live in Hindsight Cloud, which provides vector-based semantic search across all stored memories.
To start a new experiment, you fill in a name, an optional description, the target column name, and upload a CSV file. Clicking Launch Agent Workflow triggers the full six-node LangGraph pipeline — no further input is needed until the result page appears.
Architecture: The Memory Loop
The workflow is a linear LangGraph state graph with six nodes. Each node receives the shared GraphState and returns updated fields. The graph runs synchronously per experiment request.
Dataset Profiling → Memory Recall → Planner → Experiment Execution → Reflection → Memory Persistence
Here is the workflow definition exactly as it exists in the repository:
# backend/app/graph/workflow.py
from langgraph.graph import StateGraph, START, END
def build_workflow():
workflow = StateGraph(GraphState)
workflow.add_node("dataset_profiling", node_dataset_profiling)
workflow.add_node("memory_recall", node_memory_recall)
workflow.add_node("planner", node_planner)
workflow.add_node("experiment_execution", node_experiment_execution)
workflow.add_node("reflection", node_reflection)
workflow.add_node("memory_persistence", node_memory_persistence)
workflow.add_edge(START, "dataset_profiling")
workflow.add_edge("dataset_profiling", "memory_recall")
workflow.add_edge("memory_recall", "planner")
workflow.add_edge("planner", "experiment_execution")
workflow.add_edge("experiment_execution","reflection")
workflow.add_edge("reflection", "memory_persistence")
workflow.add_edge("memory_persistence", END)
return workflow.compile()
The sequencing is intentional. Memory recall happens after profiling so the query can include exact feature names and row counts. The planner receives both the current profile and the recalled memories before it chooses a model. Persistence happens last, only after a successful reflection.
Hindsight: The Memory Layer
Hindsight is a cloud memory API built for AI agents. It stores text content in a vector-indexed memory bank and retrieves semantically relevant entries on demand. For ML Memory, Hindsight serves as the long-term episodic memory that outlives any single experiment run.
The integration is minimal and direct. Here is the complete client wrapper from the repository:
# backend/app/memory/hindsight_client.py
from hindsight_client import Hindsight
class MLMemoryClient:
def __init__(self):
self.bank_id = settings.HINDSIGHT_BANK_ID
self.client = Hindsight(
base_url=settings.HINDSIGHT_BASE_URL,
api_key=settings.HINDSIGHT_API_KEY
)
def store_memory(self, content: str) -> dict:
result = self.client.retain(bank_id=self.bank_id, content=content)
return result.model_dump() if hasattr(result, "model_dump") else result.__dict__
def query_memory(self, query: str, limit: int | None = 2) -> list:
results = self.client.recall(bank_id=self.bank_id, query=query)
normalized = []
for item in results:
if hasattr(item, "model_dump"):
normalized.append(item.model_dump())
elif isinstance(item, dict):
normalized.append(item)
else:
normalized.append({"content": str(item)})
return normalized if limit is None else normalized[:limit]
Two methods. retain writes a lesson. recall retrieves the most semantically relevant ones given a query string. The limit=2 default keeps the token budget manageable for the Groq inference step during live experiments; the Memory Center UI passes limit=None to retrieve all stored lessons for browsing.
For background on why agent memory matters in this kind of system, Vectorize has a useful overview.
The Recall Step: Querying with Context
The memory recall node doesn't send a generic query. It constructs one using the actual profiled feature names and row count:
# backend/app/graph/nodes.py — node_memory_recall
def node_memory_recall(state: GraphState) -> GraphState:
service = MemoryService()
profile = state.get("profile", {})
features = ", ".join(profile.get("feature_types", {}).keys())
query = (
f"Dataset with target '{state['target_column']}', "
f"features ({features}), and {profile.get('row_count')} rows."
)
memories = service.query_memories(query)
return {"recalled_memories": memories}
This specificity matters. Earlier versions used a generic query like "classification task" and retrieved irrelevant lessons from unrelated integration tests. The feature-aware query anchors the semantic search to the actual data structure being processed.
First Experiment: Learning From a Small Churn Dataset
The first experiment was a customer churn classification task on a 20-row synthetic dataset. The planner had no memories to draw on — it noted this explicitly in its rationale and selected a decision tree with max_depth=3 as a conservative baseline for a tiny, fully numeric dataset.
The evaluation metrics came back very high: near-perfect accuracy, precision, recall, and F1 on the test split. On a 20-row dataset evaluated with an 80/20 train-test split, that means roughly 4 test samples. The numbers look impressive on paper and mean almost nothing in practice.
This is exactly where the reflection node earns its place. Its prompt explicitly forbids certain patterns of reasoning:
"Never claim that a model avoids overfitting or generalizes well based only on a tiny dataset or perfect metrics. Your lesson MUST explicitly include: EXACT dataset size (rows and features), evaluation method, key metrics, model configuration, and any limitations."
The reflection acknowledged the dataset size, flagged the single-split evaluation as insufficient for drawing generalization conclusions, and recommended cross-validation for any future runs on similarly small datasets. That lesson was written to Hindsight via client.retain(...).
Second Experiment: Recalled Knowledge as Context
When I ran a second churn experiment — this time explicitly comparing logistic regression — the recall step returned the lesson from the first run. The planner's prompt now contained real prior evidence: a memory stating that a shallow decision tree had achieved near-perfect metrics on this exact dataset, along with the caveat that 20 rows and a single split made those metrics unreliable.
The planner used this. Critically, it didn't simply repeat the same model choice. It selected logistic regression as appropriate for the comparison, and its rationale referenced the recalled memory. The plan's memory_refs field contained the Hindsight memory ID that influenced the decision.
This is the key architectural point: the recalled memory becomes a field in the LLM prompt alongside the current profile. The planner's structured output schema enforces a valid model_type, model_params, rationale, and list of memory_refs. The model is free to agree or disagree with the recalled lesson — it is context, not a hard constraint.
A New Dataset: Loan Approval and Context-Aware Planning
The third experiment used a synthetic loan approval dataset: 200 rows, 6 input features, binary target column approved. This was structurally different from the churn experiments — more rows, different feature names, a new domain.
The recall step queried Hindsight with a description of the new dataset. The results returned lessons from the churn experiments about small datasets, metric reliability, and validation strategy. The planner prompt instructs the model to distinguish dataset-specific lessons from generic ones, so these churn lessons were treated as general guidance rather than prescriptions.
The planner selected Logistic Regression with C=1, max_iter=1000, solver=lbfgs. The evaluation on the 40-row test split returned:
- Accuracy: 80.0%
- F1 Score: 77.8%
- Precision: 87.5%
- Recall: 70.0%
- ROC-AUC: 89.75%
Storing a Lesson: The Persistence Node and Memory Trace
After evaluation, the reflection node synthesized a lesson. The persistence node then assembled the full context — plan, metrics, and lesson — and called retain:
# backend/app/graph/nodes.py — node_memory_persistence
def node_memory_persistence(state: GraphState) -> GraphState:
service = MemoryService()
content = (
f"Experiment {state['experiment_id']} on {state['target_column']}. "
f"Plan: {json.dumps(state['plan'])}. "
f"Metrics: {json.dumps(state['metrics'])}. "
f"Lesson: {state['reflection']}"
)
service.client.store_memory(content)
return {"persistence_status": True}
The stored string gives Hindsight's vector index rich semantic content. Future recalls on similar datasets surface this full context — not just the lesson sentence but the specific model, hyperparameters, and metrics that produced it.
The experiment detail page exposes all three stages of the memory loop as the Hindsight Memory Trace: the memories recalled before planning, the AI Planner's decision rationale, and the lesson learned and stored after reflection.
The lesson stored for the loan approval experiment (visible in Figure 4) reads:
"Using a logistic regression (C=1, max_iter=1000, lbfgs) on a very small dataset (200 rows, 7 features) with a simple train/test split yielded 0.80 accuracy, 0.875 precision, 0.70 recall, 0.898 ROC AUC, but the lack of cross-validation or external test set limits confidence in generalization; future work should employ k-fold CV or a hold-out validation set to assess stability."
This is the kind of precise, evidence-bound lesson that becomes genuinely useful as context for a future experiment on the same domain.
The Dashboard and Memory Center
The Next.js frontend provides four views: Overview, New Experiment, Experiment History, and Memory Center. The sidebar highlights the active page using Next.js's usePathname() hook.
The Experiment History page lists every completed run in chronological order. Each row links to a detail page showing the model type, hyperparameters, evaluation metrics, and the full Hindsight Memory Trace for that run.
The Memory Center calls POST /api/memories/recall with a broad query to surface all stored Hindsight memories. Each card shows the memory type, its relevance score from the semantic index, and the full lesson text.
What I Learned
1. Store synthesized lessons, not raw transcripts.
Storing the full LLM response or raw metrics JSON produces recall results that are too noisy. The useful signal is a concise, structured lesson that names the dataset, the model, the metrics, and the limitation. That is what the reflection node is for.
2. Recalled memory should inform, not dictate.
The most valuable property of this architecture is that the planner can disagree with past lessons. Memory provides context; the LLM decides what to do with it. Hard-coding memory into a rule would eliminate the system's ability to reason about novel situations.
3. Perfect metrics on small datasets are a warning sign.
Accuracy of 1.0 on 4 test samples is not a finding — it is a red flag that the evaluation setup is insufficient. The reflection prompt explicitly prohibits claiming good generalization from small-dataset results. This constraint was not obvious at the start; it was added after discovering the failure mode.
4. The recall query needs to match the data, not just the task.
A generic query like "classification experiment" returns whatever is most frequent in the memory bank. Encoding feature names and row count in the query string dramatically improved the semantic relevance of returned memories.
5. Persistent memory survives infrastructure failures.
During development, a test fixture called Base.metadata.drop_all() against the production SQLite database, wiping all experiment records. The Hindsight memories survived intact. Every lesson from every previous experiment remained queryable. This was an unplanned but important validation of keeping episodic memory decoupled from local storage.
Limitations and What I Would Improve
Dataset size and validation. All experiments used small synthetic datasets. The ML engine uses a single 80/20 stratified split; cross-validation would produce more reliable estimates on small datasets.
Memory relevance at scale. As the memory bank grows, recalls for a new query become noisier. The current approach caps workflow recalls at two memories to manage token budget. A memory consolidation step or explicit relevance filtering would help at larger scale.
Experiment metadata consistency. Early experiments stored column_count including the target column; later experiments corrected this to feature_count (columns minus one). The discrepancy exists in the stored Hindsight memories from early runs.
Linear workflow with no branching. The LangGraph graph is a chain. If the Groq API is unavailable or the dataset fails validation, the workflow returns an error state rather than retrying individual nodes or attempting fallbacks.
No model artifact persistence. The trained scikit-learn model is discarded after evaluation. Persisting model artifacts would allow reuse and comparison without retraining.
Conclusion
The ML pipeline that can recall its own history is not a solved problem — this project is a proof of concept. But the architectural pattern is sound. Separating ephemeral experiment state (metrics, CSVs, model artifacts) from persistent distilled lessons (stored in Hindsight) means that knowledge accumulates independently of the infrastructure that generated it.
The practical consequence: the second time the agent encounters a familiar dataset, it does not start from zero. It starts from where the last experiment left off.
That is a small thing in isolation. Compounded over dozens of experiments, across multiple datasets and model architectures, it becomes a system that gets meaningfully smarter with use — which is, arguably, what intelligence looks like in practice.
References
- Hindsight on GitHub — open-source memory SDK for AI agents
- Hindsight Documentation — API reference and integration guides
- What Is Agent Memory — Vectorize — conceptual background on persistent agent memory
- LangGraph — state graph orchestration framework
- scikit-learn — ML pipeline and estimators
- Groq — LLM inference API





Top comments (0)