I Stopped Re-Explaining My Project to AI
How giving TeamForge a real memory layer with Hindsight turned a chatbot into an engineer that remembers why the project looks the way it does.
TeamForge — the AI engineer that builds with your squad.
The frustrating part of using an AI assistant on a real software project isn't usually that it can't write code. It's that I keep having to explain the project again.
Which database owns bookings? Why did I choose a modular monolith? Who owns the frontend task? Which API is the real one? Why was a particular architecture rejected?
I built TeamForge around those questions, and Hindsight became the memory layer that lets the project keep its context instead of treating every conversation like a fresh start.
The Problem Was Bigger Than Chat History
TeamForge started as an engineering decision system.
Given a problem statement, a team, a time budget, available skills, tools, constraints, and project state, it reasons through a development lifecycle: evaluate candidate problems, estimate feasibility, choose an SDLC approach, recommend an architecture, decompose work, assign tasks, identify bottlenecks, assess risk, and help developers resolve implementation questions.
That creates a problem that a normal chatbot doesn't solve particularly well.
The useful context isn't concentrated in one conversation. It is distributed across decisions and events:
The team has a particular set of skills.
A problem statement was selected for specific reasons.
An architecture was accepted after alternatives were rejected.
Individual modules own particular tables.
Tasks were assigned to specific developers.
Some tasks are blocked by dependencies.
API contracts define how those modules communicate.
Developers ask questions while implementation progresses.
Later decisions depend on what happened earlier.
A relational database is good at storing the current state. It is much less convenient for answering questions such as:
Why did we choose this architecture?
or:
What did we already decide about the booking workflow?
Those are questions about accumulated project experience. That is where I use Hindsight.
Hindsight on GitHub gives me a memory layer built around retaining information, recalling relevant memories, and reflecting over accumulated knowledge. Hindsight documentation describes the same three core operations: retain, recall, and reflect.
I use that distinction deliberately. The database answers what is true now. Hindsight helps answer what we learned, decided, and discussed along the way.
I Treat Project Memory as a Separate System
The first architectural decision was not to dump the entire TeamForge database into an LLM. That would make the system difficult to reason about and difficult to control.
Instead, I separate three kinds of information.
- Authoritative Project State
This stays in the application database. Users, teams, tasks, project IDs, assignments, schemas, statuses, and other structured entities remain queryable through normal application logic.
If a developer asks who owns a task, I don't want an embedding search deciding the answer. I query the database.
- Project Memory
This is where Hindsight fits. I retain useful events and context such as:
architecture decisions
rejected alternatives and their reasons
implementation discussions
important debugging discoveries
team handoffs
project-specific conventions
decisions made during planning
relevant conversations
The important detail is that I don't think of these as "chat logs." They are memories about the project.
- The Current Question
When a developer asks the Project Brain something, I combine the live project state with relevant recalled memory. That gives the assistant two different kinds of context.
Current state tells it what exists. Memory tells it why the project got there.
The Integration Is Intentionally Small
Hindsight's Python client exposes the core operations directly. A basic memory write looks like this:
from hindsight_client import Hindsight
memory = Hindsight(base_url=HINDSIGHT_URL)
memory.retain(
bank_id=f"project:{project.id}",
content=conversation_or_decision,
context="project engineering discussion",
tags=[f"project:{project.id}"],
)
The project ID matters. I don't want a decision from Project A appearing in the context for Project B simply because the wording happens to be similar. So the memory boundary follows the project boundary.
For larger systems, I also use Hindsight's document IDs and tags to make memory sources and visibility explicit. The API supports metadata, entities, tags, timestamps, and document IDs during retention.
That turns memory from an amorphous "AI context" feature into something I can actually reason about.
Recall Happens Before the Model Answers
The other half of the integration is retrieval. When a developer asks:
Why are we using a modular monolith instead of microservices?
I don't send that question directly to the model and hope it remembers. I first retrieve project-specific memory.
memories = memory.recall(
bank_id=f"project:{project.id}",
query=user_question,
)
context = "\n".join(
result.text
for result in memories.results
)
Hindsight's recall pipeline is particularly useful here because it isn't just a single vector similarity lookup. Its recall process combines semantic similarity, keyword retrieval, graph traversal, and temporal signals before producing a ranked result set.
That matters for engineering conversations. A question about an architecture decision might not contain the exact words used when the decision was made. The original discussion might have said:
We ruled out microservices because the team has four developers and a 48-hour delivery window.
A later question might simply be:
Why aren't these services separate?
Those two pieces of text are related without being identical.
Retain and Recall, in Action
This isn't theoretical — here is the memory layer running: a fact is retained during architecture planning, then recalled later with high confidence when the question comes back around.
The same pattern runs across a full pipeline execution — requirements analysis handing off to architecture, architecture persisting its decisions to memory, and the backend and frontend agents picking up the resulting contract:
And on boot, the memory layer initializes alongside the rest of the stack — it isn't bolted on afterward:
Then I Combine Memory With Live State
This is the part I care about most. Memory alone is not enough.
Suppose the Project Brain remembers an old API contract. If the API has since changed, returning the old memory as if it were current truth is worse than having no memory at all.
So the answer pipeline looks roughly like this:
The database remains authoritative for current entities and state. Hindsight provides historical and experiential context. The model sits above both.
That separation is important because it prevents me from treating memory as a replacement for application state.
A Concrete Example: Architecture Decisions
One of the most useful things TeamForge can explain is an architecture decision.
The architecture engine can recommend a modular monolith for a small team working under a short delivery window. It can also record rejected alternatives.
That creates two different questions. The first is straightforward:
What architecture are we using?
The application database can answer that. The second is more interesting:
Why did we reject microservices?
That answer depends on the reasoning recorded during planning. The Project Brain retrieves the earlier discussion and combines it with the current architecture record:
Current state
Architecture = Modular Monolith
Team size = 4
Time budget = 48 hours
Relevant memory
Microservices were considered but rejected because the project would add deployment, service discovery, networking, and operational overhead without enough functional benefit for the available delivery window.
The model can then explain the decision without inventing a rationale. That distinction is the entire point of the memory layer.
The Same Approach Works for Developer Handoffs
Another problem appeared when different developers worked on different parts of the project. A frontend developer might ask:
How do I connect this button to Ethan's backend?
A generic coding assistant can produce a perfectly reasonable fetch() call. But TeamForge knows more.
The active project has a real project UUID, actual tasks, real assignments, and actual API contracts. So the Project Brain combines recalled project context with current state and produces something closer to:
const response = await fetch(
${API_BASE}/api/v1/projects/${projectId}/tasks,
{
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": Bearer ${userToken}
},
body: JSON.stringify(payload)
}
);
The important part isn't the fetch() syntax. The important part is that the assistant knows which project it is talking about. That is the difference between a coding assistant and a project-aware engineering assistant.
Memory Also Changes How Conversations Accumulate
I don't want every sentence to become permanent knowledge. That creates noise.
The useful pattern is to retain information that has future value: decisions, stable project facts, discoveries, constraints, and meaningful implementation discussions. For example:
User: Why did we reject microservices?
A: The project is using a modular monolith because the current team and delivery constraints don't justify the operational overhead of distributed services.
User: Was that always the plan?
A: No. Earlier planning considered microservices, but the architecture decision rejected them because deployment, service discovery, and cross-service coordination would have added complexity without helping the current scope.
The second answer is useful because the assistant connects the current architecture with the history behind it. Over time, that becomes more valuable than another pile of chat transcripts.
Reflection Is Useful for Higher-Level Questions
Hindsight also exposes a reflection operation for synthesizing knowledge from accumulated memories. That gives me another layer above simple retrieval.
Recall is useful for:
What did we decide about authentication?
Reflection becomes more interesting for:
What patterns have emerged in how this project makes architecture decisions?
Those are different workloads. I don't need a synthesized observation every time a developer asks where an endpoint lives. But for project retrospectives, recurring engineering patterns, and longer-term context, reflection turns many individual memories into a more durable understanding.
That is one reason I prefer a dedicated memory system over implementing "vector search plus prompt stuffing" myself. The retrieval mechanism is part of the architecture rather than an afterthought.
What I Learned
- Current state and historical context should not share one source of truth.
This was the most important design decision. A database should answer questions about current state. A memory system should provide context about what happened and why. Trying to make either system do both jobs creates ambiguity.
- Project-scoped memory matters more than generic memory.
A memory about one project is not automatically useful to another. Using project-specific memory boundaries makes retrieval more predictable and reduces accidental context leakage.
- The best AI answers are often constrained by non-AI systems.
The model is only one part of the Project Brain. The database provides facts. The task graph provides dependencies. The architecture engine provides decisions. The API contracts provide interfaces. Hindsight provides historical context. The model turns those inputs into an explanation — which is much easier to trust than asking a model to reconstruct the whole project from a prompt.
- Memory needs a lifecycle.
Retaining everything is not the same as having useful memory. I have to decide what deserves to persist, how it is scoped, how it is tagged, and how it should interact with newer project state. Hindsight's retain, recall, and reflect model gives me explicit primitives for that lifecycle rather than forcing everything through one retrieval function.
- The real payoff is continuity.
The most useful result isn't that the assistant remembers a sentence from last week. It's that I can ask a question today that assumes the project has a history — that an architecture was proposed, alternatives were rejected, tasks were assigned, implementation happened, problems were discovered, and decisions changed.
That is much closer to how experienced engineers actually work.
The Larger Idea
I originally thought of the Project Brain as an AI assistant that could answer engineering questions. I now think of it differently.
The valuable part is not the chat interface. It is the combination of current project state, project memory, engineering rules, and an LLM.
Hindsight gives me the memory layer needed to make that combination persistent. Without it, every conversation starts from the same blank page. With it, the assistant accumulates context about the system it is helping build.
That changes the question from:
Can an AI write this piece of code?
to something more useful:
Can an AI understand enough of this project's history to help me make the next engineering decision?
For me, that is the more interesting problem.
Hindsight — Vectorize
https://hindsight.vectorize.io/Hindsight GitHub Repository — Vectorize
https://github.com/vectorize-io/hindsightWhat is Agent Memory? — Vectorize
https://vectorize.io/what-is-agent-memory
TeamForge GitHub
The TeamForge project is open source and available on GitHub:







Top comments (1)
thats amazing