DEV Community

Cover image for What Changed After Adding Hindsight to Agent Execution
Vudatha Dhruva Sai
Vudatha Dhruva Sai

Posted on

What Changed After Adding Hindsight to Agent Execution

The night the frontend guessed

It was a Tuesday night, six days before the demo, and the Event Booking Platform was supposed to be nearly done. The Architect agent had settled the design that afternoon: a modular monolith, with booking state owned by the booking module. Four developers and six days did not leave room for microservices, and the team had said so out loud.

Then the Backend agent started its task and proposed splitting the booking service out on its own. It had not seen the argument. The Architect's output had moved on to the next step, and the reasons went with it.

The Frontend agent had it worse. It got a task description and nothing else, so it did what a capable model does with nothing else: it wrote a confident, tidy fetch() call to an endpoint that did not exist, with a payload nobody had defined and no auth header. It looked right. It was invented.

Liam, who owned the booking API, found it the next morning when the integration test failed on a schema mismatch. It was the same kind of failure the team had hit the week before, and nobody, human or agent, remembered fixing it.

Nothing in that story is a model failing to reason. Every agent did its job well. The problem was the handoff: each step passed along its latest output and dropped the project's history. That night is why the Project Brain exists, and it is what the rest of this article is about.

The first thing I noticed after wiring persistent memory into TeamForge was not that the agents got smarter. It was that the frontend agent stopped inventing API payloads.

Before that, my execution layer looked reasonable on paper. Each agent received a task and the latest output from the previous step, and produced a result. In practice every handoff lost something: why we rejected microservices, which schema the backend had settled on, what broke the last time we integrated. This article covers how I fixed that with Hindsight, and where I chose not to use memory at all.

Figure 1. The TeamForge landing page.

Figure 1. The TeamForge landing page.

What TeamForge does

TeamForge takes a problem statement, team composition, skills, time constraints, and tool subscriptions, and turns them into an engineering plan. Specialized agents then execute that plan. The backend evaluates feasibility, recommends an SDLC, selects an architecture, decomposes work into a task dependency graph, assigns tasks by skill, recommends tools, and scores risks.

Figure 2. A squad and its active missions. Each mission is a project with its own plan and pipeline.

Figure 2. A squad and its active missions. Each mission is a project with its own plan and pipeline.

The planning flow is deliberately linear:

Problem Statement
        ↓
Feasibility → SDLC → Architecture
        ↓
Tasks + Dependencies → Assignment
        ↓
Tools + Risks
        ↓
Agent Execution
        ↓
Validation + Delivery
Enter fullscreen mode Exit fullscreen mode

The design decision that shaped everything else is that agents run after those decisions exist. The decision engines are deterministic. The problem-statement evaluator, for example, scores each candidate on explicit weighted dimensions and exposes the arithmetic:

(10.0 × 30% Skill) + (9.2 × 35% Time)
+ (9.5 × 25% Demo) + (9.2 × 10% Sponsor)
= 9.5 / 10
Enter fullscreen mode Exit fullscreen mode

An accepted architecture is required before task decomposition, and task dependencies are explicit edges rather than something a model infers. My working rule: let code decide what can be decided explicitly; let agents handle interpretation, generation, and execution around that state.

Each mission then runs through a build pipeline of Requirements, Architecture, Database, Backend API, Frontend UI, Testing, and Deployment, with a Project Brain memory log alongside it.

Figure 3. A mission with its execution pipeline, task checklist, Project Brain memory log, and agent roster.

Figure 3. A mission with its execution pipeline, task checklist, Project Brain memory log, and agent roster.

System structure: planning, memory, execution

The system has three layers on top of an authoritative state store.

  1. Decision and planning layer. The deterministic engines (feasibility, SDLC, architecture, task and dependency, skill assignment, tool recommendation, risk) produce a structured project plan.

  2. Project Brain. The persistent memory layer, built on Hindsight (the open-source Hindsight GitHub repository is what I used as the memory layer here). It stores decisions, context, task history, code and API context, learnings, and team conversations.

  3. Agent and execution layer. Architect, Backend, Frontend, Testing, DevOps, and Documentation agents, scheduled through RocketRide as the workflow orchestrator.

PostgreSQL sits underneath as the source of current truth for projects, tasks, assignments, risks, tool configuration, and execution status.

Figure 4. TeamForge system architecture. Hindsight sits between planning and execution; PostgreSQL remains authoritative for current state.

Figure 4. TeamForge system architecture. Hindsight sits between planning and execution; PostgreSQL remains authoritative for current state.

The technical story: stateless handoffs

Consider a typical execution chain: Architect, then Backend, then Frontend, then Test, then DevOps.

A standard agent loop passes the latest output to the next step. That is not project memory, and the gap becomes obvious after a few iterations. A downstream agent may need to know that:

  • the team rejected microservices for this project,
  • a particular teammate owns the booking API,
  • a database decision was made for concurrency reasons,
  • an earlier implementation failed because of a specific schema mismatch,
  • an API contract changed, or a task was blocked by another task yesterday.

I tried the obvious approaches first. A growing prompt is fragile and expensive. Raw chat history is the wrong abstraction: it records what was said, not what the project learned. I needed memory that treats a project as something that evolves over time.

Where Hindsight comes in

That is where I integrated Hindsight. At this point in the project I used the Hindsight GitHub repository as the memory layer itself and as the reference for its retain, recall, and reflect model: store experience, retrieve what is relevant, and reason over what has accumulated, instead of replaying transcripts.

To decide how the Project Brain should behave, I followed the Hindsight documentation, which draws the same line between persistent memory and ordinary prompt context. For the broader argument on why agents need a memory layer that is separate from the prompt, I relied on the Vectorize explanation of agent memory.

Hindsight is wired in at the boundary between planning and execution. The Project Brain calls retain when a decision or outcome is worth keeping, and recall before an agent starts a task:

hindsight.retain(
    session_id="prj-99a3b",
    fact="Team prefers a Modular Monolith architecture "
         "to avoid DevOps overhead.",
    context="Architecture Planning"
)

hindsight.recall(
    session_id="prj-99a3b",
    query="What architecture did we agree on?",
    top_k=1
)
Enter fullscreen mode Exit fullscreen mode

Figure 5. A retain call followed by a recall against the Project Brain. Recall returns the stored fact with a confidence score.

Figure 5. A retain call followed by a recall against the Project Brain. Recall returns the stored fact with a confidence score.

The retain and recall calls above follow the usage described in the Hindsight documentation. The confidence score is a retrieval signal, not a correctness guarantee, and I do not treat it as one. More on that under limitations.

Before and after: the frontend/backend handoff

Before. Each step was an isolated transaction. The frontend agent worked from the task description alone, and a generic model will happily produce a plausible fetch() call. Plausible is not the same as correct: the endpoint, the request schema, and the auth requirement were all guesses.

After. The same pipeline, run for an Event Booking Platform, now looks like this:

[ARCHITECT] Designing Schema (PostgreSQL) + Tech Stack (Next.js/FastAPI)...
[HINDSIGHT] Persisting architectural decisions to vector memory... DONE
[BACKEND]   Recalling schema from Hindsight Core...
[BACKEND]   Generating REST API routes and Pydantic models...
[FRONTEND]  Awaiting API contract...
[BACKEND]   API Contract generated -> /api/v1/openapi.json
[FRONTEND]  Consuming OpenAPI spec. Scaffolding React components...
Enter fullscreen mode Exit fullscreen mode

Figure 6. Pipeline run for an Event Booking Platform. Architecture is retained once, recalled by the backend, and the frontend consumes the generated contract.

Figure 6. Pipeline run for an Event Booking Platform. Architecture is retained once, recalled by the backend, and the frontend consumes the generated contract.

The Architect's decisions are retained once. The Backend agent recalls the schema instead of re-deriving it. The Frontend agent blocks on a real contract and consumes the OpenAPI spec rather than guessing a payload. Each stage leaves something for the next:

Architect → retain: "Modular monolith. Booking state belongs
                     to the booking module."
Backend   → retain: "Booking endpoint uses the accepted schema
                     and row-level locking."
Frontend  → recall: actual request/response contract,
                    not an invented payload
Test      → recall: "Previous integration exposed a schema
                     mismatch; add a contract test first."
Enter fullscreen mode Exit fullscreen mode

The locking decision in that chain shows up in the backend as a row-level lock on the booking record:

record = (
    db.query(BookingRecord)
    .filter(BookingRecord.id == item_id)
    .with_for_update()
    .first()
)
Enter fullscreen mode Exit fullscreen mode

When a developer asks the Project Brain how to connect a React button to the backend, the answer is built from the real project contract, not from syntax alone:

const response = await fetch(`${API_URL}/api/v1/...`, {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Authorization": `Bearer ${token}`
  },
  body: JSON.stringify(payload)
});
Enter fullscreen mode Exit fullscreen mode

The fetch() syntax is the least interesting part. What matters is where API_URL comes from, which endpoint it targets, what schema it expects, whether authentication is required, and who owns that API. Hindsight adds one more input: if the team already hit an integration problem here, that experience can be surfaced the next time a similar task appears.

Splitting state from memory

I kept PostgreSQL as the source of truth for explicit state, and used Hindsight for the experience around it.

PostgreSQL holds facts:

Task status  = blocked
Task owner   = Liam
Architecture = Modular Monolith
Risk score   = 6
Enter fullscreen mode Exit fullscreen mode

Hindsight holds reasoning and experience:

"Microservices were rejected because the team had four
 developers and a six-day delivery window."

"Frontend integration previously failed because the request
 schema did not match the backend contract."
Enter fullscreen mode Exit fullscreen mode

I do not want the memory layer to become a second source of truth for transactional data. I want it to remember why things are the way they are. That boundary matches the memory-versus-retrieval distinction laid out on the Vectorize agent memory page, which is where I took the framing from.

At query time the Project Brain combines current project state, current API/task/architecture data, and relevant prior memory into one contextual answer. The orchestrator also assigns different models to different roles, so memory is the shared context that lets them cooperate:

The build plan the system generates for a project is also grounded in that state. It names the stack per stage and breaks execution into phases and steps that can be assigned to teammates:

What went wrong, and what is still unsolved

Dead end: memory as a bigger context window. My first instinct was to treat persistent memory as an ever-growing prompt. It does not work. A frontend agent building a booking form does not need every conversation the team has ever had; it needs booking contracts, prior integration failures, and current architecture decisions. Switching to selective recall with a small top_k fixed both relevance and noise. The Hindsight docs make the same point: memory should be retrieved when it is relevant, not replayed wholesale.

Limitation: stale memory. If the database says an API changed and memory remembers the old contract, memory must lose. I enforce that structurally: current state is authority, recalled memory is context. Getting that precedence wrong would be worse than having no memory at all.

Limitation: retention is a judgment call. Deciding what counts as useful experience is the painful part. Retain too little and handoffs still drop context; retain too much and recall gets noisy. I currently retain at handoff boundaries and after validated outcomes rather than after every agent message, and I am still tuning that boundary.

Not solved: confidence is not truth. A high confidence score means the retrieved text matched the query well. It does not mean the fact is still current.

Results

I do not have benchmark numbers to quote, so I will stick to observable behavior:

  • The backend recalls the accepted schema instead of re-deriving it.
  • The frontend agent waits for a generated OpenAPI contract instead of inventing a payload.
  • Architecture decisions, including rejected alternatives and their reasons, survive across agents and sessions.
  • A failure that is retained can be recalled when a similar task appears, so the same mistake is less likely to repeat.

Lessons I would reuse

  1. Keep transactional state and agent memory separate. PostgreSQL owns current truth; Hindsight owns experience. Merging them creates unclear ownership.

  2. Put memory at handoffs, not everywhere. Context is lost when one component finishes and another starts. That is where retain and recall pay off.

  3. Never let memory override explicit state. Memory is context. The database is authority.

  4. Recall selectively. Persistent memory is not a longer prompt. Retrieve what is relevant to the current task.

  5. Make important decisions deterministic. When inputs, weights, and rejected alternatives are explicit, agents can explain and execute decisions instead of inventing them.

Where I ended up

The architecture I kept is small:

TeamForge
├── Decision Engines → decide
├── Project State    → define current truth
├── Hindsight        → remember experience
├── Agents           → execute specialized work
└── Validation       → verify the result
Enter fullscreen mode Exit fullscreen mode

Adding Hindsight did not replace any of these pieces. It gave the execution layer continuity. If agents are going to work on software projects across many tasks and sessions, memory has to be part of the execution architecture rather than an afterthought.

References

  1. Hindsight — Vectorize https://hindsight.vectorize.io/
  2. Hindsight GitHub Repository — Vectorize https://github.com/vectorize-io/hindsight
  3. What is Agent Memory? — Vectorize https://vectorize.io/what-is-agent-memory ## TeamForge GitHub

The TeamForge project is open source and available on GitHub:

https://github.com/DHRUVASAI/teamforge-ai

If you're interested in the implementation, architecture, or Project Brain integration, feel free to explore the repository and contribute.

Top comments (1)

Collapse
 
pokuri_lahari_c9a5f9f053d profile image
Pokuri Lahari •

thats amazing