DEV Community

Cover image for I Built a Web Agent That Remembers How I Build
Satya Harsha Yadavalli
Satya Harsha Yadavalli

Posted on

I Built a Web Agent That Remembers How I Build

I Built a Web Agent That Remembers How I Build

The first time I generated a website with an AI agent, the output looked good. The frustrating part came later: every new session meant explaining the same design decisions again.

I wanted to change that. Instead of treating each website request as an isolated generation task, I built a web development agent that can retain what I teach it, recall those decisions later, and use previous outcomes when generating the next site.

The memory layer is powered by Hindsight.

The problem wasn't generation

Generating HTML is no longer the interesting part.

A modern coding model can produce a landing page, add a form, choose a color palette, and wire up interactions from a short prompt. The harder problem is continuity.

Suppose I tell an agent:

I prefer black, white, and gold. Keep the interface minimal, use subtle hover effects, and avoid excessive animation.

It can follow those instructions for the current request. But if I start another project later, a stateless agent has no reason to know that those choices matter to me.

That creates a strange workflow: the more I use the agent, the more I have to repeat myself.

I built my system around the opposite idea: a website agent should accumulate useful experience.

How the system is structured

The application accepts both text and voice instructions. An orchestration layer then coordinates several specialized stages:

User: voice / text
        |
        v
Orchestration Agent
        |
        +---- Hindsight Recall
        |
        v
+-----------------------+
| Copywriter Agent      |
| Design Agent          |
| Code Generator Agent  |
+-----------+-----------+
            |
            v
       Generated HTML
            |
            v
+-----------------------+
| UI/UX Critic          |
| Functionality Critic  |
| Performance Critic    |
| Code Quality Critic   |
+-----------+-----------+
            |
       pass / fail
            |
      +-----+-----+
      |           |
    pass        issues
      |           |
      v           v
   Preview     Hindsight
                Retain
                  |
                Reflect
                  |
                  v
             Code revision
Enter fullscreen mode Exit fullscreen mode

The important detail is that Hindsight is not placed at the end as an audit log. It participates in the generation process.

Before the copywriter, designer, and code generator run, the orchestration layer recalls relevant memories and injects them into the generation context.

Turning feedback into durable memory

I did not want to store every sentence a user says. Most commands are temporary.

Instead, the agent extracts preferences that are likely to remain useful across projects.

For example, the preference-learning layer looks for instructions such as:

rules = [
    (("no gradient", "no gradients", "without gradient", "avoid gradients"),
     "User prefers interfaces without gradients."),

    (("black, white and gold", "black white gold"),
     "User prefers a black, white, and gold visual palette."),

    (("minimal", "minimalist", "clean and minimal"),
     "User prefers a clean, minimal interface with restrained visual clutter."),

    (("no excessive animation", "avoid excessive animation"),
     "User prefers subtle, purposeful animation rather than excessive motion."),

    (("hover effects", "interactive hover"),
     "User prefers interactive UI states and meaningful hover feedback.")
]
Enter fullscreen mode Exit fullscreen mode

Those extracted preferences are then retained as long-term memories:

for pref in preferences:
    self.retain(
        content=pref,
        tags=["user_preference", "design_preference", "long_term_memory"],
        metadata={
            "type": "user_preference",
            "workflow_id": workflow_id,
            "learning": "explicit_user_instruction"
        }
    )
Enter fullscreen mode Exit fullscreen mode

This distinction matters. A transient request such as "make this heading larger" does not necessarily deserve to become a permanent rule. A repeated design preference does.

Recall happens before generation

The next part was more important than storing the memory.

At the beginning of a new generation workflow, the orchestrator recalls architectural knowledge and user preferences:

user_preferences = hindsight_service.recall_user_preferences(
    category=business_data.get("category", ""),
    current_instruction=initial_prompt or "",
    max_results=8
)

memory_context = hindsight_service.format_memory_context(user_preferences)
Enter fullscreen mode Exit fullscreen mode

That context is then passed downstream:

memory_instructions = (
    f"{initial_prompt or ''}\n\n"
    "LONG-TERM USER PREFERENCES RECALLED FROM HINDSIGHT:\n"
    f"{memory_context}\n\n"
    "Treat these as persistent preferences unless the current user "
    "instruction explicitly overrides them."
)
Enter fullscreen mode Exit fullscreen mode

The copywriter and designer therefore receive more than the current prompt. They receive the current prompt plus relevant experience from previous interactions.

That is the difference between remembering something and actually using memory.

The second website is where the idea becomes visible

The most useful test is not asking the agent to remember something immediately after learning it.

The useful test is a new project.

I can teach the agent that I prefer a restrained black, white, and gold visual system, minimal layouts, subtle hover states, and limited animation.

Then I can start a different website without repeating those preferences.

The agent recalls the stored preferences and includes them in the generation context.

The interaction becomes:

Project 1
  User teaches preferences
        |
        v
     RETAIN
        |
        v
  Long-term memory
        |
        |
Project 2
  New request only
        |
        v
     RECALL
        |
        v
Generation uses previous preferences
Enter fullscreen mode Exit fullscreen mode

That is the behavior I wanted from the beginning. The second project is not simply another independent generation.

Memory also captures failure

The other useful part of the architecture is the critic loop.

After the first HTML is generated, four critics evaluate it across UI/UX, functionality, performance/mobile behavior, and code quality.

If the result fails the configured evaluation threshold, the issues are retained:

hindsight_service.retain(
    content=(
        f"Critic evaluation rejected code in iteration {iteration} "
        f"due to: {issues_summary}"
    ),
    tags=["critic_rejection", "mistake_guard"],
    metadata={
        "type": "experience",
        "iteration": str(iteration)
    }
)
Enter fullscreen mode Exit fullscreen mode

The agent then asks Hindsight to reflect on those issues and produce actionable repair instructions:

reflection_patch = hindsight_service.reflect(
    query=f"How to heal these specific code flaws: {issues_summary}",
    context=(
        f"Business: {business_data.get('name')}, "
        f"Category: {business_data.get('category')}"
    )
)
Enter fullscreen mode Exit fullscreen mode

The code generator receives that reflection and gets another chance to produce the site.

This gives the system two different forms of learning:

  • User learning: remember what the developer wants.
  • System learning: remember what went wrong and what patterns worked.

Cloud memory, with a resilient local path

The application is designed to use Hindsight Cloud when credentials are configured:

self.client = Hindsight(
    base_url=settings.hindsight_base_url,
    api_key=settings.hindsight_api_key,
)
Enter fullscreen mode Exit fullscreen mode

The project also contains a local persistence path for development and environments where cloud credentials are unavailable. Learned memories are written to hindsight_memory.json, allowing the application to preserve its learning across backend restarts.

This fallback was useful during development because it let me test the full retain/recall workflow without making the entire application dependent on a network connection.

In a deployed setup, the intended memory layer is Hindsight rather than the local JSON file.

What changed in the way I think about agents

The biggest lesson was that memory should not be treated as a database attached to an agent.

A database answers:

What information did we store?

A useful memory system needs to answer:

What should the agent remember, when should it recall it, and how should that memory change its next decision?

That led to three design rules for this project.

1. Store decisions, not noise

If every interaction becomes memory, recall becomes less useful. I found it more valuable to extract durable preferences and meaningful generation outcomes.

2. Recall before the expensive work

Memory has the most leverage when it changes the inputs to generation. Recalling preferences after the website is already generated is too late.

3. Make failure part of the learning loop

A failed critic evaluation is not just an error message. It is experience that can become a future constraint or repair strategy.

The part I still care about most

The most interesting moment is not when the agent generates the first website.

It is when I start a second project, intentionally leave out the design instructions I used before, and watch the agent bring those decisions back.

That changes the mental model from:

Prompt → Generate → Done
Enter fullscreen mode Exit fullscreen mode

to:

Interact → Learn → Recall → Generate → Evaluate → Reflect → Learn again
Enter fullscreen mode Exit fullscreen mode

That is a much more useful foundation for an AI development agent.

For more on the memory layer, see the Hindsight GitHub repository, the Hindsight documentation, and Vectorize's overview of agent memory.

Top comments (0)