DEV Community

poojitha chougani
poojitha chougani

Posted on

What if your deployment agent could remember every failure it ever fixed

πŸš€ Try PipelineSage Live

Deployment #1057 of our payment service failed with a database migration timeout after 30 seconds. Twenty days earlier, deployment #1017 had failed the same way, someone had fixed it, and the fix was written down.

Nobody remembered it.

I built PipelineSage to close that gap. It reads a failed deployment, retrieves relevant past incidents using Hindsight, and recommends a fix grounded in that history. When a human confirms that the fix worked, the outcome is written back to Hindsight so the next failure can learn from it.

What the system does


The flow is simple:

  1. A CI/CD failure arrives with its service, environment, commit, and error.
  2. PipelineSage queries Hindsight for relevant past incidents.
  3. Retrieved memories are filtered, deduplicated, and ranked.
  4. The current failure and relevant memories are sent to an LLM (openai/gpt-oss-120b on Groq).
  5. The agent produces historical evidence, diagnosis, and a recommended fix.
  6. A human reviews the recommendation.
  7. If the fix works, the confirmed outcome is written back to Hindsight.

The project is split into agent/, memory/, services/, and a Streamlit interface.

I deliberately didn't build my own memory layer. The Hindsight documentation already covers the difficult parts of storing and retrieving agent memories. My integration is a small wrapper around it.

Memory is a source of truth

Memory introduces a problem that stateless agents don't have: the agent can retrieve the wrong precedent and confidently present it as history.

So I designed PipelineSage around one rule:

Whatever the agent says about history must be traceable to something Hindsight actually returned.

Writing memories for a cold reader

Every incident is retained with fixed fields:

def retain_incident(self, incident):
    content = f"""
DevOps pipeline incident.

Deployment: #{incident['deployment_id']}
Service: {incident['service']}
Environment: {incident['environment']}
Status: {incident['status']}

Failure:
{incident['error']}

Root cause:
{incident.get('root_cause', 'Not yet confirmed.')}

Resolution:
{incident.get('resolution', 'Not yet resolved.')}

Outcome:
{incident.get('outcome', 'No outcome recorded.')}

Related historical incident:
{incident.get('related_historical_incident', 'Not specified.')}
"""
    return self.client.retain(
        bank_id=self.bank_id,
        content=content.strip(),
    )
Enter fullscreen mode Exit fullscreen mode

The defaults are intentional. An unresolved incident is stored as Not yet resolved. rather than left blank.

I also store a Related historical incident. This creates a simple trail showing which earlier incident influenced a later resolution.

Recall returns candidates, not answers

Retrieval uses Hindsight's recall:

result = self.client.recall(
    bank_id=self.bank_id,
    query=query,
    max_tokens=4096,
    budget="mid",
)
Enter fullscreen mode Exit fullscreen mode

I run multiple queries using the incident's service and error text, then merge and clean the results.

One important safeguard is excluding the current deployment from its own history:

# Do not use the current deployment as history.
if deployment_id and (
    f"deployment #{deployment_id}".lower() in text
    or f"deployment: #{deployment_id}".lower() in text
):
    continue
Enter fullscreen mode Exit fullscreen mode

I also had to deal with near-duplicate memories. Five retrieved memories do not necessarily mean five independent pieces of evidence. Several can be different versions of the same incident.

My first ranking approach also had keyword bonuses, including a bonus for "500" because I already knew the answer I wanted. That was effectively an answer key, not a retrieval test.

I replaced it with general signals such as:

  • Same service
  • Successful documented outcome
  • Similar failure text
  • Penalties for unrelated services or failure classes

The retrieval should work for failures the system has never seen before.

Prompt rules are not guarantees

The system prompt tells the model that Hindsight memories are the source of truth:

Never invent historical deployments, fixes, outcomes, numbers,
batch sizes, timeout values, configuration values, or
infrastructure changes.

If a historical successful resolution contains an exact value
such as "500 records", preserve that value exactly.

If historical evidence is insufficient, clearly state that.
Enter fullscreen mode Exit fullscreen mode

Temperature is set to 0.1, and the output is separated into historical evidence, diagnosis, and recommended fix.

But prompts are not enough.

What happened on #1057

For deployment #1057, the agent retrieved memories pointing to #1017.

The historical incident described a payment-service migration that exceeded the 30-second limit and succeeded after being split into batches of 500 records.

PipelineSage recommended:

  • Process the rows in batches of 500
  • Test in staging
  • Re-run the migration

The important part was that the 500-record value came from the retrieved memory.

But the model also added that the batch size had been β€œproven to keep each migration step within the 30-second timeout.”

That claim was not present in the memory.

This exposed an important limitation: a prompt can tell the model not to invent facts, but it cannot guarantee compliance.

The next improvement I want is a mechanical verification step that flags numeric or timing claims in the recommendation when they cannot be found in the retrieved memories.

Closing the memory loop

When the recommended fix works, a human confirms it.

PipelineSage then writes the outcome back to Hindsight with:

  • Status: SUCCESS
  • Resolution: Process historical transaction rows in batches of 500 records
  • Outcome: Human-confirmed success
  • Related historical incident: #1017

This human gate is important.

If the agent automatically wrote its own recommendation into memory, an unverified guess could become a future "documented precedent." Human confirmation keeps the memory based on what actually happened rather than what the model predicted.

What I learned

Write memories for a cold reader. Fixed fields and explicit unresolved states make historical incidents easier to interpret later.

Treat recall as candidate generation. Exclude the current event, expect duplicates, and don't mistake multiple memories for independent evidence.

Don't let tests contain the answer. If your query already names the fix you're trying to retrieve, you're not really testing retrieval.

Prompts need mechanical backstops. The model can still make unsupported claims even when the prompt explicitly prohibits them.

Gate memory writes on human-verified outcomes. Reading from memory is one thing; writing new facts into memory is where mistakes can compound.

I started with memory because the same failures keep costing the same hour.

What surprised me is how little of the work was about storage, and how much was about deciding what the agent is allowed to believe.

Learn More

Top comments (0)