DEV Community

Jaividhyarthi
Jaividhyarthi

Posted on

I Gave My Deploy Pipeline Hindsight Memory. Outages Stopped Repeating.

The change that broke checkout was two lines long. It fixed a rounding bug in a GST helper, passed review, passed CI, and went out at 17:50 on a Friday. For the next 18 minutes customers were charged stale prices.

The code was fine. The timing wasn't. Every Friday at 17:45 the team runs a catalog price-sync job, and any checkout deploy in that window restarts pods mid-sync and warms the cache with old prices. It had happened twice before. Nobody connected the three incidents, because the knowledge lived in three different post-mortems.

That's the problem I built Butterflaw to solve: a deployment agent that remembers every deploy a team has shipped, predicts which new ones will break production, and learns from its own wrong predictions.Why an LLM alone can't do this

My first instinct was to paste a deployment description into an LLM and ask "will this break?" It's good at sounding right. A 1,400-line refactor? "High risk." A minor dependency bump? "Low risk."

But the failures that actually hurt a team are team-specific:

pg-driver 3.2 looks like a routine bump, but it turns on prepared-statement caching, which breaks behind this team's PgBouncer.
A two-line checkout fix is harmless on Wednesday and an outage on Friday evening.
The scary 1,400-line refactor is actually safe, because this team ships dark behind feature flags.

No model can know that from the diff. It's history. So the real question was: how do I give the agent history in a way that gets better with use?

The design: same model, one difference

Butterflaw runs two agents side by side on every deployment:

With memory: it recalls similar past deployments from Hindsight agent memory and passes them to the LLM.
Without memory: the same LLM (openai/gpt-oss-120b on Groq), the same prompt, but no history.

The only variable is memory, which makes it easy to see what memory is actually worth instead of guessing.

def predict(bank_id: str, dep: dict) -> dict:
    recalled = memory.recall(bank_id, describe(dep))
    with_memory = _judge(dep, recalled)
    without_memory = _judge(dep, [])
    return {"with_memory": with_memory, "without_memory": without_memory, "recalled": recalled}
Enter fullscreen mode Exit fullscreen mode

The learning loop: retain what actually happened

Prediction is half of it. What makes it learn is what happens after the deploy ships. Butterflaw retains the real outcome and also a review of its own prediction:

if prediction["will_fail"] == dep["failed"]:
    text += f" Butterflaw predicted {predicted}, which was correct."
else:
    text += (f" Butterflaw predicted {predicted}, which was WRONG. Lesson: "
             + ("changes like this are dangerous for this team even though they look harmless."
                if dep["failed"] else
                "changes like this are safe for this team; do not over-warn on them."))

Enter fullscreen mode Exit fullscreen mode

That text goes into Hindsight with the real deploy timestamp and metadata:

memory.retain(bank_id, outcome_memory(dep, prediction), datetime.fromisoformat(dep["when"]),
              context="deployment_outcome",
              metadata={"deploy_id": dep["id"], "services": ",".join(dep["services"]),
                        "failed": str(dep["failed"]).lower()})
Enter fullscreen mode Exit fullscreen mode

The next time a similar change shows up, recall() brings back not just "this happened" but "I got this wrong last time, and here's why."

I didn't write any retrieval logic. Hindsight runs semantic, keyword, entity-graph and temporal search in parallel, and that turned out to matter:

Keyword search catches pg-driver exactly, even when a summary calls it a "routine library update."
Temporal search is what connects a "Friday 17:40 checkout deploy" to earlier Friday-evening incidents. Plain vector similarity on the change text would never find it; the texts share nothing but the time.

The "Ask Butterflaw" box uses reflect(). When it warns you, you can ask why, and Hindsight reasons over the whole memory bank rather than only the top few recalled items.

Watching it learn

To test this honestly, I replayed 24 deployments from an e-commerce team's August–September history in chronological order. The history is synthetic but realistic, with team-specific failure patterns planted in it. Both agents predict before seeing each outcome; then the outcome is retained.

In my run, the memory agent was right on 17 of 24 deployments (71%). The memoryless agent, with the same model and prompt, got 11 of 24 (46%). On every deployment where the two disagreed, the memory agent was the one that was right: it caught six incidents the baseline called safe, and the baseline never caught one that the memory agent missed.

The clearest example is the pg-driver pattern. The first bump (INC-139) fooled both agents. After that outcome was retained, the memory agent caught the next three bumps in three different services, citing the earlier incident each time. The memoryless agent called all four "routine" and missed every one. It always will.

The part I didn't expect is where memory wasn't enough. On the Friday-evening checkout deploys, the memory agent caught the second one (INC-147), then missed the next two. It recalled the right incidents both times, and then argued itself out of them: "this hotfix only adds a null check… lowering the failure likelihood." Retrieval worked; the judgment didn't. Memory gives the model the evidence, but the model still has to weigh it, and an LLM that is inclined to reason about what changed will discount a pattern that is about when it shipped.

Two more honest limits: the last deployment, switching a JSON serializer to orjson, failed in a way nobody had seen before, and both agents missed it. And accuracy didn't climb smoothly across the run (75% in the first half, 67% in the second), because the later half contained the Friday cases above plus that novel failure. The gap between the two agents is the result that held up.

What I learned
Compare memory against a no-memory baseline. Same model, same prompt. Without that comparison you can't tell whether memory helps or the model just sounds confident.
Store the agent's mistakes, not only the facts. Retaining "I predicted SAFE and was wrong" is what changes the next prediction.
Time is a feature. Some failures are about when something shipped, and temporal retrieval finds patterns that text similarity misses.
Recall is not the same as judgment. In several misses the right incident was in the prompt and the model still talked itself out of it. The next version will turn repeated patterns into hard rules instead of hoping the model weighs them correctly.
Be honest about novelty. Memory cannot predict a failure mode that has never happened, and the UI should say so.
What's next

A GitHub Actions hook that runs a preflight on every pull request and retains the CI and deploy result automatically, so the memory builds itself.

Code: https://github.com/Jaividhyarthi/Butterflaw

If you're building agents that should improve over time, have a look at Hindsight on GitHub, its documentation, and Vectorize's explainer on what agent memory is.
Built with help from @codein (Code.in).

Watch it

Top comments (0)