

The problem that kept bugging me
We've all been there. A deployment breaks at 2 AM. You spend an hour figuring out what went wrong. You fix it. Then three months later, the same thing breaks again — and you've completely forgotten how you fixed it last time. The answer is buried in some Slack thread or a doc no one reads anymore.
That frustrated me. Not the failures — those happen. What bothered me was losing the knowledge of how we fixed them. So I built something called PipelineSage. It's a simple tool that helps diagnose CI/CD pipeline failures by remembering what happened before and what worked.
What it does
When a deployment fails, PipelineSage looks back at similar failures from the past. It pulls up what went wrong before and how it got fixed. Then it uses that history to suggest a fix for the current problem.
If the fix works, you tell it. If it doesn't, you tell it that too. Either way, it saves the result. Next time something similar breaks, it already knows what to try.
That's really the whole idea. Learn from past failures. Don't start from scratch every time.
How memory works under the hood
Most AI tools forget everything after each conversation. They have no memory. I didn't want that. I wanted something that actually gets better over time.
I used Hindsight for the memory part. It's an agent memory tool that does two things really well:
- RETAIN — save something to memory
- RECALL — find relevant stuff from memory later
Here's what the code looks like. It's really short:
class HindsightMemory:
def retain(self, incident_text: str) -> dict:
response = requests.post(
f"{self.base_url}/v1/banks/{self.bank_id}/retain",
headers={"Authorization": f"Bearer {self.api_key}"},
json={"content": incident_text}
)
return response.json()
def recall(self, query: str, top_k: int = 3) -> list:
response = requests.post(
f"{self.base_url}/v1/banks/{self.bank_id}/recall",
headers={"Authorization": f"Bearer {self.api_key}"},
json={"query": query, "top_k": top_k}
)
return response.json().get("memories", [])
Just two functions. One to save, one to search. Hindsight handles all the complicated stuff like indexing and matching behind the scenes.
How the agent puts it all together
When a failure comes in, the agent grabs the error details, searches memory for similar past incidents, and then asks the LLM to come up with a diagnosis based on both the current problem and what worked before:
class PipelineSageAgent:
def diagnose(self, deployment: dict) -> dict:
failure_query = (
f"Service: {deployment['service']} "
f"Error: {deployment['error_message']} "
f"Environment: {deployment['environment']}"
)
historical = self.memory.recall(failure_query, top_k=3)
prompt = self._build_prompt(deployment, historical)
response = self.llm_client.chat.completions.create(
model=os.getenv("GROQ_MODEL", "openai/gpt-oss-120b"),
messages=[{"role": "user", "content": prompt}]
)
return {
"diagnosis": response.choices[0].message.content,
"historical_context": historical
}
The important thing here is that the LLM isn't guessing on its own. It's looking at real history before answering.
A real example
Say deployment #1057 of payment-service fails because a database migration times out. PipelineSage searches its memory and finds deployment #1017 — same service, same kind of error. That time, the fix was to break the migration into batches of 500 records. It worked.
So instead of giving a vague suggestion like "try increasing the timeout," it recommends the batch approach — because that's what actually solved the problem before.
Once the engineer confirms it worked again, PipelineSage saves that result too. Now it has two examples to pull from next time.
What I learned along the way
Keep memories clean. Dumping raw logs into memory made search results messy. Short, clear summaries work much better — what broke, what fixed it, did it work.
Let humans confirm the fix. Just because a deployment passed doesn't mean the suggested fix was the reason. A quick "yes that worked" or "no it didn't" from a person keeps the memory honest.
History makes the AI way more useful. Without memory, the LLM gives generic advice. With memory, it gives advice based on things that actually happened. That's a big difference.
Keep the memory layer simple. I didn't build my own database or write custom search logic. Hindsight gave me save and search out of the box, and that was enough.
It gets smarter on its own. The more incidents it processes, the better it gets — without changing any code. The knowledge lives in the memory, not in the model.
What's next
Right now PipelineSage handles the basic loop: something breaks, it recalls similar problems, suggests a fix, and saves the outcome. The next steps are hooking it up directly to CI/CD platforms like GitHub Actions or Jenkins so failures come in automatically, and letting teams share memory so one team's fix can help another team with the same issue.
The bigger takeaway is simple. If you're building any kind of AI tool that deals with repeating problems, give it memory. It changes everything. The code for Hindsight is open source, and the docs explain how to get started.
Top comments (0)