The 2 AM problem
Every data team has lived this. A pipeline fails in the middle of the night. Somebody gets paged, stares at a stack trace, and thinks, "I'm sure we've seen this before." Somebody probably has. They just aren't online.
The frustrating part is how repetitive it is. Upstream teams rename columns at month-end. A Synapse SQL pool gets paused right before the hourly job. A CRM export gets cut off halfway and nulls spike. The fix for each of these is known, but it lives in one person's head or buried in an old Slack thread.
For the Hack with Hyderabad hackathon, I wanted to build something that fixes that: an agent that actually remembers.
The idea
PipeMind is an on-call agent for data teams. It does three things:
- Stores every incident, what was tried, and whether the fix worked.
- Recalls similar past incidents when something new breaks.
- Uses that history to diagnose the failure, preferring fixes that worked and avoiding ones that didn't.
The detail I care about most is that it remembers failures, not just successes. A generic LLM will happily say "increase the timeout." PipeMind can say "we tried that in June and it didn't help, because the SQL pool was paused."
How it works
The stack is small on purpose:
- Hindsight (by Vectorize) is the memory layer.
-
Groq runs the reasoning, using
openai/gpt-oss-120b. - Streamlit is the interface.
Hindsight has three operations. If you know SQL, they map neatly:
| Hindsight | SQL analogy | What PipeMind uses it for |
|---|---|---|
retain |
INSERT |
Save an incident and its outcome |
recall |
SELECT ... WHERE relevant |
Fetch similar past incidents |
reflect |
GROUP BY across history |
Summarise fragile pipelines and recurring causes |
When you click Diagnose, this happens:
- You enter the pipeline, date and error message.
- PipeMind calls
recalland gets the most relevant past incidents. - Those incidents plus the new error go into a prompt that says: prefer fixes that worked, never repeat one that failed.
- Groq returns a diagnosis, and a "Memories used" panel shows exactly what was recalled.
- You click Fix worked or Fix did not work, and that outcome is retained. That feedback loop is what makes it learn instead of just retrieve.
The data
A memory demo is only convincing if the history looks real, so I built 48 incidents across 6 pipelines on Airflow, Azure Data Factory and dbt. They include recurring patterns like month-end schema drift and paused SQL pools, plus plenty of "the first fix failed, the second one worked" pairs, so the agent has real experience to learn from.
What changed with memory on
I added a Memory ON/OFF toggle and a side-by-side mode, so the same failure is answered with and without history.
Here is the memory graph Hindsight built from the incidents:
What I learned
1. Failures are as valuable as successes. Storing "this fix did not work" is what stops the agent repeating bad advice.
2. Realistic data matters more than clever code. Believable pipeline names, error messages and dates make the whole demo feel real.
3. Check your assumptions early. My connectivity test caught that one of the fallback Groq models had been deprecated. A fallback that silently doesn't exist is worse than no fallback, so I swapped it out. I'm glad I tested it before the demo.
What's next
- Slack alerts so the diagnosis lands where on-call already is
- Real webhooks from Airflow and Azure Data Factory
- Separate memory banks per team
Try it
- Code: 👉 https://github.com/amaresh0812/PipeMind
If your team has its own "only Priya knows how to fix this" stories, I'd love to hear them.


Top comments (0)