Day 1/100.
For the past two months, I’ve been building an open-source tool for on-call engineers in my spare time. When a Kubernetes alert fires, it automatically gathers context and sends a hypothesis about the "why" to Telegram, complete with evidence and a way to verify if the hypothesis is incorrect. By design, it doesn't touch anything in the cluster. Yesterday, I ran it on a real EKS cluster for the first time; the cost was 13 cents. I measured accuracy honestly: the local model got 5 out of 7 right on tricky scenarios, while the cloud model got 6 out of 6.
I’m asking for one of two things: if you use Kubernetes and handle on-call duties, I’d love a 25-minute chat about your most recent overnight incident. Alternatively, you could install and break it, it’s a single helm install command (link in the comments).
⭐ I’m not asking for star on GitHub, I’m asking for your any opinion.

Top comments (3)
Gathering context plus a falsifiable hypothesis, and never touching the cluster, is the right boundary for this — the moment an on-call bot can act, every false positive becomes an incident of its own. And the numbers are the interesting part: 5/7 on the local model versus 6/6 on the cloud one over seven scenarios is a small sample, but it's an honest one, published instead of a claim about accuracy. Was 13c for that first real run mostly embedding and context gathering, or mostly retries?
The part I'd push on is the verification path. You already send a way to mark a hypothesis wrong, which most tools skip — what I'd want to know is whether those markings come back into the eval set, because that's the only realistic way to grow past seven cases without hand-writing fixtures. And when three alerts fire at once at 03:00, does it correlate them into a single hypothesis, or send three and make you do the joining?
Thank you, the questions are exactly the right ones.
The 13 cents were neither. That was the AWS bill for the run: one NAT gateway hour ($0.052), 0.74 GB of image pulls through it ($0.039), and the EKS control plane minutes. The model on that run was the stub on purpose (no model on the cluster, no API key), so the AI cost was exactly $0.00. On the local cluster the 7B model costs nothing and the only retries were on 429/503 from the free cloud tier, which are free too. Breakdown is in docs/EKS-RUN.md.
Verification: today /wrong records the verdict, and /scenario exports that incident as a replay scenario with the engineer's cause as the expected answer. What is still manual is getting it into the eval set; the next step is grading the current model against every verdict in the store, so the set grows from real incidents instead of fixtures. That is the top of the backlog now, partly because of this comment.
Three alerts at 03:00: if it is the same alert on many pods, it is one incident and one model call (shipped this week, thresholds at 2/5/10). If it is three different alerts on one namespace, today you get three messages and do the joining yourself. Correlating those into one hypothesis is the second item on the list.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.