DEV Community

harshini bhupathiraju
harshini bhupathiraju

Posted on

Designing a Zero-Friction AI Dashboard for 3 AM PagerDuty Alerts

When a Sev-1 alert fires at 3 AM, the last thing a Site Reliability Engineer wants to do is have a conversational chat with an AI. They don't need a friendly greeting; they need root causes and safe, executable commands.

For our hackathon project, On-Call Hero, my teammates and I built an incredible backend using Groq and the Hindsight memory graph to stop AI hallucinations. My job was to build the frontend. I needed to translate that powerful stateful backend into a UI that an exhausted engineer could actually trust.

Here is how I built a fast, reactive SRE dashboard using Streamlit, and why UX in DevOps requires strict predictability over conversational fluff.

The UX Problem: Chatbots Belong in Customer Service, Not DevOps

Most AI agents default to a chat interface. But during a database outage, cognitive load is already maxed out. If an engineer has to type, "Can you analyze this Redis memory spike?" the tool is already adding friction.

I designed the On-Call Hero dashboard to be proactive, not reactive. Using Streamlit, I built a simulated PagerDuty integration. The moment an alert drops, the dashboard intercepts it and immediately passes the alert payload to the backend inference engine. Zero typing required.

Designing the Side-by-Side "Stateless vs. Stateful" View

Because we built this for a hackathon, the UI needed to instantly prove the value of our architecture to the judges. I decided the best way to show the power of Hindsight's memory graph was an A/B test built directly into the UI.

I used Streamlit's column layout to render the AI's responses side-by-side.


python
import streamlit as st
import json

# Simulated incoming PagerDuty Alert
st.header("🚨 Active Incident: `prod-cluster-3`")
st.warning(
    "PagerDuty: [SEV-1] Redis memory usage at 94% and climbing. "
    "Repeated OOM kills detected."
)

col1, col2 = st.columns(2)

with col1:
    st.subheader("🔴 Without Memory (Stateless)")
    # Renders the hallucinated, generic runbook
    st.markdown(stateless_llm_output)

with col2:
    st.subheader("🟢 With Hindsight Memory")
    # Renders the strict, graph-grounded JSON
    st.code(stateful_hindsight_output, language="json")'''




On the left, the UI renders the stateless model's output as raw markdown. It looks impressive at first glance, but it's full of hallucinated, dangerous commands.

On the right, the UI renders the response from our Hindsight-backed inference engine. Because the backend was heavily constrained to only output JSON based on historical memory, I could parse it cleanly into UI metric cards.

Parsing the JSON Payload Safely

When you force an LLM to return JSON, you have to build your frontend defensively. If the LLM misses a comma, the entire Streamlit app will throw a traceback error.

I built a strict parser to handle the response from the Groq backend:
def render_resolution_card(raw_json_string):
    try:
        data = json.loads(raw_json_string)

        # If Hindsight memory found no prior match
        if data.get("matched_incident_id") == "None":
            st.info("No historical precedent found. Escalating to human on-call.")
            return

        st.success(f"Historical Match Found: {data['matched_incident_id']}")
        st.metric(label="Root Cause", value=data["root_cause"])

        st.write("Safe Remediation Command:")
        st.code(data["recommended_fix_command"], language="bash")

    except json.JSONDecodeError:
        st.error("UI Error: Backend returned malformed payload.")
)
Predictability is the Best Feature

By combining Streamlit's fast prototyping with strict JSON parsing, we turned a raw LLM into a structured software tool. The UI doesn't say "Hello!" or guess at solutions. It either reads the Hindsight graph and returns the exact command that worked last time, or it safely admits it doesn't know.

If you are building AI tools for developers, skip the chatbot. Give them the data they need, formatted perfectly, exactly when they need it.
Enter fullscreen mode Exit fullscreen mode

Top comments (0)