DEV Community

James Joyner for DevOps AI ToolKit

Posted on

The AI Prompts I Actually Use On-Call (Copy-Paste)

I'm not here to tell you AI replaces on-call judgment. It doesn't. But a good large language model is a genuinely excellent rubber duck with a pattern-matching memory of every stack trace ever posted online — and when you're staring at an error at 2 a.m., that's worth a lot.

The trick is having the prompt ready before the incident, so you're not composing one while production burns. Here are the ones I actually keep on hand. Copy-paste, swap in your details.

One rule before we start: sanitize before you paste. Strip secrets, tokens, internal hostnames, customer data. Treat every prompt like it's going in a public pastebin, because functionally it is.

1. Triage an unknown error

You are a senior SRE. Here is an error from production:

<paste the error + the 20 lines of log around it>

Context: <service, language, runtime, what it does>.
Give me:
1. The 3 most likely root causes, ranked, with why.
2. The single fastest command to confirm or rule out #1.
3. What NOT to do (actions that look tempting but could make it worse).
Be concise. Assume I know the tools.
Enter fullscreen mode Exit fullscreen mode

The "what NOT to do" line is the one that's saved me. LLMs are eager to suggest force-flags; asking for the footguns up front surfaces them.

2. Translate a symptom into a query

I'm running Prometheus + Grafana. In plain English the symptom is:
"<e.g. checkout latency spikes every ~10 minutes for about 30 seconds>".

Write 3 PromQL queries that would help me confirm or localize this,
from broad to specific. Explain what each one would show and what
result would point to which cause.
Enter fullscreen mode Exit fullscreen mode

Turning "it feels slow sometimes" into three concrete PromQL queries is exactly the gap that eats MTTR.

3. Draft the postmortem while it's fresh

Write a blameless postmortem draft from these notes. Blameless means
focus on systems and contributing factors, never individuals.

Timeline + notes:
<paste your rough timeline>

Produce: summary, impact, timeline (as a table), root cause,
contributing factors, what went well, action items (each with a
concrete owner-role and a specific, testable definition of done).
Flag anything in my notes that's an assumption vs. a confirmed fact.
Enter fullscreen mode Exit fullscreen mode

That last line — assumptions vs. facts — catches the thing that quietly turns a postmortem into fiction.

4. Review a risky change before you run it

Review this Terraform plan output for anything destructive or
surprising before I apply it. Call out: resource replacements,
deletions, anything that would cause downtime, and any change to
security groups / IAM / public access. Rank by blast radius.

<paste `terraform plan` output>
Enter fullscreen mode Exit fullscreen mode

Great for the 200-line plan where the one scary line is hiding in the middle.

5. Summarize a wall of logs

Here are ~500 lines of logs from a failing service. Group them into
distinct error patterns, tell me how many times each occurs, order by
likely severity, and quote one representative line per group.

<paste logs>
Enter fullscreen mode Exit fullscreen mode

Turns a scroll-of-doom into a ranked list of "here are the 4 things actually going wrong."


Why these work

Notice the shared shape: give it a role, give it real context, and constrain the output. "Fix my error" gets you a wikipedia article. "You're a senior SRE, here's the log, give me 3 ranked causes and the one command to confirm the top one" gets you something you can act on in ten seconds.

I keep a much larger, categorized set of these — triage, tuning, security reviews, IaC, incident comms, postmortems — organized by tool and difficulty, every one with a worked example. It's free and copy-paste:

If you've got a prompt that's earned its place in your on-call kit, drop it in the comments — always looking for ones better than mine.

Top comments (0)