<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vitaliy Dmitriev</title>
    <description>The latest articles on DEV Community by Vitaliy Dmitriev (@vitaliy_dmitriev).</description>
    <link>https://dev.to/vitaliy_dmitriev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4163505%2Fcf22ba5c-74df-4ba1-951a-f0f3ded6e32e.jpg</url>
      <title>DEV Community: Vitaliy Dmitriev</title>
      <link>https://dev.to/vitaliy_dmitriev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vitaliy_dmitriev"/>
    <language>en</language>
    <item>
      <title>I let an AI agent make an infrastructure change</title>
      <dc:creator>Vitaliy Dmitriev</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:59:37 +0000</pubDate>
      <link>https://dev.to/vitaliy_dmitriev/i-let-an-ai-agent-make-an-infrastructure-change-lbd</link>
      <guid>https://dev.to/vitaliy_dmitriev/i-let-an-ai-agent-make-an-infrastructure-change-lbd</guid>
      <description>&lt;h1&gt;
  
  
  I let an AI agent make an infrastructure change
&lt;/h1&gt;

&lt;p&gt;An AI agent recently took an infrastructure ticket from "please create these queues" to deployed and tested.&lt;/p&gt;

&lt;p&gt;It also:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;found a bug I would probably have shipped;&lt;/li&gt;
&lt;li&gt;briefly broke a consumer;&lt;/li&gt;
&lt;li&gt;confidently gave me the wrong explanation for some infrastructure versus code drift.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All three were useful.&lt;/p&gt;

&lt;p&gt;The first one showed me what agents can do.&lt;/p&gt;

&lt;p&gt;The other two showed me what they need before you let them touch infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ticket
&lt;/h2&gt;

&lt;p&gt;Typical BAU ticket:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Create a few message queues, configure dead-lettering and permissions, and make them available to the application.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Normally this is a small DevOps task.&lt;/p&gt;

&lt;p&gt;But infrastructure changes are rarely isolated. The broker was shared, there were several environments, existing applications depended on it, and some configuration was generated from code.&lt;/p&gt;

&lt;p&gt;So I decided not to do the change myself.&lt;/p&gt;

&lt;p&gt;I gave it to an AI coding agent running in a terminal.&lt;/p&gt;

&lt;p&gt;Its permissions were deliberately asymmetric:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It could read almost everything. It could change almost nothing without my approval.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My rules were simple:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;production is read-only;&lt;/li&gt;
&lt;li&gt;non-production changes require approval;&lt;/li&gt;
&lt;li&gt;changes must be reversible;&lt;/li&gt;
&lt;li&gt;don't guess when something is ambiguous.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent followed a fixed workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;explore
  ↓
plan + pre-mortem
  ↓
independent review
  ↓
questions to human
  ↓
implement
  ↓
non-production rollout
  ↓
test
  ↓
review
  ↓
handoff
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This turned out to matter more than the model itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first useful thing: it didn't start coding
&lt;/h2&gt;

&lt;p&gt;The agent first explored the environment.&lt;/p&gt;

&lt;p&gt;That immediately found several things I would have been tempted to skip.&lt;/p&gt;

&lt;h3&gt;
  
  
  The ticket was ambiguous
&lt;/h3&gt;

&lt;p&gt;The non-production broker was shared by several environments.&lt;/p&gt;

&lt;p&gt;The queue names in the ticket didn't contain an environment identifier, so blindly creating them would have made different test environments share the same queues.&lt;/p&gt;

&lt;p&gt;The agent stopped and asked which environments actually needed the change.&lt;/p&gt;

&lt;p&gt;Good catch.&lt;/p&gt;

&lt;h3&gt;
  
  
  There was already a pattern
&lt;/h3&gt;

&lt;p&gt;Myself and other team members had implemented something very similar before.&lt;/p&gt;

&lt;p&gt;Instead of inventing a new configuration, the agent found the existing implementation and used it as a reference.&lt;/p&gt;

&lt;p&gt;This is one of the least glamorous but most valuable things an agent can do:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;look around before inventing things.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  It found a potential time bomb
&lt;/h3&gt;

&lt;p&gt;The planned queue configuration had a dangerous default. Under the expected failure scenario, messages could accumulate in memory while the consumer was unavailable. Enough messages could eventually trigger the broker's memory protection and affect unrelated applications. The agent suggested adding a size limit and changing the storage policy.&lt;/p&gt;

&lt;p&gt;That was useful.&lt;/p&gt;

&lt;p&gt;But &lt;em&gt;I didn't trust the suggestion&lt;/em&gt; just because it sounded reasonable.&lt;/p&gt;

&lt;p&gt;I asked for an independent review.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second agent found the real problem
&lt;/h2&gt;

&lt;p&gt;A separate agent reviewed the plan. Its first finding was marked &lt;strong&gt;CRITICAL&lt;/strong&gt;. There was a regular expression containing an escaped dot.&lt;/p&gt;

&lt;p&gt;Something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;amq\.default
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Perfectly valid regex. Except this value eventually ended up inside JSON. And &lt;code&gt;\.&lt;/code&gt; isn't a valid JSON escape.&lt;/p&gt;

&lt;p&gt;The interesting part was when this would fail. Not during deployment. Not during the usual validation.&lt;/p&gt;

&lt;p&gt;The broker only read that particular configuration during a restart.&lt;/p&gt;

&lt;p&gt;So the change could have been deployed successfully, tested successfully, and then caused a failure weeks later when somebody restarted a node.&lt;/p&gt;

&lt;p&gt;That's exactly the kind of infrastructure bug I want automation to catch.&lt;/p&gt;

&lt;p&gt;The fix was tiny.&lt;/p&gt;

&lt;p&gt;The important change was in the process:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;don't validate the template. Validate the actual artefact the system will consume.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Render it. Parse it. Do that for every environment.&lt;/p&gt;

&lt;p&gt;The reviewer also found a few other problems in the planned tests and configuration.&lt;/p&gt;

&lt;p&gt;This was a good reminder that an agent doesn't have to produce the final answer to be useful.&lt;/p&gt;

&lt;p&gt;Sometimes &lt;em&gt;its best job&lt;/em&gt; is simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Here are five things you haven't thought about."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Then we changed the infrastructure
&lt;/h2&gt;

&lt;p&gt;After I answered its questions, the agent implemented the change. It didn't import an entire configuration file.&lt;/p&gt;

&lt;p&gt;It made targeted changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;create these objects
change these permissions
leave everything else alone
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That sounds obvious, but it's an important property for infrastructure automation.&lt;/p&gt;

&lt;p&gt;A command that says "make the whole configuration look like this" can have a very different blast radius from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;create X
create Y
modify Z
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then it tested the change using the real application identity.&lt;/p&gt;

&lt;p&gt;It tested the happy path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;publish
→ consume
→ fail
→ retry
→ retry
→ retry
→ dead-letter
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It also tested the negative cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Could the application access things it shouldn't?&lt;/li&gt;
&lt;li&gt;Could it create things it shouldn't?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The answer needed to be no.&lt;/p&gt;

&lt;p&gt;And then came one of my favourite tests. The agent deliberately filled the queue until the configured limit was reached.&lt;/p&gt;

&lt;p&gt;If you configure a limit but never actually hit it, you don't really know whether the limit works. A limit you haven't triggered is still partly a theory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I let it restart the nodes
&lt;/h2&gt;

&lt;p&gt;The infrastructure change also needed a rolling restart.&lt;/p&gt;

&lt;p&gt;Before doing that, the agent checked for configuration differences between the repository and the live system. It also checked whether there were messages that could be affected by the restart. There was one queue containing messages that would not survive the restart.&lt;/p&gt;

&lt;p&gt;So it asked me andI approved it.&lt;/p&gt;

&lt;p&gt;The restart began, and everything looked great: all nodes were healthy, replicas were synchronized, no broker alarms, the dashboard was green.&lt;/p&gt;

&lt;p&gt;Then the agent found a problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The dashboard was green. The application wasn't.
&lt;/h2&gt;

&lt;p&gt;One test consumer had stopped consuming, while it's Pod was healthy.&lt;/p&gt;

&lt;p&gt;The cluster was healthy, the monitoring dashboard was mostly green, BUT the application was still broken.&lt;/p&gt;

&lt;p&gt;The restart had caused the application to reconnect and redeclare its queue. The queue already existed with slightly different settings from what the application expected.&lt;/p&gt;

&lt;p&gt;The broker rejected the declaration and the application stopped consuming, while its health check didn't notice.&lt;/p&gt;

&lt;p&gt;This lasted for about ten minutes. No customer impact, because it was a test environment.&lt;/p&gt;

&lt;p&gt;But it exposed a very important problem with the original verification:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I was checking server health, not client health.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The agent had done what I asked, the cluster was healthy, but that wasn't enough. :(&lt;/p&gt;

&lt;p&gt;So we changed the procedure - before every disruptive operation, we now capture the important client-side state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;consumers
connections
restart counts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After the operation, we compare it. Cluster health tells you that the cluster is healthy, but it doesn't tell you that your applications are happy.&lt;/p&gt;

&lt;p&gt;Those are two different questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  So we changed the rules
&lt;/h2&gt;

&lt;p&gt;Each failure became a new rule.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Problem&lt;/th&gt;
&lt;th&gt;New rule&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A restart broke a consumer&lt;/td&gt;
&lt;td&gt;Notify the human immediately when user-facing behaviour changes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cluster was healthy but client was broken&lt;/td&gt;
&lt;td&gt;Compare client-side state before and after disruptive operations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Drift" was based on current Git state only&lt;/td&gt;
&lt;td&gt;Check Git history before calling something drift&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Review repeated facts from the plan&lt;/td&gt;
&lt;td&gt;Reviewers must independently verify important facts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A plausible statement had no evidence&lt;/td&gt;
&lt;td&gt;Every important fact must have a command, query or source behind it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A configuration file could fail only after restart&lt;/td&gt;
&lt;td&gt;Validate the exact artefact consumed by the system&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is probably the most important part of the whole experiment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The agent is allowed to be wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The system around the agent shouldn't make the same mistake twice.&lt;/p&gt;

&lt;p&gt;The workflow itself now contains these checks, so I don't have to remember them next time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for a business
&lt;/h2&gt;

&lt;p&gt;If you're thinking about using AI to automate DevOps or infrastructure work, I wouldn't start by giving an agent production access. I'd start much more boringly.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Give it read access first
&lt;/h3&gt;

&lt;p&gt;Let it investigate.&lt;/p&gt;

&lt;p&gt;Let it correlate logs and metrics.&lt;/p&gt;

&lt;p&gt;Let it find configuration.&lt;/p&gt;

&lt;p&gt;Let it prepare a change.&lt;/p&gt;

&lt;p&gt;You can learn a lot about the agent without giving it the ability to break anything.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Gate writes with credentials
&lt;/h3&gt;

&lt;p&gt;This is important.&lt;/p&gt;

&lt;p&gt;Don't rely on:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Don't touch production."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's a prompt.&lt;/p&gt;

&lt;p&gt;Instead:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The production credential literally cannot modify production.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's a control.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Keep approval where the blast radius begins
&lt;/h3&gt;

&lt;p&gt;The agent can prepare the change.&lt;/p&gt;

&lt;p&gt;The human approves it.&lt;/p&gt;

&lt;p&gt;You can gradually automate more once you understand the failure modes.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Make the agent ask questions efficiently
&lt;/h3&gt;

&lt;p&gt;One message with five numbered questions is useful.&lt;/p&gt;

&lt;p&gt;Five separate interruptions are annoying.&lt;/p&gt;

&lt;p&gt;A good agent should come back with something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. These four environments need the queues. Agree?
2. Existing pattern A looks closest. Use it?
3. This policy may affect memory usage. Apply it?
4. One queue contains messages that won't survive restart. Proceed?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That turns the human into &lt;em&gt;a decision maker&lt;/em&gt; instead of a human API.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Turn mistakes into guardrails
&lt;/h3&gt;

&lt;p&gt;This is where I think agentic automation becomes really interesting.&lt;/p&gt;

&lt;p&gt;A mistake shouldn't just produce:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The AI made a mistake."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It should produce:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"What rule can we add so neither the AI nor the next engineer makes this mistake again?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's how the system gets better.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually changed for me
&lt;/h2&gt;

&lt;p&gt;The agent didn't replace the DevOps engineer.&lt;/p&gt;

&lt;p&gt;It changed what the DevOps engineer was doing.&lt;/p&gt;

&lt;p&gt;Instead of spending the morning:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;read ticket
find config
search Git
check cluster
write YAML
run commands
test
write documentation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I spent much more of the time asking:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What do we actually need to change?

What could go wrong?

How will we know it worked?

What must never be allowed?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those are much more valuable questions to automate around.&lt;/p&gt;

&lt;p&gt;And there is an interesting business implication here.&lt;/p&gt;

&lt;p&gt;You don't necessarily need an AI that can autonomously run your entire infrastructure. You need an AI that can take a well-defined piece of infrastructure work and move it through a &lt;strong&gt;safe, repeatable process&lt;/strong&gt;, while escalating the decisions that actually require human judgement.&lt;/p&gt;

&lt;p&gt;That's a much smaller problem.&lt;/p&gt;

&lt;p&gt;And, in my experience, a much more useful one.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you're starting with AI in DevOps
&lt;/h2&gt;

&lt;p&gt;Three habits are worth stealing even if you never use an AI agent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prove your limits.&lt;/strong&gt; If you configure a timeout, retry count or size limit, trigger it deliberately and see what happens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check clients, not only servers.&lt;/strong&gt; "The cluster is healthy" does not mean "the application works."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent didn't replace judgement. It made the places where judgement was missing much easier to see.&lt;/p&gt;

&lt;p&gt;And that's probably the part of this experiment I found most valuable.&lt;/p&gt;




&lt;p&gt;I'm experimenting with AI-assisted infrastructure automation, particularly around DevOps and on-prem environments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Would you trust an AI agent to make a production infrastructure change if the permissions, approval gates and rollback path were designed correctly?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'm curious where other teams draw that line.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>automation</category>
      <category>sre</category>
    </item>
    <item>
      <title>Don't put an LLM where a bash script will do</title>
      <dc:creator>Vitaliy Dmitriev</dc:creator>
      <pubDate>Tue, 06 Oct 2026 08:58:11 +0000</pubDate>
      <link>https://dev.to/vitaliy_dmitriev/dont-put-an-llm-where-a-bash-script-will-do-728</link>
      <guid>https://dev.to/vitaliy_dmitriev/dont-put-an-llm-where-a-bash-script-will-do-728</guid>
      <description>&lt;h1&gt;
  
  
  Don't put an LLM where a bash script will do
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Or: how I almost gave a language model a login, a service account and a cron job to check four numbers.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Recently I was debugging a production Prometheus exporter that was restarting during the night. It wasn't crashing in the usual sense. It was getting killed by Kubernetes because its liveness probe timed out.&lt;/p&gt;

&lt;p&gt;The reason was fairly simple: the exporter depended on another service that became slow when the nightly jobs started. The probe had a short timeout, so Kubernetes eventually decided the exporter was dead and restarted it.&lt;/p&gt;

&lt;p&gt;The fix was boring: increase the probe timeout, adjust the monitoring timeout, deploy the change.&lt;/p&gt;

&lt;p&gt;The interesting question was what happens after the deployment.&lt;/p&gt;

&lt;p&gt;The fix only really matters at 3 a.m. That's when the nightly load happens. I didn't want to wake up at 3 a.m. just to check whether the exporter survived. So I needed some kind of automated verification.&lt;/p&gt;

&lt;h2&gt;
  
  
  The obvious AI solution
&lt;/h2&gt;

&lt;p&gt;I had been using an AI agent during the investigation and it was actually quite useful. It helped correlate the restarts with the nightly load, look through metrics and logs, find a problem with one of my assumptions, and prepare the ticket and merge requests.&lt;/p&gt;

&lt;p&gt;So the next idea seemed reasonable:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Run the agent every morning for a few days and ask it whether the fix worked.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent could query the Prometheus, look at the deployment and send me a message. Then I started thinking about what would actually be required to run an agent unattended.&lt;/p&gt;

&lt;p&gt;It would need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a dedicated non-root user;&lt;/li&gt;
&lt;li&gt;its own authentication to the model provider;&lt;/li&gt;
&lt;li&gt;read-only access to the monitoring system;&lt;/li&gt;
&lt;li&gt;some kind of K8S credentials if it needed its data;&lt;/li&gt;
&lt;li&gt;a restricted tool set;&lt;/li&gt;
&lt;li&gt;a timeout and a way to stop it;&lt;/li&gt;
&lt;li&gt;logs;&lt;/li&gt;
&lt;li&gt;and probably some tests around the whole thing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's a lot of infrastructure for a question that turned out to be:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Did the exporter restart?&lt;/li&gt;
&lt;li&gt;Was it up during the whole observation period?&lt;/li&gt;
&lt;li&gt;What was the maximum successful scrape duration?&lt;/li&gt;
&lt;li&gt;Did the alert fire?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's four PromQL queries and four comparisons.&lt;/p&gt;

&lt;p&gt;At that point I stopped.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boring solution
&lt;/h2&gt;

&lt;p&gt;I wrote a shell script instead. It was around 80 lines, most of which were notification and error handling. The actual check was tiny.&lt;/p&gt;

&lt;p&gt;Something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;query&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    curl &lt;span class="nt"&gt;-sk&lt;/span&gt; &lt;span class="nt"&gt;--max-time&lt;/span&gt; 25 &lt;span class="nt"&gt;-G&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROM&lt;/span&gt;&lt;span class="s2"&gt;/api/v1/query"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
        &lt;span class="nt"&gt;--data-urlencode&lt;/span&gt; &lt;span class="s2"&gt;"query=&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; |
        jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'
            select(.status == "success") |
            .data.result[0].value[1] // "none"
        '&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

check&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nb"&gt;local &lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="nb"&gt;local &lt;/span&gt;&lt;span class="nv"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$2&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="nb"&gt;local &lt;/span&gt;&lt;span class="nv"&gt;condition&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$3&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

    &lt;span class="nb"&gt;local &lt;/span&gt;value
    &lt;span class="nv"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;query &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$query&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        errors+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="s2"&gt;: query failed"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$value&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s2"&gt;"none"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        errors+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="s2"&gt;: no data"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    jq &lt;span class="nt"&gt;-en&lt;/span&gt; &lt;span class="nt"&gt;--arg&lt;/span&gt; v &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$value&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
        &lt;span class="s1"&gt;'($v | tonumber) as $v | '&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$condition&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null &lt;span class="o"&gt;||&lt;/span&gt;
        failures+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="s2"&gt; = &lt;/span&gt;&lt;span class="nv"&gt;$value&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

check restarts &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'...'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'$v == 0'&lt;/span&gt;

check &lt;span class="nb"&gt;uptime&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'...'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'$v == 1'&lt;/span&gt;

check scrape_time &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'...'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'$v &amp;lt; 25'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The real script had a few more checks and notification handling, but the principle was the same.&lt;/p&gt;

&lt;p&gt;There were three possible results:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CONFIRMED&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Everything was within the expected limits.&lt;/p&gt;

&lt;p&gt;No notification during the observation period. On the last day, send one short confirmation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;REGRESSED&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Something went wrong.&lt;/p&gt;

&lt;p&gt;Send me a message with the relevant result and the rollback procedure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;INCONCLUSIVE&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The monitoring query failed, there was no data, or the thing being checked no longer existed.&lt;/p&gt;

&lt;p&gt;Send a message too.&lt;/p&gt;

&lt;p&gt;This last case is important:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Silence should mean "checked and everything is fine", not "the checker failed".&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So I also wrote one JSONL record for every run. Nothing fancy, just enough to know that the check actually happened.&lt;/p&gt;

&lt;h3&gt;
  
  
  One useful detail
&lt;/h3&gt;

&lt;p&gt;The script also checked that the deployment being verified was still the one I intended to verify.&lt;/p&gt;

&lt;p&gt;I used an identifier associated with the newly deployed workload rather than simply asking Prometheus about "the exporter".&lt;/p&gt;

&lt;p&gt;That matters because deployments can change while your verification job is waiting.&lt;/p&gt;

&lt;p&gt;If somebody rolls back or deploys another version, the original workload disappears. In that case the script reports &lt;code&gt;INCONCLUSIVE&lt;/code&gt; instead of accidentally verifying the new deployment.&lt;/p&gt;

&lt;p&gt;It's a small detail, but it prevents a particularly annoying kind of false positive.&lt;/p&gt;

&lt;h2&gt;
  
  
  systemd instead of cron
&lt;/h2&gt;

&lt;p&gt;I could have used cron, but systemd made the rest of this surprisingly simple.&lt;/p&gt;

&lt;p&gt;The service looked roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[Unit]&lt;/span&gt;
&lt;span class="py"&gt;ConditionPathExists&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;!/etc/example/PAUSE&lt;/span&gt;

&lt;span class="nn"&gt;[Service]&lt;/span&gt;
&lt;span class="py"&gt;Type&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;oneshot&lt;/span&gt;
&lt;span class="py"&gt;User&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;example-verify&lt;/span&gt;
&lt;span class="py"&gt;ExecStart&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;/usr/local/libexec/example-verify/verify.sh&lt;/span&gt;
&lt;span class="py"&gt;TimeoutStartSec&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;300&lt;/span&gt;

&lt;span class="py"&gt;StateDirectory&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;example-verify&lt;/span&gt;
&lt;span class="py"&gt;LoadCredential&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;notifier.env:/path/to/notifier.env&lt;/span&gt;

&lt;span class="py"&gt;ProtectSystem&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;strict&lt;/span&gt;
&lt;span class="py"&gt;ProtectHome&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;yes&lt;/span&gt;
&lt;span class="py"&gt;PrivateTmp&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;yes&lt;/span&gt;
&lt;span class="py"&gt;NoNewPrivileges&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the timer contained the few dates on which I wanted the check to run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[Timer]&lt;/span&gt;
&lt;span class="py"&gt;OnCalendar&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;
&lt;span class="py"&gt;OnCalendar&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;
&lt;span class="py"&gt;OnCalendar&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;

&lt;span class="py"&gt;Persistent&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There were a couple of nice things here.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;LoadCredential=&lt;/code&gt; lets systemd provide the notification credential to the service without copying it into the script's environment or configuration.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;ConditionPathExists=!PAUSE&lt;/code&gt; gives me a simple kill switch. Create the file and systemd doesn't start the service.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And because the timer had a finite number of scheduled runs, it didn't become another forgotten cron job that would still be running six months later ;)&lt;/p&gt;

&lt;p&gt;There was also a small systemd gotcha. I initially used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="py"&gt;RuntimeMaxSec&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;300&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;because that's what I wanted: kill the process after five minutes. But it appeared that if service is &lt;code&gt;Type=oneshot&lt;/code&gt;, that isn't the right setting for limiting its startup/run phase.&lt;/p&gt;

&lt;p&gt;The correct option here was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="py"&gt;TimeoutStartSec&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;300&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;systemd-analyze verify&lt;/code&gt; happily caught it. It's worth running, costs nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the LLM actually helped
&lt;/h2&gt;

&lt;p&gt;This isn't an anti-LLM story, on the contrary the LLM was useful during the investigation.&lt;/p&gt;

&lt;p&gt;The interesting part was figuring out &lt;strong&gt;what should be checked&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example, one of the initial metrics I was looking at seemed to show that the exporter had enough headroom, while second review showed that this conclusion was wrong.&lt;/p&gt;

&lt;p&gt;The apparent maximum duration was based on successful scrapes. When a scrape exceeds the timeout, it can fail instead of giving you a nice metric saying "this scrape took 27 seconds". So a graph showing a maximum of 9 seconds doesn't necessarily mean the exporter never took longer than 9 seconds. It can mean that 9 seconds was the longest successful scrape.&lt;/p&gt;

&lt;p&gt;The failed scrapes were the important part. That kind of investigation is where an LLM is useful:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;correlate different signals;&lt;/li&gt;
&lt;li&gt;suggest what to look at;&lt;/li&gt;
&lt;li&gt;challenge assumptions;&lt;/li&gt;
&lt;li&gt;generate queries;&lt;/li&gt;
&lt;li&gt;review a proposed solution;&lt;/li&gt;
&lt;li&gt;write the boring documentation and ticket text.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My current rule
&lt;/h2&gt;

&lt;p&gt;I ended up with a fairly simple rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If the pass/fail condition can be written as a comparison, use a script. / If you still need to figure out what should be compared, an LLM can be useful.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The important part is not "bash good, AI bad". It's rather about using the right amount of machinery.&lt;/p&gt;

&lt;p&gt;An LLM is useful when there is ambiguity and you need judgement. A script is useful when the decision has already been made and needs to be repeated reliably (not "re-decided each and every time").&lt;/p&gt;

&lt;p&gt;In this case, using an agent would have meant building a small security boundary around a system whose job was basically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;query
query
query
query
compare
notify
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The shell script was easier to understand, easier to secure and easier to debug.&lt;/p&gt;

&lt;p&gt;And it was considerably cheaper ;)&lt;/p&gt;

&lt;h2&gt;
  
  
  One last surprise
&lt;/h2&gt;

&lt;p&gt;After three nights the check reported exactly what I wanted: no restarts and scrape times comfortably below the new timeout. So I deleted the timer, the service, the script and the temporary user.&lt;/p&gt;

&lt;p&gt;Except I didn't.&lt;/p&gt;

&lt;p&gt;The cleanup command appeared to run successfully but didn't remove the files. The reason was my shell configuration.&lt;/p&gt;

&lt;p&gt;My interactive shell had an &lt;code&gt;rm -i&lt;/code&gt; alias. The non-interactive environment inherited it. Without a terminal to answer the confirmation prompt, &lt;code&gt;rm -i&lt;/code&gt; effectively did nothing.&lt;/p&gt;

&lt;p&gt;That was a good reminder that automation doesn't necessarily run in a clean environment. Your shell configuration is part of the environment too.&lt;/p&gt;

&lt;p&gt;If you use interactive-only aliases, make sure they are actually limited to interactive shells:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="nv"&gt;$-&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt;i&lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nt"&gt;-z&lt;/span&gt; &lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;AGENT_SHELL&lt;/span&gt;&lt;span class="k"&gt;:-}&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;alias rm&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'rm -i'&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use whatever environment variable or detection mechanism makes sense for your agent.&lt;/p&gt;

&lt;p&gt;The broader lesson was probably the same as the one from the beginning:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use the clever tool to figure out what needs to be checked. Use the boring tool to check it.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>ai</category>
      <category>kubernetes</category>
      <category>monitoring</category>
    </item>
  </channel>
</rss>
