<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vitaliy</title>
    <description>The latest articles on DEV Community by Vitaliy (@vi7al).</description>
    <link>https://dev.to/vi7al</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4163505%2Fcf22ba5c-74df-4ba1-951a-f0f3ded6e32e.jpg</url>
      <title>DEV Community: Vitaliy</title>
      <link>https://dev.to/vi7al</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vi7al"/>
    <language>en</language>
    <item>
      <title>Don't put an LLM where a bash script will do</title>
      <dc:creator>Vitaliy</dc:creator>
      <pubDate>Tue, 06 Oct 2026 08:58:11 +0000</pubDate>
      <link>https://dev.to/vi7al/dont-put-an-llm-where-a-bash-script-will-do-728</link>
      <guid>https://dev.to/vi7al/dont-put-an-llm-where-a-bash-script-will-do-728</guid>
      <description>&lt;h1&gt;
  
  
  Don't put an LLM where a bash script will do
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Or: how I almost gave a language model a login, a service account and a cron job to check four numbers.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Recently I was debugging a production Prometheus exporter that was restarting during the night. It wasn't crashing in the usual sense. It was getting killed by Kubernetes because its liveness probe timed out.&lt;/p&gt;

&lt;p&gt;The reason was fairly simple: the exporter depended on another service that became slow when the nightly jobs started. The probe had a short timeout, so Kubernetes eventually decided the exporter was dead and restarted it.&lt;/p&gt;

&lt;p&gt;The fix was boring: increase the probe timeout, adjust the monitoring timeout, deploy the change.&lt;/p&gt;

&lt;p&gt;The interesting question was what happens after the deployment.&lt;/p&gt;

&lt;p&gt;The fix only really matters at 3 a.m. That's when the nightly load happens. I didn't want to wake up at 3 a.m. just to check whether the exporter survived. So I needed some kind of automated verification.&lt;/p&gt;

&lt;h2&gt;
  
  
  The obvious AI solution
&lt;/h2&gt;

&lt;p&gt;I had been using an AI agent during the investigation and it was actually quite useful. It helped correlate the restarts with the nightly load, look through metrics and logs, find a problem with one of my assumptions, and prepare the ticket and merge requests.&lt;/p&gt;

&lt;p&gt;So the next idea seemed reasonable:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Run the agent every morning for a few days and ask it whether the fix worked.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent could query the Prometheus, look at the deployment and send me a message. Then I started thinking about what would actually be required to run an agent unattended.&lt;/p&gt;

&lt;p&gt;It would need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a dedicated non-root user;&lt;/li&gt;
&lt;li&gt;its own authentication to the model provider;&lt;/li&gt;
&lt;li&gt;read-only access to the monitoring system;&lt;/li&gt;
&lt;li&gt;some kind of K8S credentials if it needed its data;&lt;/li&gt;
&lt;li&gt;a restricted tool set;&lt;/li&gt;
&lt;li&gt;a timeout and a way to stop it;&lt;/li&gt;
&lt;li&gt;logs;&lt;/li&gt;
&lt;li&gt;and probably some tests around the whole thing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's a lot of infrastructure for a question that turned out to be:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Did the exporter restart?&lt;/li&gt;
&lt;li&gt;Was it up during the whole observation period?&lt;/li&gt;
&lt;li&gt;What was the maximum successful scrape duration?&lt;/li&gt;
&lt;li&gt;Did the alert fire?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's four PromQL queries and four comparisons.&lt;/p&gt;

&lt;p&gt;At that point I stopped.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boring solution
&lt;/h2&gt;

&lt;p&gt;I wrote a shell script instead. It was around 80 lines, most of which were notification and error handling. The actual check was tiny.&lt;/p&gt;

&lt;p&gt;Something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;query&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    curl &lt;span class="nt"&gt;-sk&lt;/span&gt; &lt;span class="nt"&gt;--max-time&lt;/span&gt; 25 &lt;span class="nt"&gt;-G&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$PROM&lt;/span&gt;&lt;span class="s2"&gt;/api/v1/query"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
        &lt;span class="nt"&gt;--data-urlencode&lt;/span&gt; &lt;span class="s2"&gt;"query=&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; |
        jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'
            select(.status == "success") |
            .data.result[0].value[1] // "none"
        '&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

check&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="nb"&gt;local &lt;/span&gt;&lt;span class="nv"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="nb"&gt;local &lt;/span&gt;&lt;span class="nv"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$2&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="nb"&gt;local &lt;/span&gt;&lt;span class="nv"&gt;condition&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$3&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

    &lt;span class="nb"&gt;local &lt;/span&gt;value
    &lt;span class="nv"&gt;value&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;query &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$query&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        errors+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="s2"&gt;: query failed"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$value&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s2"&gt;"none"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        errors+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="s2"&gt;: no data"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    jq &lt;span class="nt"&gt;-en&lt;/span&gt; &lt;span class="nt"&gt;--arg&lt;/span&gt; v &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$value&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
        &lt;span class="s1"&gt;'($v | tonumber) as $v | '&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$condition&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null &lt;span class="o"&gt;||&lt;/span&gt;
        failures+&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$name&lt;/span&gt;&lt;span class="s2"&gt; = &lt;/span&gt;&lt;span class="nv"&gt;$value&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

check restarts &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'...'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'$v == 0'&lt;/span&gt;

check &lt;span class="nb"&gt;uptime&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'...'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'$v == 1'&lt;/span&gt;

check scrape_time &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'...'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s1"&gt;'$v &amp;lt; 25'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The real script had a few more checks and notification handling, but the principle was the same.&lt;/p&gt;

&lt;p&gt;There were three possible results:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CONFIRMED&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Everything was within the expected limits.&lt;/p&gt;

&lt;p&gt;No notification during the observation period. On the last day, send one short confirmation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;REGRESSED&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Something went wrong.&lt;/p&gt;

&lt;p&gt;Send me a message with the relevant result and the rollback procedure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;INCONCLUSIVE&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The monitoring query failed, there was no data, or the thing being checked no longer existed.&lt;/p&gt;

&lt;p&gt;Send a message too.&lt;/p&gt;

&lt;p&gt;This last case is important:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Silence should mean "checked and everything is fine", not "the checker failed".&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So I also wrote one JSONL record for every run. Nothing fancy, just enough to know that the check actually happened.&lt;/p&gt;

&lt;h3&gt;
  
  
  One useful detail
&lt;/h3&gt;

&lt;p&gt;The script also checked that the deployment being verified was still the one I intended to verify.&lt;/p&gt;

&lt;p&gt;I used an identifier associated with the newly deployed workload rather than simply asking Prometheus about "the exporter".&lt;/p&gt;

&lt;p&gt;That matters because deployments can change while your verification job is waiting.&lt;/p&gt;

&lt;p&gt;If somebody rolls back or deploys another version, the original workload disappears. In that case the script reports &lt;code&gt;INCONCLUSIVE&lt;/code&gt; instead of accidentally verifying the new deployment.&lt;/p&gt;

&lt;p&gt;It's a small detail, but it prevents a particularly annoying kind of false positive.&lt;/p&gt;

&lt;h2&gt;
  
  
  systemd instead of cron
&lt;/h2&gt;

&lt;p&gt;I could have used cron, but systemd made the rest of this surprisingly simple.&lt;/p&gt;

&lt;p&gt;The service looked roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[Unit]&lt;/span&gt;
&lt;span class="py"&gt;ConditionPathExists&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;!/etc/example/PAUSE&lt;/span&gt;

&lt;span class="nn"&gt;[Service]&lt;/span&gt;
&lt;span class="py"&gt;Type&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;oneshot&lt;/span&gt;
&lt;span class="py"&gt;User&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;example-verify&lt;/span&gt;
&lt;span class="py"&gt;ExecStart&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;/usr/local/libexec/example-verify/verify.sh&lt;/span&gt;
&lt;span class="py"&gt;TimeoutStartSec&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;300&lt;/span&gt;

&lt;span class="py"&gt;StateDirectory&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;example-verify&lt;/span&gt;
&lt;span class="py"&gt;LoadCredential&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;notifier.env:/path/to/notifier.env&lt;/span&gt;

&lt;span class="py"&gt;ProtectSystem&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;strict&lt;/span&gt;
&lt;span class="py"&gt;ProtectHome&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;yes&lt;/span&gt;
&lt;span class="py"&gt;PrivateTmp&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;yes&lt;/span&gt;
&lt;span class="py"&gt;NoNewPrivileges&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;yes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the timer contained the few dates on which I wanted the check to run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[Timer]&lt;/span&gt;
&lt;span class="py"&gt;OnCalendar&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;
&lt;span class="py"&gt;OnCalendar&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;
&lt;span class="py"&gt;OnCalendar&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;

&lt;span class="py"&gt;Persistent&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There were a couple of nice things here.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;LoadCredential=&lt;/code&gt; lets systemd provide the notification credential to the service without copying it into the script's environment or configuration.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;ConditionPathExists=!PAUSE&lt;/code&gt; gives me a simple kill switch. Create the file and systemd doesn't start the service.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And because the timer had a finite number of scheduled runs, it didn't become another forgotten cron job that would still be running six months later ;)&lt;/p&gt;

&lt;p&gt;There was also a small systemd gotcha. I initially used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="py"&gt;RuntimeMaxSec&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;300&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;because that's what I wanted: kill the process after five minutes. But it appeared that if service is &lt;code&gt;Type=oneshot&lt;/code&gt;, that isn't the right setting for limiting its startup/run phase.&lt;/p&gt;

&lt;p&gt;The correct option here was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="py"&gt;TimeoutStartSec&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;300&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;systemd-analyze verify&lt;/code&gt; happily caught it. It's worth running, costs nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the LLM actually helped
&lt;/h2&gt;

&lt;p&gt;This isn't an anti-LLM story, on the contrary the LLM was useful during the investigation.&lt;/p&gt;

&lt;p&gt;The interesting part was figuring out &lt;strong&gt;what should be checked&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example, one of the initial metrics I was looking at seemed to show that the exporter had enough headroom, while second review showed that this conclusion was wrong.&lt;/p&gt;

&lt;p&gt;The apparent maximum duration was based on successful scrapes. When a scrape exceeds the timeout, it can fail instead of giving you a nice metric saying "this scrape took 27 seconds". So a graph showing a maximum of 9 seconds doesn't necessarily mean the exporter never took longer than 9 seconds. It can mean that 9 seconds was the longest successful scrape.&lt;/p&gt;

&lt;p&gt;The failed scrapes were the important part. That kind of investigation is where an LLM is useful:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;correlate different signals;&lt;/li&gt;
&lt;li&gt;suggest what to look at;&lt;/li&gt;
&lt;li&gt;challenge assumptions;&lt;/li&gt;
&lt;li&gt;generate queries;&lt;/li&gt;
&lt;li&gt;review a proposed solution;&lt;/li&gt;
&lt;li&gt;write the boring documentation and ticket text.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My current rule
&lt;/h2&gt;

&lt;p&gt;I ended up with a fairly simple rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If the pass/fail condition can be written as a comparison, use a script. / If you still need to figure out what should be compared, an LLM can be useful.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The important part is not "bash good, AI bad". It's rather about using the right amount of machinery.&lt;/p&gt;

&lt;p&gt;An LLM is useful when there is ambiguity and you need judgement. A script is useful when the decision has already been made and needs to be repeated reliably (not "re-decided each and every time").&lt;/p&gt;

&lt;p&gt;In this case, using an agent would have meant building a small security boundary around a system whose job was basically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;query
query
query
query
compare
notify
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The shell script was easier to understand, easier to secure and easier to debug.&lt;/p&gt;

&lt;p&gt;And it was considerably cheaper ;)&lt;/p&gt;

&lt;h2&gt;
  
  
  One last surprise
&lt;/h2&gt;

&lt;p&gt;After three nights the check reported exactly what I wanted: no restarts and scrape times comfortably below the new timeout. So I deleted the timer, the service, the script and the temporary user.&lt;/p&gt;

&lt;p&gt;Except I didn't.&lt;/p&gt;

&lt;p&gt;The cleanup command appeared to run successfully but didn't remove the files. The reason was my shell configuration.&lt;/p&gt;

&lt;p&gt;My interactive shell had an &lt;code&gt;rm -i&lt;/code&gt; alias. The non-interactive environment inherited it. Without a terminal to answer the confirmation prompt, &lt;code&gt;rm -i&lt;/code&gt; effectively did nothing.&lt;/p&gt;

&lt;p&gt;That was a good reminder that automation doesn't necessarily run in a clean environment. Your shell configuration is part of the environment too.&lt;/p&gt;

&lt;p&gt;If you use interactive-only aliases, make sure they are actually limited to interactive shells:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="nv"&gt;$-&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt;i&lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nt"&gt;-z&lt;/span&gt; &lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;AGENT_SHELL&lt;/span&gt;&lt;span class="k"&gt;:-}&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;alias rm&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'rm -i'&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use whatever environment variable or detection mechanism makes sense for your agent.&lt;/p&gt;

&lt;p&gt;The broader lesson was probably the same as the one from the beginning:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use the clever tool to figure out what needs to be checked. Use the boring tool to check it.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>ai</category>
      <category>kubernetes</category>
      <category>monitoring</category>
    </item>
  </channel>
</rss>
