<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Justyn Larry</title>
    <description>The latest articles on DEV Community by Justyn Larry (@justyn_larry_e12a0d9779f4).</description>
    <link>https://dev.to/justyn_larry_e12a0d9779f4</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2596002%2Ffdfef0d8-b625-4804-adf5-f3fab1c34777.jpg</url>
      <title>DEV Community: Justyn Larry</title>
      <link>https://dev.to/justyn_larry_e12a0d9779f4</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/justyn_larry_e12a0d9779f4"/>
    <language>en</language>
    <item>
      <title>Grafana Agent vs Alloy: What Changed and Why</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Wed, 05 Aug 2026 12:45:59 +0000</pubDate>
      <link>https://dev.to/irinobservability/grafana-agent-vs-alloy-what-changed-and-why-2f90</link>
      <guid>https://dev.to/irinobservability/grafana-agent-vs-alloy-what-changed-and-why-2f90</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Grafana Agent reached End-of-Life on November 1, 2025 and has been replaced by Grafana Alloy. Alloy consolidates Agent's Static mode, Flow mode, and Kubernetes Operator into a single collector built on the OpenTelemetry Collector while maintaining native support for Prometheus and Loki. If you're using Flow mode, migration is relatively straightforward. If you're using Static mode, the migration process will involve reviewing and testing the converted configuration. Before switching over, verify relabeling rules, recheck resource usage, and confirm that Prometheus and Loki are receiving the same data and labels as before. If you're still running Promtail, it's worth migrating both to Alloy at the same time since Promtail is also End-of-Life.&lt;/p&gt;

&lt;p&gt;If you deployed Grafana Agent a couple of years ago, there's a good chance you haven't thought about it since. It quietly collects metrics, ships logs, and generally stays out of the way.&lt;/p&gt;

&lt;p&gt;What you may not realize is that Grafana Agent reached End-of-Life on November 1, 2025. That includes Static mode, Flow mode, and the Kubernetes Operator. Grafana Labs has stopped creating bug fixes, security patches, and official support. If you're still running it, your collection layer is probably still performing normally, but is now unsupported.&lt;/p&gt;

&lt;p&gt;That doesn't necessarily mean it will stop working tomorrow, plenty of unsupported software continues running for years. It does mean you're taking on the risk yourself, especially as the rest of your monitoring stack continues to evolve.&lt;/p&gt;

&lt;p&gt;This article covers why Grafana Labs replaced Agent with Alloy, what actually changes during the migration, and where people tend to run into problems.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Grafana Agent was deprecated
&lt;/h2&gt;

&lt;p&gt;One of the biggest issues with Grafana Agent is that it was essentially three agents, not one product:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Static mode, which used YAML and looked similar to Prometheus.&lt;/li&gt;
&lt;li&gt;Flow mode, which introduced a component-based configuration using River.&lt;/li&gt;
&lt;li&gt;The Kubernetes Operator, which managed Agent deployments inside Kubernetes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each had its own documentation, examples, bugs, and learning curve. Before setting it up, you had to first decide which version of Agent to deploy. Grafana Labs solved that by consolidating everything into Grafana Alloy.&lt;/p&gt;

&lt;p&gt;Rather than maintaining multiple collectors, Alloy becomes the single collection agent moving forward. It's built on the OpenTelemetry Collector while keeping first-class support for Prometheus and Loki, and replaces Promtail altogether. Instead of separate binaries for different use cases, everything lives in one collector. It's similar enough to Flow mode that if you were already using it, this probably feels like a rebranding exercise, but if you were using Static mode, it's a much bigger change.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alloy isn't just Grafana Agent with a new name
&lt;/h2&gt;

&lt;p&gt;It's easy to think of Alloy as Agent 2.0, but it's more complex than that. Grafana Labs is betting on OpenTelemetry becoming the long-term standard for telemetry collection. Rather than building around Prometheus first and adding OpenTelemetry later, Alloy starts with the OpenTelemetry Collector and layers Prometheus scraping, Loki log collection, and Grafana integrations on top.&lt;/p&gt;

&lt;p&gt;The result is one collector that can ingest almost anything while still feeling familiar if you're already running a Prometheus-based stack. For most users, that means fewer moving pieces and a single collector to manage, and reduced maintenance overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually changes during migration
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Static mode
&lt;/h3&gt;

&lt;p&gt;Static mode users have the biggest leap to make, since Agent Static mode used YAML with instance-based configuration. Alloy uses &lt;code&gt;.alloy&lt;/code&gt; configuration files built from components connected together into pipelines. Instead of one large configuration file describing everything, you explicitly define how data flows from one component to another.&lt;/p&gt;

&lt;p&gt;Grafana Labs provides a conversion utility:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;alloy convert &lt;span class="nt"&gt;--source-format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;static
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It does a good job handling the mechanical translation, but I wouldn't trust the output without reviewing it carefully. If you've accumulated years of custom scrape jobs, relabeling rules, and integrations, expect to spend some time validating the converted configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Flow mode
&lt;/h3&gt;

&lt;p&gt;The process of transitioning from Flow mode to Alloy is much smoother. Flow already introduced the component model that Alloy uses, and most of the migration comes down to renamed components and minor syntax updates. If you're already comfortable with Flow, this is probably an afternoon project instead of a week-long one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Understanding the component model
&lt;/h2&gt;

&lt;p&gt;Even if you're coming from Static mode, the component model ends up making a lot of sense once you get used to it. Instead of one giant configuration file, Alloy builds pipelines from individual components.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;prometheus.scrape&lt;/code&gt; collects metrics.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;prometheus.remote_write&lt;/code&gt; forwards those metrics.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;loki.source.file&lt;/code&gt; reads log files.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;loki.write&lt;/code&gt; sends those logs to Loki.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each component has a specific job, and when writing the &lt;code&gt;.alloy&lt;/code&gt; file, you explicitly connect them together.&lt;/p&gt;

&lt;p&gt;Consequently, it feels more like building a pipeline than writing a traditional configuration file. Once you understand how the pieces fit together, it's generally easier to troubleshoot because you can see exactly where data is flowing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Things to verify before switching over
&lt;/h2&gt;

&lt;p&gt;Most migrations typically go smoothly, but the problems usually show up in the little details.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verify relabeling
&lt;/h3&gt;

&lt;p&gt;The conversion tool will translate your &lt;code&gt;relabel_configs&lt;/code&gt;, but it's dangerous to assume the behavior is identical.&lt;/p&gt;

&lt;p&gt;In Alloy, relabeling happens inside a &lt;code&gt;prometheus.relabel&lt;/code&gt; component as part of a pipeline. If your configuration depended on rule ordering, make sure the converted pipeline preserves that behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  Recheck resource usage
&lt;/h3&gt;

&lt;p&gt;Don't assume Alloy has the same resource requirements as Agent. If you're combining Prometheus scraping with OTLP receivers or adding log collection into the same instance, memory usage can change enough that your old container limits are no longer appropriate.&lt;/p&gt;

&lt;h3&gt;
  
  
  If you're still running Promtail, migrate both together
&lt;/h3&gt;

&lt;p&gt;Promtail reached End-of-Life on March 2, 2026, and Alloy replaces that as well.&lt;/p&gt;

&lt;p&gt;Instead of migrating Agent now and Promtail later, consider doing both at the same time. Alloy can collect metrics and logs from the same process, which simplifies your collection layer considerably. I wrote &lt;a href="https://www.irinobservability.com/blog/migrate-node-exporter-to-grafana-alloy" rel="noopener noreferrer"&gt;an article discussing the transition from node_exporter to Alloy&lt;/a&gt;, and &lt;a href="https://www.irinobservability.com/blog/promtail-to-alloy-logs" rel="noopener noreferrer"&gt;another that covers transitioning from Promtail to Alloy&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Validate what reaches your backends
&lt;/h3&gt;

&lt;p&gt;It's easy to verify that Alloy is running, it's much more important to verify that Prometheus and Loki are receiving the same data they were before. Pay particular attention to labels. A missing or renamed label won't trigger an alert, but it can quietly break dashboards, recording rules, and alerting logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is migrating worth it?
&lt;/h2&gt;

&lt;p&gt;If you're running a simple Flow deployment, it probably is. It's a relatively straightforward migration. If you're running years of Static mode configuration across dozens or hundreds of servers, this is an intense project. You'll need to plan it, test it, and validate the results.&lt;/p&gt;

&lt;p&gt;The alternative, though, is continuing to run unsupported infrastructure. That means no security updates, no bug fixes, and no guarantee that future versions of Prometheus, Loki, or the surrounding ecosystem will continue working the way they do today. For a handful of servers, that's manageable, but for larger environments, it quickly becomes technical debt that only gets more expensive to deal with later.&lt;/p&gt;

&lt;p&gt;Migrating to Alloy isn't exciting work, but it's one of those maintenance tasks that's easier to do on your schedule rather than during an outage.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>infrastructure</category>
      <category>monitoring</category>
      <category>sre</category>
    </item>
    <item>
      <title>A Dead Man's Switch for Your Monitoring Stack</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Wed, 29 Jul 2026 12:37:21 +0000</pubDate>
      <link>https://dev.to/irinobservability/a-dead-mans-switch-for-your-monitoring-stack-2335</link>
      <guid>https://dev.to/irinobservability/a-dead-mans-switch-for-your-monitoring-stack-2335</guid>
      <description>&lt;p&gt;Your monitoring catches problems on everything except itself. Here is how an always-firing Watchdog alert plus an external heartbeat check turns silence into a signal, so you find out when your own alerting dies.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;A monitoring system can't reliably monitor its own failure, so use a dead man's switch. Create an always-firing Prometheus Watchdog alert and route it to an independent external heartbeat service. As long as the monitoring pipeline is working, the Watchdog continuously refreshes the heartbeat. If Prometheus, Alertmanager, or the delivery path fails, the heartbeat stops and the external service alerts you through a separate channel. The key is independence: the system responsible for detecting that your monitoring is down must not depend on the monitoring stack itself.&lt;/p&gt;

&lt;p&gt;One of the traps of creating alerts on a monitoring stack is the hidden assumption that the mechanism evaluating the alert is running properly and has the ability to evaluate it. Prometheus watches your hosts and Alertmanager delivers the warnings. But what watches Prometheus? If something goes wrong and the monitoring stack fails in the middle of the night, no alerts are going out but there is definitely a problem. That is the failure mode that you should be most concerned about, because it is the one your monitoring cannot report on.&lt;/p&gt;

&lt;p&gt;The fix is an old idea with a grim name: a dead man's switch. A train's dead man's switch stops the train when the operator stops holding it down. The safe state requires continuous positive action, while the absence of that action is what triggers the response. Applied to monitoring, it means building one alert whose silence is itself the alarm.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step one: an alert that always fires
&lt;/h2&gt;

&lt;p&gt;This feels backwards the first time you see it, and it took me a little time to get it right. Basically, you create an alert with a condition that is always true, so it fires constantly, forever, on purpose. In the Prometheus world this is conventionally called Watchdog.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;alert&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Watchdog&lt;/span&gt;
  &lt;span class="na"&gt;expr&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;vector(1)&lt;/span&gt;
  &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;watchdog&lt;/span&gt;
  &lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Monitoring&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pipeline&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;is&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;alive"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;vector(1)&lt;/code&gt; is always true, so as long as Prometheus is evaluating rules at all, this alert is firing. It flows through Alertmanager just like any real alert. On its own it does nothing useful, it is just noise you would never want a human to see, which is the problem I initially ran into with my Alerts Dashboard. Its value is entirely in what its absence means. If the Watchdog stops firing, that tells you the evaluation-and-delivery pipeline itself has broken, which is the one thing no ordinary alert could ever tell you.&lt;/p&gt;

&lt;p&gt;Because it fires constantly, you route it deliberately away from human channels. I give it its own &lt;code&gt;severity: watchdog&lt;/code&gt; label and filter that severity out of every dashboard panel and human-facing route, so it never lands in anyone's Slack. It exists to be watched by a machine, not a person.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step two: something outside the stack has to watch for the silence
&lt;/h2&gt;

&lt;p&gt;Here is the crucial part that a lot of people miss on the first pass. The Watchdog alert cannot be checked by the same stack that produces it. If Prometheus and Alertmanager are what would notice the Watchdog going quiet, then when they die they take the noticer down with them, and that's no better than not having it.&lt;/p&gt;

&lt;p&gt;So the watcher has to live somewhere else entirely, on infrastructure that fails independently. The best way to do this is a hosted heartbeat service. Alertmanager is configured to deliver the constantly-firing Watchdog to an external endpoint, effectively petting the switch on a schedule. That external service expects a ping at a regular interval and checks to make sure that if the ping does not arrive on time, it alerts you through a path that does not touch your monitoring stack at all.&lt;/p&gt;

&lt;p&gt;The logic is inverted, and that's what threw me off when I first set it up. Your monitoring alerts you when a real problem occurs. The heartbeat service alerts you when an &lt;em&gt;expected&lt;/em&gt; ping fails to arrive — the mirror image of that. If there's a real problem, you get a message. If monitoring itself dies, the pings stop, and the external service messages you about the silence. The two together cover the gap each leaves on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the separation is the entire point
&lt;/h2&gt;

&lt;p&gt;It is tempting to cut the corner and have the stack monitor its own health with an internal check. That gives you a comforting green panel that says everything is fine right up until the moment it cannot say anything at all, because the same outage that broke your alerting also broke the panel reporting on your alerting. A self-check is only ever as trustworthy as the system running it, and the system running it is what you also need to be keeping an eye on.&lt;/p&gt;

&lt;p&gt;The dead man's switch works because it refuses that shortcut. The signal path for "your monitoring is down" runs entirely outside your monitoring, completely independent. The cost is one deliberately-useless alert and one small external service, and in exchange you close the single scariest gap in any observability setup: the silent death of the thing you were counting on to break the silence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this looks like running
&lt;/h2&gt;

&lt;p&gt;In my stack the Watchdog fires continuously, carries its own &lt;code&gt;severity: watchdog&lt;/code&gt; label so it stays out of every human channel, and gets delivered outward to a heartbeat check on a schedule. I also surface it as a single internal panel that reads "pipeline healthy" while the Watchdog is present and flips to "broken" the instant it goes absent, using a &lt;code&gt;count()&lt;/code&gt; guarded with &lt;code&gt;or vector(0)&lt;/code&gt; so the panel shows a real red state rather than an empty no-data tile when the alert stops arriving. That panel is for me, during normal working hours. The external heartbeat is what actually wakes me up, because it is the only part of the chain that will still be running when the rest of it is not.&lt;/p&gt;

&lt;p&gt;If you run your own monitoring, this is worth the hour it takes to set up. Have you built a dead man's switch into your stack, and if so, did you route it to a hosted heartbeat service or roll your own external checker? I am always curious how other people close this particular gap.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>monitoring</category>
      <category>sre</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Migrating from Promtail to Alloy for Log Collection</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Thu, 23 Jul 2026 14:44:33 +0000</pubDate>
      <link>https://dev.to/irinobservability/migrating-from-promtail-to-alloy-for-log-collection-18j2</link>
      <guid>https://dev.to/irinobservability/migrating-from-promtail-to-alloy-for-log-collection-18j2</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Promtail reached end of life in March 2026, making Alloy the supported replacement for shipping logs to Loki. The migration is mostly a matter of replacing Promtail's scrape configuration with &lt;code&gt;loki.source.*&lt;/code&gt;, &lt;code&gt;loki.process&lt;/code&gt;, and &lt;code&gt;loki.write&lt;/code&gt; components. The biggest pitfall isn't the configuration itself—it's assuming how a host writes logs. Detect whether a machine uses the systemd journal or log files before configuring Alloy, or you may end up silently collecting nothing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In this &lt;a href="https://www.irinobservability.com/blog/migrate-node-exporter-to-grafana-alloy" rel="noopener noreferrer"&gt;article&lt;/a&gt; I discussed moving host metrics from a scraped &lt;code&gt;node_exporter&lt;/code&gt; to a push-based Alloy agent, and followed it up with a piece comparing &lt;a href="https://www.irinobservability.com/blog/prometheus-agent-vs-alloy" rel="noopener noreferrer"&gt;Alloy against Prometheus Agent mode&lt;/a&gt;. Both of those articles discussed metrics and how they move between the host and client nodes. This post covers the log-shipping portion of the same story, and it has a deadline attached that the metrics migration didn't.&lt;/p&gt;

&lt;p&gt;Promtail reached end of life on March 2, 2026. It continues to function, but it no longer receives bug fixes or security updates. If you are shipping logs to Loki with Promtail today, you are running an unmaintained agent, and the replacement Grafana points you at is the same Alloy you may already be running for metrics. If you're already running Alloy for metrics, the log migration is mostly adding components to a config that already exists. If you're not, this article will strengthen the case for consolidating onto Alloy rather than running a separate, now-unmaintained log agent.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             Promtail

  Journal/File Logs
         │
         ▼
     Promtail
         │
         ▼
       Loki


             Alloy

 Journal/File Logs
         │
         ▼
  loki.source.*
         │
         ▼
   loki.process
         │
         ▼
    loki.write
         │
         ▼
        Loki
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What actually changes
&lt;/h2&gt;

&lt;p&gt;Promtail is a standalone binary with its own YAML config: &lt;code&gt;scrape_configs&lt;/code&gt;, &lt;code&gt;clients&lt;/code&gt;, and a positions file, built for one job. Alloy replaces that YAML configuration with a component-based pipeline. The host produces log entries, Alloy processes them if required, and forwards them to Loki. There are three component types, all wired together with &lt;code&gt;forward_to&lt;/code&gt; references.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Promtail&lt;/th&gt;
&lt;th&gt;Alloy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Journal scrape&lt;/td&gt;
&lt;td&gt;&lt;code&gt;loki.source.journal&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;File scrape&lt;/td&gt;
&lt;td&gt;&lt;code&gt;loki.source.file&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pipeline stages&lt;/td&gt;
&lt;td&gt;&lt;code&gt;loki.process&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;clients&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;loki.write&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Between the source and destination, &lt;code&gt;loki.process&lt;/code&gt; handles parsing, filtering, label manipulation, and other pipeline stages that previously lived inside Promtail.&lt;/p&gt;

&lt;h2&gt;
  
  
  The journal case
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;source&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;journal&lt;/span&gt; &lt;span class="s2"&gt;"systemd_journal"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;forward_to&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;add_labels&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The source reads the journal and hands each line to a &lt;code&gt;loki.process&lt;/code&gt; component named &lt;code&gt;add_labels&lt;/code&gt;, which is where I attach the same tenant, cluster, environment, and role labels that ride along with the metrics from that host to help maintain separation between clients. You can read more about my tenant isolation model &lt;a href="https://www.irinobservability.com/blog/three-layer-tenant-isolation" rel="noopener noreferrer"&gt;here&lt;/a&gt;. Labeling logs and metrics identically at the edge is what lets me line them up later in Grafana, and it is worth getting consistent from the first host rather than fixing it in queries forever after.&lt;/p&gt;

&lt;h2&gt;
  
  
  The file-based case
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="nx"&gt;local&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;file_match&lt;/span&gt; &lt;span class="s2"&gt;"logs"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;path_targets&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"__path__"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/syslog"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"__path__"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/auth.log"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"__path__"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/messages"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"__path__"&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"/var/log/secure"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;source&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;file&lt;/span&gt; &lt;span class="s2"&gt;"log_scrape"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;targets&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;local&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;file_match&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;logs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;targets&lt;/span&gt;
  &lt;span class="nx"&gt;forward_to&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;loki&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;add_labels&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;local.file_match&lt;/code&gt; resolves glob patterns into concrete targets, and &lt;code&gt;loki.source.file&lt;/code&gt; tails whatever it finds. I list both Debian-family paths and RHEL-family paths because which files exist depends on the distro, and missing paths are simply skipped. In practice I detect which log method a host actually uses at install time and inject only the matching block.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trap: don't guess the log method by distro
&lt;/h2&gt;

&lt;p&gt;I learned this lesson the hard way. I assumed Debian meant &lt;code&gt;/var/log/syslog&lt;/code&gt; and RHEL meant &lt;code&gt;journald&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I thought the obvious way to decide between journal and file collection was to base it on distro family. That doesn't cover edge cases. Plenty of Debian systems have &lt;code&gt;rsyslog&lt;/code&gt; disabled or absent and log only to the journal, while some hosts have been configured differently by whoever set them up. Alloy starts cleanly, reports itself healthy, and quietly ships &lt;strong&gt;no logs&lt;/strong&gt; because it is watching the wrong source. The first indication anything is wrong is often during an incident, when the logs you expected simply aren't there.&lt;/p&gt;

&lt;p&gt;What actually works is an empirical check. Determine whether the host is actively writing journal entries or log files, then configure Alloy accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Migrating without a gap
&lt;/h2&gt;

&lt;p&gt;Stand up Alloy's log collection alongside Promtail rather than cutting over immediately. Both can ship to the same Loki, and the extra resource overhead is negligible. You will get some duplicate lines during the overlap, which is a much better failure mode than a hole in your logs. Confirm in Grafana that the Alloy-sourced logs are arriving with the correct labels, then remove Promtail.&lt;/p&gt;

&lt;p&gt;Loki deduplicates identical log entries that share the same labels. If Alloy adds even one different label, the same log line becomes part of a different stream and both copies remain visible. Match your label set before tearing the old agent down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this leaves you
&lt;/h2&gt;

&lt;p&gt;If you already migrated metrics to Alloy, adding log collection is just a handful of additional components. If you are still on Promtail, the March 2026 end-of-life date makes this migration worth prioritizing. The migration itself is straightforward. The only genuinely dangerous part is the silent-failure trap, and that's avoidable if you verify how each host actually writes logs instead of assuming based on its distro.&lt;/p&gt;

&lt;p&gt;If you've already made this move, I'm curious whether you ran into the journal-versus-file mismatch too, or whether your fleet was uniform enough that the distro heuristic held.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>infrastructure</category>
      <category>monitoring</category>
    </item>
    <item>
      <title>Monitoring Docker Containers with Grafana Alloy and cAdvisor</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Tue, 21 Jul 2026 12:29:15 +0000</pubDate>
      <link>https://dev.to/irinobservability/monitoring-docker-containers-with-grafana-alloy-and-cadvisor-146p</link>
      <guid>https://dev.to/irinobservability/monitoring-docker-containers-with-grafana-alloy-and-cadvisor-146p</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;\&lt;br&gt;
Host metrics tell you whether a server is healthy. cAdvisor tells you&lt;br&gt;
which container isn't. This article shows how to integrate cAdvisor&lt;br&gt;
into a Grafana Alloy push architecture, avoid the common cardinality&lt;br&gt;
trap, and build a few alerts that catch real problems.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Container Monitoring
&lt;/h2&gt;

&lt;p&gt;In my previous posts, I compared &lt;a href="https://www.irinobservability.com/blog/prometheus-agent-vs-alloy" rel="noopener noreferrer"&gt;Prometheus Agent Mode with Grafana&lt;br&gt;
Alloy&lt;/a&gt; and walked through &lt;a href="https://www.irinobservability.com/blog/migrate-node-exporter-to-grafana-alloy" rel="noopener noreferrer"&gt;migrating from &lt;code&gt;node_exporter&lt;/code&gt; to Alloy&lt;/a&gt;. Both&lt;br&gt;
focused on the agent responsible for shipping host metrics and logs&lt;br&gt;
upstream.&lt;/p&gt;

&lt;p&gt;The natural next step is monitoring Docker containers.&lt;/p&gt;

&lt;p&gt;One of my servers runs eighteen Docker containers spread across multiple&lt;br&gt;
Compose files. Host-level CPU and memory metrics can tell me the server&lt;br&gt;
is healthy, but they cannot tell me &lt;em&gt;which&lt;/em&gt; container is consuming all&lt;br&gt;
of the memory or unexpectedly restarting.&lt;/p&gt;

&lt;p&gt;cAdvisor solves that problem. Originally developed by Google, it reads&lt;br&gt;
container resource usage directly from Linux cgroups and namespaces&lt;br&gt;
without requiring any instrumentation inside the containers themselves.&lt;br&gt;
In this article I'll show how it fits cleanly into a push-based Alloy&lt;br&gt;
architecture.&lt;/p&gt;
&lt;h2&gt;
  
  
  Deployment
&lt;/h2&gt;

&lt;p&gt;cAdvisor runs as its own container with several read-only mounts so it&lt;br&gt;
can inspect Docker and the host's cgroup state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--name&lt;/span&gt; cadvisor &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--privileged&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--volume&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/:/rootfs:ro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--volume&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/var/run:/var/run:ro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--volume&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/sys:/sys:ro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--volume&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/var/lib/docker:/var/lib/docker:ro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--volume&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/dev/disk:/dev/disk:ro &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--publish&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&amp;lt;port&amp;gt;:8080 &lt;span class="se"&gt;\&lt;/span&gt;
  gcr.io/cadvisor/cadvisor:latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Although &lt;code&gt;--privileged&lt;/code&gt; often raises eyebrows, every mounted volume is&lt;br&gt;
read-only. cAdvisor is observing host state, not modifying it.&lt;/p&gt;

&lt;p&gt;I also recommend avoiding a hardcoded &lt;code&gt;8080&lt;/code&gt; mapping. It is one of the&lt;br&gt;
most frequently occupied ports on development machines and small&lt;br&gt;
servers. My installer probes for an available port and falls back to&lt;br&gt;
&lt;code&gt;9338&lt;/code&gt;, then reports the selected port back so Alloy can generate the&lt;br&gt;
correct scrape target automatically.&lt;/p&gt;
&lt;h2&gt;
  
  
  Wiring it into Alloy
&lt;/h2&gt;

&lt;p&gt;Rather than exposing cAdvisor to a central Prometheus server, Alloy&lt;br&gt;
scrapes it locally over &lt;code&gt;localhost&lt;/code&gt; and forwards those metrics using the&lt;br&gt;
same &lt;code&gt;remote_write&lt;/code&gt; pipeline that already carries host metrics and logs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    Docker Host
┌──────────────────────────────────────────┐
│                                          │
│  Docker Containers                       │
│        │                                 │
│        ▼                                 │
│    cAdvisor                              │
│        │ localhost scrape                │
│        ▼                                 │
│   Grafana Alloy                          │
│        │ remote_write                    │
└────────┼─────────────────────────────────┘
         │
         ▼
┌───────────────────────────────┐
│ Central Monitoring            │
│                               │
│ Prometheus                    │
│ Loki                          │
│ Grafana                       │
│ Alertmanager                  │
└───────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From the backend's perspective, container metrics arrive exactly like&lt;br&gt;
host metrics: already labeled, already authenticated, and without&lt;br&gt;
requiring any inbound connectivity to the client.&lt;/p&gt;

&lt;p&gt;This becomes especially valuable for clients behind CGNAT or dynamic&lt;br&gt;
residential IP addresses. Once a push-based pipeline exists, adding&lt;br&gt;
another local exporter is simply another scrape target---not another&lt;br&gt;
monitoring system.&lt;/p&gt;
&lt;h2&gt;
  
  
  Avoiding the Cardinality Trap
&lt;/h2&gt;

&lt;p&gt;One detail that many getting-started guides overlook is that cAdvisor&lt;br&gt;
exports metrics for every cgroup it can see, not just your Docker&lt;br&gt;
containers.&lt;/p&gt;

&lt;p&gt;Without filtering, dashboards become cluttered with unnamed&lt;br&gt;
infrastructure cgroups while your active series count grows for little&lt;br&gt;
benefit.&lt;/p&gt;

&lt;p&gt;The simplest fix is to consistently filter container metrics using a&lt;br&gt;
PromQL label matcher:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;container_cpu_usage_seconds_total{name!=""}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Alternatively, you can drop unnamed series at scrape time using metric&lt;br&gt;
relabeling in Alloy or Prometheus.&lt;/p&gt;
&lt;h2&gt;
  
  
  Three Useful Alerts
&lt;/h2&gt;

&lt;p&gt;Restart loops:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;changes(container_start_time_seconds{name!=""}[1h]) &amp;gt; 3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;High sustained CPU:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;rate(container_cpu_usage_seconds_total{name!="",name!="POD"}[5m]) * 100 &amp;gt; 80
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Container memory as a percentage of total host memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(container_memory_usage_bytes{name!=""}
 / on(instance, tenant) group_left node_memory_MemTotal_bytes) * 100 &amp;gt; 85
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That final query uses a &lt;code&gt;group_left&lt;/code&gt; join to compare a container-level&lt;br&gt;
metric against host memory, producing a percentage that's much easier to&lt;br&gt;
reason about than raw bytes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Fits
&lt;/h2&gt;

&lt;p&gt;In Irin, Docker monitoring is implemented as an optional module, but the&lt;br&gt;
architecture described here works with any push-based monitoring stack.&lt;br&gt;
Once you've adopted a push-based agent, adding exporters like cAdvisor&lt;br&gt;
becomes incremental work.&lt;/p&gt;

&lt;p&gt;The difficult part isn't collecting the metrics---it's deciding which&lt;br&gt;
ones are worth keeping.&lt;/p&gt;

&lt;p&gt;If you're already running cAdvisor, I'd be interested to hear whether&lt;br&gt;
you've run into the unnamed cgroup problem or found a filtering strategy&lt;br&gt;
that works even better.&lt;/p&gt;

</description>
      <category>monitoring</category>
      <category>docker</category>
      <category>containers</category>
      <category>cadvisor</category>
    </item>
    <item>
      <title>Why LLM Decisions Should Be Deterministic</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Wed, 15 Jul 2026 14:07:14 +0000</pubDate>
      <link>https://dev.to/irinobservability/why-llm-decisions-should-be-deterministic-2i7p</link>
      <guid>https://dev.to/irinobservability/why-llm-decisions-should-be-deterministic-2i7p</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt;  I originally treated deterministic boundaries around LLMs as a consistency mechanism. I now think their real value is auditability. If the system's decisions are made by deterministic code rather than the model, every decision has a reproducible implementation. The LLM can explain that decision to humans, but it should never be the source of the decision itself.&lt;/p&gt;

&lt;p&gt;In two previous posts I argued for keeping narration and decision-making separate. &lt;a href="https://www.irinobservability.com/blog/adding-llm-narration" rel="noopener noreferrer"&gt;One covered monthly reporting&lt;/a&gt;. &lt;a href="https://www.irinobservability.com/blog/llm-narrates-code-decides" rel="noopener noreferrer"&gt;The other covered a real-time alert annotator&lt;/a&gt;. Both focused on consistency. This post argues that consistency is only the visible benefit; auditability is the deeper one. By auditability, I mean that a third party can inspect how a decision was reached and reproduce it from the implementation rather than relying on an after-the-fact explanation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Deterministic Layer
&lt;/h2&gt;

&lt;p&gt;The alert annotator classifies every alert into one of eight values, enforced in Python against a fixed set, not requested in a prompt. If the model's output does not match one of the eight strings exactly, the field falls back to &lt;code&gt;unknown&lt;/code&gt; rather than accepting whatever the model produced. That validation step is more important than the enum itself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ALLOWED_CAUSES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memory_pressure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cpu_saturation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;disk_pressure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;service_unavailability&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;network_issue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;configuration_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;external_dependency&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unknown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;resolve_cause&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_output&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;cause&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model_output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;cause&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ALLOWED_CAUSES&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unknown&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;cause&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Given the same input, this function always produces the same output. That means every classification is reproducible from the source code alone.&lt;/p&gt;

&lt;p&gt;This is the shape of the check, not a copy of the production code, but the principle is that you should never trust an external system's output by letting it pass through unchecked. This is not an LLM-specific idea, it is the same discipline that should apply to a third-party webhook payload or an API response before it touches your data model. The model just happens to be the least predictable external system I integrate with, which makes the validation step the most visibly valuable.&lt;/p&gt;

&lt;p&gt;I originally wrote about this as a consistency fix, and failed to discuss its full value. Before the enum existed, the same alert produced different category strings across separate runs, which made cross-tenant pattern analysis impossible. What the validation step actually guarantees is that every classification belongs to a bounded, deterministic set of outcomes. &lt;code&gt;resolve_cause()&lt;/code&gt; is a pure function, and every valid result is reproducible from its implementation. It creates consistency in the input and output every time, and the entire decision is just sixteen lines of Python.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Traditional pipeline:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Metrics
   |
   v
Deterministic rules
   |
   |-- classification
   |
   v
LLM narration
   |
   v
Client
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;LLM-centric pipeline:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Metrics
   |
   v
LLM
   |
   |-- classification
   |-- explanation
   |
   v
Client
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important distinction isn't whether an LLM appears accurate most of the time. It's whether a reviewer can reproduce the decision months later from the same inputs. In the first architecture, the decision exists independently of the model's narration. In the second, the classification and its explanation originate from the same probabilistic process, making them difficult to separate during an audit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Industry Trends
&lt;/h2&gt;

&lt;p&gt;It turns out a great deal of current AI governance work is aimed at exactly this property, approached from the opposite direction. A large part of the LLM governance tooling market exists because production language model behavior does not offer the guarantee classical software does. The same prompt can produce different answers across runs, so teams are building trace-level evidence systems, output evaluators, and audit logging specifically to reconstruct what a probabilistic system did and why after the fact.&lt;/p&gt;

&lt;p&gt;The regulatory backdrop makes the stakes explicit rather than abstract. Under the EU AI Act, certain classes of LLM application are treated as high-risk and come with mandatory logging and human-oversight requirements, with an explicit standard that records must be sufficient to reconstruct the system's operation, not just gesture at it. NIST's AI Risk Management Framework leans on the same assumption from a different angle, treating a reliable and auditable record of system behavior as a prerequisite for its governance functions.&lt;/p&gt;

&lt;p&gt;None of that applies to the projects that I'm working on directly. Advisory infrastructure monitoring for small server fleets is not a high-risk category, and I am not building toward EU AI Act compliance. But the underlying problem is the same problem, just at a different scale and a different level of legal consequence. If a decision matters enough that someone might reasonably ask "why did the system do that," the honest answer needs to survive more scrutiny than a model's own account of itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-Reporting Is Not an Explanation
&lt;/h2&gt;

&lt;p&gt;You can ask a language model why it produced a given output, and it will give you a fluent, plausible answer, but it will not give you the actual causal mechanism. The explanation is generated during a new inference over the conversation history, not by inspecting the internal computation that produced the original answer. The model generates a plausible explanation rather than retrieving the causal process behind its earlier output. The model will make a guess and present it as a first-hand account.&lt;/p&gt;

&lt;p&gt;Recent AI governance research states the practical consequence of this directly, stating that systems that need genuinely reproducible decisions should not rely on a probabilistic layer alone, they should sit a deterministic enforcement layer underneath it as a secondary safeguard. That is a formal way of describing something a lot of people building on LLMs arrive at independently once they hit production. I arrived at it because an enum kept coming back spelled three different ways, before I had even thought about what the industry was doing. Other people are arriving at it because a regulator asked for a record that would hold up. We're both coming to the same conclusion, but the reasons behind it are different.&lt;/p&gt;

&lt;p&gt;Classical software doesn't explain itself, we inspect the implementation. A pure function doesn't tell us why it returned &lt;code&gt;cpu_saturation&lt;/code&gt;, we read the code that produced that output. LLMs invert that relationship. They readily generate explanations, but those explanations are themselves model outputs rather than evidence of the underlying computation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Does Narration Actually Do?
&lt;/h2&gt;

&lt;p&gt;None of this makes the narration layer pointless, it just refines its allowed job. The prose the model writes for a client report, or the plain-English line attached to an alert, is still probabilistic. I can log exactly what it said. I cannot claim a rigorous account of why it phrased something one way over another, but I do not need one, because that text is not load-bearing. It explains a decision, but it doesn't make one. If the narration layer disappeared entirely tomorrow, every report and every alert would still carry a correct, if blunter, classification, because the classification does not depend on the prose existing.&lt;/p&gt;

&lt;p&gt;That is the actual shape of the boundary. Not "the model is untrustworthy so keep it away from important things," which is too broad to be useful, but "know exactly which outputs in your pipeline need to be reproducible, and make sure none of those outputs pass through a step you cannot validate against a fixed answer."&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Is Heading
&lt;/h2&gt;

&lt;p&gt;Before the narrative layer ships to a client, it needs a short, plain statement of what the model is permitted to produce, how its output is labeled in the report, and what happens when it is wrong or unavailable. If a system needs to justify a decision months later, the justification should be found in deterministic code and recorded inputs, not in asking the model what it thinks it did. The model narrates a state that was already decided; it never gets to decide the state itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Related Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.irinobservability.com/blog/adding-llm-narration" rel="noopener noreferrer"&gt;Adding an LLM Narration Layer to Prometheus and Grafana&lt;/a&gt;. Where this series started, on the monthly report pipeline.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.irinobservability.com/blog/llm-narrates-code-decides" rel="noopener noreferrer"&gt;The LLM Narrates. The Code Decides.&lt;/a&gt; The harder real-time case, on the alert annotator.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Prometheus Agent Mode vs Grafana Alloy: Choosing the Right Push Agent in 2026</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Tue, 14 Jul 2026 12:36:22 +0000</pubDate>
      <link>https://dev.to/irinobservability/prometheus-agent-mode-vs-grafana-alloy-choosing-the-right-push-agent-in-2026-52m2</link>
      <guid>https://dev.to/irinobservability/prometheus-agent-mode-vs-grafana-alloy-choosing-the-right-push-agent-in-2026-52m2</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; If you only collect metrics, Prometheus Agent mode is lightweight, familiar, and difficult to beat. If you collect metrics, logs, or traces together, or expect to in the future, Grafana Alloy's unified pipeline is usually worth the additional complexity.&lt;/p&gt;

&lt;p&gt;Once you've decided to &lt;a href="https://www.irinobservability.com/blog/alloy-prometheus-comparison" rel="noopener noreferrer"&gt;move from pull-based scraping to a push architecture&lt;/a&gt;, the next question is which agent should actually run on each host. In 2026, the two strongest choices are Prometheus Agent mode and Grafana Alloy. I run Alloy across my production fleet, but that doesn't automatically make it the right answer for everyone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Shift in the Monitoring Landscape
&lt;/h2&gt;

&lt;p&gt;Over the last couple of years, Grafana has consolidated both metrics and log collection into Grafana Alloy. Grafana Agent reached end of life on November 1, 2025, and Promtail followed on March 2, 2026. Neither receives security fixes anymore.&lt;/p&gt;

&lt;p&gt;The practical choice moving forward:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Prometheus Agent&lt;/th&gt;
&lt;th&gt;Grafana Alloy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Metrics&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Logs&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Traces&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Config&lt;/td&gt;
&lt;td&gt;Prometheus YAML&lt;/td&gt;
&lt;td&gt;Alloy components&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Footprint&lt;/td&gt;
&lt;td&gt;Smaller&lt;/td&gt;
&lt;td&gt;Larger&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learning curve&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Future direction&lt;/td&gt;
&lt;td&gt;Metrics agent&lt;/td&gt;
&lt;td&gt;Unified telemetry&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The table gives the short answer. The rest of this article explains where those differences actually matter in practice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prometheus Agent mode.&lt;/strong&gt; Run the Prometheus binary with the &lt;code&gt;--agent&lt;/code&gt; flag and it stops acting as a full Prometheus server. It no longer stores local TSDB blocks, evaluates alerting rules, or serves queries. Instead, it scrapes targets, buffers samples in a write-ahead log, and forwards them upstream via &lt;code&gt;remote_write&lt;/code&gt;. It is Prometheus with the storage and query layers removed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Grafana Alloy.&lt;/strong&gt; A single agent that collects metrics, logs, and traces, processes them in a component pipeline, and pushes each signal to its backend. It embeds many exporters directly, so a line like &lt;code&gt;prometheus.exporter.unix "node_exporter" {}&lt;/code&gt; gives you full node_exporter functionality without installing a separate binary.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Case for Prometheus Agent
&lt;/h2&gt;

&lt;p&gt;If you only need metrics, agent mode is hard to argue with.&lt;/p&gt;

&lt;p&gt;The configuration is the Prometheus config you are probably already familiar with. The &lt;code&gt;scrape_configs&lt;/code&gt;, relabeling, and service discovery are all the same. If your team is fluent in Prometheus YAML, there is nothing new to learn, and every Stack Overflow answer from the last decade is still applicable.&lt;/p&gt;

&lt;p&gt;The resource footprint is small and predictable. Agent mode exists specifically to reduce Prometheus to one job: scrape metrics and forward them upstream. On a constrained edge box collecting a modest number of series, it is the lighter option, and it is maintained by the Prometheus project itself. If your logs already have a home, or you don't collect them at all, adding Alloy adds complexity you won't use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Alloy Wins
&lt;/h2&gt;

&lt;p&gt;If you want to monitor logs, the landscape changes dramatically. With agent mode you need a second agent for log shipping, and the tool most people used for that was Promtail, which is now end-of-life. You would probably end up running agent mode plus Alloy, at which point you may as well run one agent instead of two.&lt;/p&gt;

&lt;p&gt;That consolidation is what sold me. On every host I monitor, one systemd service collects host metrics through the embedded node_exporter component, tails the journal for logs, and pushes both upstream over the same authenticated tunnel. One binary to install, one service to health-check, one config to manage per host. When I later added container metrics and disk health collection, those became new components in the same pipeline instead of new daemons.&lt;/p&gt;

&lt;p&gt;The pipeline model streamlines the operation on the processing side too. Labels get attached at the edge before data ever leaves the client machine: every sample arrives already tagged with tenant, cluster, environment, and role, which is what makes multi-tenant isolation by label possible. That means routing, dashboards, and alerting can all rely on the same label set without additional processing upstream. Doing the equivalent in agent mode means metric relabeling rules, and applying it to logs means a second tool entirely.&lt;/p&gt;

&lt;p&gt;Alloy has become Grafana Labs' strategic collection agent following the retirement of Grafana Agent and Promtail. It has first-class OTLP support, so when I added tracing, I was able to add the receiver as a config block instead of installing a new agent. Everything Grafana folded in from Promtail and Grafana Agent now lives here, and this is where Grafana Labs is focusing new collection features.&lt;/p&gt;

&lt;p&gt;Although Alloy is developed by Grafana Labs, it isn't tied to Grafana Cloud. It speaks standard protocols such as Prometheus &lt;code&gt;remote_write&lt;/code&gt; and OTLP, so it works just as well with self-hosted Prometheus, Loki, Tempo, Mimir, or other compatible backends.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alloy's Hidden Costs
&lt;/h2&gt;

&lt;p&gt;The configuration language is a real learning curve. Alloy configs are components wired together with &lt;code&gt;forward_to&lt;/code&gt; references, in a Terraform-like syntax. I think it is genuinely better than YAML once it clicks, because the pipeline is explicit, and you can read a config top to bottom and see exactly where data flows. The learning curve is steep, and small syntax details can create headaches. Alloy has a larger runtime footprint because it bundles a much broader telemetry pipeline, including OpenTelemetry Collector capabilities, embedded exporters, and support for multiple signal types. For metrics-only work on tiny hosts, agent mode is leaner.&lt;/p&gt;

&lt;p&gt;Fleet management is also more complicated. Alloy configs are per-host and declarative, which is great until you have dozens of them and a label schema change means touching every one. The method I used to streamline this process was to generate configs from a template, then build a sync mechanism where hosts pull updated configs on a schedule.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose the Tool That Fits
&lt;/h2&gt;

&lt;p&gt;If your infrastructure is only tracking metrics and your team is already fluent in Prometheus config, running Prometheus Agent mode is probably the right choice.&lt;/p&gt;

&lt;p&gt;If you need to track both metrics and logs or traces, or if you plan to in the future, Alloy is probably the better choice. The single-agent model pays for its learning curve quickly, especially if your business is growing and your infrastructure is expanding.&lt;/p&gt;

&lt;p&gt;If you're already running Grafana Agent or Promtail, you don't have a choice anymore, and &lt;code&gt;alloy convert&lt;/code&gt; will translate your existing config as a starting point. Treat the output as a draft and verify it against the running system, not the migration guide.&lt;/p&gt;

&lt;p&gt;I went with Alloy because I ship logs alongside metrics for every host, embedding node_exporter meant one less binary on client machines, and because edge labeling is load-bearing for how we isolate tenant data. Those reasons are specific to my environment. If I only needed metrics from a handful of systems, I would probably choose Prometheus Agent mode instead. The decision isn't really Prometheus versus Alloy. It's whether you want a dedicated metrics forwarder or a unified telemetry pipeline. Once you know which problem you're solving, the choice becomes much clearer.&lt;/p&gt;

</description>
      <category>prometheus</category>
      <category>grafana</category>
      <category>monitoring</category>
      <category>observability</category>
    </item>
    <item>
      <title>Migrating from node_exporter to Grafana Alloy, One Server at a Time</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Wed, 08 Jul 2026 12:46:19 +0000</pubDate>
      <link>https://dev.to/irinobservability/migrating-from-nodeexporter-to-grafana-alloy-one-server-at-a-time-58an</link>
      <guid>https://dev.to/irinobservability/migrating-from-nodeexporter-to-grafana-alloy-one-server-at-a-time-58an</guid>
      <description>&lt;p&gt;If you've been monitoring Linux servers for any length of time, there's a good chance &lt;strong&gt;node_exporter&lt;/strong&gt; was the first thing you installed.  It's lightweight, reliable, and exposes a huge amount of machine metrics for Prometheus to scrape.  For years, it has been the default answer.&lt;br&gt;
As your infrastructure grows, though, your monitoring stack usually grows with it.  First comes log collection. Then traces. Before long you're running &lt;code&gt;node_exporter&lt;/code&gt;, a log shipper, and maybe another telemetry agent.  Each component has its own configuration, service unit, upgrade cycle, and failure modes.  &lt;/p&gt;

&lt;p&gt;Grafana Alloy changes that by consolidating those responsibilities into a single telemetry agent.&lt;br&gt;
This post walks through migrating from &lt;code&gt;node_exporter&lt;/code&gt; to Alloy on a real fleet, one server at a time, while maintaining continuous visibility throughout the process.  These are the exact steps that survived contact with production on the Irin monitoring stack, not the idealized version that looks clean in a diagram.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;  If you're already running &lt;code&gt;node_exporter&lt;/code&gt;, don't replace it overnight.  Install Grafana Alloy alongside it, configure Alloy's built-in prometheus.exporter.unix component, verify that metrics are reaching your remote Prometheus instance, and only then retire node_exporter.&lt;br&gt;
Migrating one server at a time minimizes risk, preserves visibility, and positions your infrastructure for logs, traces, and future telemetry without deploying additional agents.&lt;/p&gt;
&lt;h2&gt;
  
  
  The real difference is the direction of travel
&lt;/h2&gt;

&lt;p&gt;Before getting started, it's worth understanding what actually changes.&lt;br&gt;
This isn't simply replacing one monitoring agent with another.&lt;br&gt;
&lt;code&gt;node_exporter&lt;/code&gt; is a server. It listens on a port, typically 9100,and waits for Prometheus to connect and scrape metrics. That means every monitored machine needs an open endpoint, network connectivity from Prometheus, firewall rules, and scrape configurations.&lt;/p&gt;

&lt;p&gt;Alloy flips that model around.&lt;/p&gt;

&lt;p&gt;Instead of waiting for Prometheus to connect, Alloy collects metrics locally and pushes them to a remote endpoint using Prometheus Remote Write.&lt;/p&gt;

&lt;p&gt;On my stack, that outbound traffic travels through a Cloudflare Tunnel. Nothing reaches into the monitored servers. There are no metrics ports exposed to the LAN, no inbound firewall rules to maintain, and no scrape network that has to remain routable.  The user’s metrics are exposed through the secure tunnel, and the monitoring stack has no access to the server.&lt;/p&gt;

&lt;p&gt;That shift is the real migration, you're not replacing a binary, you're changing the direction your telemetry flows.  Once you frame it that way, the rest of the migration makes much more sense.&lt;/p&gt;
&lt;h2&gt;
  
  
  The component that replaces node_exporter
&lt;/h2&gt;

&lt;p&gt;Alloy is configured using River, which is less like a traditional configuration file and more like a telemetry pipeline.  Each component performs one task before handing data to the next component.  As you begin to put your model together, the configuration becomes surprisingly readable.&lt;/p&gt;

&lt;p&gt;The component that replaces node_exporter is &lt;code&gt;prometheus.exporter.unix&lt;/code&gt;.&lt;br&gt;
Under the hood it's using the same collector code as &lt;code&gt;node_exporter&lt;/code&gt;, so the metrics themselves remain familiar.&lt;br&gt;
A minimal configuration looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight hcl"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Collect host metrics.&lt;/span&gt;
&lt;span class="nx"&gt;prometheus&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;exporter&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;unix&lt;/span&gt; &lt;span class="s2"&gt;"host"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Scrape those locally collected metrics.&lt;/span&gt;
&lt;span class="nx"&gt;prometheus&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;scrape&lt;/span&gt; &lt;span class="s2"&gt;"host"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;targets&lt;/span&gt;    &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;prometheus&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;exporter&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;unix&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;host&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;targets&lt;/span&gt;
  &lt;span class="nx"&gt;forward_to&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;prometheus&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;remote_write&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;default&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;receiver&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// Push metrics to Prometheus Remote Write.&lt;/span&gt;
&lt;span class="nx"&gt;prometheus&lt;/span&gt;&lt;span class="err"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;remote_write&lt;/span&gt; &lt;span class="s2"&gt;"default"&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;endpoint&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"https://metrics.example.internal/api/v1/write"&lt;/span&gt;

    &lt;span class="c1"&gt;# Production deployments typically authenticate here using&lt;/span&gt;
    &lt;span class="c1"&gt;# Basic Auth, bearer tokens, or mTLS.&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read from top to bottom, it tells a story:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The exporter gathers metrics.&lt;/li&gt;
&lt;li&gt;The scraper collects those metrics internally.&lt;/li&gt;
&lt;li&gt;The remote write component sends them to Prometheus.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of maintaining a YAML file full of scrape targets, you're wiring components together into a pipeline.  That same pattern extends naturally to logs, traces, profiling, and nearly every other telemetry signal Alloy supports.&lt;/p&gt;

&lt;p&gt;One thing to remember during migration: the metrics themselves stay the same, but some labels—particularly the &lt;code&gt;instance&lt;/code&gt; label—may change depending on how Alloy identifies the host.&lt;br&gt;
That's expected, and it's important when you start verifying your migration.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why migrate one server at a time?
&lt;/h2&gt;

&lt;p&gt;The temptation is to deploy Alloy everywhere and immediately disable node_exporter, but the best way is to do it slowly.  The safest migration pattern is a canary.  Pick a non-critical server, install Alloy beside node_exporter, and let both run simultaneously while you verify that the new telemetry path is working correctly.  The resource overhead is negligible, but the confidence you gain is enormous.&lt;/p&gt;

&lt;p&gt;Running both agents briefly means you always have a known-good monitoring path while validating the new one.  Only after you've confirmed that Alloy is producing fresh, accurate metrics should you retire &lt;code&gt;node_exporter&lt;/code&gt;.  Once you've done that successfully a few times, batching servers becomes much less stressful.&lt;/p&gt;
&lt;h2&gt;
  
  
  The migration
&lt;/h2&gt;
&lt;h3&gt;
  
  
  1. Pick a canary host
&lt;/h3&gt;

&lt;p&gt;Choose a stable server where a few minutes of metric oddities wouldn't be catastrophic, but would still be noticeable.&lt;br&gt;
Install Grafana Alloy using your distribution's package manager.&lt;br&gt;
Most Linux distributions install Alloy as a systemd service automatically, with the primary configuration file located at:&lt;br&gt;
&lt;code&gt;/etc/alloy/config.alloy&lt;/code&gt;&lt;br&gt;
Enable the service, but leave node_exporter exactly as it is.&lt;br&gt;
&lt;code&gt;sudo systemctl enable alloy&lt;/code&gt;&lt;br&gt;
&lt;code&gt;sudo systemctl start alloy&lt;/code&gt;&lt;br&gt;
At this point nothing has changed from Prometheus' perspective, you’re just running two collector concurrently.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. Configure Alloy
&lt;/h3&gt;

&lt;p&gt;Create your River configuration and point the &lt;code&gt;remote_write&lt;/code&gt; endpoint at your Prometheus receiver.&lt;br&gt;
Start Alloy and verify the service is healthy.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl status alloy &lt;span class="nt"&gt;--no-pager&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;--no-pager&lt;/code&gt; option simply prints the status and returns you to the shell instead of opening the interactive pager, making it much friendlier for automation and scripting.  Alloy also exposes a built-in UI on port 12345.  Opening it gives you a live view of every component in your telemetry pipeline.&lt;br&gt;
If &lt;code&gt;prometheus.exporter.unix&lt;/code&gt; and &lt;code&gt;prometheus.remote_write&lt;/code&gt; are healthy, your data is flowing.&lt;br&gt;
If &lt;code&gt;remote_write&lt;/code&gt; is unhealthy, the problem is usually one of three things:&lt;br&gt;
    • Incorrect endpoint URL&lt;br&gt;
    • Network connectivity&lt;br&gt;
    • Authentication&lt;br&gt;
...in roughly that order.&lt;/p&gt;
&lt;h3&gt;
  
  
  3. Verify the new telemetry
&lt;/h3&gt;

&lt;p&gt;Now both &lt;code&gt;node_exporter&lt;/code&gt; and Alloy are reporting metrics.&lt;br&gt;
First, verify Alloy locally.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:12345/metrics | &lt;span class="nb"&gt;head&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or check for a specific metric:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:12345/metrics | &lt;span class="nb"&gt;grep &lt;/span&gt;node_memory_MemAvailable_bytes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then open Prometheus or Grafana and confirm the new series is arriving. Since Alloy pushes via remote_write instead of being scraped, the standard up metric won't reflect this host anymore — up is generated per scrape target, and there's no scrape target here. Checking it after cutover will look like the host vanished, even though everything is working.&lt;br&gt;
Instead, check for freshness directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;time() - timestamp(node_uname_info{instance="canary-host"})
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A small, stable number means data is arriving on schedule. A growing number means the host has stopped pushing.&lt;br&gt;
Then compare a real metric such as available memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;node_memory_MemAvailable_bytes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't expect the labels to match perfectly — the important part is that the values agree and the new series continues updating.&lt;br&gt;
If your dashboards, recording rules, or alerting rules explicitly reference labels such as job="node_exporter" — or reference up for these hosts — now is a good time to identify them before removing the old exporter.&lt;/p&gt;
&lt;h3&gt;
  
  
  4. Retire node_exporter
&lt;/h3&gt;

&lt;p&gt;Once you're confident Alloy is working correctly, stop the old exporter.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl stop node_exporter
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl disable node_exporter
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;stop&lt;/code&gt; ends the running process.&lt;br&gt;
disable prevents it from quietly returning after the next reboot.&lt;br&gt;
Then remove the server's scrape job from Prometheus and reload the configuration.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;curl -X POST http://localhost:9090/-/reload
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reloading avoids interrupting metric collection for every other host while updating the configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Watch the canary
&lt;/h3&gt;

&lt;p&gt;Give the server at least one scrape interval under the new path.&lt;br&gt;
Verify that:&lt;br&gt;
    • Metrics remain fresh.&lt;br&gt;
    • Alerts continue behaving normally.&lt;br&gt;
    • Dashboards still populate.&lt;br&gt;
    • No recording rules broke because of label changes.&lt;/p&gt;

&lt;p&gt;Once everything looks healthy, repeat the process on the next server.&lt;br&gt;
After a handful of successful migrations, you'll have enough confidence to migrate small batches instead of individual hosts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common migration mistakes
&lt;/h2&gt;

&lt;p&gt;There are a few issues that show up repeatedly during migrations:&lt;br&gt;
    • Removing node_exporter before verifying Alloy.&lt;br&gt;
    • Forgetting to reload Prometheus after removing scrape targets.&lt;br&gt;
    • Alert rules or dashboards still referencing the old job label.&lt;br&gt;
    • Firewall rules blocking outbound Remote Write traffic.&lt;br&gt;
    • Assuming label changes won't affect existing dashboards.&lt;/p&gt;

&lt;p&gt;None of these are difficult to fix, but catching them during a canary migration is far less stressful than discovering them after migrating twenty servers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you gain
&lt;/h2&gt;

&lt;p&gt;The obvious benefit is consolidation.&lt;br&gt;
Instead of deploying separate agents for metrics, logs, and traces, Alloy provides a single telemetry pipeline that grows with your infrastructure.&lt;br&gt;
If you need logs, you can add a Loki component, for OTLP traces, add another component.  The overall architecture doesn't change.&lt;br&gt;
The less obvious benefit is security.  Once every monitored machine pushes telemetry outward, you can close your metrics ports entirely.  There's no longer a listening endpoint on every server, no scrape network to maintain, and no inbound firewall rule whose only purpose is monitoring.&lt;/p&gt;

&lt;p&gt;For anyone managing customer infrastructure—or simply trying to reduce attack surface—that's arguably the biggest improvement Alloy brings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Five years ago, exposing port 9100 across an internal network wasn't unusual, but today we're steadily moving toward zero-trust networking, outbound-only connectivity, and centralized telemetry pipelines.&lt;br&gt;
Grafana Alloy isn't compelling because it replaces node_exporter.  It's compelling because it aligns your monitoring architecture with where modern infrastructure is already heading, and strengthens your security posture by removing exposed enpoints.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Take the migration slowly.&lt;/li&gt;
&lt;li&gt;Run both agents for a while.&lt;/li&gt;
&lt;li&gt;Prove the new telemetry path.&lt;/li&gt;
&lt;li&gt;Then remove the old one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Done this way, there should never be a moment when a server isn't being watched.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>linux</category>
      <category>monitoring</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>The LLM narrates. The code decides.</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Tue, 07 Jul 2026 12:51:40 +0000</pubDate>
      <link>https://dev.to/irinobservability/the-llm-narrates-the-code-decides-33f2</link>
      <guid>https://dev.to/irinobservability/the-llm-narrates-the-code-decides-33f2</guid>
      <description>&lt;p&gt;Most of the "AI for observability" work I see right now hands the language model the judgment. I think that's backwards. Feed it the alert, feed it some metrics, ask it what's wrong, what should be done, and let it make the judgement call. Based on my experience working with language models, I decided that inverting the process provides better results.&lt;br&gt;
The short version: in my alerting pipeline, the set of allowable classifications is fixed in deterministic Python, and the model has to pick from it. The LLM's only job is to turn a structured verdict into an easily digestible sentence. It never decides whether something is bad, how bad it is, or what category of problem it is. It narrates within a decision space the code has already locked down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Instead of letting an LLM decide what's wrong with an alert, I let deterministic Python make every operational decision and restrict the model to explaining the result in plain English. The code classifies, validates, and aggregates; the LLM only narrates. That keeps the data consistent, prevents hallucinated classifications, and ensures the monitoring pipeline continues working even if the model fails.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;I run a small managed monitoring service. Alertmanager fires, a webhook lands, and historically that webhook produced a line like HighMemoryUsage on host web-vm, severity warning, which is accurate, but not terribly helpful. The person reading it still has to know what HighMemoryUsage implies, whether this host always runs hot, and whether to care. I wanted plain-English context attached to the alerts without altering the alert delivery process.&lt;/p&gt;

&lt;p&gt;The obvious move was to throw the whole alert at an LLM and ask it to explain. I tried that in the first iteration of this experiment, expecting it to be somewhat accurate, but not entirely reliable, and it did not disappoint, the model was confidently inconsistent. The same alert, fired three times, produced three different "root cause" categories. One run called a test alert a "Configuration or setup issue," the next called it "Configuration/Testing," the next something else again. If you're storing that output to do any kind of aggregation later (I am, I want to know when three different clients hit the same class of problem in the same week), free-form model output fragments into noise. Grouping on a field the model changes at random won't work.&lt;/p&gt;

&lt;p&gt;I kind of knew from the onset that I wouldn't get amazing results, and that it would be harder than it looked. So I started doing some research, and decided to flip the design. The little voice in the back of my head was right all along, don't let the model make the decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Split
&lt;/h2&gt;

&lt;p&gt;I created a pipeline that has a hard wall down the middle.&lt;br&gt;
On the deterministic side, Python does the classifying. The output is constrained to an eight-value enum: memory_pressure, cpu_saturation, disk_pressure, service_unavailability, network_issue, configuration_error, external_dependency, unknown. I aggregate on that field because it can only ever be one of eight strings. If nothing fits, the answer is unknown, which is itself a useful signal rather than a hallucinated/variant guess.&lt;br&gt;
On the narrative side, the LLM (llama3:8b, running locally on a box on my own LAN, data/network secure) must choose its classification from that fixed eight-value set, and it writes two short fields alongside it: what the alert is, in plain English, and what it means operationally. The code defines the shape of the answer; the model only fills in a slot that already exists. It is explicitly instructed not to suggest fixes and not to invent a cause, so it performs translation instead of analysis.&lt;br&gt;
The prompt returns strict JSON, grammar-constrained, so I get {what, means, likely_cause_class} every time and the enum value is validated against the allowed set on the way out. If the model returns something off-list, I capture the bug instead of storing a row.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context hydration
&lt;/h2&gt;

&lt;p&gt;A naive version of even the narration step gets you alarmist prose. When tuning the system I received a DiskFillPredicted alert, which on its face looked worth investigating. Then I looked at the host in question, which has had a flat disk-utilization baseline for months. The prediction was a rounding artifact, "your disk is about to fill" is actively misleading. The model had no way to know that from the alert alone, so it just wrote something.&lt;/p&gt;

&lt;p&gt;I fixed it by giving the model the same context a human would look at before reacting. Prior to the LLM call, Python does a fast lookup against the metrics backend for that host's recent baseline, and the prompt carries an explicit rule: if historical context is provided, weigh it over the alert's literal text. A predicted-disk-fill on a host with a stable months-long baseline is informational, not urgent.&lt;/p&gt;

&lt;p&gt;The latency budget for that lookup is five seconds. The LLM call itself takes about eighty seconds, because this runs deliberately on modest CPU-only hardware, a 4th generation Intel i7 with 16GB RAM and no GPU. That is a choice, not a constraint I am apologizing for: the whole posture of the service is that nothing leaves my LAN, so a slow local model beats a fast remote one. And the eighty seconds never reaches the person being alerted. Because the annotator rides alongside the existing path (more on that below), the raw alert lands in Slack and email instantly; the narrated version shows up as a separate annotation a minute or so later. Nothing is ever waiting on the model. Against that eighty-second call, five seconds of pre-fetch is under seven percent overhead and invisible. What mattered most was that if the metrics backend is slow or unreachable, the system fails immediately and falls back to the un-hydrated path, so the enrichment step can never block the result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fail-Closed
&lt;/h2&gt;

&lt;p&gt;That fallback instinct runs through the whole process, the ethos is that the annotator is additive, a 'nice-to-have.' Alertmanager routes to it with continue: true, so it sits alongside the existing Slack and email delivery processes, and can never block them. The webhook always returns 200, even when the LLM box is down or when the JSON is malformed, and a static fallback annotation gets used instead. The worst case is that the end-user gets a less flowery alert, but never misses one. The narration is an amenity layered on top of a delivery path that doesn't depend on it at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Advice, For What it's Worth
&lt;/h2&gt;

&lt;p&gt;"Let the model narrate, not analyze" is the easy version of the lesson, and it's true, but it isn't the hard part. The hard part is the enforcement: constraining the output to a fixed set you can validate, and keeping the model off the critical path so its failures cost you prose and never an alert. The model's value is fluency, not judgment, and fluency is the most replaceable thing in the stack. Push every actual decision into code you can test, constrain anything you'll later aggregate down to an enum, and treat the model as the last, most replaceable stage in the pipeline. If you can swap the model out tomorrow and your data stays clean, you've drawn the line in the right place.&lt;br&gt;
Mine's been running against live infrastructure for a few weeks. The prose is good, I'm sure when the hardware is upgraded and a more powerful model is put in place it will be better, but the reason I trust (tentatively) it is that the prose isn't doing the work.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>monitoring</category>
      <category>python</category>
    </item>
    <item>
      <title>Stop Relying Entirely on Uptime Kuma for Incident Response</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Thu, 25 Jun 2026 14:32:45 +0000</pubDate>
      <link>https://dev.to/irinobservability/stop-relying-entirely-on-uptime-kuma-for-incident-response-39fj</link>
      <guid>https://dev.to/irinobservability/stop-relying-entirely-on-uptime-kuma-for-incident-response-39fj</guid>
      <description>&lt;p&gt;Before I get into this, it is not a knock on Uptime Kuma. It's a genuinely amazing, easy-to-use piece of software. If you run a homelab or a small fleet and you're not using it, you probably should be. It's free, self-hosted, beautiful, and it does the thing it was built to do better than almost anything else at any price.&lt;/p&gt;

&lt;p&gt;There's always a "but," though, so before we get to it I want to spend a little time on what Uptime Kuma does well.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR:
&lt;/h2&gt;

&lt;p&gt;Uptime Kuma is excellent at telling you when a service becomes unreachable, but it cannot explain why a service is slow or unhealthy while still responding. That requires internal metrics from tools like Prometheus, Grafana, and Alloy. Reachability monitoring and systems monitoring solve different problems, and mature environments typically use both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Uptime Kuma excels
&lt;/h2&gt;

&lt;p&gt;Uptime Kuma answers one question extremely well: is this reachable? It'll ping a host, hit an HTTP endpoint and check the status code, watch a TCP port, validate a TLS cert's expiry, query a DNS record, check a keyword on a page, watch a Docker container, even poke a game server. It checks on a tight interval, shows you a clean history, and when something stops responding it fires a notification through basically any channel you can name. Ninety-plus notification integrations. Status pages you can hand to your users. Two-factor auth. A genuinely nice UI.&lt;/p&gt;

&lt;p&gt;For "tell me the moment my website, my reverse proxy, my Plex, or my Home Assistant stops answering," it's close to perfect. The interval is short, setup is measured in minutes, and there's practically no maintenance. It has earned a famously loyal userbase for a reason.&lt;/p&gt;

&lt;p&gt;I'm not here to tell you it isn't the answer, or to convince you to ditch it for something else. I still think everyone running infrastructure of any size should have something like it watching their endpoints. I have it running in a Proxmox container on my own homelab. But there's a gap I noticed while using it, and this post is about the specific moment when you ask Uptime Kuma a question it wasn't designed to answer, and what you do when that moment arrives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Growing pains
&lt;/h2&gt;

&lt;p&gt;There usually comes a time, as your homelab or business grows, when your database starts feeling slow. Queries that used to be instant are taking a little longer. It's not a real problem yet, but you can tell something is off.&lt;/p&gt;

&lt;p&gt;The Uptime Kuma dashboard is all green. The database port is answering, the HTTP healthcheck returns 200, every light on the board is on. Uptime Kuma is correctly reporting that the service is up.&lt;/p&gt;

&lt;p&gt;And it's not wrong. That's the thing. The service &lt;em&gt;is&lt;/em&gt; reachable. But "reachable" and "healthy" mean different things, and you've just walked into the space between them. If the disk that database lives on is pinned at 100% IO utilization because a backup job and a big query are fighting over it, your queries are queuing behind that contention, and from the outside the port still answers in time to pass the check. The board is green, the database is slow, and there's no contradiction.&lt;/p&gt;

&lt;p&gt;Uptime Kuma doesn't see any of that, and the reason it can't isn't a missing feature, it's the architecture. It checks your systems from the outside looking in. It has no way to see what's happening inside your servers. What are the disk, memory, CPU, and kernel actually doing?&lt;/p&gt;

&lt;p&gt;What you need at that moment is something standing &lt;em&gt;inside&lt;/em&gt; the box, reading the system from within. That's a different category of tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reachability versus internals
&lt;/h2&gt;

&lt;p&gt;There are two kinds of monitoring, and once you see the split you understand why mature setups end up running both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reachability monitoring&lt;/strong&gt; (Uptime Kuma) asks the basic questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can I get to it?&lt;/li&gt;
&lt;li&gt;Is the port open, the page loading, the cert valid, the container running?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It reports what it can see from the outside, which is exactly what you want for "is my service up and can my users reach it." It's easy, simple, and honest about what it knows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Systems monitoring&lt;/strong&gt; (the Prometheus world) asks questions that are a little more involved:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What's going on inside the machine?&lt;/li&gt;
&lt;li&gt;How busy is each CPU core?&lt;/li&gt;
&lt;li&gt;How much memory is actually available once you account for cache?&lt;/li&gt;
&lt;li&gt;What's the disk IO utilization, the queue depth, the read and write latency?&lt;/li&gt;
&lt;li&gt;How much network throughput, how many dropped packets?&lt;/li&gt;
&lt;li&gt;Is memory slowly leaking over days?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's an internal view, and it answers &lt;em&gt;why&lt;/em&gt; a service is behaving the way it is.&lt;/p&gt;

&lt;p&gt;Neither replaces the other. Reachability tells you that something is wrong. Systems metrics tell you why. The database scenario above needs both: Uptime Kuma to eventually notice if the slowness becomes an actual outage, and system metrics to explain the slowness long before it gets there.&lt;/p&gt;

&lt;h2&gt;
  
  
  The internal view
&lt;/h2&gt;

&lt;p&gt;The standard way to get the inside view on Linux is a tiny agent called node_exporter. It's a small binary that runs on the box, reads metrics straight from the kernel, and exposes them for a time-series database (Prometheus) to collect. Pair it with Grafana for dashboards, and for logs, pair Loki with a shipper. The traditional choice there was Promtail, though Grafana has since moved Promtail into long-term support and now steers you toward Grafana Alloy, which handles both metrics and logs in a single agent. (I wrote a comparison of those two separately.)&lt;/p&gt;

&lt;p&gt;With either node_exporter or Alloy running, the database scenario stops being a mystery. The exact moment things felt slow, you can pull up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Disk IO utilization on that box, and watch it pin to 100% right when the slowness started.&lt;/li&gt;
&lt;li&gt;The specific disk and the read/write split, so you can see it was the backup volume contending with queries.&lt;/li&gt;
&lt;li&gt;CPU broken out by mode, so you can rule out CPU as the cause.&lt;/li&gt;
&lt;li&gt;Memory availability over the past week, so you can see whether pressure had been building.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And if you have Loki collecting logs alongside the metrics, you can line up the disk IO spike against the log line where the backup job kicked off, and the whole story assembles itself in one view. Uptime Kuma told you the service was up. The system metrics tell you the backup job is strangling your database disk, which is what you actually need to know to fix it before it hits production.&lt;/p&gt;

&lt;p&gt;(If the PromQL behind those dashboards is unfamiliar, I wrote up the five queries you actually need to monitor a Linux server separately.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part: this is more work
&lt;/h2&gt;

&lt;p&gt;Standing up a stack to see inside your servers is not as simple as setting up Uptime Kuma, and that simplicity is a real part of why Uptime Kuma is so loved. Moving to system metrics comes with a cost.&lt;/p&gt;

&lt;p&gt;node_exporter or Alloy goes on every server, with Prometheus running somewhere to collect from them. Grafana dashboards have to be built, or imported from the community and then tweaked until they're readable instead of overwhelming. Alert rules have to be written to fire on real problems without crying wolf. Metric and log retention have to be configured. And then the whole thing needs ongoing maintenance.&lt;/p&gt;

&lt;p&gt;This is the irony nobody warns you about: you now have a monitoring stack that itself needs monitoring, which is partly why you want predictive disk alerts on the box running Prometheus.&lt;/p&gt;

&lt;p&gt;None of it is hard, exactly. But it's an ongoing process with no end, and it's a different commitment than the near-zero maintenance of an Uptime Kuma container you set up once and edit when new services come online. Uptime Kuma is the right tool for reachability and status pages, and it costs almost nothing to run. System metrics cost more, but they become relevant the moment you start wondering why services aren't behaving, even though they're still showing green.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what do you do with this
&lt;/h2&gt;

&lt;p&gt;The real takeaway here isn't a product, it's the distinction, because that understanding outlives any particular tool. Outside-in tells you something broke. Inside-out tells you why. Uptime Kuma is one of the best outside-in tools ever made, and it'll happily keep doing that job for you forever. It just wasn't built to explain the why.&lt;/p&gt;

&lt;p&gt;When you do need the why, you've got two options: run the inside-view stack yourself (node_exporter or Alloy, Prometheus, Grafana, Loki), which is completely viable and a great way to learn, or hand it to someone who runs it for you.&lt;/p&gt;

&lt;p&gt;For full disclosure, that second path is the reason I built &lt;a href="https://www.irinobservability.com/pricing.html" rel="noopener noreferrer"&gt;Irin Observability&lt;/a&gt;, a managed version of that inside-view stack for small teams and homelabs that have outgrown pure reachability checks but don't want a second full-time job maintaining a metrics pipeline. It's meant to sit &lt;em&gt;alongside&lt;/em&gt; something like Uptime Kuma, not replace it, because reachability and internals are different jobs and the mature answer is to run both.&lt;/p&gt;

&lt;p&gt;Either way, keep the green board. Just add the view from inside the box when you start asking why.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>linux</category>
      <category>monitoring</category>
      <category>productivity</category>
    </item>
    <item>
      <title>You Don't Need Kubernetes to Monitor 20 Linux VMs</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Tue, 23 Jun 2026 12:42:27 +0000</pubDate>
      <link>https://dev.to/irinobservability/you-dont-need-kubernetes-to-monitor-20-linux-vms-5af4</link>
      <guid>https://dev.to/irinobservability/you-dont-need-kubernetes-to-monitor-20-linux-vms-5af4</guid>
      <description>&lt;p&gt;If you've ever tried to set up Prometheus by following the official getting-started path, you're likely to find a path that does not follow your infrastructure model. Out of the gate, page one mentions kube-prometheus-stack. Page two wants you to install a Helm chart, and page three assumes you already have a cluster running. The documentation for monitoring plain Linux servers is in there somewhere, but you have to dig for it. When you do find it, the tone suggests you are doing something slightly old-fashioned.&lt;/p&gt;

&lt;p&gt;If that sounds like your setup, the tooling is making this harder than it actually is. Monitoring a fleet of Linux VMs is fairly simple and has been for years. It is just obscured behind documentation that would prefer to sell you something bigger.&lt;/p&gt;

&lt;p&gt;Modern infrastructure tooling has quietly decided everyone runs Kubernetes. If you don't, the assumption is that you eventually will. Meanwhile, most real-world infrastructure still runs on VMs.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Modern observability documentation often assumes you're running Kubernetes. Most small teams aren't. If you're managing a fleet of Linux VMs, node_exporter plus Prometheus gives you everything you need for infrastructure monitoring with a single lightweight agent and a straightforward deployment model. No cluster required.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  VMs are often the answer
&lt;/h2&gt;

&lt;p&gt;For most small businesses, running VMs instead of Kubernetes does not mean you failed to evolve. Most workloads under a certain scale perform better on VMs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One process per box, predictable resource limits, and the ability to ssh in and look at what's happening, which makes it easier to keep track of the infrastructure as a whole.&lt;/li&gt;
&lt;li&gt;They're cheaper, both financially and in the mental overhead of running them.&lt;/li&gt;
&lt;li&gt;Backups and snapshots are straightforward in a way stateful Kubernetes still isn't.&lt;/li&gt;
&lt;li&gt;There's no control plane that itself needs monitoring and upgrades and care.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Kubernetes solves problems that mostly pertain to companies with dozens of engineers and hundreds of services. For platforms that consist of 20 VMs, Kubernetes is the wrong tool, and being told you need it before you're allowed to have monitoring is the wrong approach.&lt;/p&gt;

&lt;h2&gt;
  
  
  What node_exporter actually is
&lt;/h2&gt;

&lt;p&gt;What you need is called node_exporter, a lightweight systemd process.&lt;/p&gt;

&lt;p&gt;It's a single Go binary, around 25 MB. It runs as one process on each VM, reads metrics from the kernel through &lt;code&gt;/proc&lt;/code&gt; and &lt;code&gt;/sys&lt;/code&gt;, and exposes them on an HTTP endpoint, normally port 9100. It's very uncomplicated: there's no daemon set, operator, sidecar, CRD, cluster, or control plane. It runs quietly in the background and answers HTTP on port 9100 with a plain-text list of numbers. You can curl it yourself and read it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://&amp;lt;localhost or IP&amp;gt;:9100/metrics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What comes back is a few hundred lines of metrics containing CPU time per core per mode, memory broken down by category, disk space per mountpoint, network bytes per interface, load, uptime, and open file handles. It tells you everything the kernel knows about the server, in a format Prometheus reads directly.&lt;/p&gt;

&lt;p&gt;The agent the big observability vendors want to install on your servers is doing this same job. It reads from &lt;code&gt;/proc&lt;/code&gt; and exposes metrics, but they've wrapped it in a config model and an update mechanism and a logo. The core of it is what node_exporter has been doing for over a decade. You are not missing out on some sophisticated technology by over-complicating your system. The simple, plain version is the technology.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting up one VM
&lt;/h2&gt;

&lt;p&gt;Here's the actual setup on a single box. Check the releases page for the current version before you run this, the version string changes.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Download the binary&lt;/span&gt;
wget https://github.com/prometheus/node_exporter/releases/download/v1.8.2/node_exporter-1.8.2.linux-amd64.tar.gz

&lt;span class="c"&gt;# Extract and install&lt;/span&gt;
&lt;span class="nb"&gt;tar &lt;/span&gt;xzf node_exporter-1.8.2.linux-amd64.tar.gz
&lt;span class="nb"&gt;sudo mv &lt;/span&gt;node_exporter-1.8.2.linux-amd64/node_exporter /usr/local/bin/

&lt;span class="c"&gt;# Run it as its own unprivileged user&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;useradd &lt;span class="nt"&gt;--no-create-home&lt;/span&gt; &lt;span class="nt"&gt;--shell&lt;/span&gt; /bin/false node_exporter

&lt;span class="c"&gt;# systemd unit&lt;/span&gt;
&lt;span class="nb"&gt;sudo tee&lt;/span&gt; /etc/systemd/system/node_exporter.service &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /dev/null &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;'
[Unit]
Description=Node Exporter
After=network.target

[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter

[Install]
WantedBy=multi-user.target
&lt;/span&gt;&lt;span class="no"&gt;EOF

&lt;/span&gt;&lt;span class="c"&gt;# Start it&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl daemon-reload
&lt;span class="nb"&gt;sudo &lt;/span&gt;systemctl &lt;span class="nb"&gt;enable&lt;/span&gt; &lt;span class="nt"&gt;--now&lt;/span&gt; node_exporter

&lt;span class="c"&gt;# Confirm it's alive&lt;/span&gt;
curl http://localhost:9100/metrics | &lt;span class="nb"&gt;head&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With these ten commands, you can have it running in under five minutes. It sits at roughly 20 MB of RAM and you'll likely forget it's there. One thing you should do is lock down port 9100. Leave it open to your monitoring server and nothing else. node_exporter exposes details about your system and it shouldn't be reachable from the public internet. It should be behind your firewall.&lt;/p&gt;

&lt;h2&gt;
  
  
  It is a little repetitive
&lt;/h2&gt;

&lt;p&gt;The same setup runs on every machine, so there are a few ways to deploy it if you have more than 5 to 10 servers to monitor. The setup is the same for almost all Linux distributions.&lt;/p&gt;

&lt;p&gt;If you're already using Ansible, the node_exporter playbook is about 30 lines and is one of the most copy-pasted snippets out there. The &lt;code&gt;cloudalchemy.node_exporter&lt;/code&gt; role does it for you with reasonable defaults if you'd rather not write your own.&lt;/p&gt;

&lt;p&gt;You can also use a shell loop over ssh if you don't want to add new tooling. Walk your hostnames, ssh in, run the commands above. Twenty boxes will probably take around ten minutes.&lt;/p&gt;

&lt;p&gt;If you spin servers up and down often using a VM image or cloud-init, you can just include node_exporter in the base image. Every new VM will show up already monitoring itself.&lt;/p&gt;

&lt;p&gt;The monitoring side is one Prometheus instance pointed at the list of servers you want to monitor:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# prometheus/prometheus.yml&lt;/span&gt;
&lt;span class="na"&gt;scrape_configs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;job_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;linux-vms'&lt;/span&gt;
    &lt;span class="na"&gt;static_configs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;targets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;vm1.example.com:9100&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;vm2.example.com:9100&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;vm3.example.com:9100&lt;/span&gt;
          &lt;span class="c1"&gt;# ...the rest of them&lt;/span&gt;
        &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For 20 boxes, that static list is genuinely fine. If you add and remove servers a lot, &lt;code&gt;file_sd_configs&lt;/code&gt; lets Prometheus pick up target changes from a file without a restart, which carries you much further. The setup isn't too much more complicated:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# prometheus/prometheus.yml&lt;/span&gt;
&lt;span class="na"&gt;scrape_configs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;job_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;linux-vms'&lt;/span&gt;
    &lt;span class="na"&gt;file_sd_configs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;files&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/etc/prometheus/file_sd/linux-vms.yml&lt;/span&gt;
        &lt;span class="na"&gt;refresh_interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;30s&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The file structure requires that you add a &lt;code&gt;file_sd&lt;/code&gt; directory to the prometheus folder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;prometheus/
├── prometheus.yml
└── file_sd/
    └── linux-vms.yml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# file_sd/linux-vms.yml&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;targets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;vm1.example.com:9100&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;vm2.example.com:9100&lt;/span&gt;
  &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
    &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;

&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;targets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;db1.example.com:9100&lt;/span&gt;
  &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;production&lt;/span&gt;
    &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;database&lt;/span&gt;

&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;targets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;staging1.example.com:9100&lt;/span&gt;
  &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;staging&lt;/span&gt;
    &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you put each server directly into &lt;code&gt;prometheus.yml&lt;/code&gt;, you have to restart Prometheus every time you add one. By putting your servers in the file under &lt;code&gt;file_sd&lt;/code&gt;, Prometheus picks them up automatically on the refresh interval. That's a little extra structure up front, so if your infrastructure is largely static it isn't really worth it. If you're constantly onboarding or removing servers, the extra layer removes a lot of the maintenance.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you can actually see
&lt;/h2&gt;

&lt;p&gt;With node_exporter on every VM and one Prometheus pulling from them, here are real questions you can answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CPU across the whole fleet for the last hour: one query over &lt;code&gt;node_cpu_seconds_total&lt;/code&gt;, split by instance.&lt;/li&gt;
&lt;li&gt;Which box is closest to full: &lt;code&gt;node_filesystem_avail_bytes&lt;/code&gt; against &lt;code&gt;node_filesystem_size_bytes&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;When vm7 last rebooted: &lt;code&gt;node_boot_time_seconds&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Which box is dropping the most packets: a rate over &lt;code&gt;node_network_receive_drop_total&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Whether memory has been slowly tightening on anything over the past week: &lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt; plotted across all instances.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything can be viewed in Grafana using queries written in PromQL. I wrote up &lt;a href="https://www.irinobservability.com/blog/five-promql-queries-linux-server" rel="noopener noreferrer"&gt;the five basic queries you need to monitor a Linux server&lt;/a&gt; separately, with each one explained in detail.&lt;/p&gt;

&lt;p&gt;That covers what a small fleet typically needs. Monitoring doesn't require Kubernetes, or giant vendors like Datadog, or agent vendors. A Go binary on each box and one instance of Prometheus and Grafana.&lt;/p&gt;

&lt;h2&gt;
  
  
  Maintenance costs
&lt;/h2&gt;

&lt;p&gt;Getting node_exporter onto 20 VMs and setting up Prometheus and Grafana is relatively easy. It's all open source and available to anyone. But most teams underestimate dashboard design, alert tuning, retention planning, and long-term maintenance. Making sure Prometheus stays healthy and the &lt;code&gt;prometheus.yml&lt;/code&gt; and &lt;code&gt;file_sd/*.yml&lt;/code&gt; files are all up to date, building functional dashboards, writing alert rules that fire on real problems without creating noise, sorting out retention, getting alerts somewhere a human will actually see them, and keeping all of it patched as each piece ships new versions: that becomes ongoing operational work somebody has to own. All of it grows in complexity with the fleet. On top of that, the monitoring stack itself can go down, which takes time and effort to troubleshoot and fix.&lt;/p&gt;

&lt;p&gt;If you like that sort of work, or you have dedicated people who can take on the additional load, node_exporter, Prometheus, and Grafana are excellent. If you have the money to spend, Datadog is a great company.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Irin comes in
&lt;/h2&gt;

&lt;p&gt;Because maintaining the monitoring stack is a burden most small businesses don't have the time or resources for, I built &lt;a href="https://www.irinobservability.com/pricing.html" rel="noopener noreferrer"&gt;Irin Observability&lt;/a&gt;. You keep your attention on running your business and keep an eye on it through dashboards and alerts that are already built and tuned. Instead of node_exporter, Irin uses Grafana Alloy as the agent. It covers the same infrastructure metrics, ships your logs, supports additional telemetry pipelines, and installs with a single bootstrap command. Instead of a pull-based model that requires you to open a port to your monitoring server, it pushes your data out through an encrypted Cloudflare tunnel. Your dashboards, alerts, and retention live on Irin's infrastructure. The only thing on your boxes is the agent, and it stays out of the way.&lt;/p&gt;

&lt;p&gt;The pitch really isn't the point, though, and I'm only scratching the surface of what node_exporter or Alloy can do. The point is that the docs may be telling you a story that isn't true for your situation. You do not need Kubernetes to watch a handful of Linux servers. You need a small binary on each box and something to scrape it. Run that something yourself or pay someone to run it, either is fine. The architecture underneath is simple no matter who operates it, and it's been sitting in plain sight the whole time under a pile of cloud-native marketing.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>linux</category>
      <category>monitoring</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>The Only 5 PromQL Queries You Really Need to Monitor a Linux Server</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Tue, 16 Jun 2026 14:45:14 +0000</pubDate>
      <link>https://dev.to/irinobservability/the-only-5-promql-queries-you-really-need-to-monitor-a-linux-server-39n6</link>
      <guid>https://dev.to/irinobservability/the-only-5-promql-queries-you-really-need-to-monitor-a-linux-server-39n6</guid>
      <description>&lt;p&gt;PromQL has its quirks, and can be difficult, but basic monitoring of a Linux server is not.  I’ve boiled it down to five queries that will give you the basic outline of how your system is performing.  This article discusses the queries for CPU, memory, disk space, disk IO, and network, with a plain explanation of how each one works. &lt;/p&gt;

&lt;p&gt;PromQL has a reputation for being intimidating, and the reputation is half-earned.  The full language is genuinely deep, with subtleties around ranges, rates, and vector matching that take a while to learn and understand.  What nobody tells you when you are starting out is that monitoring a single Linux box well does not require a comprehensive grasp of the language.  It requires about five questions, asked correctly.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;You don’t need hundreds of metrics to monitor a Linux server effectively.  Five PromQL queries covering CPU, memory, disk space, disk IO, and network traffic will catch the most common server issues.  This article explains each query, how it works, and why it belongs on your dashboard.&lt;br&gt;
These queries work with both node_exporter and Grafana Alloy and are commonly used in Grafana dashboards, Prometheus alert rules, and Linux server monitoring setups. If you're looking for practical PromQL examples rather than a full PromQL tutorial, start here.&lt;/p&gt;
&lt;h2&gt;
  
  
  Quick Reference:
&lt;/h2&gt;

&lt;p&gt;These are the exact PromQL queries used to monitor CPU usage, memory utilization, disk space, disk IO, and network throughput on Linux servers running node_exporter or Grafana Alloy.&lt;/p&gt;

&lt;p&gt;CPU Usage&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;100 - (avg by (instance)(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Memory Usage&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;100 * (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes))
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Disk Space&amp;nbsp;100&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes{fstype!~"tmpfs|overlay"} * 100)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Disk IO Saturation&amp;nbsp;rate&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(node_disk_io_time_seconds_total[5m])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Network Throughput&amp;nbsp;rate&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(node_network_receive_bytes_total{device!~"lo|veth.*"}[5m])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This article assumes you have node_exporter or Grafana Alloy running and Prometheus scraping it.  Alloy’s metrics are identical to node_exporter’s, Alloy's &lt;code&gt;prometheus.exporter.unix&lt;/code&gt; component is node_exporter under the hood, so every query below works for both.  If you are still deciding between the two or would like to learn more, we wrote a separate comparison of Alloy and node_exporter that discusses the two and when each makes sense that can be found here.&lt;/p&gt;

&lt;p&gt;Before we dive into the queries, it’s important to point out the difference between a gauge and a counter on the Grafana dashboard.  A gauge is a value that goes up and down, like memory in use right now or CPU temperature.  It shows you what’s happening now, and you read a gauge directly.  A counter only ever goes up, like total bytes received since boot or total seconds the CPU has spent working.   It’s a count over time.  You almost never read a counter directly, because "847 billion bytes since the machine booted" is useless.  The relevant question to ask yourself when looking at counters is: how fast it is climbing?  That’s what rate() tells you.  Three of the five queries discussed below are counters, and once you see why they all use rate(), the pattern makes sense and PromQL starts making a little more sense.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. CPU usage (percent busy)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When I created my first Grafana dashboard, I expected this to be the easiest query to write.  I think most people expect it to be simple and, like me, are confused when it’s not.&lt;br&gt;&lt;br&gt;
node_exporter does not expose a "CPU percent" metric, because there is no honest single number for it.  What it exposes is &lt;code&gt;node_cpu_seconds_total&lt;/code&gt;, a counter that tracks how many seconds each CPU core has spent in each mode: idle, user, system, iowait, and a few others.  The machine is always doing one of these, so the modes always add up to 100 percent of available CPU time.&lt;br&gt;
The cleanest way to ask "how busy is the CPU" is to measure how much it is not idle, so we work from the idle mode and subtract from 100. Reading the query from the inside out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;node_cpu_seconds_total{mode="idle"}&lt;/code&gt; selects just the idle counter, for every core. &lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;rate(...[5m])&lt;/code&gt; is the key piece. It looks at how that counter changed over the last 5 minutes and returns a per-second rate. For the idle counter, the rate is "idle seconds accumulated per second," which is a number between 0 and 1 per core: 1.0 means a core was fully idle, 0.0 means fully busy. &lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;avg by&lt;/code&gt; (instance) averages that across all the cores on the machine, so a 4-core box gives you one number instead of four. &lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;code&gt;* 100&lt;/code&gt; turns the 0-to-1 fraction into a percentage, and 100 - (...) flips "percent idle" into "percent busy."&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;[5m]&lt;/code&gt; window is a smoothing choice, not a magic number.  A wider window like &lt;code&gt;[5m]&lt;/code&gt; smooths out brief spikes and shows the server sustained load.  If you use a narrower window like &lt;code&gt;[1m]&lt;/code&gt; it’s twitchier and catches short bursts.  For alerting on a server, sustained load is usually what matters, which is why our own default alert fires on CPU above 80 percent for five-plus minutes rather than reacting to every momentary peak.  By extending the window from &lt;code&gt;[1m]&lt;/code&gt; to &lt;code&gt;[5m]&lt;/code&gt; you’re able to reduce noise, but can still see when there’s a problem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7t4suvu0e6smj88jzjvh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7t4suvu0e6smj88jzjvh.png" alt=" " width="800" height="363"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The production query adds label filters for multi-tenant use; the core logic is identical.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Memory usage (percent used)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;100 * (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes))
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Memory is best read as a gauge, so no &lt;code&gt;rate()&lt;/code&gt; is used for this query.  The values are read at a glance.  There is one trap worth understanding, because it’s fairly common.&lt;/p&gt;

&lt;p&gt;The naive instinct is to use &lt;code&gt;node_memory_MemFree_bytes&lt;/code&gt;, the amount of completely unused memory.  It’s a baseline metric that node_exporter provides, and it seems like it makes perfect sense to pull it directly to the panel. On a healthy Linux system, "free" memory is often very low by design.  Linux uses otherwise-idle RAM for the page cache, holding recently-read files in memory so it does not have to hit the disk again.  That memory looks "used" but is instantly reclaimable the moment a program actually needs it.  If you track and alert on low &lt;code&gt;MemFree&lt;/code&gt;, you’ll get unnecessary alerts on servers that are working as intended.&lt;br&gt;
The number you need to track is &lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt;.  The kernel calculates this for you.  It is the memory genuinely available for new programs to use, after accounting for the cache it can reclaim.&lt;br&gt;&lt;br&gt;
So the query reads: take available memory divided by total memory, which gives you the fraction available.  Subtract that from 1 to get the fraction used, and multiply by 100 for a percentage.  A good threshold for this panel is 85 percent, or when available memory drops below 15 percent.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpgux55yrbipdyduib78p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpgux55yrbipdyduib78p.png" alt=" " width="800" height="377"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The production query adds label filters for multi-tenant use; the core logic is identical.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Disk space (percent full)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;100 - (node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes{fstype!~"tmpfs|overlay"} * 100)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Disk space is also a gauge, and structurally this is the same shape as the memory query.  Take the available divided by total, and turn it into percent used.  What makes this query tricky is the label filter, because disk monitoring using unfiltered queries get noisy.&lt;br&gt;
A Linux machine reports many "filesystems" that are not real disks.  Tracking every single one would create a massive amount of noise, and make it difficult to parse out which disks are likely to cause a problem in the near future.  &lt;code&gt;tmpfs&lt;/code&gt; is memory-backed temporary storage, overlay filesystems belong to running containers, and there are others. If you monitor all of them, your "disk full" dashboard lights up over ephemeral mounts that are largely irrelevant. The filter &lt;code&gt;fstype!~"tmpfs|overlay"&lt;/code&gt; strips those out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;fstype&lt;/code&gt; is the label node_exporter attaches describing the filesystem type. &lt;/li&gt;
&lt;li&gt;
&lt;code&gt;!~ means&lt;/code&gt; "does not match this regular expression." (=~ would be "does match.") &lt;/li&gt;
&lt;li&gt;
&lt;code&gt;"tmpfs|overlay"&lt;/code&gt;is the regex: the | is an OR, so this matches either type, and !~ excludes both. &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This query leaves you with the actual disks on your server.  This is also the first time a regex-matching operator has popped up in this article.   These two operators, &lt;code&gt;=~&lt;/code&gt; and &lt;code&gt;!~&lt;/code&gt; are how to do most of the flexible filtering in PromQL.  Once you can include or exclude by pattern, you can filter metrics any way you need.&lt;/p&gt;

&lt;p&gt;One caveat: this query returns one result per mounted disk, which will show metrics for each mounted drive on your server.  A server with a separate / and /data should show you both, because either can fill independently.  Setting the threshold limit to something like 85 percent full gives you time to act before the disk is full.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fitsv6lwnaefzx1n840of.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fitsv6lwnaefzx1n840of.png" alt=" " width="799" height="384"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The production query adds label filters for multi-tenant use; the core logic is identical.  The production query uses max by (instance) rather than the simplified version described above.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Disk IO (how saturated the disk is)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;rate(node_disk_io_time_seconds_total[5m])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fourth query uses a counter, so &lt;code&gt;rate()&lt;/code&gt; returns.  This query answers a question that disk-space monitoring doesn’t address.  Your disk can have plenty of free space and still create problems because it cannot read and write fast enough to keep up with what the system is demanding of it.&lt;br&gt;
&lt;code&gt;node_disk_io_time_seconds_total&lt;/code&gt; counts the total seconds the disk spent actively busy with input/output (IO).  Because it is a counter, you wrap it in &lt;code&gt;rate(...[5m])&lt;/code&gt; to get "seconds of IO activity per second," which is effectively a utilization fraction.  A result near 1.0 means the disk was busy essentially the entire time, which tells you that the disk is saturated.  A result near 0.1 means it was busy about 10 percent of the time, with plenty of headroom.&lt;/p&gt;

&lt;p&gt;This is the metric that can help to identify where slowdowns are coming from. When a database gets sluggish, or backups drag but CPU and memory look fine, disk IO saturation is very often the culprit.  It’s the kind of problem that simple up-or-down monitoring won’t tell you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ff2g2emt3rolx8icdv9dv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Ff2g2emt3rolx8icdv9dv.png" alt=" " width="799" height="357"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The production query adds label filters for multi-tenant use; the core logic is identical.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Network throughput (bytes per second)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight prometheus"&gt;&lt;code&gt;&lt;span class="nb"&gt;rate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;node_network_receive_bytes_total&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="na"&gt;device&lt;/span&gt;&lt;span class="o"&gt;!~&lt;/span&gt;&lt;span class="s2"&gt;"lo|veth.*"&lt;/span&gt;&lt;span class="p"&gt;}[&lt;/span&gt;&lt;span class="mi"&gt;5m&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fifth query is a counter and a regex filter together, which is why I saved it for last.  If you understand this one, you’ll have a better understanding of  the pattern behind all five.&lt;br&gt;
&lt;code&gt;node_network_receive_bytes_total&lt;/code&gt; is a counter of total bytes received on each network interface since boot. &lt;code&gt;rate(...[5m])&lt;/code&gt; turns it into bytes per second, your live inbound throughput.  To watch outbound traffic, swap in &lt;code&gt;node_network_transmit_bytes_total&lt;/code&gt;, or create a second query in your panel to view the two side by side.&lt;br&gt;
The filter handles the same noise problem that the disk query does.  A Linux host has interfaces that you typically don’t need to keep an eye on: &lt;code&gt;lo&lt;/code&gt; is the loopback (the machine talking to itself), and &lt;code&gt;veth&lt;/code&gt; interfaces are the virtual ethernet links Docker and other container runtimes create, often dozens of them. &lt;code&gt;device!~"lo|veth.*"&lt;/code&gt; excludes them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;lo&lt;/code&gt; matches the loopback exactly. &lt;/li&gt;
&lt;li&gt;
&lt;code&gt;veth.*&lt;/code&gt; is a regex where &lt;code&gt;.&lt;/code&gt; means "any character" and &lt;code&gt;*&lt;/code&gt; means "zero or more of the preceding," so &lt;code&gt;veth.*&lt;/code&gt; matches &lt;code&gt;veth&lt;/code&gt; followed by anything: veth1a2b3c, vethABCD`, all of them. &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This leaves your physical or primary virtual interface(s), the one(s) carrying traffic that actually matters.  The output is in bytes per second, so if you would rather see bits per second to compare against your network provider's numbers, multiply by 8.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8okll2hcbvb6fiaocfxd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8okll2hcbvb6fiaocfxd.png" alt=" " width="800" height="449"&gt;&lt;/a&gt;&lt;br&gt;
_The production query adds label filters for multi-tenant use; the core logic is identical.  TX is shown as negative so RX and TX can share one panel without overlap.&lt;br&gt;
_&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern
&lt;/h2&gt;

&lt;p&gt;Looking over the five queries discussed here, there are only a few moving parts.  Gauges (memory, disk space) you read directly as available-over-total.  Counters (CPU, disk IO, network) you wrap in rate() to ask how fast they are climbing.  And label filters with &lt;code&gt;=~&lt;/code&gt; and &lt;code&gt;!~&lt;/code&gt; let you cut out the noise so you are watching real disks and real interfaces instead of container ephemera.  Five queries, three ideas that give you basic coverage of your servers. &lt;/p&gt;

&lt;h2&gt;
  
  
  Where this leaves you
&lt;/h2&gt;

&lt;p&gt;If you put these five on a dashboard with sensible thresholds, you have covered the large majority of what goes wrong on a single Linux server: the CPU is overworked, it runs out of memory, a disk fills up, a disk chokes, or its network saturates.  There’s always more that you can monitor, node_exporter and Alloy provide a massive amount of system metrics, but anything fancier is a refinement of these fundamentals.  You can view Irin’s System Health dashboard here to see these five queries alongside a few others.&lt;br&gt;
Going from "five queries in an expression browser" to "a real monitoring setup" is more work than it looks.  Some of the gauge queries are modified and used as time series, so you can see what’s happening over time, not just in that instant.  Prometheus needs to be set up to store the data with appropriate retention, Grafana dashboards need to be built around these queries, alert rules wired to thresholds that don’t create noise, and somewhere for the alerts to actually go.  Setup isn’t overwhelming, but it is an ongoing process to keep it running and tuned.  The monitoring system itself needs to be kept healthy, and thresholds/alerts need to be tuned to your system.  Then, there’s always the danger of over-monitoring, the first dashboard I created years ago was an endless scroll, it had EVERYTHING, which turned out to be too much.  Looking at the dashboard was overwhelming, and I couldn’t just take a glance at it to see how the system was doing, which is the goal.&lt;br&gt;
That recurring chore is the gap Irin Observability exists to fill. We ship these exact queries, pre-built dashboards, and tuned alert thresholds (the 80-percent-CPU, 15-percent-memory, 85-percent-disk defaults referenced above are ours) as a flat-rate managed service, so you get the visibility without becoming the person who maintains the monitoring stack. But whether you run it yourself or hand it off, the five queries above are the foundation either way.&lt;/p&gt;

&lt;p&gt;_Want the next step? Once you’re familiar with these, the natural follow-up is wiring node_exporter or Alloy up properly and understanding what it can and cannot tell you on its own.&lt;br&gt;
_&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Metrics Tell You Something Broke. Tracing Tells You What, Where, and Why.</title>
      <dc:creator>Justyn Larry</dc:creator>
      <pubDate>Thu, 04 Jun 2026 15:00:00 +0000</pubDate>
      <link>https://dev.to/irinobservability/metrics-tell-you-something-broke-tracing-tells-you-what-where-and-why-3j6b</link>
      <guid>https://dev.to/irinobservability/metrics-tell-you-something-broke-tracing-tells-you-what-where-and-why-3j6b</guid>
      <description>&lt;p&gt;Complacency is a killer. The monitoring stack that I built works, and it’s reliable, so leaving it alone seems like the most obvious thing to do. Focusing on marketing, documentation, taking time away from it all seem like good options, but there’s always a better way to do something, to solve a problem you didn’t realize you had.&lt;/p&gt;

&lt;p&gt;In my spare time, I look through Reddit and Dev.to for ideas or inspiration. Systems that others are using that I’m not, or that I’m not aware of. Distributed traces jumped out at me from both forums — I can tie a system event to the metrics, instead of stumbling around logs? This is a monitoring goldmine. How had I missed this?&lt;/p&gt;

&lt;h2&gt;
  
  
  WHAT EXACTLY IS DISTRIBUTED TRACING?
&lt;/h2&gt;

&lt;p&gt;For any kind of multi-step processes running on your system, distributed tracing provides a timeline of exactly what happened, and how long each step took. It’s like getting a receipt for the work showing you where time and resources were spent. Each request or job gets a trace ID, and every step records a span — a named block with a start time, end time, and any attributes you want to attach. Those spans assemble into a waterfall, and you can see at a glance where time was spent, what succeeded, and what failed.&lt;/p&gt;

&lt;p&gt;This added visibility can take a technical team from “this seems slow” to a detailed accounting of how long a process took and what the system was actually doing when the process was lagging.&lt;/p&gt;

&lt;h2&gt;
  
  
  THE ORIGINAL CORE STACK
&lt;/h2&gt;

&lt;p&gt;Irin Observability runs on Prometheus, Grafana, Loki, Grafana Alloy, and Alertmanager. I’ve built a robust monitoring stack that tracks metrics for request rates, error rates, LLM call counts, and report generation status. There are also logs flowing from all the services through Loki, so overall, I believed that the stack was well-instrumented and very readable.&lt;/p&gt;

&lt;p&gt;The alert system that I built runs through five internal services to process each alert through an alert annotator and to generate a monthly report in sequence:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An alert comes in from a client’s infrastructure&lt;/li&gt;
&lt;li&gt;The alert annotator calls a local LLM to add a plain-English explanation for a panel on one of the dashboards&lt;/li&gt;
&lt;li&gt;The annotated result gets pushed back into Loki&lt;/li&gt;
&lt;li&gt;At the end of the month, the aggregation script gathers all findings for report generation&lt;/li&gt;
&lt;li&gt;The LLM narrative layer writes a summary&lt;/li&gt;
&lt;li&gt;The report generator assembles everything into a PDF and sends it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of those steps runs in a different process. Some run as Docker containers, some as host Python scripts. When auditing the reports and something didn’t look right, I had to check the logs on the Loki Log Exporter Dashboard or grep logs across multiple services, correlate timestamps manually, and piece together what happened. This was both frustrating and time-consuming. The platform should be telling me what the problem is in addition to telling me that something is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  THE SOLUTION: OPENTELEMETRY
&lt;/h2&gt;

&lt;p&gt;OpenTelemetry (OTel) is an open source standard for collecting telemetry data — traces, metrics, and logs — from applications. It’s vendor-neutral, well-maintained, and has solid Python libraries.&lt;/p&gt;

&lt;p&gt;Grafana Tempo is an open source backend for storing and querying traces. It integrates directly with Grafana, so once it’s running you can navigate from a log line to a trace, or from a trace to the logs that were happening at the same time.&lt;/p&gt;

&lt;p&gt;Getting this running involved three parts. First, I deployed Tempo as a Docker Compose service, with a config file and a Grafana datasource. The second step was to wire up Grafana Alloy as the collector. Since Alloy is the agent already running on my servers to ship metrics and logs, I was able to add an OTLP receiver block to accept traces from internal services and forward them to Tempo — one config change, and the heartbeat API distributed the updated config files to all the monitored servers. The final step was to instrument the Python services. This is where things got a little more difficult, but it also taught me some valuable lessons.&lt;/p&gt;

&lt;h2&gt;
  
  
  THE PYTHON IMPLEMENTATION
&lt;/h2&gt;

&lt;p&gt;The OTel Python SDK has two modes. The first is auto-instrumentation, which handles the common cases automatically. If you’re running a Flask or FastAPI app, importing two libraries and calling .instrument() captures every HTTP request with no further changes. If you’re using psycopg2 for Postgres queries, one more library call and every query becomes a span.&lt;/p&gt;

&lt;p&gt;The second, manual spans, are for the logic your code owns — units of work that typical instrumentation frameworks can’t see automatically. I used these to capture the LLM call itself (duration, prompt size, whether the response parsed cleanly), each section of the aggregation script so I can see which Prometheus query is slow, and the overall per-tenant run so every trace carries a tenant name.&lt;/p&gt;

&lt;h2&gt;
  
  
  LESSONS LEARNED
&lt;/h2&gt;

&lt;p&gt;Short-lived scripts need an explicit flush.&lt;/p&gt;

&lt;p&gt;The aggregation script and report generator run once and exit. The default OTel exporter batches spans and sends them on a timer. If the process exits before the batch fires, you lose all your spans. I fixed it by adding two lines: force_flush() and shutdown() in a try/finally block before exit. I lost my first few test traces before I figured this out.&lt;br&gt;
The psycopg2-binary package breaks auto-instrumentation silently.&lt;/p&gt;

&lt;p&gt;The OTel instrumentation library checks for a package literally named psycopg2. If you installed psycopg2-binary — the same library, different distribution name — the check fails and you receive no database spans, no error message, nothing reported. The fix is one parameter: Psycopg2Instrumentor().instrument(skip_dep_check=True).&lt;br&gt;
Background tasks break parent-child trace linkage.&lt;/p&gt;

&lt;p&gt;My alert annotator returns a 200 response immediately and processes the alert in a background thread. The HTTP span closes when the response is sent, but before the real work begins, which means each alert generates two separate traces — a brief HTTP span and an orphaned processing span. The model behavior was correct, not a bug, but it looked confusing until I understood the threading model. I accepted it and correlate the two traces by alert fingerprint when necessary.&lt;/p&gt;
&lt;h2&gt;
  
  
  THE BIG DIFFERENCE
&lt;/h2&gt;

&lt;p&gt;This is where things get interesting, and how the original monitoring stack differs from its current iteration.&lt;/p&gt;

&lt;p&gt;Prior to integrating distributed tracing, I knew that the report pipeline ran. That’s it — pass/fail, true/false. If something went wrong, where did it happen, and why? What was the system state at the time of the failure? Now I can open a trace in Grafana Tempo and see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;report.generate: total duration 4m 12s
  db.get_contacts: 41ms
  aggregation.run (per tenant): 2m 18s
    aggregation.stability: 39ms
    aggregation.resources: 1.2s  (slow Prometheus query range)
    aggregation.alerts: 88ms
  llm.narrative_generation: 1m 44s
    llm.build_prompt: 12ms
    llm.call attempt 1: 119s  (timeout)
    llm.call attempt 2: 44s   (success)
    llm.parse: 3ms
  report.build_pdf: 8s
  report.send_email: 2s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That waterfall tells me that the Ollama model timed out on the first attempt and succeeded on the second. I don’t have to go digging through logs in an approximate time frame to figure out what happened. The Prometheus query for resource metrics was the slow step in aggregation. PDF build and email delivery were fast. The problem isn’t solved, but I know exactly what the problem is.&lt;/p&gt;

&lt;p&gt;Through the alert annotator, I can now see every alert as a trace. The system shows me the dedup check against Loki, the LLM call, the result push. I can filter by tenant, by alert name, by whether the LLM call succeeded. A 55-second LLM call that I used to see only as a latency spike in a Prometheus histogram is now a named span with the prompt size, the response size, and whether the JSON parsed cleanly.&lt;/p&gt;

&lt;h2&gt;
  
  
  THE IMPLICATIONS
&lt;/h2&gt;

&lt;p&gt;If you have any experience with monitoring, you have almost certainly hit the “something seems wrong but I can’t tell what” problem. The logs are probably available, you can see the metrics, but you’re stuck sifting through them in sequence trying to reconstruct what happened.&lt;/p&gt;

&lt;p&gt;Distributed tracing changes the diagnostic workflow from “search for clues” to “read the receipt.” The trace tells you what happened, in order, with timing, which virtually eliminates investigation time and lets you go directly to the problem at hand.&lt;/p&gt;

&lt;p&gt;It also changes how you think about reliability. When I see the LLM call timing out on first attempt consistently, I know to tune the timeout or check model load before it impacts the client. Being proactive in monitoring is a moving target, but it is still the goal.&lt;/p&gt;

&lt;h2&gt;
  
  
  THE TOOLCHAIN
&lt;/h2&gt;

&lt;p&gt;Everything I used is open source and self-hostable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenTelemetry Python SDK (opentelemetry-sdk, exporter packages, auto-instrumentation libraries)&lt;/li&gt;
&lt;li&gt;Grafana Tempo for trace storage and querying&lt;/li&gt;
&lt;li&gt;Grafana Alloy as the collector and forwarder&lt;/li&gt;
&lt;li&gt;Grafana for visualization, with native Tempo datasource support and log/trace correlation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you’re already running Prometheus and Grafana for metrics, adding Tempo for traces is a natural extension of the same stack. You can use the same agent, dashboards, and query interface. You’re adding one more signal type, but no new tooling paradigm.&lt;/p&gt;

&lt;p&gt;The monitoring stack I run for Irin clients is the same stack I use to observe both Irin and my private infrastructure. It’s what lets me catch instrumentation gotchas and gives me a reliable view of all of my systems. I built Irin because I believe that monitoring your system shouldn’t be a full-time job. If the monitoring stack does what it’s supposed to, you should be able to check it intermittently through the day. It should tell you at a glance if something’s wrong, and send an alert if the problem merits it. If it’s noisy, crowded, and you don’t know where to begin when there’s a problem, the system doesn’t work — and the real problems get drowned out.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>distributedsystems</category>
      <category>monitoring</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
