<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Slim</title>
    <description>The latest articles on DEV Community by Slim (@slima4).</description>
    <link>https://dev.to/slima4</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1192518%2F030c5a4c-0dfe-4f7a-a9e2-e06b76c62117.JPG</url>
      <title>DEV Community: Slim</title>
      <link>https://dev.to/slima4</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/slima4"/>
    <language>en</language>
    <item>
      <title>Uncle Bob's uml-viewer on Rust: 35 dependency cycles down to 3</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Tue, 22 Sep 2026 06:58:19 +0000</pubDate>
      <link>https://dev.to/slima4/uncle-bobs-uml-viewer-on-rust-35-dependency-cycles-down-to-3-30h1</link>
      <guid>https://dev.to/slima4/uncle-bobs-uml-viewer-on-rust-35-dependency-cycles-down-to-3-30h1</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt; Uncle Bob released &lt;a href="https://github.com/unclebob/uml-viewer" rel="noopener noreferrer"&gt;uml-viewer&lt;/a&gt;, a clickable architecture diagram that draws Dependency Rule violations in red and is meant to be driven together with a coding agent. It only parses Clojure, so I wrote a 200-line script that turns my Rust crate into its input file. The first picture showed 167 red arrows and 35 dependency cycles. Six commits and 315 files later: 109 red arrows, 3 cycles, and &lt;code&gt;app.rs&lt;/code&gt; down from 1,087 lines to 642. No behavior changed. The tool matters less than the loop it forces on you: look, point at one arrow, let the agent move code, look again. The &lt;a href="https://uptimepage.dev/blog/uml-viewer-rust-dependency-cycles?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=uml-viewer&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;original post&lt;/a&gt; has the same text with a collapsible FAQ and the source list at the end.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Uncle Bob's argument for the tool is short: agents still need supervision, and reading every line they write is the bottleneck. So supervise the structure instead of the text: draw the system, find the shape that is wrong, tell the agent to fix that shape, and check the new drawing.&lt;/p&gt;

&lt;p&gt;I wanted to know if that works on a real codebase, so I pointed it at &lt;a href="https://github.com/uptimepage/uptimepage" rel="noopener noreferrer"&gt;Uptimepage&lt;/a&gt;, about 146,000 lines of Rust in one crate. It does, with one condition: the picture is only as honest as the input file you feed it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F02687sxtvl3bgva5wowl.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F02687sxtvl3bgva5wowl.webp" alt="Two panels of five stacked layers, vocabulary at the bottom and assembly at the top. Before: five red arrows point up from storage and config into api, worker, auth and billing, and two grey loops mark cycles between api and web and between storage and quotas. After: the same modules plus dashed green boxes for request, templates, security and pagination, and every arrow points down or sideways." width="800" height="507"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Five of the 167 red arrows and two of the 35 cycles, and where the code went. Every arrow that pointed up now points down.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What uml-viewer is
&lt;/h2&gt;

&lt;p&gt;uml-viewer is an open-source desktop tool by Uncle Bob (Robert Martin) that draws a codebase as a UML-like diagram you can click. It is written in Clojure and needs the Clojure CLI and Java 21 or newer.&lt;/p&gt;

&lt;p&gt;Namespaces are components. The modules inside a namespace are the component's elements. Nesting can go as deep as your source tree does. You double-click a component to open the next level, double-click a module to see its class card, and click a function to open the source file at that line.&lt;/p&gt;

&lt;p&gt;Color comes from CRAP and mutation scores, so a red box is code with high complexity and weak tests, or code with no metrics at all, which the README counts as the worst grade. Red arrows come from the Dependency Rule: you tell the tool which namespaces sit at which architectural level, and every dependency that points from an inner level to an outer level is drawn red.&lt;/p&gt;

&lt;p&gt;The tool is built to run next to an agent. By default it opens a tmux window with Grok in the examined project and the two talk through a small mailbox directory. You can also drive it by hand from any agent session: regenerate the input file, press R in the viewer, and the diagram reloads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting a Rust crate into it
&lt;/h2&gt;

&lt;p&gt;The only parser it ships reads Clojure. That sounded like the end of the experiment, until I read what the viewer actually consumes. It never sees source code. It reads one EDN file: a list of classes with a namespace and a level, a list of edges with a from, a to and a kind, and a list of levels. That is a format any script can emit.&lt;/p&gt;

&lt;p&gt;So I asked the agent to write a generator for Rust. It came out at about 200 lines of Python and does four things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every &lt;code&gt;.rs&lt;/code&gt; file is one class. &lt;code&gt;mod.rs&lt;/code&gt; stands for its directory and &lt;code&gt;lib.rs&lt;/code&gt; is skipped.&lt;/li&gt;
&lt;li&gt;Every &lt;code&gt;use crate::…&lt;/code&gt;, &lt;code&gt;super::…&lt;/code&gt; and &lt;code&gt;self::…&lt;/code&gt; path, plus every inline &lt;code&gt;crate::a::b&lt;/code&gt; path in a function body, becomes a dependency edge to the nearest enclosing module that exists as a file.&lt;/li&gt;
&lt;li&gt;Comments are stripped first, so a path mentioned in a doc comment does not count as a dependency.&lt;/li&gt;
&lt;li&gt;A hand-written &lt;code&gt;LEVELS&lt;/code&gt; list gives each top-level module a rank.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last item is the only opinion in the script:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# inner (high level) first, matching the Dependency Rule ranks
&lt;/span&gt;&lt;span class="n"&gt;LEVELS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;domain&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;metric_names&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;storage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;security&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;net&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http_client&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;config&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
     &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;quotas&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;observability&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pagination&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pipeline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;escalation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;notifier&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;telegram&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
     &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;whatsapp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;auth&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;analytics&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
     &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;jobs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http_outbound&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;targets&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;api&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;web&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mcp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;marketing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;worker&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scheduler&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
     &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;request&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;public_status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;templates&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;oauth&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
     &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ad_hoc_dispatch&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;channels&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;app&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;router&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bootstrap&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;main&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bin&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rank 0 is vocabulary that everyone may use: domain types, error codes, text helpers. Rank 1 is infrastructure: storage, security primitives, config. Rank 2 is the services that do the work: probing, escalation, notification, billing. Rank 3 is every way into the system: HTTP handlers, HTML views, the MCP server, the marketing site. Rank 4 is assembly: the app state, the router, main.&lt;/p&gt;

&lt;p&gt;The rule is then mechanical. An edge is a violation when the module it comes from has a smaller rank than the module it points to. Same rank is allowed. Foreign crates like axum and sqlx are ovals outside the diagram and are never compared.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The picture is only as honest as the levels list&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The generator does not decide the architecture. I do, in that list. If you put &lt;code&gt;api&lt;/code&gt; in the same group as &lt;code&gt;domain&lt;/code&gt;, the diagram turns green and you have learned nothing. Write the list you want to be true, then let the red arrows show how far the code is from it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the first picture showed
&lt;/h2&gt;

&lt;p&gt;The first diagram had 167 red arrows, 35 pairs of modules that imported each other, 17 modules importing &lt;code&gt;api&lt;/code&gt;, and an &lt;code&gt;app.rs&lt;/code&gt; of 1,087 lines.&lt;/p&gt;

&lt;p&gt;Behind the numbers were shapes I half knew about and had never seen drawn:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;storage&lt;/code&gt; reached up into &lt;code&gt;api&lt;/code&gt; for error codes and dashboard read models, and into &lt;code&gt;worker&lt;/code&gt; for heartbeat state.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;api&lt;/code&gt; and &lt;code&gt;web&lt;/code&gt; imported each other. The JSON side took session and token extractors, cookies and client IP from the HTML side, and the HTML side took the heartbeat read model from the JSON handlers.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;quotas&lt;/code&gt; read accounts and organizations from &lt;code&gt;storage&lt;/code&gt;, while &lt;code&gt;storage&lt;/code&gt; embedded quota SQL fragments and read plan types from &lt;code&gt;quotas&lt;/code&gt;. A cycle in both directions.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;marketing&lt;/code&gt; and &lt;code&gt;oauth&lt;/code&gt; imported &lt;code&gt;web&lt;/code&gt; for three template filters.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;config&lt;/code&gt; imported &lt;code&gt;auth&lt;/code&gt; and &lt;code&gt;billing&lt;/code&gt; for two enums.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this was visible in the review of any single pull request. Every one of those imports was reasonable on the day it was written. A diff shows one import at a time, so the sum of them never appeared in any review.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop
&lt;/h2&gt;

&lt;p&gt;Every round had the same six steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Regenerate the EDN file and reload the viewer.&lt;/li&gt;
&lt;li&gt;Pick one red arrow. Hover it to see which module pairs it bundles.&lt;/li&gt;
&lt;li&gt;Tell the agent one thing: what must not import what, and where the shared piece should live.&lt;/li&gt;
&lt;li&gt;The agent moves the code and runs the compiler and the tests.&lt;/li&gt;
&lt;li&gt;Regenerate. Check that the arrow is gone and count the new ones.&lt;/li&gt;
&lt;li&gt;Run the full suite, review the diff, commit.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The instruction in step 3 is the part that makes this work. It looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;storage must not import api. The error codes and the dashboard read models
it takes from there are crate-wide vocabulary. Move the codes to error::codes
and the read models to domain::metrics, and point every reader at the new
path. No behavior changes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is one rule, one violation, and one place to put the result. The agent does not have to guess what "cleaner" means. It has a named arrow to remove, and the next diagram says whether it did.&lt;/p&gt;

&lt;p&gt;I ran this with Claude Code, but nothing in the loop depends on it. Codex, Cursor, Grok, or any agent that can edit files and run a test suite gets the same instruction and the same picture afterwards.&lt;/p&gt;

&lt;p&gt;Six rounds took the crate from 167 red arrows to 109:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Round&lt;/th&gt;
&lt;th&gt;Red arrows I pointed at&lt;/th&gt;
&lt;th&gt;Where the code went&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;storage&lt;/code&gt; into &lt;code&gt;api&lt;/code&gt; and &lt;code&gt;worker&lt;/code&gt;, &lt;code&gt;config&lt;/code&gt; into &lt;code&gt;auth&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;error codes to &lt;code&gt;error::codes&lt;/code&gt;, 12 read models to &lt;code&gt;domain::metrics&lt;/code&gt;, provider enums to &lt;code&gt;domain::credential&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;storage&lt;/code&gt; into &lt;code&gt;auth&lt;/code&gt; and &lt;code&gt;api&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;token hashing, HMAC and SHA-256 helpers to &lt;code&gt;security&lt;/code&gt;, redaction scrubbers to &lt;code&gt;security::redaction&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;api&lt;/code&gt; and &lt;code&gt;web&lt;/code&gt; into each other&lt;/td&gt;
&lt;td&gt;extractors, cookies, client IP and host resolution to a new &lt;code&gt;request&lt;/code&gt; module that imports neither&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;storage&lt;/code&gt; and &lt;code&gt;quotas&lt;/code&gt; into each other&lt;/td&gt;
&lt;td&gt;SQL fragments to &lt;code&gt;storage::count_sql&lt;/code&gt;, plan types to &lt;code&gt;domain::quota&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;marketing&lt;/code&gt;, &lt;code&gt;oauth&lt;/code&gt; and &lt;code&gt;api&lt;/code&gt; into &lt;code&gt;web&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;template filters and formatters to &lt;code&gt;templates&lt;/code&gt;, &lt;code&gt;app.rs&lt;/code&gt; split into &lt;code&gt;config::boot&lt;/code&gt;, &lt;code&gt;observability::readiness&lt;/code&gt;, &lt;code&gt;targets::status&lt;/code&gt; and &lt;code&gt;net&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;public_status&lt;/code&gt; and &lt;code&gt;web&lt;/code&gt; into &lt;code&gt;api&lt;/code&gt;, &lt;code&gt;security&lt;/code&gt; into &lt;code&gt;http_client&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;page envelopes to &lt;code&gt;pagination&lt;/code&gt;, &lt;code&gt;Cipher::from_config&lt;/code&gt; into &lt;code&gt;security&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And the totals:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F01eoys6zyo3qxvwsb13m.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F01eoys6zyo3qxvwsb13m.webp" alt="Four paired horizontal bars, red for before and green for after: Dependency Rule violations 167 to 109, module pairs importing each other 35 to 3, modules importing api 17 to 4, lines in app.rs 1,087 to 642." width="799" height="413"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Six commits, 315 files, same test suite before and after. Violations down 35 percent, cycles down 91 percent.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;After&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Dependency Rule violations&lt;/td&gt;
&lt;td&gt;167&lt;/td&gt;
&lt;td&gt;109&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Module pairs importing each other&lt;/td&gt;
&lt;td&gt;35&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Modules importing &lt;code&gt;api&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lines in &lt;code&gt;app.rs&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;1,087&lt;/td&gt;
&lt;td&gt;642&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Files touched&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;315&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Behavior changes&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I regenerated the graph at every one of the six commits to see what each round removed:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fphmfr98vocwisvvr93cw.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fphmfr98vocwisvvr93cw.webp" alt="A line chart over seven points from start to round six. Module pairs importing each other fall 35, 28, 23, 19, 16, 8, 3. Modules importing api fall 17, 11, 9, 6, 6, 6, 4. Under each round a short label names what moved." width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Cycles and api importers per round. The first two rounds did most of the work on the red count; rounds three to six were about the cycles.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The shape surprised me. The red arrow count stopped moving after round two. Rounds three to six barely touched it, but they took the cycles from 23 down to 3, because a cycle between two modules at the same level is not a Dependency Rule violation at all. The rule catches arrows that point up and says nothing about two modules on the same level that import each other, so you need both counts.&lt;/p&gt;

&lt;p&gt;Not every move survived. In a later round the agent moved the health-check paths into &lt;code&gt;observability&lt;/code&gt;, and a coupling test on the marketing module said no, because that module is only allowed to reach a short list of leaf modules. Another move put the strict JSON parser under &lt;code&gt;request&lt;/code&gt;, then had to come back because that parser reads the OpenAPI document. Both reverts took minutes, because the diagram and the test said so before the commit did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I stopped at 91
&lt;/h2&gt;

&lt;p&gt;After the six commits, 109 red arrows were left. Most pointed into &lt;code&gt;app&lt;/code&gt;, but ten did not: a sampler and a silence job that lived under &lt;code&gt;observability&lt;/code&gt; but reached into the scheduler and the notifier, two background jobs reading &lt;code&gt;public_status&lt;/code&gt;, the HTTP metrics layer, and the rate-limit middleware in &lt;code&gt;quotas&lt;/code&gt; reading the app state. That middleware was also the third of the three cycles. One more round moved each of those into the module that owns it and left 91 red arrows and 2 cycles.&lt;/p&gt;

&lt;p&gt;Every one of the 91 is the same shape. Handlers read the app state, and the app state is built from the modules those handlers live in. &lt;code&gt;app&lt;/code&gt; imports &lt;code&gt;api&lt;/code&gt; and &lt;code&gt;web&lt;/code&gt; to mount them; &lt;code&gt;api&lt;/code&gt; and &lt;code&gt;web&lt;/code&gt; import &lt;code&gt;app&lt;/code&gt; to get the state.&lt;/p&gt;

&lt;p&gt;Fixing it is possible. Split the state into per-handler sub-states and hand each handler only its slice. I counted what that would touch: about 90 files of plumbing that make the code harder to read, not easier. I stopped because I could name every remaining arrow, and at that point the diagram had told me everything it was going to.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Stop when every arrow has a name&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Drive the count down until every arrow that is left has a name and a reason, then stop. The 91 arrows left in my crate are all "this handler reads the app state", and a diagram that shows them is more honest than one that hides them behind 90 files of indirection.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Keeping it fixed
&lt;/h2&gt;

&lt;p&gt;Two things stop the graph from drifting back.&lt;/p&gt;

&lt;p&gt;The marketing module has a test that is an allow-list: the exact set of leaf modules it may import. A new reach into the app fails the build. It used to be a deny-list, which only catches the mistakes you already thought of.&lt;/p&gt;

&lt;p&gt;The other is a habit, not a test. Regressions arrive with feature commits, not with refactors. After every push I regenerate the graph and read the list of red arrows that do not point into &lt;code&gt;app&lt;/code&gt;. If the list is empty, the feature stayed in its layer. If it is not, the offending edge is one instruction away from gone, before the next feature builds on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this beats "refactor the codebase"
&lt;/h2&gt;

&lt;p&gt;Two months ago I wrote about &lt;a href="https://uptimepage.dev/blog/map-your-codebase-for-ai-agents?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=uml-viewer&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;mapping this codebase for humans and AI agents&lt;/a&gt;. The lesson then was that a model is good at shape and bad at numbers, so I had to check every count it produced by hand.&lt;/p&gt;

&lt;p&gt;This is the same lesson, applied to refactoring. Here the agent never counts anything; the generator does. The agent gets one rule with one violation and a picture that says pass or fail after every change. It is the same reason a failing test is a better instruction than a paragraph of requirements: the acceptance criterion exists before the work starts, and it is not the agent that judges it.&lt;/p&gt;

&lt;p&gt;What I did not get is the color. CRAP and mutation scores come from Uncle Bob's Clojure tooling, and there is no Rust equivalent wired in yet. The README is clear that a box with no metrics is painted as the worst grade, so the fill on my boxes means nothing until coverage and mutation numbers exist for Rust. The arrows alone were worth the setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do it yourself
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Install the Clojure CLI and Java 21 or newer, then clone &lt;a href="https://github.com/unclebob/uml-viewer" rel="noopener noreferrer"&gt;uml-viewer&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Write a generator for your language. One class per module, one dependency edge per import, output as EDN with &lt;code&gt;:hierarchical true&lt;/code&gt;. The generated &lt;code&gt;examples/uml-viewer.edn&lt;/code&gt; in the repo is a complete example of the format. In Rust, imports are &lt;code&gt;use&lt;/code&gt; paths; in TypeScript they are &lt;code&gt;import&lt;/code&gt; statements; in Python, &lt;code&gt;import&lt;/code&gt; and &lt;code&gt;from&lt;/code&gt;. A regex over each file is enough to start.&lt;/li&gt;
&lt;li&gt;Write the &lt;code&gt;:levels&lt;/code&gt; list by hand, inner layer first. Describe the architecture you want; the red arrows will show where the code differs.&lt;/li&gt;
&lt;li&gt;Run the viewer against your file. On a fresh start it waits for its companion agent; the &lt;code&gt;--restart&lt;/code&gt; flag restores the last view without spawning one, which is what you want when you drive it from your own agent session.&lt;/li&gt;
&lt;li&gt;Pick one red arrow. Write one instruction that names the two modules and the new home. Regenerate. Repeat.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaways&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;uml-viewer reads a plain EDN file, not source code. A 200-line script gets any language in.&lt;/li&gt;
&lt;li&gt;The levels list is the only opinion in the input, so write the architecture you want and let the red show the distance.&lt;/li&gt;
&lt;li&gt;Give the agent one rule, one violation and one destination per instruction. The next diagram is the acceptance test.&lt;/li&gt;
&lt;li&gt;Rust will not tell you two modules import each other. Build the graph and count the mutual pairs yourself.&lt;/li&gt;
&lt;li&gt;Stop when every remaining red arrow has a name. Mine are all "handler reads app state", and that is fine.&lt;/li&gt;
&lt;li&gt;Turn the result into an allow-list test, and regenerate the graph after every feature push, because that is when regressions arrive.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;The six commits are public in the &lt;a href="https://github.com/uptimepage/uptimepage/compare/b791c262...cee12c44" rel="noopener noreferrer"&gt;Uptimepage repository&lt;/a&gt; if you want to read what moved and why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Robert C. Martin, &lt;a href="https://github.com/unclebob/uml-viewer" rel="noopener noreferrer"&gt;uml-viewer&lt;/a&gt; on GitHub. The README documents the EDN format, the &lt;code&gt;:levels&lt;/code&gt; rule and the companion agent mailbox.&lt;/li&gt;
&lt;li&gt;Robert C. Martin, &lt;a href="https://blog.cleancoder.com/uncle-bob/2012/08/13/the-clean-architecture.html" rel="noopener noreferrer"&gt;The Clean Architecture&lt;/a&gt;, The Clean Code Blog, August 2012. The Dependency Rule.&lt;/li&gt;
&lt;li&gt;Uptimepage, &lt;a href="https://github.com/uptimepage/uptimepage" rel="noopener noreferrer"&gt;source repository&lt;/a&gt;, AGPL. Commits &lt;code&gt;bda59342&lt;/code&gt; through &lt;code&gt;cee12c44&lt;/code&gt; are the six rounds described here.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I build &lt;a href="https://uptimepage.dev/?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=uml-viewer&amp;amp;utm_content=cta" rel="noopener noreferrer"&gt;Uptimepage&lt;/a&gt; because I wanted uptime checks that say why something failed instead of "transport error": HTTP with per-phase timing, TLS and domain expiry, ping and TCP, heartbeats for background jobs, and browser flows, from several regions. It is AGPL-3.0 open source, which is why the six commits above are public and you can check every number in this post against them.&lt;/p&gt;

&lt;p&gt;Now the question for the comments: what draws the module graph for your language, and does it show the cycles or only the layer violations? My rounds three to six only moved the cycle count, and I would have missed that with a tool that colors arrows but never counts pairs.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rust</category>
      <category>architecture</category>
      <category>programming</category>
    </item>
    <item>
      <title>ICMP vs TCP vs UDP: the difference, explained for developers</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Sun, 13 Sep 2026 09:49:11 +0000</pubDate>
      <link>https://dev.to/slima4/icmp-vs-tcp-vs-udp-the-difference-explained-for-developers-49g6</link>
      <guid>https://dev.to/slima4/icmp-vs-tcp-vs-udp-the-difference-explained-for-developers-49g6</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://uptimepage.dev/blog/icmp-vs-tcp-vs-udp" rel="noopener noreferrer"&gt;Uptimepage blog&lt;/a&gt;. I build an uptime monitor in Rust, and its ping, TCP and DNS checks are these three protocols with a timer on them, so I had to get the differences straight.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt; ICMP, TCP and UDP all travel inside IP packets, but they ask different questions. ICMP echo (ping) asks "does this address answer at all?" and needs no port. TCP asks "is a program listening on this port?" and gets a clear yes, a clear no, or silence. UDP asks nothing by itself: you only learn something if the program on the other side chooses to reply. A closed port answers differently on each one, and that difference is most of what a network check can and cannot tell you. The &lt;a href="https://uptimepage.dev/blog/icmp-vs-tcp-vs-udp?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=icmp-tcp-udp&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;original post&lt;/a&gt; has the same text with a collapsible FAQ and the RFC list at the end.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  One address, three ways to knock
&lt;/h2&gt;

&lt;p&gt;Every packet on the internet is an IP packet. IP knows how to move a packet from one address to another and nothing else. One byte in the IP header says what is inside, and &lt;a href="https://www.iana.org/assignments/protocol-numbers/protocol-numbers.xhtml" rel="noopener noreferrer"&gt;IANA keeps the list&lt;/a&gt;: ICMP is protocol 1, TCP is protocol 6, UDP is protocol 17. The three are siblings: same addresses, same routers, and a different promise about what happens once the packet arrives.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnyoeeaad36lhumosejz3.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnyoeeaad36lhumosejz3.webp" alt="ICMP echo, TCP three-way handshake and UDP datagram drawn as three lanes between you and a host, each ending with what silence means on that protocol." width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Three ways to knock on one address. The last row of each lane is what silence means there.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;ICMP&lt;/th&gt;
&lt;th&gt;TCP&lt;/th&gt;
&lt;th&gt;UDP&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Specification&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.rfc-editor.org/rfc/rfc792" rel="noopener noreferrer"&gt;RFC 792&lt;/a&gt; (1981), &lt;a href="https://www.rfc-editor.org/rfc/rfc4443" rel="noopener noreferrer"&gt;RFC 4443&lt;/a&gt; for IPv6&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.rfc-editor.org/rfc/rfc9293" rel="noopener noreferrer"&gt;RFC 9293&lt;/a&gt; (2022)&lt;/td&gt;
&lt;td&gt;
&lt;a href="https://www.rfc-editor.org/rfc/rfc768" rel="noopener noreferrer"&gt;RFC 768&lt;/a&gt; (1980)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IP protocol number&lt;/td&gt;
&lt;td&gt;1 (58 for ICMPv6)&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ports&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;source and destination&lt;/td&gt;
&lt;td&gt;source and destination&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Handshake&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;three packets before any data&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delivery guarantee&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;td&gt;ordered, complete, retransmitted&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Header&lt;/td&gt;
&lt;td&gt;8 bytes for an echo&lt;/td&gt;
&lt;td&gt;20 bytes minimum&lt;/td&gt;
&lt;td&gt;8 bytes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical users&lt;/td&gt;
&lt;td&gt;ping, traceroute, error reports&lt;/td&gt;
&lt;td&gt;HTTP/1.1 and HTTP/2, TLS, SSH, SQL, SMTP&lt;/td&gt;
&lt;td&gt;DNS, NTP, QUIC and HTTP/3, video, games&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  ICMP: the network talking about itself
&lt;/h2&gt;

&lt;p&gt;ICMP is not a transport. RFC 792 says it plainly: ICMP "is actually an integral part of IP, and must be implemented by every IP module." It exists so that routers and hosts can report on delivery: this destination is unreachable, that packet ran out of hops. ICMP messages have a type and a code but no ports, because nothing sits on top of ICMP waiting for them. The IP stack consumes them itself.&lt;/p&gt;

&lt;p&gt;The one ICMP message everybody has typed is the echo. &lt;code&gt;ping&lt;/code&gt; sends an echo request (type 8, or 128 in ICMPv6) and the far host's kernel sends back an echo reply (type 0, or 129). No program had to be running. &lt;a href="https://www.rfc-editor.org/rfc/rfc1122" rel="noopener noreferrer"&gt;RFC 1122&lt;/a&gt;, the host requirements standard, makes it a rule: "Every host MUST implement an ICMP Echo server function that receives Echo Requests and sends corresponding Echo Replies."&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ ping -c 1 uptimepage.dev
PING uptimepage.dev (204.168.246.94): 56 data bytes
64 bytes from 204.168.246.94: icmp_seq=0 ttl=50 time=96.069 ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A reply proves two things: the address belongs to a machine that is switched on, and packets can travel there and back. It also gives you the round-trip time for free.&lt;/p&gt;

&lt;p&gt;That is all a ping proves. It says nothing about whether a web server, a database or anything else is running, and a host can answer every ping while every service on it is dead.&lt;/p&gt;

&lt;p&gt;Two things about ICMP surprise people. The first is that no reply does not mean down. Firewalls drop ICMP all the time. On AWS a fresh security group blocks it, and the &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security-group-rules-reference.html" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; says so: "To ping your instance, you must add one of the following inbound ICMP rules." Before you rely on ping for a host, confirm the host answers a ping at all.&lt;/p&gt;

&lt;p&gt;The second is that sending ICMP needs privilege. A raw socket needs &lt;code&gt;CAP_NET_RAW&lt;/code&gt;. Linux has an unprivileged alternative, the ICMP datagram socket, but the &lt;a href="https://docs.kernel.org/networking/ip-sysctl.html" rel="noopener noreferrer"&gt;kernel documentation&lt;/a&gt; says the default &lt;code&gt;ping_group_range&lt;/code&gt; is "1 0", "meaning, that nobody (not even root) may create ping sockets." Docker opens that range in every container that gets its own network namespace, which is why ping works there without anyone thinking about it; a container on &lt;code&gt;--network host&lt;/code&gt; inherits the host's setting instead. I learned the limits of that the slow way: the same probe binary that pinged fine in Docker could not open an ICMP socket at all on a Firecracker microVM, until I gave it &lt;code&gt;CAP_NET_RAW&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  TCP: a promise before the first byte
&lt;/h2&gt;

&lt;p&gt;TCP is a connection. Before either side sends a byte of data, the two ends run the three-way handshake: SYN, SYN-ACK, ACK. RFC 9293, which replaced the 1981 specification in 2022, calls it "the procedure used to establish a connection." It costs one round trip, and that round trip pays for every promise TCP makes afterwards: bytes arrive in order, lost segments are sent again, the receiver is not flooded (flow control), and the network is not flooded (congestion control). To the program it looks like a stream. You write bytes on one side and read the same bytes, in the same order, on the other.&lt;/p&gt;

&lt;p&gt;This is what HTTP/1.1, HTTP/2, TLS, SSH, PostgreSQL, MySQL, SMTP and most of what you deploy run on.&lt;/p&gt;

&lt;p&gt;The handshake is also a good probe, because a SYN gets one of three answers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;SYN-ACK.&lt;/strong&gt; A program is listening. The connect succeeds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RST.&lt;/strong&gt; Nothing is listening. RFC 9293 says that if the connection does not exist, "a reset is sent in response to any incoming segment except another reset." Your connect fails at once with "connection refused".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing.&lt;/strong&gt; A firewall dropped the SYN, or the host is gone. The kernel keeps retrying. On Linux the default &lt;code&gt;tcp_syn_retries&lt;/code&gt; is 6, and the kernel documentation puts the final timeout for an active connection attempt at 131 seconds.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here are all three from my laptop. The last one has a three second limit; without it macOS waits 75 seconds.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ nc -zv uptimepage.dev 443
Connection to uptimepage.dev port 443 [tcp/https] succeeded!

$ nc -zv 127.0.0.1 9
nc: connectx to 127.0.0.1 port 9 (tcp) failed: Connection refused

$ nc -zv -G 3 uptimepage.dev 8443
nc: connectx to uptimepage.dev port 8443 (tcp) failed: Operation timed out
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://nmap.org/book/man-port-scanning-basics.html" rel="noopener noreferrer"&gt;Nmap's documentation&lt;/a&gt; gives these states the names everyone uses: open, closed and filtered.&lt;/p&gt;

&lt;p&gt;What a successful handshake proves: a process accepted a connection on that port. Not that it is healthy. A database that accepts connections and then rejects every query still passes. That is why a TCP check belongs on things that have no HTTP surface (SSH, SMTP, a database port) and an HTTP check belongs on things that do.&lt;/p&gt;

&lt;h2&gt;
  
  
  UDP: a message, nothing more
&lt;/h2&gt;

&lt;p&gt;UDP is the smallest transport there is. RFC 768 fits on three pages. It adds a source port, a destination port, a length and a checksum to an IP packet, 8 bytes in total, and stops. The RFC says: "The protocol is transaction oriented, and delivery and duplicate protection are not guaranteed." There is no handshake, no acknowledgement, no ordering and no retransmission, and no congestion control either. &lt;a href="https://www.rfc-editor.org/rfc/rfc8085" rel="noopener noreferrer"&gt;RFC 8085&lt;/a&gt; opens by saying UDP "has no inherent congestion control mechanisms" and then spends more than fifty pages telling application designers how to add their own.&lt;/p&gt;

&lt;p&gt;Why anyone uses it: latency and control. A DNS lookup is one packet out and one packet back, no handshake. &lt;a href="https://www.rfc-editor.org/rfc/rfc1035" rel="noopener noreferrer"&gt;RFC 1035&lt;/a&gt; capped a DNS answer over UDP at 512 bytes. EDNS(0) (&lt;a href="https://www.rfc-editor.org/rfc/rfc6891" rel="noopener noreferrer"&gt;RFC 6891&lt;/a&gt;) later let a client advertise a bigger buffer, so most answers fit in one datagram today; one that still does not fit is cut short with the TC bit set, and the client asks again over TCP (&lt;a href="https://www.rfc-editor.org/rfc/rfc7766" rel="noopener noreferrer"&gt;RFC 7766&lt;/a&gt;). NTP, most game traffic and real-time video use UDP because a late packet is worth less than the next one, so waiting for a retransmission is the wrong trade. And QUIC, the transport under HTTP/3, is built entirely on UDP. &lt;a href="https://www.rfc-editor.org/rfc/rfc9000" rel="noopener noreferrer"&gt;RFC 9000&lt;/a&gt; states that "QUIC packets are carried in UDP datagrams" and gives the reason as "to better facilitate deployment in existing systems and networks". In practice that means QUIC does its own handshake, loss recovery and congestion control in the application, and can change them without waiting for every operating system to update its TCP stack.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ dig @1.1.1.1 uptimepage.dev A +noall +answer +stats
; &amp;lt;&amp;lt;&amp;gt;&amp;gt; DiG 9.10.6 &amp;lt;&amp;lt;&amp;gt;&amp;gt; @1.1.1.1 uptimepage.dev A +noall +answer +stats
; (1 server found)
;; global options: +cmd
uptimepage.dev.     3600    IN  A   204.168.246.94
;; Query time: 125 msec
;; SERVER: 1.1.1.1#53(1.1.1.1)
;; WHEN: Wed Sep 02 11:40:44 EEST 2026
;; MSG SIZE  rcvd: 59
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One datagram out, one datagram back carrying a 59-byte answer (67 bytes with the UDP header), and no handshake before it. Over TCP the same lookup costs two round trips, one for the handshake and one for the query, which is why DNS defaults to UDP.&lt;/p&gt;

&lt;p&gt;Now the probing problem. Send a UDP datagram to a port and what comes back?&lt;/p&gt;

&lt;p&gt;If nothing listens there, the host's IP stack should answer with an ICMP error. RFC 1122 again: "If a datagram arrives addressed to a UDP port for which there is no pending LISTEN call, UDP SHOULD send an ICMP Port Unreachable message." So "closed" is detectable, and it is ICMP that detects it.&lt;/p&gt;

&lt;p&gt;If something listens, you get whatever that program chooses to send. A DNS server sends an answer. A syslog receiver sends nothing, ever. So silence means open, or filtered, or dead, and Nmap has a state for this, &lt;code&gt;open|filtered&lt;/code&gt;. Its documentation explains: "This occurs for scan types in which open ports give no response."&lt;/p&gt;

&lt;p&gt;The consequence is that a generic "UDP port check" does not exist. To check a UDP service you have to speak its protocol: send a real DNS query and read the answer, send a real NTP request and check the timestamp. This is why monitoring tools have a DNS check and a TCP check but rarely a UDP check.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same closed door, three answers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;ICMP echo&lt;/th&gt;
&lt;th&gt;TCP SYN&lt;/th&gt;
&lt;th&gt;UDP datagram&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Service up&lt;/td&gt;
&lt;td&gt;echo reply&lt;/td&gt;
&lt;td&gt;SYN-ACK&lt;/td&gt;
&lt;td&gt;the program's reply, or nothing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Port closed, host up&lt;/td&gt;
&lt;td&gt;echo reply (ICMP has no ports)&lt;/td&gt;
&lt;td&gt;RST, "connection refused"&lt;/td&gt;
&lt;td&gt;ICMP port unreachable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Firewall drops the packet&lt;/td&gt;
&lt;td&gt;silence&lt;/td&gt;
&lt;td&gt;silence, retried for about two minutes&lt;/td&gt;
&lt;td&gt;silence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Host gone&lt;/td&gt;
&lt;td&gt;silence, or an unreachable error from the last router&lt;/td&gt;
&lt;td&gt;silence, or that same ICMP error surfacing as "no route to host"&lt;/td&gt;
&lt;td&gt;silence, or that same ICMP error&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read the table by columns. TCP is the only one where "closed" and "filtered" look different on their own, UDP borrows ICMP to say "closed", and ICMP cannot see ports at all. Silence is the ambiguous answer on all three, which is why no check should conclude "down" from one silent probe sent from one place. Two or three consecutive misses, seen from more than one network, is the honest bar. I wrote about that in &lt;a href="https://uptimepage.dev/blog/stop-false-uptime-alerts?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=icmp-tcp-udp&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;how to stop false uptime alerts&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for a monitor
&lt;/h2&gt;

&lt;p&gt;I build &lt;a href="https://uptimepage.dev/?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=icmp-tcp-udp&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;Uptimepage&lt;/a&gt;, and its check types map onto these questions, so this is how the theory turns into a config.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ping sends one ICMP echo with a 32-byte payload, the size Windows &lt;code&gt;ping&lt;/code&gt; uses (Unix &lt;code&gt;ping&lt;/code&gt; sends 56, as the transcript above shows), so middleboxes see ordinary diagnostic traffic. It tries the first IPv4 and the first IPv6 address of the host and splits the timeout between them. Silence for the whole budget is down, because an echo has no way to refuse. Use it for routers, gateways and hosts that expose no service.&lt;/li&gt;
&lt;li&gt;TCP runs the handshake and closes the connection. Accepted is up, refused is down, and a timeout is recorded as an error with the time it took. An error counts as a failed check for alerting exactly like down, so a firewalled port still opens an incident once the confirmation threshold is met. Use it for databases, brokers, SSH and mail, where "a process accepted the connection" is the most you can assert from outside.&lt;/li&gt;
&lt;li&gt;DNS sends a real query and checks the answer's content, so it is the UDP check for the one UDP service almost everyone depends on. It can point at a specific resolver, which matters when you want that server's view and not your cache's.&lt;/li&gt;
&lt;li&gt;HTTP is a TCP handshake, then TLS, then a request, and it records the time spent in each phase: DNS, connect, TLS and time to first byte. The connect phase is the handshake from this article, measured inside a bigger check.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Pick the check that asks the question you care about, and remember what each one cannot see. There is a fuller version of that decision in the &lt;a href="https://uptimepage.dev/docs/monitor-types?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=icmp-tcp-udp&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;monitor types documentation&lt;/a&gt;, and a story-shaped version in &lt;a href="https://uptimepage.dev/blog/osi-layers?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=icmp-tcp-udp&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;the mystery of the "down" website&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is ICMP a transport protocol like TCP and UDP?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. ICMP is a control protocol that belongs to IP itself, and RFC 792 says every IP module must implement it. It has no ports and carries no application data. It reports on delivery: a destination is unreachable, a packet ran out of hops, and the echo pair that ping uses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does ping work while my website is down?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because ping only proves that the host's IP stack answers and that packets can travel there and back. The kernel sends the echo reply, so no web server, database or other program has to be running. A host can answer every ping while every service on it is dead.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does a closed TCP port fail at once but a firewalled one hangs?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A closed port sends back a TCP reset, so the connect fails immediately with connection refused. A firewall that drops the SYN sends nothing, and the kernel keeps retrying. On Linux the default is six retries, and the kernel documentation puts the final timeout at 131 seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I check a UDP port the way I check a TCP port?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No, because an open UDP port sends nothing back unless the program behind it chooses to answer. A closed port is detectable, since the host should reply with an ICMP port unreachable message, but open and filtered both look like silence. To check a UDP service you have to speak its protocol, for example send a real DNS query.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If UDP is unreliable, why does HTTP/3 use it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;UDP itself gives no delivery guarantee, and QUIC adds its own on top. QUIC runs inside UDP datagrams and does its own handshake, retransmission and congestion control in the application, which lets it evolve without waiting for every operating system to update its TCP stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc792" rel="noopener noreferrer"&gt;RFC 792: Internet Control Message Protocol&lt;/a&gt;, September 1981.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc4443" rel="noopener noreferrer"&gt;RFC 4443: ICMPv6 for the Internet Protocol Version 6&lt;/a&gt;, March 2006.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc768" rel="noopener noreferrer"&gt;RFC 768: User Datagram Protocol&lt;/a&gt;, August 1980.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc9293" rel="noopener noreferrer"&gt;RFC 9293: Transmission Control Protocol&lt;/a&gt;, August 2022.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc1122" rel="noopener noreferrer"&gt;RFC 1122: Requirements for Internet Hosts, Communication Layers&lt;/a&gt;, October 1989. Sections 3.2.2.6 and 4.1.3.1.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc8085" rel="noopener noreferrer"&gt;RFC 8085: UDP Usage Guidelines&lt;/a&gt;, March 2017.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc1035" rel="noopener noreferrer"&gt;RFC 1035: Domain Names, Implementation and Specification&lt;/a&gt;, November 1987. Section 4.2.1.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc6891" rel="noopener noreferrer"&gt;RFC 6891: Extension Mechanisms for DNS (EDNS(0))&lt;/a&gt;, April 2013.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc7766" rel="noopener noreferrer"&gt;RFC 7766: DNS Transport over TCP, Implementation Requirements&lt;/a&gt;, March 2016.&lt;/li&gt;
&lt;li&gt;IETF, &lt;a href="https://www.rfc-editor.org/rfc/rfc9000" rel="noopener noreferrer"&gt;RFC 9000: QUIC, a UDP-Based Multiplexed and Secure Transport&lt;/a&gt;, May 2021.&lt;/li&gt;
&lt;li&gt;IANA, &lt;a href="https://www.iana.org/assignments/protocol-numbers/protocol-numbers.xhtml" rel="noopener noreferrer"&gt;Assigned Internet Protocol Numbers&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Linux kernel, &lt;a href="https://docs.kernel.org/networking/ip-sysctl.html" rel="noopener noreferrer"&gt;IP sysctl documentation&lt;/a&gt;, &lt;code&gt;tcp_syn_retries&lt;/code&gt; and &lt;code&gt;ping_group_range&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Nmap, &lt;a href="https://nmap.org/book/man-port-scanning-basics.html" rel="noopener noreferrer"&gt;Port scanning basics&lt;/a&gt;, the six port states.&lt;/li&gt;
&lt;li&gt;Amazon Web Services, &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security-group-rules-reference.html" rel="noopener noreferrer"&gt;Security group rules for different use cases&lt;/a&gt;, rules for ping/ICMP.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I build &lt;a href="https://uptimepage.dev/?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=icmp-tcp-udp&amp;amp;utm_content=cta" rel="noopener noreferrer"&gt;Uptimepage&lt;/a&gt; because I wanted these checks for my own services: ping, TCP, DNS and HTTP with per-phase timing, plus TLS and domain expiry, heartbeats for background jobs, and browser flows, from several regions. It is AGPL-3.0 open source, so you can run it on your own server if you prefer.&lt;/p&gt;

&lt;p&gt;Now the question for the comments: do you still run ping checks on anything, and what for? I keep them for routers, gateways and hosts that expose no service, and almost nothing else. If a ping check has earned its keep for you somewhere I have not listed, I would like to hear where.&lt;/p&gt;

</description>
      <category>networking</category>
      <category>devops</category>
      <category>monitoring</category>
      <category>programming</category>
    </item>
    <item>
      <title>Your domain can expire while your uptime monitor stays green</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Tue, 04 Aug 2026 14:15:27 +0000</pubDate>
      <link>https://dev.to/slima4/your-domain-can-expire-while-your-uptime-monitor-stays-green-2b2e</link>
      <guid>https://dev.to/slima4/your-domain-can-expire-while-your-uptime-monitor-stays-green-2b2e</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt; One HTTP check on your homepage is not enough to know your site works. It cannot tell you when your domain expires, because an expired domain still returns 200 OK from a parking page. Mail dies at the same time, just as quietly, and you keep losing money while every monitor stays green. The fix is not one clever check. It is a few plain ones, starting with the two that fail on a known date: domain and TLS expiry.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feplinmvbiko94hkbeiua.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feplinmvbiko94hkbeiua.webp" alt="A server topped with a dark error cross, while two monitoring checks beside it still show green success ticks." width="800" height="347"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Two checks green, the server behind them is down. A status check cannot tell the difference.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I keep seeing new products treat reliability as something that takes care of itself. You ship the site, you watch it load once, and you assume it will keep loading. Hosted once, up forever. That is the most expensive assumption a young team can make, because nobody is watching on the day it stops being true.&lt;/p&gt;

&lt;p&gt;So ask one honest question. Is one HTTP check on your homepage enough to know your whole site works?&lt;/p&gt;

&lt;p&gt;For most teams it is not. Here is a real case that shows why.&lt;/p&gt;

&lt;h2&gt;
  
  
  The site was gone, but it answered 200 OK
&lt;/h2&gt;

&lt;p&gt;I was looking at a startup's site last week and it would not load right. Not an error, not a timeout, just the wrong page. The domain had expired two days earlier.&lt;/p&gt;

&lt;p&gt;When a domain lapses, the registrar does not switch the site off. It points the domain at its own parking page, the "this domain may be for sale" one. That page is a normal web page. It returns &lt;code&gt;200 OK&lt;/code&gt; and loads fast. It stays that way for a while too, because an expired domain is not deleted the next morning.&lt;/p&gt;

&lt;p&gt;So the monitor sees a healthy site. Green light, no alert, everyone asleep. What could be wrong?&lt;/p&gt;

&lt;p&gt;Your domain expired. You knew the date once, on the day you bought it a year or two ago. You told yourself you would renew closer to the time, and then you forgot. The card on file lapsed, or the reminder went to an inbox you no longer read.&lt;/p&gt;

&lt;p&gt;Now your customers cannot get in. They panic. They email to ask whether you are still in business. And they get no answer, because the mail died with the domain too.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why the mail dies with it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The registration carries your MX records, so when it lapses, mail to every address on the domain starts bouncing. The website looks up and the inbox goes silent, and the two failures do not obviously connect. That is why the angry email never reaches you.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faa18fg4spy67syhmm8bc.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faa18fg4spy67syhmm8bc.webp" alt="An illustration of a server disconnected from its monitor, the link between them broken." width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The front page still answers. Behind it, nothing is running.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;And the meter keeps running. While the page says 200, your ads keep sending paid clicks to a parking lander, signups never arrive, and recovery is not fully in your hands. Two days down is not a rounding error either, because &lt;a href="https://uptimepage.dev/blog/is-98-uptime-good?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=domain-expiry&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;downtime adds up faster than the percentages suggest&lt;/a&gt;. Once a domain enters redemption after expiry, getting it back can cost far more than a renewal, and the registry sets the clock, not you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The clock after expiry
&lt;/h2&gt;

&lt;p&gt;For a gTLD like a .com, the sequence is roughly this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Grace period.&lt;/strong&gt; The registrar parks the domain. The site is gone but the name still answers, usually for around 30 days, and a normal renewal still works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redemption, 30 days.&lt;/strong&gt; Now it really goes dark. You can still get it back, but through a restore fee that is many times the renewal price.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pending delete, about 5 days.&lt;/strong&gt; Nothing you can do. At the end it drops to whoever is waiting.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;ICANN's &lt;a href="https://www.icann.org/resources/pages/domain-name-renewal-expiration-faqs-2018-12-07-en" rel="noopener noreferrer"&gt;renewal and expiration FAQ&lt;/a&gt; and its &lt;a href="https://www.icann.org/resources/pages/grace-2013-05-03-en" rel="noopener noreferrer"&gt;grace period notes&lt;/a&gt; have the exact stages. The practical takeaway is that the first month, the one where an uptime monitor is happily reporting 200, is the cheap month. Everything after that costs money or the domain.&lt;/p&gt;

&lt;h2&gt;
  
  
  My own uptime check would miss it too
&lt;/h2&gt;

&lt;p&gt;I build an uptime monitor, so here is the uncomfortable part. My own HTTP check would have called that dead site up too. A 200 is a 200. A status-code check cannot tell a real page from a parking page, and it should not pretend that it can.&lt;/p&gt;

&lt;p&gt;That is the trap. One HTTP check on the homepage, the most common setup there is, is blind by design. The site answers, the code is green, and the part that pays your bills is broken behind it.&lt;/p&gt;

&lt;p&gt;You can bolt on body matching, and it helps a bit. Assert that the response contains a string only your app renders, and a parking page fails it. But a keyword check is still guessing at the cause after the fact. The registration date is a fact you can read months in advance, so read it.&lt;/p&gt;

&lt;h2&gt;
  
  
  One check is not enough, set up several
&lt;/h2&gt;

&lt;p&gt;So how do you avoid reading about your own outage in a customer's email? Not with one cleverer check. With a few plain ones, each watching a different way to fail. Start with the two that hand you a date instead of a surprise.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Domain expiry.&lt;/strong&gt; Reads the registration record and warns you weeks before the day it lapses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;TLS certificate expiry.&lt;/strong&gt; The same shape, a full outage with the date printed on it in advance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then cover the failures a homepage check never sees: a TCP check on your database port, a DNS check that reads the answer from outside your own network, and a heartbeat for the backup job that dies without a sound. The full list is in &lt;a href="https://uptimepage.dev/blog/do-i-need-an-uptime-monitor?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=domain-expiry&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;do I need an uptime monitor&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Six checks you trust beat one that only ever watches the front door.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to set up domain expiry with an alert
&lt;/h2&gt;

&lt;p&gt;Two minutes now buys you weeks of warning later. The shape is the same in any tool that supports it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Add a monitor and choose the &lt;strong&gt;Domain expiry&lt;/strong&gt; type.&lt;/li&gt;
&lt;li&gt;Enter the registered domain, for example &lt;code&gt;yourbusiness.com&lt;/code&gt;, not a subdomain.&lt;/li&gt;
&lt;li&gt;Set the thresholds in days. A common choice is a warning at 30 days and a critical alert at 7.&lt;/li&gt;
&lt;li&gt;Attach a notification channel you actually read. Telegram, Slack, PagerDuty, SMS, a webhook, or email all work. Pick the one that reaches you at night.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;From then on it checks once a day, warns you when the date crosses your line, and reminds you as it gets close. No page-watching, no false green.&lt;/p&gt;

&lt;p&gt;If you would rather not add a service for it, a cron job that runs &lt;code&gt;whois&lt;/code&gt; and greps the expiry date will get you most of the way. The point is the date, not the tool. What you must not do is leave the only reminder in a registrar email going to a founder's old inbox.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaways&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One HTTP check on the homepage is not enough. It stays green while the domain, the certificate, DNS, or a background job is broken.&lt;/li&gt;
&lt;li&gt;An expired domain serves a parking page that returns 200 OK, so status-code monitors report it as healthy.&lt;/li&gt;
&lt;li&gt;Email dies at the same time and even more quietly, because the MX records lapse with the registration.&lt;/li&gt;
&lt;li&gt;You keep losing money the whole time: paid clicks to a dead page, lost signups, and a slow, costly recovery.&lt;/li&gt;
&lt;li&gt;Add the checks that fail on a known date first, domain and TLS expiry, then the rest.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Common questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why does an expired domain still return 200 OK?&lt;/strong&gt; The registrar parks it on a lander page instead of taking it offline. The old server is gone, but the parking page is a real web page that answers normally, so a check that only reads the status code sees success.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long before an expired domain is gone for good?&lt;/strong&gt; You usually get a month or two. Grace period, then a 30-day redemption where the site is dark but recoverable for a fee, then a short pending-delete stage before it is released to anyone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What else breaks when a domain expires?&lt;/strong&gt; Your email, at the same moment and more quietly. The MX records go with the registration, so mail starts bouncing while you are still looking at a page that loads.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is a keyword check enough instead?&lt;/strong&gt; It catches the parking page, which is better than nothing, but only once you are already down. The expiry date is knowable months ahead. Use both if you like, but do not skip the date.&lt;/p&gt;

&lt;p&gt;I build &lt;a href="https://uptimepage.dev/?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=domain-expiry&amp;amp;utm_content=cta" rel="noopener noreferrer"&gt;Uptimepage&lt;/a&gt; because I wanted checks like this for my own services: domain and TLS expiry, HTTP, TCP, ping and DNS, heartbeats for background jobs, and browser flows for logins, from several regions. It is AGPL-3.0 open source, so you can run it on your own server if you prefer.&lt;/p&gt;

&lt;p&gt;Now the question I actually want answered in the comments: what is the dumbest green check you have been burned by? I have seen a health endpoint that returned 200 with a hardcoded &lt;code&gt;{"status":"ok"}&lt;/code&gt; while the database behind it was unreachable for six hours. I suspect yours is better.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>monitoring</category>
      <category>dns</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How I stop one bad probe from waking you at 3 a.m.</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Tue, 28 Jul 2026 14:55:48 +0000</pubDate>
      <link>https://dev.to/slima4/how-i-stop-one-bad-probe-from-waking-you-at-3-am-1k7l</link>
      <guid>https://dev.to/slima4/how-i-stop-one-bad-probe-from-waking-you-at-3-am-1k7l</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4yv8hh94rv26d9rmzfy8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4yv8hh94rv26d9rmzfy8.png" alt=" " width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt; The most common false alert in uptime monitoring is one probe location having a bad network day. I treated that as a design requirement from day one, so the checking pipeline has two gates: a region only counts as down after it fails the same check twice in a row, and an incident only opens when enough regions agree. A region that goes silent leaves the vote instead of counting as down. One bad location cannot page you. The &lt;a href="https://uptimepage.dev/blog/stop-false-uptime-alerts?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=false-alerts&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;original post&lt;/a&gt; carries the three figures below as live widgets you can drive yourself.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The alert that was nothing
&lt;/h2&gt;

&lt;p&gt;The phone buzzes at 3 a.m. "Your API is down." You get up, open the laptop, and everything is green. The site was fine the whole time. One monitoring server, somewhere far away, had a bad network moment and sent an alert about nothing.&lt;/p&gt;

&lt;p&gt;This is the false alert everyone in monitoring knows. It costs you sleep, and then it costs you something worse: the next alert feels less serious. The day a real outage comes, you look at your phone and think "probably nothing again."&lt;/p&gt;

&lt;p&gt;I knew this failure mode before I wrote the first line of the scheduler, so it became a requirement, on the same level as "checks must run on time." The rule I started from: a single bad location must never be able to page a customer. The whole checking pipeline, from how probes report to how incidents open, is shaped by that rule.&lt;/p&gt;

&lt;p&gt;The shape it took is two gates. A failure has to pass both before anyone gets paged.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate one: fail twice, from the same place
&lt;/h2&gt;

&lt;p&gt;One failed check proves very little. Networks lose packets, routers restart, and sometimes a DNS answer arrives a second too late. All of that can make one check fail while your site is healthy.&lt;/p&gt;

&lt;p&gt;So a region only counts as down after it fails the same check twice in a row. Not two failures somewhere in the system: two failures from that one region, back to back. A single blip resets to zero on the next good check.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqecsqpxg7r245f00zwrw.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqecsqpxg7r245f00zwrw.gif" alt="Two rows of check results. In the top row a single failed check sits between passing ones, the counter resets and nothing happens. In the bottom row two failures land back to back, so that region counts as down and earns one vote." width="720" height="357"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;One blip resets the counter. Two failures in a row from the same region earn that region a single vote.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The count is a setting on each monitor. Two is the default. For a very sensitive check you can set it to one, and for a noisy target you can raise it. Checks run on the monitor's own interval, so with a one-minute check the second failure arrives about a minute after the first. That minute buys you a lot of silence for a very small delay.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gate two: the vote
&lt;/h2&gt;

&lt;p&gt;Passing gate one gives a region exactly one vote. Nothing more.&lt;/p&gt;

&lt;p&gt;Say you check from five regions and one of them has a bad ISP day. It fails twice in a row and votes "down." The other four keep passing. One vote against four is not enough, so nothing happens and nobody gets paged.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu0outz33y2xgbg0d56u2.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu0outz33y2xgbg0d56u2.gif" alt="Two votes across five regions. In the first, one region is down and four are up, there is no majority and nobody is paged. In the second, three of the five are down, that is a majority and an incident opens." width="600" height="345"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;One region against four cannot page you. Three of five can.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The default rule is a majority: more than half of the reporting regions have to agree. You can change it per monitor. "Any" opens an incident on the first confirmed region, which is useful when you want the earliest possible signal and can accept some noise. "All" waits until every region agrees, for a service that only matters when it is unreachable everywhere. Or you pick a fixed number, like two regions out of whatever you assign.&lt;/p&gt;

&lt;p&gt;A monitor checked from a single region behaves the same under every rule. The vote only starts to protect you when you add a second location, and it gets better with a third.&lt;/p&gt;

&lt;h2&gt;
  
  
  Silence is not failure
&lt;/h2&gt;

&lt;p&gt;Here is the part that took the most care to get right.&lt;/p&gt;

&lt;p&gt;The probe regions push their results to the control plane, the brain of the system. The brain never calls out to ask "are you alive?" during a vote. It just counts the results that arrived in the last few check cycles.&lt;/p&gt;

&lt;p&gt;That gives silence a clear meaning. A region that lost its own connection cannot send anything, so it simply is not in the vote. It is not a down vote and not an up vote; the region is out until it reports again. The majority recalculates over the regions that still speak. Five regions where one goes dark becomes a vote of four.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fslg2hll5bmb4v6jmrdph.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fslg2hll5bmb4v6jmrdph.gif" alt="Five probe regions reporting to the control plane. Four send results over solid arrows, one sends nothing and is drawn as an empty dashed circle. The majority is now counted over four regions." width="720" height="344"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A silent region is not a down vote and not an up vote. It leaves the vote until it reports again.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I think this is the only honest reading. A region that cannot reach the brain is telling you nothing about your site. Treating its silence as "down" would turn every probe outage on the monitoring side into a fake incident on yours.&lt;/p&gt;

&lt;p&gt;And it works in both directions. Missing data never opens an incident, and it never closes one. An open incident only closes when the down votes fall below the threshold and at least one region shows a real run of passing checks. Recovery needs proof, the same way failure does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who watches the watchers
&lt;/h2&gt;

&lt;p&gt;There is one gap left. If every region that covers your monitor goes dark, the vote has nobody in it. Your site could burn down and no incident would open, because no data means no votes.&lt;/p&gt;

&lt;p&gt;For that case there is a separate signal, one level below the checks. Every probe sends a small check-in to the brain on its own schedule, a heartbeat that has nothing to do with your monitors. When the last live probe covering a monitor goes stale, you get a different message: "NO DATA: monitoring interrupted, no check results received." It is honest about what it knows: the service cannot see your site right now. It does not claim your site is down, because it has no idea.&lt;/p&gt;

&lt;p&gt;When probing returns, you get a "monitoring RESUMED, receiving check results again" note, and the vote picks up where it left off.&lt;/p&gt;

&lt;p&gt;One more detail I like. If a large share of all monitors goes silent at the same time, that pattern is almost never a thousand customer problems. It is one problem, mine. In that case the system alarms me and holds the customer notices, so my bad day does not become spam on yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for your setup
&lt;/h2&gt;

&lt;p&gt;Three practical points follow from this, whatever tool you run.&lt;/p&gt;

&lt;p&gt;Check from more than one region. The vote cannot protect you with a single location. Two regions with the majority rule means both have to agree, which already kills the classic false alert. Three gives you two out of three, which is the right balance for most services.&lt;/p&gt;

&lt;p&gt;Leave the confirmation count at two unless you have a reason. It is the difference between "a packet got lost" and "this endpoint is failing."&lt;/p&gt;

&lt;p&gt;Pick the rule to match the monitor. A payment API deserves majority or even any. An internal tool nobody uses at night can wait for all. The setting is per monitor, so you do not have to choose one policy for everything.&lt;/p&gt;

&lt;p&gt;The goal of all this machinery is boring: when your phone buzzes, it is real. Everything else is plumbing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is a false alert in uptime monitoring?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A false alert says your site is down when it is not. The usual cause is a problem near the probe, not near your site: a bad network path, a busy datacenter, a DNS hiccup on the monitoring side. Your site answered fine for every real user the whole time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many regions need to agree before an incident opens?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;By default, more than half of the regions that are reporting results. A region only joins the down side after it fails the same check twice in a row. You can change the rule per monitor: any single region, all regions, or a fixed count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens when a probe region goes offline?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It leaves the vote. A region that sends no results is not counted as down and not counted as up. The vote recalculates over the regions that still report. Missing data alone never opens an incident.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does a silent region close my open incident?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. Closing needs real proof: the down votes must fall below the threshold and at least one region must show a run of passing checks. Silence is not proof of recovery, so an open incident stays open until real results come back.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaways&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One failed check from one place proves nothing. Confirm it from the same region before that region gets a say.&lt;/li&gt;
&lt;li&gt;Give each confirmed region one vote, then require agreement. A majority of reporting regions is a good default.&lt;/li&gt;
&lt;li&gt;A silent region must leave the vote, not join the down side. Otherwise every probe outage becomes a fake incident.&lt;/li&gt;
&lt;li&gt;Missing data must not close an incident either. Recovery needs a real run of passing checks.&lt;/li&gt;
&lt;li&gt;Cover the all-dark case with a separate heartbeat and a message that says "we cannot see your site," not "your site is down."&lt;/li&gt;
&lt;li&gt;Two regions kill the classic false alert. Three is the sweet spot for most services.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;The three figures above are live widgets in the &lt;a href="https://uptimepage.dev/blog/stop-false-uptime-alerts?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=false-alerts&amp;amp;utm_content=cta" rel="noopener noreferrer"&gt;original post&lt;/a&gt;: you can change the confirmation count and the vote rule and watch where the line falls. The whole project is &lt;a href="https://github.com/uptimepage/uptimepage" rel="noopener noreferrer"&gt;open source&lt;/a&gt;, so the scheduler and the incident writer are there to read.&lt;/p&gt;

&lt;p&gt;How does your setup handle a probe that goes quiet? I have seen tools count silence as down, and I would like to know if anyone has a better reading than "it leaves the vote."&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>monitoring</category>
      <category>programming</category>
    </item>
    <item>
      <title>How I mapped my codebase for humans and AI agents</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Thu, 23 Jul 2026 14:17:23 +0000</pubDate>
      <link>https://dev.to/slima4/how-i-mapped-my-codebase-for-humans-and-ai-agents-1i4p</link>
      <guid>https://dev.to/slima4/how-i-mapped-my-codebase-for-humans-and-ai-agents-1i4p</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt; I asked an AI model to turn my codebase into three things: a one-page summary for me, a JSON file for the next AI agent, and an interactive map you can click. It worked well, but only after one boring step: check every number against the code. The model got the shape right and several counts wrong. You can &lt;a href="https://uptimepage.dev/architecture?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=codebase-map&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;see the live map&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;The interactive map itself, with one flow lit up across the system. Open the &lt;a href="https://uptimepage.dev/architecture?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=codebase-map&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;live version&lt;/a&gt; and click any flow.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A big codebase does not fit in your head. Mine is about 146,000 lines of Rust across 31 modules. When you open a project that size, the first hour is just finding where things are.&lt;/p&gt;

&lt;p&gt;AI agents have the same problem. Every time an agent starts a task, it reads many files to learn how the system fits together. It does this from zero, every time. That is slow, and it costs money.&lt;/p&gt;

&lt;p&gt;So I tried something. I asked an AI model to read the whole codebase and write three things. One for me, one for the next agent, and one for anyone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three files
&lt;/h2&gt;

&lt;p&gt;The first file is a one-page summary for a human. It lists the rules that must always hold, the main parts and what each one does, the path a request takes, and the traps that waste time. It is the page I wish existed on my first day.&lt;/p&gt;

&lt;p&gt;The second file is a JSON map for the next AI agent. It is not written to be pretty. It lists the invariants with the file or test that enforces each one, a short recipe for each common task, the known traps, and the key files. When the next agent starts a task, it reads this first and skips an hour of searching.&lt;/p&gt;

&lt;p&gt;The third file is an interactive map for everyone. It shows the parts as boxes in columns. You pick a flow, like "a scheduled check" or "an agent login", and the path lights up across the boxes with numbered steps. It is live here: &lt;a href="https://uptimepage.dev/architecture?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=codebase-map&amp;amp;utm_content=inline" rel="noopener noreferrer"&gt;the architecture map&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7uygsaibtths4ihn4pza.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7uygsaibtths4ihn4pza.webp" alt="A scattered codebase of many small files on the left flows through an amber arrow into a single JSON map file marked with a green source dot, which then branches into a human-readable page and an interactive node map." width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The three files. The JSON map is the source of truth, and the human page and the interactive map are both built from it.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;One file is the source of truth&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The JSON map holds every fact. The other two files are only views of it. Build the map first, and the human page and the interactive map cannot disagree.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The method, not one magic prompt
&lt;/h2&gt;

&lt;p&gt;There is no single prompt that does this well. The result comes from four steps, in order.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2qvanrh5pq35srlgt99r.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2qvanrh5pq35srlgt99r.webp" alt="Four steps left to right joined by an amber line: explore in parallel across several small files, build one large JSON map file with a green source dot, check the numbers in it with a magnifying glass, then render the views into two output cards. The map is drawn largest as the anchor." width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The four steps. Everything hangs off step two, the map.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;First, explore in parallel. One agent cannot read 146,000 lines in one go. I let several agents read different parts at the same time, then joined their notes. This is faster and it covers more.&lt;/p&gt;

&lt;p&gt;Second, build the machine map first. The JSON for the agent is the source of truth. The human page and the interactive map are only views of it. So I build the JSON first and put every fact in one place.&lt;/p&gt;

&lt;p&gt;Third, check every number in that map, before you build anything else. This is the step people skip, and it is the most important one. You verify once, at the source, so a wrong count cannot spread into the other files.&lt;/p&gt;

&lt;p&gt;Fourth, render the views from the map. The human page and the interactive map both come from the same JSON, so they cannot disagree. Each view has one clear reader: the page is for a new engineer, and the map is for a visitor who has never seen the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers the AI got wrong
&lt;/h2&gt;

&lt;p&gt;The model was good at structure. It found the parts, the flows, and the rules. But it guessed numbers, and some guesses were wrong.&lt;/p&gt;

&lt;p&gt;It said there were 13 alert channels. The real number is 14.&lt;/p&gt;

&lt;p&gt;It said there were about 80 error codes. The real number is 155.&lt;/p&gt;

&lt;p&gt;It said there were 19 blog posts. The real number is 18.&lt;/p&gt;

&lt;p&gt;None of these are small. If you publish them, you look careless, and the next agent trusts a wrong map. I only found them because I checked each count against the source code.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Check the numbers, at the source&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An AI model is strong at shape and words. It is weak at exact numbers. Verify every count in the map, once, before you build anything from it. Then check the ones that matter yourself.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The prompts
&lt;/h2&gt;

&lt;p&gt;Here are the three prompts, in the order I ran them. Build the map first, check it, then make the views. Change the details for your own project, and keep the "verify against the source" line in the first one.&lt;/p&gt;

&lt;p&gt;For the machine map, build this first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Read my whole codebase. Write one JSON file for the next AI agent that
will add a feature. Include the invariants, and for each one name the file
or test that enforces it. Add a short recipe for each common task: the
goal, then the files to touch in order. Add the known traps and the key
files. Keep every path exact. Verify every count against the source code.
Do not invent numbers. This file is data for a machine, not prose.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the human summary, built from the map:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;From that verified JSON, write one self-contained HTML page that explains
the system to a new engineer. Include the rules that must always hold,
the main parts and what each does, the path a request takes, the path data
takes, and the traps that waste time. Rank the parts by size. Do not add
any number that is not already in the JSON.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the interactive map, built from the same map:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;From the same JSON, build one self-contained interactive HTML page. Show
the parts as boxes in columns. Show each flow as a numbered path that lights
up across the boxes when I click it. Keep all styles and scripts in the
page. No build step.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why this is worth doing
&lt;/h2&gt;

&lt;p&gt;The human page saved me time the next week. The map helps me explain the system in one screen. And the JSON is the part I did not expect to like. The next agent that touches this code reads a map first, so it starts from step one instead of step zero.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The next agent starts from step one, not step zero&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A map file hands the next AI agent the rules, the tasks, and the traps up front. It reads one small file instead of re-reading the whole codebase every time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key takeaways&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Build the machine map (JSON) first. The human page and the interactive map are only views of it, so they cannot drift.&lt;/li&gt;
&lt;li&gt;An AI model is strong at structure and weak at numbers. Verify every count at the source before you build anything from it.&lt;/li&gt;
&lt;li&gt;Write a map file for the next AI agent, so it starts from step one instead of re-reading the whole codebase.&lt;/li&gt;
&lt;li&gt;Explore in parallel. Several agents reading different parts cover more than one agent reading everything.&lt;/li&gt;
&lt;li&gt;Give each file one clear reader. It keeps the writing simple.&lt;/li&gt;
&lt;li&gt;The model got real counts wrong, like 13 channels when the real number was 14. Do not publish a number you did not check.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you want to see the output, open the &lt;a href="https://uptimepage.dev/architecture?utm_source=devto&amp;amp;utm_medium=community&amp;amp;utm_campaign=codebase-map&amp;amp;utm_content=cta" rel="noopener noreferrer"&gt;interactive map&lt;/a&gt;. The whole project is open source, so you can read the real code behind every box.&lt;/p&gt;

&lt;p&gt;If you have mapped a codebase for an AI agent: which format did the agent actually use, and how many numbers did the model get wrong on your run? Mine missed three. I would like to know if that is typical.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>programming</category>
      <category>documentation</category>
    </item>
    <item>
      <title>How to write incident status updates that build trust</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Tue, 21 Jul 2026 13:42:58 +0000</pubDate>
      <link>https://dev.to/slima4/how-to-write-incident-status-updates-that-build-trust-7d4</link>
      <guid>https://dev.to/slima4/how-to-write-incident-status-updates-that-build-trust-7d4</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt; During an outage your status page has a second job: holding trust while the fix is still in progress. Write plainly. Say what customers feel first, say what you are doing, and promise a time for the next update. Then keep that promise, even when there is no news. The four stages are investigating, identified, monitoring, and resolved.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The idea of an incident commander did not come from software. It came from wildfires. In 1970s California, a run of huge fires pushed fire chiefs to study why large responses fell apart. The cause surprised them. It was not too few firefighters or too little water. It was communication. Different teams used different words and no one shared a plan. The fix was a simple system with clear roles and clear language, &lt;a href="https://en.wikipedia.org/wiki/Incident_Command_System" rel="noopener noreferrer"&gt;now known as the Incident Command System&lt;/a&gt;, and on-call engineers still use a version of it today.&lt;/p&gt;

&lt;p&gt;Your status page is the communication part of that system. When something breaks, the repair happens in your code. The trust happens on the status page. This post is about the second part: how to write updates that keep people calm while the first part is still in progress.&lt;/p&gt;

&lt;h2&gt;
  
  
  The words cost real money
&lt;/h2&gt;

&lt;p&gt;An outage is expensive before you write a single word. In &lt;a href="https://itic-corp.com/itic-2024-hourly-cost-of-downtime-report/" rel="noopener noreferrer"&gt;ITIC's 2024 survey&lt;/a&gt; of more than 1,000 companies, over 90% of mid-size and large firms said one hour of downtime costs them more than $300,000. For 41% of them, one hour costs between $1 million and $5 million. &lt;a href="https://uptimeinstitute.com/resources/research-and-reports/annual-outage-analysis-2024" rel="noopener noreferrer"&gt;Uptime Institute's outage research&lt;/a&gt; points the same way: about one in five recent outages cost more than $1 million, and more than half cost over $100,000.&lt;/p&gt;

&lt;p&gt;You do not need to be a bank for this to matter. If your store makes $2,000 a day, one bad hour during a sale can cost more than a slow afternoon, because sales are not spread evenly across the day. The size of the number changes. The shape of the problem does not.&lt;/p&gt;

&lt;p&gt;Here is the part teams forget. Some of that cost is the downtime itself. The rest is trust, and trust is where writing helps. Many providers offer an SLA, a promise to stay up with a penalty if they miss it. The penalty is almost always a service credit, which is money back on your bill. A credit refunds your invoice, not your customer's bad afternoon. Good updates cannot bring the service back. They can protect the relationship while it is down.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four stages of an incident
&lt;/h2&gt;

&lt;p&gt;Most status pages, including the big public ones, use the same four stages. They come from a shared standard, so a customer who has read one status page already understands yours.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;What it means&lt;/th&gt;
&lt;th&gt;What to write&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Investigating&lt;/td&gt;
&lt;td&gt;You know something is wrong. You do not know why yet.&lt;/td&gt;
&lt;td&gt;Name the impact and say you are looking into it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Identified&lt;/td&gt;
&lt;td&gt;You found the cause and are working on a fix.&lt;/td&gt;
&lt;td&gt;Say what is wrong in plain words and that a fix is coming.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monitoring&lt;/td&gt;
&lt;td&gt;The fix is live. You are watching to be sure it holds.&lt;/td&gt;
&lt;td&gt;Say the fix is in and you are checking that it works.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resolved&lt;/td&gt;
&lt;td&gt;The service is back and stable.&lt;/td&gt;
&lt;td&gt;Say it is over, and thank people for waiting.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The order matters. Do not jump to identified because you have a hunch. If you say you found the cause and then change your story an hour later, every update after that is worth less. Move a stage forward only when the facts move with you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7e2rllfk6qtmfa3qgboe.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7e2rllfk6qtmfa3qgboe.webp" alt="The four incident stages in order: investigating, identified, monitoring, resolved, each with a sample status message." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The rules of a good update
&lt;/h2&gt;

&lt;p&gt;Five rules cover almost every message.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Start with impact.&lt;/strong&gt; Say what the customer feels before you talk about servers or databases. "Some payments are failing" is more useful than "we are seeing elevated 5xx errors on the API gateway." Only the first one means anything to a customer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Share facts, not guesses.&lt;/strong&gt; Say what you have confirmed. If something is still unknown, say that too. Writing "we do not yet know the cause" is fine. Inventing a cause to sound in control is not. This rule earns its place under pressure: the &lt;a href="https://uptimeinstitute.com/resources/research-and-reports/annual-outage-analysis-2025" rel="noopener noreferrer"&gt;Uptime Institute's 2025 report&lt;/a&gt; found that about 40% of organizations had a major outage caused by human error in the last three years, and 85% of those came from someone not following a procedure. An incident is exactly when tired people guess.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Promise the next update.&lt;/strong&gt; Give a real time, like "next update by 15:00 UTC" or "another update within 30 minutes." This one line does more than any apology. It tells people they can close the tab and get on with their day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Add a timestamp and a time zone.&lt;/strong&gt; "In 30 minutes" means nothing if the reader does not know when you posted. Every update should carry a clear time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep it plain.&lt;/strong&gt; Short sentences. No blame, no jargon, no backstory. A worried customer reads fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  A bad update and a better one
&lt;/h2&gt;

&lt;p&gt;Here is a bad update of the kind you have all seen:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;We are aware of an issue and our team is working hard to resolve it as quickly as possible. We apologize for any inconvenience.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It sounds polite and says nothing. What is broken? Who is affected? When will I hear more? A better version answers those questions:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Investigating: some customers cannot log in. Sign-ups and password resets are affected too. Our team is looking into the cause now. Next update by 14:30 UTC.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Same length. Only one of them lets a customer decide what to do next.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fht2289fiy1fn5qtp1vj4.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fht2289fiy1fn5qtp1vj4.webp" alt="A vague incident update next to a specific one that names the impact and a next-update time." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How often to post
&lt;/h2&gt;

&lt;p&gt;Set the rhythm in your first message and keep it. Every 30 to 60 minutes is a good starting point for a serious outage. The exact number matters less than the promise. If you said 30 minutes, post at 30 minutes, even when the only news is "still working, no change yet." Silence after a promised time reads as "they have lost control." A boring on-time update beats an exciting late one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Planned maintenance is the same skill, calmer
&lt;/h2&gt;

&lt;p&gt;Maintenance uses three stages: scheduled, in progress, and completed. The difference is that you have time to write it well in advance. Tell people what will happen, when it will happen, and whether they need to do anything. Mark it in progress while the work runs, and completed when it is done. The tone is calmer because nothing is on fire, but the shape is the same: impact, timing, next update.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mistake almost everyone makes
&lt;/h2&gt;

&lt;p&gt;It is not the wording. It is the second update.&lt;/p&gt;

&lt;p&gt;The first update is easy. The outage just happened, and everyone is paying attention. Then the fix takes longer than expected, the team goes quiet, and the page sits at "investigating" for two hours. From the outside, a frozen status page looks the same as a dead company.&lt;/p&gt;

&lt;p&gt;The fix is a habit, not a talent. Decide the next-update time in your first message, set a timer, and post again when it rings. If there is no news, that is the update: no change, still working, next check in 30 minutes. The point of the promise is that people stop refreshing and trust you to come back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write the message, then give it a home
&lt;/h2&gt;

&lt;p&gt;You can write all of this by hand during an outage. Most people write worse under stress, which is the worst possible time to start from a blank box. That is why we built a free &lt;a href="https://uptimepage.dev/tools/incident-update-generator" rel="noopener noreferrer"&gt;incident update generator&lt;/a&gt;: pick a stage, fill in a few fields, and it writes a clean message you can paste into any status page. It runs in your browser and stores nothing you type.&lt;/p&gt;

&lt;p&gt;The message still needs somewhere to live. A status page is where customers look when your product will not load, so it should be honest and easy to reach. If you are choosing an uptime number to promise, the &lt;a href="https://uptimepage.dev/tools/uptime-sla-calculator" rel="noopener noreferrer"&gt;uptime SLA calculator&lt;/a&gt; turns a percentage into real minutes, and &lt;a href="https://uptimepage.dev/blog/is-98-uptime-good" rel="noopener noreferrer"&gt;is 98% uptime good&lt;/a&gt; shows how small percentages hide large amounts of downtime. For the page itself, &lt;a href="https://uptimepage.dev/blog/status-page-you-cant-fake" rel="noopener noreferrer"&gt;a status page you cannot fake&lt;/a&gt; covers how to keep it worth trusting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What should the first incident update say?&lt;/strong&gt; Say what the customer feels, say you are looking into it, and give a time for the next update. You do not need the cause yet. Waiting for the root cause before posting is the most common mistake.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the difference between investigating and identified?&lt;/strong&gt; Investigating means the cause is not confirmed. Identified means you have confirmed the cause and are working on a fix. Do not move to identified based on a hunch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Should you apologize in an incident update?&lt;/strong&gt; One short, honest line is fine. Skip the long apology. A clear next-update time does more for trust than three sentences of sorry.&lt;/p&gt;

&lt;p&gt;If you have run incidents: what update cadence actually held up for you under a long outage, and did anyone ever complain that you posted too often? I have never seen that complaint, but I would like to know if it exists.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>monitoring</category>
      <category>writing</category>
    </item>
    <item>
      <title>Why I chose Rust over Go for an uptime monitor</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Mon, 20 Jul 2026 14:46:30 +0000</pubDate>
      <link>https://dev.to/slima4/why-i-chose-rust-over-go-for-an-uptime-monitor-lo8</link>
      <guid>https://dev.to/slima4/why-i-chose-rust-over-go-for-an-uptime-monitor-lo8</guid>
      <description>&lt;p&gt;I build Uptimepage, an &lt;a href="https://uptimepage.dev" rel="noopener noreferrer"&gt;open-source uptime monitor&lt;/a&gt; and status page written in Rust. People ask why Rust and not Go, since Go is the usual pick for this kind of network service. Here is the honest answer. It is not that Rust wins everywhere. It is that one part of this job made the choice for me.&lt;/p&gt;

&lt;p&gt;The product is one promise: tell you fast and honestly when your site is slow or down. The numbers I show you, like your p99 response time, have to be clean. If my own code adds random delay, I blur the exact signal you pay for. So the runtime under the prober matters more here than it would for a normal app.&lt;/p&gt;

&lt;h2&gt;
  
  
  Garbage collection shows up in the tail
&lt;/h2&gt;

&lt;p&gt;Go has a garbage collector. It is fast, and most apps never feel it. But it still has to do work to free memory, and that work can land inside the millisecond timings I report. Run tens of thousands of checks at once and a pause at the wrong moment lifts a p99 number. At that point I am measuring my own runtime, not your server.&lt;/p&gt;

&lt;p&gt;Rust has no garbage collector. Memory is freed at a point I can see in the code. There is no background pause I did not write. For a tool that sells timing, that control is worth the extra work Rust asks for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that looks like in memory
&lt;/h2&gt;

&lt;p&gt;Here are some real numbers, with a warning attached. They come from a load test on a developer laptop, not a production server. I use them to catch a slowdown between two versions of my code, not to plan capacity. A real server does better, a small box does worse. Treat them as a floor.&lt;/p&gt;

&lt;p&gt;In one run, a single machine held 50,000 checks in flight and peaked at 933 MiB of memory. That is under one gigabyte for fifty thousand live checks. My running server uses about 42 MiB while watching its monitors, and it sits quiet when there is nothing to do. That kind of density is normal for a service with no garbage collector, and it means one small box covers a lot of monitors. I go deeper on the prober and the throughput numbers in &lt;a href="https://uptimepage.dev/blog/building-an-uptime-monitor-in-rust" rel="noopener noreferrer"&gt;the build story&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug that never compiles
&lt;/h2&gt;

&lt;p&gt;Speed is only half of it. The other half is a class of bug that Rust refuses to build.&lt;/p&gt;

&lt;p&gt;Picture many workers writing to one shared map of results at the same time. In Go this compiles and runs. Sometimes it is fine. Sometimes two goroutines write at once, you get corrupt data or a crash, and it only happens under load, which is the worst time to find out:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="c"&gt;// compiles fine, races at runtime&lt;/span&gt;
&lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="k"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt;&lt;span class="p"&gt;{}&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="k"&gt;range&lt;/span&gt; &lt;span class="n"&gt;checks&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;go&lt;/span&gt; &lt;span class="k"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt; &lt;span class="n"&gt;Check&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ID&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Latency&lt;/span&gt; &lt;span class="c"&gt;// data race&lt;/span&gt;
    &lt;span class="p"&gt;}(&lt;/span&gt;&lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Go gives you good tools for this. There is a race detector, a sync.Mutex, and channels. But remembering to reach for them is on you.&lt;/p&gt;

&lt;p&gt;In Rust the same concurrent write does not compile. The compiler stops you until the shared map is wrapped in a lock:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight rust"&gt;&lt;code&gt;&lt;span class="c1"&gt;// will not compile unless the shared map is a Mutex&lt;/span&gt;
&lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nn"&gt;Mutex&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nn"&gt;HashMap&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;new&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="nn"&gt;std&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nn"&gt;thread&lt;/span&gt;&lt;span class="p"&gt;::&lt;/span&gt;&lt;span class="nf"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;(|&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt; &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="n"&gt;checks&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;let&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="nf"&gt;.spawn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;move&lt;/span&gt; &lt;span class="p"&gt;||&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="c1"&gt;// each worker locks only for its own insert&lt;/span&gt;
            &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="nf"&gt;.lock&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="nf"&gt;.unwrap&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="nf"&gt;.insert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="py"&gt;.id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="py"&gt;.latency&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each worker holds the lock only for its own insert, so the writes stay safe without blocking the others for long. For code that runs day and night across many machines, "the compiler will not let you ship the race" removes a whole set of late-night bugs before they exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Go is the better pick
&lt;/h2&gt;

&lt;p&gt;None of this makes Go a bad choice. Often it is the better one. Go is faster to learn. It builds in seconds. Its standard library for network services is excellent, and a new engineer can be useful in days. If I were building a normal web service, or something I had to ship this week, Go would be on the table and might win.&lt;/p&gt;

&lt;p&gt;Rust asks more from you first. The compiler argues with you. The build is slower. You spend time on things Go would just handle. I take that trade because this job is narrow and it rewards tight control over memory and timing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The compiler as a safety net
&lt;/h2&gt;

&lt;p&gt;Heavy code review is how large teams catch data races and memory bugs. Rust gives you a lot of that for free. The compiler turns those mistakes into build errors, so they never reach production and never wake anyone up. Every change gets a strong, automatic check before it ships, which is a big part of why I trust the service to run unattended.&lt;/p&gt;

&lt;p&gt;It is also why Uptimepage is open source and self-hostable. You can read the code, run it on your own hardware, and export your data whenever you want, so you are never locked into a single vendor. You can start from &lt;a href="https://github.com/uptimepage/uptimepage" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest version
&lt;/h2&gt;

&lt;p&gt;The honest version is not "Rust beats Go." It is "for a tool that lives or dies by timing and runs at high concurrency, Rust fit better." Pick the language for the job in front of you. Mine happened to be a job that Rust is very good at.&lt;/p&gt;

&lt;p&gt;If you have shipped a heavily concurrent network service, did you reach for Go or Rust, and did the GC ever actually show up in your tail latencies? Curious where people land when the timing is the product.&lt;/p&gt;

</description>
      <category>rust</category>
      <category>go</category>
      <category>monitoring</category>
      <category>devops</category>
    </item>
    <item>
      <title>Is 98% uptime good? It allows 7.3 days of downtime a year</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Fri, 17 Jul 2026 13:26:37 +0000</pubDate>
      <link>https://dev.to/slima4/is-98-uptime-good-it-allows-73-days-of-downtime-a-year-4dd3</link>
      <guid>https://dev.to/slima4/is-98-uptime-good-it-allows-73-days-of-downtime-a-year-4dd3</guid>
      <description>&lt;p&gt;&lt;em&gt;Cover photo by &lt;a href="https://unsplash.com/@theblowup" rel="noopener noreferrer"&gt;the blowup&lt;/a&gt; on Unsplash.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I build an uptime monitor, so this question lands in my inbox a lot: is 98% uptime good?&lt;/p&gt;

&lt;p&gt;For a public website or a paid API, no. For an internal tool or a side project, it is fine. The difference is one division away.&lt;/p&gt;

&lt;p&gt;98% looks like a top grade because school taught us that 98 out of 100 is excellent. Uptime does not grade like school. The whole scale for public services lives between 99% and 100%, and serious targets differ only in the digits after the decimal. On that scale, 98% sits at the bottom.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turn the percentage into time
&lt;/h2&gt;

&lt;p&gt;The allowed failure at 98% is 2%. Two percent of a year is 7.3 days. Two percent of a 30-day month is 14.4 hours. If your shop makes $2,000 a day, 98% uptime means you accept about $14,600 of closed-door time per year and the target still counts as met.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fooi41uhc3xx4mxhi4qy4.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fooi41uhc3xx4mxhi4qy4.webp" alt="One year at 98% uptime drawn as 52 weeks of day cells: four short red outage runs totalling 7.3 days scattered through a green year." width="800" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here is the same math for every common target. The last column prices the downtime for that $2,000-a-day shop.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Uptime&lt;/th&gt;
&lt;th&gt;Per day&lt;/th&gt;
&lt;th&gt;Per 30-day month&lt;/th&gt;
&lt;th&gt;Per year&lt;/th&gt;
&lt;th&gt;Lost sales per year&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;2.4 hours&lt;/td&gt;
&lt;td&gt;3 days&lt;/td&gt;
&lt;td&gt;36.5 days&lt;/td&gt;
&lt;td&gt;$73,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;95%&lt;/td&gt;
&lt;td&gt;1.2 hours&lt;/td&gt;
&lt;td&gt;36 hours&lt;/td&gt;
&lt;td&gt;18.3 days&lt;/td&gt;
&lt;td&gt;$36,500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;98%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;28.8 minutes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14.4 hours&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7.3 days&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$14,600&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;99%&lt;/td&gt;
&lt;td&gt;14.4 minutes&lt;/td&gt;
&lt;td&gt;7.2 hours&lt;/td&gt;
&lt;td&gt;3.7 days&lt;/td&gt;
&lt;td&gt;$7,300&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;99.5%&lt;/td&gt;
&lt;td&gt;7.2 minutes&lt;/td&gt;
&lt;td&gt;3.6 hours&lt;/td&gt;
&lt;td&gt;1.8 days&lt;/td&gt;
&lt;td&gt;$3,650&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;99.9%&lt;/td&gt;
&lt;td&gt;1.4 minutes&lt;/td&gt;
&lt;td&gt;43 minutes&lt;/td&gt;
&lt;td&gt;8.8 hours&lt;/td&gt;
&lt;td&gt;$730&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;99.95%&lt;/td&gt;
&lt;td&gt;43 seconds&lt;/td&gt;
&lt;td&gt;21.6 minutes&lt;/td&gt;
&lt;td&gt;4.4 hours&lt;/td&gt;
&lt;td&gt;$365&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;99.99%&lt;/td&gt;
&lt;td&gt;8.6 seconds&lt;/td&gt;
&lt;td&gt;4.3 minutes&lt;/td&gt;
&lt;td&gt;52.6 minutes&lt;/td&gt;
&lt;td&gt;$73&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;99.999%&lt;/td&gt;
&lt;td&gt;0.9 seconds&lt;/td&gt;
&lt;td&gt;26 seconds&lt;/td&gt;
&lt;td&gt;5.3 minutes&lt;/td&gt;
&lt;td&gt;$7&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two jumps on this table do most of the work in real contracts. From 98% to 99.9%, the allowed downtime drops from 14.4 hours a month to 43 minutes. From 99.9% to 99.99%, it drops from 43 minutes to 4.3 minutes, and that second jump usually costs about ten times more engineering than the first while saving the shop $657 a year.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the downtime matters
&lt;/h2&gt;

&lt;p&gt;98% per month is 14.4 hours, but the number says nothing about how those hours land.&lt;/p&gt;

&lt;p&gt;Thirty minutes of planned maintenance every night at 03:00 adds up to 98% and most users never notice. One 14-hour outage on the day of your product launch is also 98%. Same score, very different month.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiov47zlspdeu5bzsmf3j.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiov47zlspdeu5bzsmf3j.webp" alt="Two 30-day bars that both score 98% uptime: thin red ticks every night versus one 14.4-hour red block on launch day." width="800" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So a single uptime percentage is a summary, not the full story. When someone quotes you a number, also ask about the longest single outage and when it happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  When 98% is enough
&lt;/h2&gt;

&lt;p&gt;Plenty of systems can live at 98% and nobody gets hurt:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An internal wiki. People retry after lunch.&lt;/li&gt;
&lt;li&gt;A staging environment. Downtime there is often planned.&lt;/li&gt;
&lt;li&gt;A batch job that builds reports at night. It has hours of slack before anyone reads the output.&lt;/li&gt;
&lt;li&gt;A home server on a residential connection. Your power company already decided your uptime for you.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The shared pattern: when these go down, nobody loses money and nobody loses trust. Paying for more nines there is waste.&lt;/p&gt;

&lt;h2&gt;
  
  
  When it is not
&lt;/h2&gt;

&lt;p&gt;A checkout page, a paid API, a login service. Here 98% fails twice. First the direct cost: 14.4 hours a month of failed requests and support tickets. Second the trust cost, which is larger and slower. A customer who hits your outage twice in one week does not check your uptime report. They remember that your service is the one that breaks.&lt;/p&gt;

&lt;p&gt;One more trap: an SLA is not uptime. Uptime is the measured number. An SLA is a contract promise with a penalty, and the penalty is almost always a service credit. If a provider misses its 99.9% SLA, you get part of your bill back. Your customers get nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to aim for instead
&lt;/h2&gt;

&lt;p&gt;99.9% is the default target for customer-facing services for a reason. 43 minutes a month is enough room for a bad deploy and a couple of small failures, and a small team can hit it without heroics. Above that, each nine costs roughly ten times more and most users cannot feel the difference.&lt;/p&gt;

&lt;p&gt;If you want to run your own numbers, I keep a free &lt;a href="https://uptimepage.dev/tools/uptime-sla-calculator" rel="noopener noreferrer"&gt;uptime SLA calculator&lt;/a&gt; and an &lt;a href="https://uptimepage.dev/tools/error-budget-calculator" rel="noopener noreferrer"&gt;error budget calculator&lt;/a&gt; on our site; both work without signup.&lt;/p&gt;

&lt;p&gt;What target do you actually run in production, and did you pick it or inherit it? I am curious how many teams measured before they promised.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>monitoring</category>
      <category>beginners</category>
    </item>
    <item>
      <title>The status page you can't fake: measured uptime, not published</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Wed, 15 Jul 2026 13:50:26 +0000</pubDate>
      <link>https://dev.to/slima4/the-status-page-you-cant-fake-measured-uptime-not-published-ckp</link>
      <guid>https://dev.to/slima4/the-status-page-you-cant-fake-measured-uptime-not-published-ckp</guid>
      <description>&lt;p&gt;A status page is the one dashboard a company publishes about its own service. It is also the one place where the company has a reason to look good. That is a problem. If the page can be edited to look better than reality, it stops being useful. So when you build or choose a status page, ask one thing: can someone hide a real outage on it?&lt;/p&gt;

&lt;p&gt;The short answer: a status page you can trust builds its uptime bar from real checks, not from the incidents a person chose to publish.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two layers, and only one is yours to edit
&lt;/h2&gt;

&lt;p&gt;A good status page has two parts. The first part is the measured data: the green and red timeline, the 90-day bar, the uptime percent. The second part is incidents: the notes a person writes to explain what broke, what they are doing, and when it is fixed. On the page they sit next to each other and look the same. They are not the same, and mixing them is where trust gets lost.&lt;/p&gt;

&lt;p&gt;The measured data answers one question: what did the checks see? The incident notes answer a different one: what does the team want to say about it? The first is a fact. The second is a story. A status page you can trust lets the team write the story, but keeps them away from the facts.&lt;/p&gt;

&lt;p&gt;The story layer also includes the postmortem, the write-up you post after an outage. A postmortem is honesty you add on purpose. You explain what broke and why, because you choose to. The bar works the other way. It shows the failure on its own, whether you write anything or not. You control the story. You do not control the facts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trap: a bar that only shows what you published
&lt;/h2&gt;

&lt;p&gt;The common mistake is to build the uptime bar from incidents. It feels natural. You already open an incident when something breaks, so why not color the timeline from incidents too? Now the bar turns red only where an incident exists and is marked public.&lt;/p&gt;

&lt;p&gt;The problem comes on the day an outage has no public incident. Maybe no one published it. Maybe a setting was wrong. Maybe the monitor was added to the page after the outage, so its incident was saved as private and never checked again. In every case the timeline shows green over a real red day, and it does this quietly. The uptime number goes up. The customer sees 100 percent over a week they remember as broken.&lt;/p&gt;

&lt;p&gt;One version of this trap is easy to build by accident, so it is worth explaining. You decide "is this incident public?" once, at the moment the incident opens, based on whether the monitor was on a public page right then. Then you save that answer and never look again. In the code it looks like a live check. It is really a photo taken one time. Move the monitor onto the page a day later, and its past outages stay hidden, because the photo was taken before the monitor was there. Any yes or no mark that is set once and then trusted forever has this problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix is a rule, not a switch
&lt;/h2&gt;

&lt;p&gt;The bar should come from the measured data. It needs only one rule to stay calm: wait for confirmation before you count downtime. You do not want one failed check from one place to turn a whole day red, because networks are noisy and one bad checker is not an outage. So you wait for agreement: more than one region failing at the same time, for more than one check in a row. That one rule is enough to keep a short blip from becoming a red day.&lt;/p&gt;

&lt;p&gt;One thing matters here: the same rule feeds both your alerts and your uptime bar. If the rule that wakes your on-call person is the same rule that colors the timeline, the two can never tell different stories. The moment you add a second rule just for the bar, it will drift from the first, and the page will disagree with itself. On-call gets paged, but the public history says everything was fine.&lt;/p&gt;

&lt;p&gt;Publishing stays where it belongs, on the notes. The team decides whether to write an incident, what to say, and when to post the all-clear. They do not decide whether last Tuesday was down. The checks already decided that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the status page you already have
&lt;/h2&gt;

&lt;p&gt;You can check any status page, including your own, in a few minutes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Add a monitor to a page after it has already had an outage. Does the history show the outage, or does it start clean from the day you added the monitor?&lt;/li&gt;
&lt;li&gt;Take a real incident and unpublish it. Does the bar keep the red day, or does the day turn green?&lt;/li&gt;
&lt;li&gt;Make one region fail for one second. Does the whole day go red, or does the bar stay calm?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A page you can trust shows the outage in the first test, keeps the red in the second, and stays calm in the third. A page that fails these is usually not lying on purpose. It just built its timeline on top of what people chose to publish, and that always has holes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this leaves you
&lt;/h2&gt;

&lt;p&gt;This is how I built it into &lt;a href="https://uptimepage.dev" rel="noopener noreferrer"&gt;Uptimepage&lt;/a&gt;: the 90-day bar and each status light come from confirmed downtime, measured across regions with a confirmation rule, not from what someone chose to publish. Incidents and postmortems are the layer you write by hand. It is AGPL and open source on &lt;a href="https://github.com/uptimepage/uptimepage" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;, and there is a &lt;a href="https://uptimepage.dev/blog/building-an-uptime-monitor-in-rust" rel="noopener noreferrer"&gt;longer write-up on the internals&lt;/a&gt; if you want the Rust side.&lt;/p&gt;

&lt;p&gt;You should not be able to fake your uptime. You should not be able to fake it by accident either. The bar is a measurement. Keep it one.&lt;/p&gt;

&lt;p&gt;How does your status page compute its uptime bar, from checks or from published incidents? Curious what people run and whether anyone has been bitten by the frozen-flag version of this.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>sre</category>
      <category>monitoring</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Error budgets, explained: SLOs, burn rate, and when to stop shipping</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Mon, 13 Jul 2026 14:11:33 +0000</pubDate>
      <link>https://dev.to/slima4/error-budgets-explained-slos-burn-rate-and-when-to-stop-shipping-27a0</link>
      <guid>https://dev.to/slima4/error-budgets-explained-slos-burn-rate-and-when-to-stop-shipping-27a0</guid>
      <description>&lt;h2&gt;
  
  
  The idea
&lt;/h2&gt;

&lt;p&gt;An uptime target has a second number hidden inside it. Say you promise 99.9% uptime. You are also saying that 0.1% is allowed to fail. Put that 0.1% into real time, and that is your error budget: the downtime you can have before you break the promise.&lt;/p&gt;

&lt;p&gt;This changes how you look at downtime. It stops being a mistake to feel bad about and becomes a budget you can spend: on a risky deploy, on a slow service you depend on, or on a migration. When the budget runs low you slow down. When it is healthy you can move fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  The formula
&lt;/h2&gt;

&lt;p&gt;An SLO is your reliability target, for example 99.9%. The gap to 100% is the failure you are allowed. Multiply that gap by the length of the window, and you get real time you can spend.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;error budget = window x (1 - SLO)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A 99.9% SLO over a 30-day month is 2,592,000 seconds times 0.001. That is 43 minutes and 12 seconds. That is the whole budget for the month, shared across every incident, not a fresh 43 minutes each day.&lt;/p&gt;

&lt;p&gt;So three nines is not "never go down". It is a 43-minute budget each month. Every extra nine costs about ten times more engineering, for downtime that most users never notice. This is why Google says that 100% is the wrong target for almost every service.&lt;/p&gt;

&lt;h2&gt;
  
  
  Burn rate is the speed
&lt;/h2&gt;

&lt;p&gt;The total tells you how much you can spend. It does not tell you how fast. Two services can both sit at 99.9% for the month: one loses the budget slowly, the other loses it all in a single bad hour. Burn rate tells them apart.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;burn rate = (1 - measured) / (1 - SLO)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A burn rate of 1x spends the whole window exactly. 2x spends it in half the time. Below 1x, you finish the month with budget left. Above 1x, the rate tells you the deadline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;budget runs out in = window / burn rate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At 2x, a 30-day budget is gone in fifteen days. At 14.4x it is gone in about two days. That same 14.4x spends 2% of the budget in one hour, which is the level most fast alerts are set to.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fipp0rxzzl2709nfrdnix.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fipp0rxzzl2709nfrdnix.webp" alt="Burn-down chart at 99.0% measured against a 99.9% SLO: an amber line hits zero after about a tenth of the month, labelled gone in 3d." width="799" height="302"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Turning burn rate into alerts
&lt;/h2&gt;

&lt;p&gt;One threshold is not enough. It either alerts too late or it sends too many false alarms. The common fix uses two windows: a long one to confirm the problem is real, and a short one to clear the alert quickly once you fix it. Both have to be burning for the alert to fire. For a 30-day budget, these are the usual settings:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fast page: 2% of the budget in 1 hour (with a 5-minute short window). This is a 14.4x burn.&lt;/li&gt;
&lt;li&gt;Page: 5% in 6 hours (30-minute short window). This is a 6x burn.&lt;/li&gt;
&lt;li&gt;Slow ticket: 10% in 3 days (6-hour short window). This is a 1x burn.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fast page catches a sudden outage. The slow ticket catches a slow problem that would still use up the whole month if nobody looked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule is the point
&lt;/h2&gt;

&lt;p&gt;The math is the easy part. The value comes from a rule you agree on before anything breaks. The rule is simple. When the budget runs out, risky launches stop, and the team works on reliability until the budget grows back. While the budget is healthy, you ship and you take the risk.&lt;/p&gt;

&lt;p&gt;One more rule keeps it fair. If you never spend your budget, your SLO is too strict, and you are paying for reliability that no user asked for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;I put the formulas into a free, no sign-up &lt;a href="https://uptimepage.dev/tools/error-budget-calculator" rel="noopener noreferrer"&gt;error budget calculator&lt;/a&gt;: enter an SLO and your measured availability, and it shows budget spent, budget left, burn rate, and a burn-down chart. There is also an &lt;a href="https://uptimepage.dev/tools/uptime-sla-calculator" rel="noopener noreferrer"&gt;uptime SLA calculator&lt;/a&gt; for the full downtime-per-nine table.&lt;/p&gt;

&lt;p&gt;What SLO and burn-rate thresholds do you run in production? Curious how others pick them.&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>reliability</category>
      <category>monitoring</category>
    </item>
    <item>
      <title>ClickHouse system tables ate my disk (and the fix)</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Thu, 09 Jul 2026 15:13:13 +0000</pubDate>
      <link>https://dev.to/slima4/clickhouse-system-tables-ate-my-disk-and-the-fix-827</link>
      <guid>https://dev.to/slima4/clickhouse-system-tables-ate-my-disk-and-the-fix-827</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A monitoring alert said I was dropping check results. The disk was 100% full. My actual data was 20 MB. ClickHouse had quietly written 12 GB of logs about itself, mostly &lt;code&gt;system.text_log&lt;/code&gt; and &lt;code&gt;system.trace_log&lt;/code&gt;. Those same logs also burn CPU at idle. The fix is a few lines of ClickHouse config that disable the noisy logs and slow the metrics collector. Full config is below.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The alert
&lt;/h2&gt;

&lt;p&gt;One morning a critical alert fired: &lt;code&gt;UptimepageResultsLost&lt;/code&gt;. Its description is blunt: storage write failures or dropped results, checks run but results are not persisting. Then it did something worse than fire once. It cleared, fired again, cleared, and fired again, over and over.&lt;/p&gt;

&lt;p&gt;I run Uptimepage, an uptime monitoring service. "Results not persisting" means the one thing customers pay for, recording whether their sites are up, might be failing. So it had my full attention.&lt;/p&gt;

&lt;p&gt;The good news first: no data was lost. The write path retries, and every failed write was caught by a retry. I keep a counter for results that actually get dropped, and it stayed at zero the whole time. But something was clearly wrong underneath.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real signal: a full disk
&lt;/h2&gt;

&lt;p&gt;The first real signal came from Postgres, which had crashed and restarted a couple of minutes earlier:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FATAL: could not write lock file "postmaster.pid": No space left on device
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The disk was 100% full. Postgres could not write its lock file, crashed, and recovered on its own through WAL replay. ClickHouse, &lt;a href="https://uptimepage.dev/blog/postgres-vs-clickhouse-uptime-monitor" rel="noopener noreferrer"&gt;where I store raw check results&lt;/a&gt;, was rejecting inserts with its own version of the same complaint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Code: 243. DB::Exception: Cannot reserve 1.00 MiB, not enough space:
While executing WaitForAsyncInsert. (NOT_ENOUGH_SPACE)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So the alert was a symptom. The real problem was a full disk. That reframes the question: I do not store much, so what filled it?&lt;/p&gt;

&lt;h2&gt;
  
  
  Leak one: old Docker images
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;docker system df&lt;/code&gt; gave the first half of the answer:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2qai3wbkmvs4chrtqc14.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2qai3wbkmvs4chrtqc14.webp" alt="Docker disk usage broken down: images 16.2 GB, volumes 13.8 GB (Postgres and ClickHouse data), build cache 3.8 GB, containers 2.5 GB"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;88 Docker images, only two of them in use. Every deploy pulls a fresh image, and nothing pruned the old ones, so they piled up for weeks. A &lt;code&gt;docker image prune -af&lt;/code&gt; reclaimed about 10 GB and took the disk off the ceiling.&lt;/p&gt;

&lt;p&gt;That stopped the bleeding. But 13.8 GB of volumes is a lot for a service whose data I thought was tiny. That number turned out to be the real story.&lt;/p&gt;

&lt;h2&gt;
  
  
  Leak two: ClickHouse logging about itself
&lt;/h2&gt;

&lt;p&gt;I went into ClickHouse and asked the obvious question, which table is big:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="k"&gt;database&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;formatReadableSize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bytes_on_disk&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;size&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;rows&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="k"&gt;system&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parts&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;active&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="k"&gt;database&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="k"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bytes_on_disk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The answer stopped me:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Database&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Rows&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;system&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;11.8 GiB&lt;/td&gt;
&lt;td&gt;577,340,927&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;monitor&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;19.7 MiB&lt;/td&gt;
&lt;td&gt;1,813,884&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;monitor&lt;/code&gt; is my data: every check result and rollup I keep. About 20 MB. The &lt;code&gt;system&lt;/code&gt; database, ClickHouse's own diagnostic tables, was 11.8 GiB. Nearly all of the storage was ClickHouse logging about itself.&lt;/p&gt;

&lt;p&gt;Breaking &lt;code&gt;system&lt;/code&gt; down by table:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fanojqdlwd6lrdhlj4d4n.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fanojqdlwd6lrdhlj4d4n.webp" alt="ClickHouse system tables ranked by on-disk size: text_log 5.29 GiB, trace_log 3.02 GiB, part_log 1.11 GiB and smaller logs, next to the actual data table at 2.8 MiB"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two tables did most of the damage. &lt;code&gt;system.text_log&lt;/code&gt; (5.29 GiB) is a copy of the server's own log output written into a table. &lt;code&gt;system.trace_log&lt;/code&gt; (3.02 GiB) is the query profiler, which samples running queries. Both are handy when you are actively debugging ClickHouse. Neither is worth multiple gigabytes when I am not. And &lt;code&gt;text_log&lt;/code&gt; is off by default, so something in my setup had switched it on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;Two parts: reclaim the space now, and stop it coming back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reclaim now.&lt;/strong&gt; The system log tables are throwaway diagnostics, not real data. Truncate them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;TRUNCATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="k"&gt;system&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text_log&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;TRUNCATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;IF&lt;/span&gt; &lt;span class="k"&gt;EXISTS&lt;/span&gt; &lt;span class="k"&gt;system&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;trace_log&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="c1"&gt;-- and the rest of the system.*_log tables&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That gave back 12 GB at once and took the disk from 74% down to 41%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stop the regrowth.&lt;/strong&gt; Truncating is a one-time cleanup. Without a cap, the logs fill right back up. The durable fix is ClickHouse config, added under &lt;code&gt;/etc/clickhouse-server/config.d/&lt;/code&gt;. I watch ClickHouse through Grafana, not these tables. So I disable almost all of them and keep only &lt;code&gt;query_log&lt;/code&gt; and &lt;code&gt;part_log&lt;/code&gt;, both bounded, for the rare hands-on debugging session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight xml"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;clickhouse&amp;gt;&lt;/span&gt;
    &lt;span class="c"&gt;&amp;lt;!-- Sample async metrics every 60s, not every second. --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;asynchronous_metrics_update_period_s&amp;gt;&lt;/span&gt;60&lt;span class="nt"&gt;&amp;lt;/asynchronous_metrics_update_period_s&amp;gt;&lt;/span&gt;

    &lt;span class="c"&gt;&amp;lt;!-- Disable the log tables. remove="1" on an absent table is a no-op,
         so this list is safe to paste as-is across ClickHouse versions. --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;asynchronous_metric_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;asynchronous_insert_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;backup_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;error_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;crash_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;metric_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;query_metric_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;query_thread_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;query_views_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;session_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;text_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;trace_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;opentelemetry_span_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;zookeeper_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;processors_profile_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;latency_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;background_schedule_pool_log&lt;/span&gt; &lt;span class="na"&gt;remove=&lt;/span&gt;&lt;span class="s"&gt;"1"&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;

    &lt;span class="c"&gt;&amp;lt;!-- Keep query_log and part_log, bounded, for on-hand debugging. --&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;query_log&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;ttl&amp;gt;&lt;/span&gt;event_date + INTERVAL 3 DAY DELETE&lt;span class="nt"&gt;&amp;lt;/ttl&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;max_size_rows&amp;gt;&lt;/span&gt;1048576&lt;span class="nt"&gt;&amp;lt;/max_size_rows&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/query_log&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;part_log&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;ttl&amp;gt;&lt;/span&gt;event_date + INTERVAL 3 DAY DELETE&lt;span class="nt"&gt;&amp;lt;/ttl&amp;gt;&lt;/span&gt;
        &lt;span class="nt"&gt;&amp;lt;max_size_rows&amp;gt;&lt;/span&gt;1048576&lt;span class="nt"&gt;&amp;lt;/max_size_rows&amp;gt;&lt;/span&gt;
    &lt;span class="nt"&gt;&amp;lt;/part_log&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;/clickhouse&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;remove="1"&lt;/code&gt; disables a system log completely. The &lt;code&gt;&amp;lt;ttl&amp;gt;&lt;/code&gt; element on &lt;code&gt;query_log&lt;/code&gt; and &lt;code&gt;part_log&lt;/code&gt; adds a TTL and keeps the table's default partitioning, so you do not have to restate the whole engine. ClickHouse picks up the change after a restart. Altinity's "system tables ate my disk" note covers the same ground and is worth a read. If you would rather keep the diagnostics, give every log a short &lt;code&gt;&amp;lt;ttl&amp;gt;&lt;/code&gt; instead of disabling it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Bonus: lower CPU, not just disk&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those log tables are written constantly, and on a small server that steady write load shows up as CPU. A long-running ClickHouse issue tracks the server using noticeable CPU at zero load, &lt;a href="https://github.com/ClickHouse/ClickHouse/issues/60016" rel="noopener noreferrer"&gt;#60016&lt;/a&gt;. People there report dropping from 40 to 70% CPU down to about 1.5% after disabling the logs and slowing the async-metrics collector. So those two settings, the &lt;code&gt;asynchronous_metrics_update_period_s&lt;/code&gt; line and the &lt;code&gt;remove="1"&lt;/code&gt; block, pay off twice: less disk and less CPU.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The real lesson: alert on the cause, not the effect
&lt;/h2&gt;

&lt;p&gt;Here is the part that stings. I had an alert for "results are being dropped." I did not have an alert for "the disk is filling up." So a full disk is a slow, predictable problem that builds over days. But it only reached me as a sudden downstream symptom, after Postgres had already crashed once.&lt;/p&gt;

&lt;p&gt;A downstream alert like "results lost" is not a substitute for watching the resource that actually runs out. I added the missing one: a plain host disk-space alert on &lt;code&gt;node_filesystem_avail_bytes&lt;/code&gt;, firing at 80% and 90% used, well before anything starts failing. That is the alert that should have caught this on day one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key takeaways&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ClickHouse system log tables are unbounded by default and can dwarf your real data. Mine were 11.8 GB against 20 MB of actual data.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;system.text_log&lt;/code&gt; and &lt;code&gt;system.trace_log&lt;/code&gt; are the usual offenders. &lt;code&gt;text_log&lt;/code&gt; is off by default, so check whether something enabled it.&lt;/li&gt;
&lt;li&gt;Cap them in config: &lt;code&gt;remove="1"&lt;/code&gt; to disable a log, or &lt;code&gt;&amp;lt;ttl&amp;gt;&lt;/code&gt; plus &lt;code&gt;&amp;lt;max_size_rows&amp;gt;&lt;/code&gt; to bound the ones you keep. Truncate to reclaim space right away.&lt;/li&gt;
&lt;li&gt;It is not just disk. The same logs burn CPU at idle on small servers. Disabling them, plus &lt;code&gt;asynchronous_metrics_update_period_s = 60&lt;/code&gt;, took reporters in ClickHouse issue #60016 from 40 to 70% CPU down to about 1.5%.&lt;/li&gt;
&lt;li&gt;Anything that pulls artifacts on a schedule, Docker images in my case, needs matching cleanup or it becomes a slow disk leak.&lt;/li&gt;
&lt;li&gt;Alert on the cause (disk space), not only the effect (dropped writes). The cause gives you days of warning; the effect gives you minutes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;This happened on &lt;a href="https://uptimepage.dev/blog/clickhouse-system-tables-filled-disk" rel="noopener noreferrer"&gt;Uptimepage&lt;/a&gt;, the uptime monitor I run and dogfood.&lt;/p&gt;

</description>
      <category>clickhouse</category>
      <category>database</category>
      <category>observability</category>
      <category>devops</category>
    </item>
    <item>
      <title>Postgres vs ClickHouse? I use both. 4 tricks from the split.</title>
      <dc:creator>Slim</dc:creator>
      <pubDate>Thu, 09 Jul 2026 05:55:03 +0000</pubDate>
      <link>https://dev.to/slima4/postgres-vs-clickhouse-i-use-both-4-tricks-from-the-split-4420</link>
      <guid>https://dev.to/slima4/postgres-vs-clickhouse-i-use-both-4-tricks-from-the-split-4420</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR.&lt;/strong&gt; My uptime monitor keeps config and incidents in Postgres and check results in ClickHouse. The split is one rule: does a row ever change? Config gets edited, so it wants Postgres transactions and constraints. A check result is written once and never touched again, so it goes to ClickHouse, where the right codec, a careful sort key, and a per-row TTL make billions of rows cheap. Four tricks and a bonus below, useful even if you only ever run one database.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every monitor I run writes a row every 20 seconds, from every region, and never stops. One monitor is about 4,300 rows a day, per region. Multiply by every monitor on the platform and the rows only ever go up.&lt;/p&gt;

&lt;p&gt;That stream would slowly crush a normal Postgres table, and watching it is the whole product. So the check results do not live in Postgres. They live in ClickHouse. Everything else, the monitors and incidents and teams, lives in Postgres. The interesting part is the line between them.&lt;/p&gt;

&lt;p&gt;People ask why not one database. "Postgres vs ClickHouse" is the wrong question, because the two are not fighting over the same job. Here is the line I draw, and the one rule under all of it: split your data by how it is written, not by which engine is faster.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Does the row change? That picks the database
&lt;/h2&gt;

&lt;p&gt;Forget benchmarks for a second. One question sorts a table into one store or the other: after you write a row, will you ever change it?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;one monitor, two stores:

  you edit a monitor  -&amp;gt;  Postgres   (targets, incidents, team)
                          the row changes, has constraints, lives in a transaction

  a probe checks it   -&amp;gt;  ClickHouse (check_results)
                          one row, written once, never updated
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A monitor is a row that changes. You toggle it off, edit the URL, change the interval. So it lives in Postgres:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;targets&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;id&lt;/span&gt;            &lt;span class="n"&gt;UUID&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="n"&gt;gen_random_uuid&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;          &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;check_spec&lt;/span&gt;    &lt;span class="n"&gt;JSONB&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;interval_secs&lt;/span&gt; &lt;span class="nb"&gt;INTEGER&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;CHECK&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;interval_secs&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;enabled&lt;/span&gt;       &lt;span class="nb"&gt;BOOLEAN&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="k"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;updated_at&lt;/span&gt;    &lt;span class="n"&gt;TIMESTAMPTZ&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;CHECK&lt;/code&gt; means a 3-second interval can never reach the table, no matter which part of the app tried to write it. This is Postgres doing the thing it is best at: a small set of rows that must stay correct while many callers change them at once.&lt;/p&gt;

&lt;p&gt;A check result is the opposite. It is written once when a probe finishes, and then it never changes. Nothing ever updates it, so it does not need any of that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;check_results&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;org_id&lt;/span&gt;      &lt;span class="n"&gt;UUID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;target_id&lt;/span&gt;   &lt;span class="n"&gt;UUID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;region&lt;/span&gt;      &lt;span class="n"&gt;LowCardinality&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;String&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="nb"&gt;timestamp&lt;/span&gt;   &lt;span class="nb"&gt;DateTime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'UTC'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;CODEC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DoubleDelta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ZSTD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;      &lt;span class="n"&gt;Enum8&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'up'&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'down'&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'degraded'&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'error'&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;duration_ms&lt;/span&gt; &lt;span class="n"&gt;UInt32&lt;/span&gt; &lt;span class="n"&gt;CODEC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;T64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ZSTD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;ENGINE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;MergeTree&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No &lt;code&gt;updated_at&lt;/code&gt;, no foreign key, no transaction. Just an append-only stream that grows forever. That is the shape ClickHouse is built for and the shape that slowly hurts Postgres.&lt;/p&gt;

&lt;p&gt;Notice the types too. &lt;code&gt;status&lt;/code&gt; is an &lt;code&gt;Enum8&lt;/code&gt;, one byte on disk, not the string &lt;code&gt;"up"&lt;/code&gt;. &lt;code&gt;region&lt;/code&gt; is &lt;code&gt;LowCardinality&lt;/code&gt;, stored once in a dictionary and referenced by a small id instead of repeating the text on every row. On a table that only ever grows, a byte saved per row is a byte saved times billions.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The right codec turns a day of timestamps into almost nothing
&lt;/h2&gt;

&lt;p&gt;Once the results are in a column store, the type on each column is most of the compression, and the default setting wastes a lot of space.&lt;/p&gt;

&lt;p&gt;Look at that timestamp again. A monitor checks every 20 seconds, so for one monitor a day of timestamps is a run of 20, 20, 20, 20, over and over. &lt;code&gt;DoubleDelta&lt;/code&gt; stores the change in the change. For a steady interval the first delta is a constant 20 and the second delta is zero, so each row after the first packs down to about a bit. The codec is not a ClickHouse invention: it comes from Facebook's Gorilla time-series paper, built for exactly this, measurements taken at a steady rate.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="nb"&gt;timestamp&lt;/span&gt;   &lt;span class="nb"&gt;DateTime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'UTC'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="n"&gt;CODEC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DoubleDelta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ZSTD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
&lt;span class="n"&gt;duration_ms&lt;/span&gt; &lt;span class="n"&gt;UInt32&lt;/span&gt;           &lt;span class="n"&gt;CODEC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;T64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ZSTD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
&lt;span class="n"&gt;dns_ms&lt;/span&gt;      &lt;span class="k"&gt;Nullable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;UInt16&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;CODEC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;T64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ZSTD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I also store the timestamp at whole seconds on purpose. The smallest interval is 20 seconds, so no two checks for one monitor ever land in the same second, and the sub-second detail lives in &lt;code&gt;duration_ms&lt;/code&gt; where it belongs. Whole seconds that step by a fixed amount compress far harder than milliseconds would.&lt;/p&gt;

&lt;p&gt;The latency columns get &lt;code&gt;T64&lt;/code&gt; instead, because they are small integers that stay in a small range, a different shape from a steady clock. &lt;code&gt;T64&lt;/code&gt; crops the unused high bits off a block of values, so a &lt;code&gt;UInt32&lt;/code&gt; that never climbs past a few thousand milliseconds stops paying for all 32 bits. That is the trick worth copying, and it is not the specific codec names. It is that a timestamp ticking by a fixed step and a latency staying in a small range are two different shapes, and telling the database which is which does the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Put the tenant in the sort key, not the partition
&lt;/h2&gt;

&lt;p&gt;I got this one wrong first, so learn it from my mistake instead of your own.&lt;/p&gt;

&lt;p&gt;This is a multi-tenant app, so every row carries an &lt;code&gt;org_id&lt;/code&gt; and almost every query filters by it. My first schema partitioned by org, one slot per customer, because it felt tidy. ClickHouse turned that into a flood of tiny parts, the merges could not keep up, and startup got slower the more customers I had. I ripped it out. A partition is not a folder for tidiness. It is a physical unit ClickHouse merges and expires, and you want few large ones, not many small ones.&lt;/p&gt;

&lt;p&gt;Here is what it should be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;ENGINE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;MergeTree&lt;/span&gt;
&lt;span class="k"&gt;PARTITION&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;toYYYYMMDD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;target_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;org_id&lt;/code&gt; leads the &lt;code&gt;ORDER BY&lt;/code&gt;, so a per-org query walks the sorted primary index and reads only that org's slice. But it stays out of &lt;code&gt;PARTITION BY&lt;/code&gt;. The rule I follow now: partition by something low-cardinality that you also delete by, here the day, and put the high-cardinality tenant key in the sort order. Sort key answers "find this org fast". Partition answers "drop old data cheaply", which is the next trick.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Make retention a number in a row, not a schema change
&lt;/h2&gt;

&lt;p&gt;Different plans keep history for different lengths of time. The clumsy way is a migration or a cleanup job per plan. The clean way is to make the retention window a column:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;ttl_days&lt;/span&gt; &lt;span class="n"&gt;UInt16&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt; &lt;span class="n"&gt;CODEC&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ZSTD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;ENGINE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;MergeTree&lt;/span&gt;
&lt;span class="k"&gt;PARTITION&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;toYYYYMMDD&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;target_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;TTL&lt;/span&gt; &lt;span class="nb"&gt;timestamp&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;toIntervalDay&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ttl_days&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each row carries its own &lt;code&gt;ttl_days&lt;/code&gt;, stamped from the org's plan at the moment it is written. A free plan keeps 30 days, a paid plan keeps more, and changing that needs no &lt;code&gt;ALTER&lt;/code&gt;, no migration, no backfill. The next write just stores a different number. The daily partitions line up with the TTL, so ClickHouse expires old data by dropping whole parts, close to free, and a whole column of the same &lt;code&gt;30&lt;/code&gt; compresses away to nothing. The flexibility costs no space.&lt;/p&gt;

&lt;p&gt;One more piece makes reading that history cheap. A materialized view rolls raw checks into per-minute and per-hour summaries as they land:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="n"&gt;MATERIALIZED&lt;/span&gt; &lt;span class="k"&gt;VIEW&lt;/span&gt; &lt;span class="n"&gt;check_results_1m&lt;/span&gt;
&lt;span class="n"&gt;ENGINE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;AggregatingMergeTree&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;target_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;minute&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;SELECT&lt;/span&gt;
    &lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;target_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;toStartOfMinute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;timestamp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="k"&gt;minute&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;countState&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;                &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;total_checks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;countIfState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'up'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;up_checks&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;avgState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;duration_ms&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;       &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;avg_duration_ms&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;check_results&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="n"&gt;org_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;target_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;minute&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So a dashboard showing 30 days across a thousand monitors reads a few thousand minute-buckets, not millions of raw rows. Recent views read the minute rollup, long history reads an hour rollup, and the raw rows underneath expire on their own TTL. The firehose is there when you need to drill into one bad minute, and left alone the rest of the time.&lt;/p&gt;

&lt;h2&gt;
  
  
  One more, because everyone hits it: make inserts idempotent
&lt;/h2&gt;

&lt;p&gt;Here is the trick I wish someone had told me first. My agents batch check results and send them to ClickHouse. A network blip after the server commits but before my side gets the ack means the batch retries and sends the exact same block again. Without protection, that double-counts every row in it, and your uptime numbers quietly drift.&lt;/p&gt;

&lt;p&gt;ClickHouse has a fix built in, but for a plain (non-Replicated) MergeTree it is off until you turn it on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;SETTINGS&lt;/span&gt; &lt;span class="n"&gt;index_granularity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
         &lt;span class="n"&gt;non_replicated_deduplication_window&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The server hashes each inserted block and remembers the last 1,000 hashes. A retry sends an identical block, the hash matches, and the server drops it instead of appending it again. The retry becomes safe to do blindly, so the client stays simple: send, and if unsure, send again. Make the window bigger than the most blocks a single retry could resend, and you stop having to worry about duplicate writes.&lt;/p&gt;

&lt;h2&gt;
  
  
  So, one database or two?
&lt;/h2&gt;

&lt;p&gt;For almost every app, one. If your data fits in Postgres and stays fast, a second database is a cost you should not pay: two schemas, two clients, two things to back up and think about.&lt;/p&gt;

&lt;p&gt;Reach for the second store only when one table stops looking like the rest of your tables. For me that table was &lt;code&gt;check_results&lt;/code&gt;. It is written once and never edited, it grows without end, and every question I ask it is a summary over a time range. That is a different shape from my monitors and my incidents, so it wanted a different database.&lt;/p&gt;

&lt;p&gt;The rule I would give my past self: do not split by "which database is faster". Split by how the data is written. Rows that change and must stay correct want Postgres. An append-only stream you only ever summarize wants a column store like ClickHouse. Most of the tricks above are that one idea pushed down into the schema.&lt;/p&gt;

&lt;p&gt;Uptimepage is open source, AGPL-3.0, and both schemas are in the repo: &lt;a href="https://github.com/uptimepage/uptimepage" rel="noopener noreferrer"&gt;github.com/uptimepage/uptimepage&lt;/a&gt;. The probe that writes those rows is &lt;a href="https://uptimepage.dev/blog/http-prober-in-rust-no-reqwest" rel="noopener noreferrer"&gt;its own post&lt;/a&gt;, and the wider build story, one binary and two databases, is &lt;a href="https://uptimepage.dev/blog/building-an-uptime-monitor-in-rust" rel="noopener noreferrer"&gt;here&lt;/a&gt;. Or &lt;a href="https://uptimepage.dev" rel="noopener noreferrer"&gt;start free on the hosted tier&lt;/a&gt; and point a check at something.&lt;/p&gt;

</description>
      <category>rust</category>
      <category>database</category>
      <category>postgres</category>
      <category>clickhouse</category>
    </item>
  </channel>
</rss>
