<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Saxon Nicholls</title>
    <description>The latest articles on DEV Community by Saxon Nicholls (@saxonnicholls).</description>
    <link>https://dev.to/saxonnicholls</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146055%2Fa5c66bd8-3b5c-4045-b9cf-c17e763b04c1.png</url>
      <title>DEV Community: Saxon Nicholls</title>
      <link>https://dev.to/saxonnicholls</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/saxonnicholls"/>
    <language>en</language>
    <item>
      <title>Logs, Errors, Code, Versions: Why Agentic Debugging Needs All Four</title>
      <dc:creator>Saxon Nicholls</dc:creator>
      <pubDate>Sun, 27 Sep 2026 22:59:46 +0000</pubDate>
      <link>https://dev.to/saxonnicholls/logs-errors-code-versions-why-agentic-debugging-needs-all-four-5bpl</link>
      <guid>https://dev.to/saxonnicholls/logs-errors-code-versions-why-agentic-debugging-needs-all-four-5bpl</guid>
      <description>&lt;p&gt;&lt;strong&gt;Title:&lt;/strong&gt;&lt;br&gt;
Logs, Errors, Code, Versions: Why Agentic Debugging Needs All Four&lt;/p&gt;

&lt;p&gt;Give an AI coding agent your code and your error message, and it debugs like&lt;br&gt;
someone standing on a two-legged stool: upright for a moment, then guessing.&lt;br&gt;
The third leg — logs — tells it what actually happened, not just what should&lt;br&gt;
have happened and what didn't. Most agent setups never hand it over.&lt;/p&gt;

&lt;p&gt;This post is about why that third leg matters, why it needs to be kept&lt;br&gt;
separate from a fourth thing most people don't think of as "debugging data"&lt;br&gt;
at all — versions — and why the failure modes in all of this rhyme with each&lt;br&gt;
other more than you'd expect.&lt;/p&gt;
&lt;h2&gt;
  
  
  1. Collection and interpretation are not the same job
&lt;/h2&gt;

&lt;p&gt;In December 2021, a single defect in &lt;code&gt;log4j-core&lt;/code&gt; — present from the 2.0&lt;br&gt;
betas through 2.14.1 — became one of the worst vulnerabilities the industry&lt;br&gt;
has seen. The mechanism is worth sitting with: log4j didn't just &lt;em&gt;write&lt;/em&gt;&lt;br&gt;
log messages, it &lt;em&gt;interpreted&lt;/em&gt; them. A JNDI lookup embedded in a string you&lt;br&gt;
logged would be resolved and executed, at log-write time, by the logging&lt;br&gt;
library itself. Something as mundane as logging a User-Agent header became&lt;br&gt;
a remote code execution path, because the thing recording your data and the&lt;br&gt;
thing acting on it were the same code, in the same process, on the same&lt;br&gt;
write.&lt;/p&gt;

&lt;p&gt;That's the general lesson, not just a log4j postmortem: &lt;strong&gt;when the layer&lt;br&gt;
that records your data can also act on content embedded in that data, a&lt;br&gt;
hostile string stops being text and becomes a code path.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The architectural answer is not "write better regex" — it's separation.&lt;br&gt;
A collector's only job should be to move bytes verbatim: capture a line,&lt;br&gt;
relay it, store it. No template substitution, no lookups, no execution&lt;br&gt;
triggered by what's &lt;em&gt;in&lt;/em&gt; the line. Interpretation — reading the data,&lt;br&gt;
summarizing it, having an LLM reason over it, deciding whether to alert —&lt;br&gt;
happens later, out of band, over data that's already at rest, as a read,&lt;br&gt;
never as a side effect of collection itself. A hostile string in a log line&lt;br&gt;
can still be alarming to read. It should never be able to make the&lt;br&gt;
&lt;em&gt;collector&lt;/em&gt; do anything.&lt;/p&gt;

&lt;p&gt;This isn't a claim that any particular architecture is immune to every future&lt;br&gt;
bug — it's a claim about which failure class you've structurally ruled out&lt;br&gt;
by not putting an interpreter in the write path.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. Code + error is half the picture — and versions are a third thing agents miss
&lt;/h2&gt;

&lt;p&gt;Most "AI debugs your code" workflows hand an agent two things: the source&lt;br&gt;
and the stack trace. That's pattern-matching against a crash report. It's&lt;br&gt;
not wrong, it's just incomplete — code says what &lt;em&gt;should&lt;/em&gt; happen, the error&lt;br&gt;
says what &lt;em&gt;didn't&lt;/em&gt;, and neither says what &lt;em&gt;did&lt;/em&gt;: which service actually saw&lt;br&gt;
the request first, what the retry logic did before it gave up, what the&lt;br&gt;
previous line was a second before the timeout. That's what logs are for,&lt;br&gt;
and it's why an MCP server that lets an agent call &lt;code&gt;tail_logs&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;search_logs&lt;/code&gt;, and &lt;code&gt;wait_for&lt;/code&gt; directly changes what "debug this" actually&lt;br&gt;
means — the agent reads evidence instead of asking you to paste output into&lt;br&gt;
the chat.&lt;/p&gt;

&lt;p&gt;But there's a fourth axis that code, error, and logs &lt;em&gt;together&lt;/em&gt; still miss:&lt;br&gt;
&lt;strong&gt;what changed underneath you.&lt;/strong&gt; A bug that appears on a Tuesday, on a&lt;br&gt;
machine where a transitive dependency silently auto-upgraded on Monday&lt;br&gt;
night, looks — from code, error, and logs alone — like a mystery. It isn't&lt;br&gt;
a mystery. It's a version that moved. An agent that can correlate "this&lt;br&gt;
broke" against "this is what actually changed in the dependency graph&lt;br&gt;
around that time" is doing real diagnosis instead of pattern-matching&lt;br&gt;
against a stack trace it's seen shaped like this before.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. The devil is in the interleaving
&lt;/h2&gt;

&lt;p&gt;Take five independent processes — five agents, five services, doesn't&lt;br&gt;
matter — each logging into its own place: one's own stdout, a file nobody&lt;br&gt;
tails, nothing at all. Two of them touch the same resource within moments&lt;br&gt;
of each other. Reconstructing what happened means stitching together clocks&lt;br&gt;
that don't agree, from logs that were never meant to be read side by side.&lt;/p&gt;

&lt;p&gt;The bug isn't in any one process's log. It's in the order things actually&lt;br&gt;
happened &lt;em&gt;across&lt;/em&gt; them — and that order is exactly the thing you lose the&lt;br&gt;
moment each stream is captured and read separately. This is the same shape&lt;br&gt;
of problem, at smaller scale, as the version story below: something that is&lt;br&gt;
completely invisible when you look at the parts in isolation, and only&lt;br&gt;
visible when you look at them interleaved.&lt;/p&gt;

&lt;p&gt;Concretely: logging that interleaves by arrival time — not grouped by&lt;br&gt;
source, not reassembled after the fact — means the story reads in the order&lt;br&gt;
it actually happened. That sounds like a small implementation detail. It's&lt;br&gt;
the whole difference between "here are five log files, good luck" and&lt;br&gt;
"here's what happened."&lt;/p&gt;
&lt;h2&gt;
  
  
  4. Why "the manifest was clean" isn't the same as "it works"
&lt;/h2&gt;

&lt;p&gt;Dependabot (and tools like it) answer one question well: is this &lt;em&gt;one&lt;/em&gt;&lt;br&gt;
declared version of this &lt;em&gt;one&lt;/em&gt; package known-bad, against a public advisory&lt;br&gt;
database? That's real, useful, and worth having — we mirror the same class&lt;br&gt;
of advisory feed ourselves for exactly that question, because there's no&lt;br&gt;
reason to reinvent it.&lt;/p&gt;

&lt;p&gt;What that check &lt;em&gt;cannot&lt;/em&gt; see is two versions that each pass it individually&lt;br&gt;
and still fail together at runtime. We tested this rather than argue it:&lt;br&gt;
&lt;strong&gt;197 real combinations of packages, installed and actually run.&lt;/strong&gt; For 12&lt;br&gt;
of the 25 pairs that then failed, the package manager's own compatibility&lt;br&gt;
check — the same class of check Dependabot runs — was clean. Nothing&lt;br&gt;
declared the problem. Only running it did.&lt;/p&gt;

&lt;p&gt;And "known-bad" has a subtler trap: it's easy to assume "latest version" and&lt;br&gt;
"safe version" are the same thing. They aren't always. Try this yourself —&lt;br&gt;
it's public data, not our claim:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;lodash@4.17.21
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That's not an old, abandoned version. It's lodash's actual &lt;em&gt;last-ever&lt;/em&gt;&lt;br&gt;
release. It still carries CVE-2021-23337 (a command-injection issue in&lt;br&gt;
&lt;code&gt;template()&lt;/code&gt;), because upstream's position was that &lt;code&gt;template&lt;/code&gt; was never&lt;br&gt;
meant for untrusted input, and the release train that would have carried a&lt;br&gt;
fix never shipped. "I'm on the latest version" and "I checked once at&lt;br&gt;
install time" both quietly stop being true the moment you stop checking —&lt;br&gt;
dependencies auto-upgrade, base images roll forward, and a scan is a&lt;br&gt;
snapshot, not a subscription. A lockfile tells you what you &lt;em&gt;meant&lt;/em&gt; to&lt;br&gt;
install. It doesn't tell you what's actually running, and it doesn't check&lt;br&gt;
itself again tomorrow.&lt;/p&gt;

&lt;p&gt;That's what we mean by version vigilance: not a one-time scan, but matching&lt;br&gt;
what's &lt;em&gt;actually installed and running&lt;/em&gt;, continuously, against real&lt;br&gt;
published advisory and end-of-life data — and being honest about the&lt;br&gt;
difference between "we checked this specific range" and "we're guessing."&lt;br&gt;
We'd rather say "we don't know" than imply a coverage we don't have.&lt;/p&gt;
&lt;h2&gt;
  
  
  How super-log actually addresses each of these
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Collection and interpretation, kept separate.&lt;/strong&gt; The SDKs are&lt;br&gt;
zero-dependency and MIT — their only job is to capture a line and relay it&lt;br&gt;
to your local hub. Nothing in that path parses, templates, or executes&lt;br&gt;
content found &lt;em&gt;inside&lt;/em&gt; a log line. Reading and reasoning over what's been&lt;br&gt;
captured — including anything an LLM does with it over MCP — happens&lt;br&gt;
afterward, as a read against data already at rest, never as a side effect&lt;br&gt;
of writing it. That's the structural answer to "the collector shouldn't be&lt;br&gt;
able to act on what it collects." This code is easy to analyse and completely transparent being published on GitHub.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The third leg, today; the fourth, in progress.&lt;/strong&gt; The MCP server&lt;br&gt;
(&lt;code&gt;tail_logs&lt;/code&gt;, &lt;code&gt;search_logs&lt;/code&gt;, &lt;code&gt;wait_for&lt;/code&gt;, &lt;code&gt;stream_guide&lt;/code&gt;) is live and free —&lt;br&gt;
one command, and an agent can read your actual logs instead of asking you&lt;br&gt;
to paste them in. Correlating "this broke" against "this is what changed in&lt;br&gt;
your dependency graph" is the part we're building now, not shipped yet — a&lt;br&gt;
version-tracking layer that reads what's actually installed and running,&lt;br&gt;
not just your manifest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Interleaving, solved at the collection layer.&lt;/strong&gt; This one's live and&lt;br&gt;
it's the core of the product: every stream you point at the hub — your app,&lt;br&gt;
your GPU, your build, your agents, however many of them — lands on one&lt;br&gt;
bench, ordered by arrival, not grouped by source and reconciled afterward.&lt;br&gt;
The story reads in the order it actually happened because it was captured&lt;br&gt;
that way, not reconstructed from clocks that don't agree.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Version vigilance — what's free today, what Cloud adds.&lt;/strong&gt; The open&lt;br&gt;
bench already logs when anything on your machine changes version, for free,&lt;br&gt;
whether or not you ever pay us anything. What we're building on top: reading&lt;br&gt;
that timeline and telling you what a change &lt;em&gt;means&lt;/em&gt; — when a runtime stops&lt;br&gt;
getting security fixes, which versions you actually run have a published&lt;br&gt;
CVE against them, where two machines differ by the one version that&lt;br&gt;
explains your bug. Free tells you what you have and when it changed. Cloud&lt;br&gt;
will tell you what's wrong with it. &lt;/p&gt;

&lt;p&gt;Try the part that's live:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add super-log &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; @super-log/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
        &lt;div class="c-embed__cover"&gt;
          &lt;a href="https://super-log.com/" class="c-link align-middle" rel="noopener noreferrer"&gt;
            &lt;img alt="" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsuper-log.com%2Fbrand%2Fsocial-1280x640.png" height="auto" class="m-0"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="c-embed__body"&gt;
        &lt;h2 class="fs-xl lh-tight"&gt;
          &lt;a href="https://super-log.com/" rel="noopener noreferrer" class="c-link"&gt;
            super-log — every stream, one bench
          &lt;/a&gt;
        &lt;/h2&gt;
          &lt;p class="truncate-at-3"&gt;
            Your bench's journal, saved off the machine, readable by your team and its agents.
          &lt;/p&gt;
        &lt;div class="color-secondary fs-s flex items-center"&gt;
            &lt;img alt="favicon" class="c-embed__favicon m-0 mr-2 radius-0" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsuper-log.com%2Fbrand%2Ffavicon.svg"&gt;
          super-log.com
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;




&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/saxonnicholls" rel="noopener noreferrer"&gt;
        saxonnicholls
      &lt;/a&gt; / &lt;a href="https://github.com/saxonnicholls/super-log" rel="noopener noreferrer"&gt;
        super-log
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      One hub for every log stream you have — devices, servers, containers, browsers, chains, apps
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/saxonnicholls/super-log/image/README/1788338555134.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fsaxonnicholls%2Fsuper-log%2FHEAD%2Fimage%2FREADME%2F1788338555134.png" alt="1788338555134"&gt;&lt;/a&gt;# super-log&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/saxonnicholls/super-log/actions/workflows/ci.yml" rel="noopener noreferrer"&gt;&lt;img src="https://github.com/saxonnicholls/super-log/actions/workflows/ci.yml/badge.svg" alt="ci"&gt;&lt;/a&gt;
&lt;a href="https://github.com/saxonnicholls/super-log/LICENSE" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/84ba0b50ad44e854f0382b3a99afaef96f3d4db9e861686a3297ccd3bd397de7/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e63652d4d49542d626c75652e737667" alt="licence: MIT"&gt;&lt;/a&gt;
&lt;a href="https://github.com/saxonnicholls/super-log/releases" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/b66c4b63d5287c73b1f606220be5585da24f7798ffb1941eccafd99ff90a9c7c/68747470733a2f2f696d672e736869656c64732e696f2f6769746875622f762f7461672f7361786f6e6e6963686f6c6c732f73757065722d6c6f673f6c6162656c3d72656c6561736526736f72743d73656d766572" alt="release"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://mcptoplist.com/server/com.super-log%2Fsuper-log" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/c101deed634015eeea8f2dbeb6521d74240485e07f7e51f756d60cce08fe1ad2/68747470733a2f2f6d6370746f706c6973742e636f6d2f62616467652f636f6d2e73757065722d6c6f6725324673757065722d6c6f672e737667" alt="MCP Toplist"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One hub for every log stream you have — devices, servers, containers
browsers, chains and apps.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This project is a consolidation of a patchwork of tools I have used, in
one form or another, over the last fifteen years — the log mergers, port
watchers, build wrappers, ad-hoc proxies and one-off scripts every
long-running bench accumulates — rebuilt here as one coherent thing, on
one wire protocol, with one screen.&lt;/p&gt;
&lt;p&gt;Free and self-hosted, forever. It &lt;strong&gt;collects and consolidates&lt;/strong&gt;; analysis
is a separate, cleaner concern — hand the consolidated stream to
&lt;a href="https://super-log.com" rel="nofollow noopener noreferrer"&gt;super-log.com&lt;/a&gt; for real-time LLM analysis and team
features, or to your own store. See
&lt;a href="https://github.com/saxonnicholls/super-log#collection-is-not-analysis--and-that-is-the-whole-design" rel="noopener noreferrer"&gt;Collection is not analysis&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/saxonnicholls/super-log/assets/bench-overview.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fsaxonnicholls%2Fsuper-log%2FHEAD%2Fassets%2Fbench-overview.png" alt="Twelve streams interleaved on one screen"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Twelve producers on one screen, interleaved by arrival: C++ through both
SN_LOG and spdlog, Rust, Go, Python, Swift, Fortran, a POSIX shell script
two React Native devices, Metal GPU work reporting real bandwidth, and a
live Binance WebSocket.&lt;/em&gt;…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/saxonnicholls/super-log" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


</description>
      <category>logging</category>
      <category>security</category>
      <category>ai</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
