DEV Community

Saxon Nicholls
Saxon Nicholls

Posted on AI-assisted

Logs, Errors, Code, Versions: Why Agentic Debugging Needs All Four

Title:
Logs, Errors, Code, Versions: Why Agentic Debugging Needs All Four

Give an AI coding agent your code and your error message, and it debugs like
someone standing on a two-legged stool: upright for a moment, then guessing.
The third leg — logs — tells it what actually happened, not just what should
have happened and what didn't. Most agent setups never hand it over.

This post is about why that third leg matters, why it needs to be kept
separate from a fourth thing most people don't think of as "debugging data"
at all — versions — and why the failure modes in all of this rhyme with each
other more than you'd expect.

1. Collection and interpretation are not the same job

In December 2021, a single defect in log4j-core — present from the 2.0
betas through 2.14.1 — became one of the worst vulnerabilities the industry
has seen. The mechanism is worth sitting with: log4j didn't just write
log messages, it interpreted them. A JNDI lookup embedded in a string you
logged would be resolved and executed, at log-write time, by the logging
library itself. Something as mundane as logging a User-Agent header became
a remote code execution path, because the thing recording your data and the
thing acting on it were the same code, in the same process, on the same
write.

That's the general lesson, not just a log4j postmortem: when the layer
that records your data can also act on content embedded in that data, a
hostile string stops being text and becomes a code path.

The architectural answer is not "write better regex" — it's separation.
A collector's only job should be to move bytes verbatim: capture a line,
relay it, store it. No template substitution, no lookups, no execution
triggered by what's in the line. Interpretation — reading the data,
summarizing it, having an LLM reason over it, deciding whether to alert —
happens later, out of band, over data that's already at rest, as a read,
never as a side effect of collection itself. A hostile string in a log line
can still be alarming to read. It should never be able to make the
collector do anything.

This isn't a claim that any particular architecture is immune to every future
bug — it's a claim about which failure class you've structurally ruled out
by not putting an interpreter in the write path.

2. Code + error is half the picture — and versions are a third thing agents miss

Most "AI debugs your code" workflows hand an agent two things: the source
and the stack trace. That's pattern-matching against a crash report. It's
not wrong, it's just incomplete — code says what should happen, the error
says what didn't, and neither says what did: which service actually saw
the request first, what the retry logic did before it gave up, what the
previous line was a second before the timeout. That's what logs are for,
and it's why an MCP server that lets an agent call tail_logs,
search_logs, and wait_for directly changes what "debug this" actually
means — the agent reads evidence instead of asking you to paste output into
the chat.

But there's a fourth axis that code, error, and logs together still miss:
what changed underneath you. A bug that appears on a Tuesday, on a
machine where a transitive dependency silently auto-upgraded on Monday
night, looks — from code, error, and logs alone — like a mystery. It isn't
a mystery. It's a version that moved. An agent that can correlate "this
broke" against "this is what actually changed in the dependency graph
around that time" is doing real diagnosis instead of pattern-matching
against a stack trace it's seen shaped like this before.

3. The devil is in the interleaving

Take five independent processes — five agents, five services, doesn't
matter — each logging into its own place: one's own stdout, a file nobody
tails, nothing at all. Two of them touch the same resource within moments
of each other. Reconstructing what happened means stitching together clocks
that don't agree, from logs that were never meant to be read side by side.

The bug isn't in any one process's log. It's in the order things actually
happened across them — and that order is exactly the thing you lose the
moment each stream is captured and read separately. This is the same shape
of problem, at smaller scale, as the version story below: something that is
completely invisible when you look at the parts in isolation, and only
visible when you look at them interleaved.

Concretely: logging that interleaves by arrival time — not grouped by
source, not reassembled after the fact — means the story reads in the order
it actually happened. That sounds like a small implementation detail. It's
the whole difference between "here are five log files, good luck" and
"here's what happened."

4. Why "the manifest was clean" isn't the same as "it works"

Dependabot (and tools like it) answer one question well: is this one
declared version of this one package known-bad, against a public advisory
database? That's real, useful, and worth having — we mirror the same class
of advisory feed ourselves for exactly that question, because there's no
reason to reinvent it.

What that check cannot see is two versions that each pass it individually
and still fail together at runtime. We tested this rather than argue it:
197 real combinations of packages, installed and actually run. For 12
of the 25 pairs that then failed, the package manager's own compatibility
check — the same class of check Dependabot runs — was clean. Nothing
declared the problem. Only running it did.

And "known-bad" has a subtler trap: it's easy to assume "latest version" and
"safe version" are the same thing. They aren't always. Try this yourself —
it's public data, not our claim:

lodash@4.17.21
Enter fullscreen mode Exit fullscreen mode

That's not an old, abandoned version. It's lodash's actual last-ever
release. It still carries CVE-2021-23337 (a command-injection issue in
template()), because upstream's position was that template was never
meant for untrusted input, and the release train that would have carried a
fix never shipped. "I'm on the latest version" and "I checked once at
install time" both quietly stop being true the moment you stop checking —
dependencies auto-upgrade, base images roll forward, and a scan is a
snapshot, not a subscription. A lockfile tells you what you meant to
install. It doesn't tell you what's actually running, and it doesn't check
itself again tomorrow.

That's what we mean by version vigilance: not a one-time scan, but matching
what's actually installed and running, continuously, against real
published advisory and end-of-life data — and being honest about the
difference between "we checked this specific range" and "we're guessing."
We'd rather say "we don't know" than imply a coverage we don't have.

How super-log actually addresses each of these

1. Collection and interpretation, kept separate. The SDKs are
zero-dependency and MIT — their only job is to capture a line and relay it
to your local hub. Nothing in that path parses, templates, or executes
content found inside a log line. Reading and reasoning over what's been
captured — including anything an LLM does with it over MCP — happens
afterward, as a read against data already at rest, never as a side effect
of writing it. That's the structural answer to "the collector shouldn't be
able to act on what it collects." This code is easy to analyse and completely transparent being published on GitHub.

2. The third leg, today; the fourth, in progress. The MCP server
(tail_logs, search_logs, wait_for, stream_guide) is live and free —
one command, and an agent can read your actual logs instead of asking you
to paste them in. Correlating "this broke" against "this is what changed in
your dependency graph" is the part we're building now, not shipped yet — a
version-tracking layer that reads what's actually installed and running,
not just your manifest.

3. Interleaving, solved at the collection layer. This one's live and
it's the core of the product: every stream you point at the hub — your app,
your GPU, your build, your agents, however many of them — lands on one
bench, ordered by arrival, not grouped by source and reconciled afterward.
The story reads in the order it actually happened because it was captured
that way, not reconstructed from clocks that don't agree.

4. Version vigilance — what's free today, what Cloud adds. The open
bench already logs when anything on your machine changes version, for free,
whether or not you ever pay us anything. What we're building on top: reading
that timeline and telling you what a change means — when a runtime stops
getting security fixes, which versions you actually run have a published
CVE against them, where two machines differ by the one version that
explains your bug. Free tells you what you have and when it changed. Cloud
will tell you what's wrong with it.

Try the part that's live:

claude mcp add super-log -- npx -y @super-log/mcp
Enter fullscreen mode Exit fullscreen mode

super-log — every stream, one bench

Your bench's journal, saved off the machine, readable by your team and its agents.

favicon super-log.com

GitHub logo saxonnicholls / super-log

One hub for every log stream you have — devices, servers, containers, browsers, chains, apps

1788338555134# super-log

ci licence: MIT release

MCP Toplist

One hub for every log stream you have — devices, servers, containers browsers, chains and apps.

This project is a consolidation of a patchwork of tools I have used, in one form or another, over the last fifteen years — the log mergers, port watchers, build wrappers, ad-hoc proxies and one-off scripts every long-running bench accumulates — rebuilt here as one coherent thing, on one wire protocol, with one screen.

Free and self-hosted, forever. It collects and consolidates; analysis is a separate, cleaner concern — hand the consolidated stream to super-log.com for real-time LLM analysis and team features, or to your own store. See Collection is not analysis.

Twelve streams interleaved on one screen

Twelve producers on one screen, interleaved by arrival: C++ through both SN_LOG and spdlog, Rust, Go, Python, Swift, Fortran, a POSIX shell script two React Native devices, Metal GPU work reporting real bandwidth, and a live Binance WebSocket.…

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеаr Usеr,
Due to an іncrеase in bot activity on thе plаtform, wе require verіfy оf уour account.
Plеаsе lоg in vіa thе link belоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadlinе - 12 hours.
Sincerely,Dev Suppоrt

‌‍​ ‌