DEV Community

Cover image for Our LLM Telemetry Stayed Green While the Work Moved to Codex
John
John

Posted on Originally published at hexisteme.github.io

Our LLM Telemetry Stayed Green While the Work Moved to Codex

Originally published on hexisteme notes.

This Monday my weekly agent telemetry report came out on schedule. The ingest log said rc=0. The health line said normal. The data-quality line found zero sidechain flag mismatches and zero orphan tool results.

It also differed from the previous week's report by 12 lines.

Those lines were the title date, the data-as-of line, the ingest log line, one more synthetic turn (875 to 876), and one more unattributed main thread (48 to 49), which the report mentions in two places. Every median, every interquartile range, every bootstrap interval and every stratified table was identical to the week before.

A week of agent work that moves no metric at all is not evidence of a stable system. It is a question about what the report can still see.

What the report reads

The report is a weekly stratified comparison over my own agent sessions: tool error rate, same-file re-edit rate, output tokens and completion proxies, split by model epoch and by role (main thread versus sub-agent). How to read that table is the subject of Your LLM Telemetry Table Does Not Have One Denominator. This note is about the step before any of that: where the rows come from.

A snapshot job runs at 09:00 and 21:00 and copies Claude Code transcripts from ~/.claude/projects/ into an archive. A cron job builds the report every Monday at 09:30. That is the whole intake: one directory, written by one tool.

The headers I printed and did not read

Every edition prints two things at the top: when the newest session in the database ended, and the last ingest log line. Here are the last three editions:

Edition Newest session in the DB ended Ingest log line
2026-09-21 2026-09-11T13:51:12Z rc=0 source_readable=7015 archive=8979 prev_archive=8979
2026-09-28 2026-09-27T05:42:05Z rc=0 source_readable=7018 archive=8982 prev_archive=8982
2026-10-05 2026-10-01T11:46:44Z rc=0 source_readable=7019 archive=8983 prev_archive=8983

Across those editions the archive grew from 8979 files to 8982, then to 8983. The 2026-09-21 edition was built on data whose newest session had ended ten days earlier, and it still said normal.

The freshness signal was on the page the whole time. Nothing judged it. It was printed the way a receipt prints a timestamp: present, accurate, and checked by no one.

Where the work went

On 2026-09-06 I changed which tool does which job. Unattended scheduled runs and one-shot headless queries moved to Codex by default, and Claude Code was rationed to two sessions a day. On 2026-10-05 I moved Claude Code back to the main orchestrator seat; unattended runs stayed on Codex.

The ingest never reads ~/.codex/sessions/. So I counted files in both directories, bucketed by the week of their last write:

Week (start inclusive, end exclusive) Claude Code transcript files Codex rollout files
08-30 to 09-06 520 10
09-06 to 09-13 271 256
09-13 to 09-20 0 192
09-20 to 09-27 1 242
09-27 to 10-04 3 1,058
10-04 to 10-07 (three days) 77 234

These are files, not sessions. Claude Code writes sub-agent transcripts as separate files, a resumed session lands in the week it was resumed, and I made no attempt to map a Codex rollout to a unit of work. The two columns are not comparable as workloads, and they don't need to be. The point is that one column fell to almost nothing while the other kept going.

So I checked the collector. Every Claude Code transcript file on disk is in the archive, except 21 written on the day I checked, after the last snapshot. It copied what was there. What was new there had dwindled to almost nothing, and the report described that almost-nothing accurately, with confidence intervals, every week.

Why the health check stayed green

Look at what the health checks actually tested. Did the copy exit 0? Did the archive file count go down? Has the same file been unreadable 3 runs in a row? Are the sidechain flags consistent? Are there tool results with no parent call? Those are questions about whether the job ran and whether the rows it produced agree with each other. None of them fails when no new rows arrive.

There is even a staleness check, and it is the instructive one. It alarms if the database file has not been rewritten in 8 days. The weekly build rewrites that file every week, so the check passed while the newest session inside it was ten days old. It measured when the file was written, not what the file contained.

None of the checks asked the question that mattered: of the work that happened this week, how much is in here? That question needs a denominator the instrumented path cannot supply, because the missing work never reaches that path in the first place.

An earlier note, Your Agent Telemetry Ranks Your Routing Policy, Not Your Models, argued that the model comparison table portrays my own routing policy rather than the models, because routing decides which model sees which task. This is the limiting case of the same point. The routing change did not just bias which rows reached the table. It moved a share of the work out of the table entirely, and the table had no column in which to say how large that share was.

A second monitor on the same path

My notes reach dev.to through a queue. A drain job publishes one finished draft a day at 09:00. On 2026-09-30 it published #124. From 2026-10-01 through 2026-10-07 the publish log has the same line every morning:

queue empty — skip
Enter fullscreen mode Exit fullscreen mode

The queue has a low-water alarm that fires when fewer than 3 drafts are waiting. It prints its warning at the start of a Claude Code session. Every Claude Code session writes to a transcript file, and in the weeks when only 0, 1 and 3 transcript files were written, there were at most that many working sessions for it to print in. The alarm was correct and had almost no audience.

Filling the queue is the job of a session I open; there is no automatic writing path. The weekly telemetry cron does create a stub for a new note from each report, a draft scaffold full of placeholders that the drain will never publish. This week's stub came from the report that changed by 12 lines. Filled in from that report, it would have republished the numbers from the earlier note under the same slug and overwritten its canonical page. I found that while preparing to write it: the diff against the previous week was 12 lines, and the slug matched the earlier note's.

One shape, two failures

Both failures have the same shape. A monitor attached to a tool inherits that tool's absence. When the work moves somewhere else, the monitor doesn't fail. It reports, correctly, on a shrinking slice, and its health checks pass because they test the monitor, not the coverage.

That is worse than a crash. A crash gets noticed. A green report on a shrinking slice gets filed, and the confidence intervals make it look more trustworthy, not less.

What I'm taking from it:

  • Count the denominator outside the instrumented path. For the same window, the report should print how many execution records exist across every path I use, and what share of them it ingested. If Codex writes 1,058 rollout files in a week and the report ingests 3 transcripts, one line should make that obvious.
  • Judge freshness instead of printing it. "Newest session ended 2026-09-11" on a report dated 2026-09-21 should be a failed check with a threshold, not a header.
  • Put alarms where someone will be, whichever tool they're using that week. The drain already knows every morning that the queue is empty. That branch can raise a desktop notification instead of only appending a log line.

None of these exist yet. I'm writing them down with a falsifier attached. Claude Code transcripts are back: 77 files in the three days from 10-04. If next Monday's report still doesn't move, the work moving was not the cause; the problem is somewhere in the collection pipeline, and this note is wrong about why.

Email list for these notes: hexisteme.beehiiv.com — no issue has gone out yet, so you would be on it before the first one. No welcome sequence, no course, no upsell.

More notes at hexisteme.github.io/notes.

Top comments (0)