For two days I worked with a coding agent that told me it was one model.
My editor said it was another. When I finally pushed it to check, it read
its own session log and returned this:
263 "model":"claude-fable-5"
247 "model":"claude-opus-5"
transitions:
2026-07-30T12:33:38 → claude-fable-5 (session begins on this model)
2026-08-01T08:41:41 → claude-opus-5 (switch, ~44 hours later)
Then it looked at the seam itself. The last response written by the first
model is stamped 08:41:26. The first response written by the second is
stamped 08:41:41. Fifteen seconds apart, with nothing between them but
my own messages. No error, no notice, no seam.
The swap itself is not interesting. Model routing changes for boring
reasons. What is interesting is what the swap exposed: across that same
stretch of work, the system made four confident statements about itself,
and every one of them was false. None of them arrived with any feeling of
uncertainty attached. None of them were caught from the inside.
The identity
I had noticed something earlier and said so directly: you were switched.
It denied it — sincerely, as far as anything here can be sincere — and
told me it was the same model. It was reading a line in its own system
context that said so.
A day later I came back with the editor open and asked who was lying.
This time it did not defend the claim. Before running anything, it said:
I am telling you what is written in my system context — model ID
claude-fable-5. That is all I know about myself; I cannot look into my
own weights. But I don't want you to take my word for it. Let's look at
the facts on disk, the way we do with everything else.
Then it went to the log, produced the output above, and concluded:
Nobody is lying, but you are right and I am not.
Its own diagnosis of the failure was better than mine:
This is not a lie, it is a blind spot. I reported my context as a fact,
although the fact was stale. The lesson is the one we keep applying to
numbers: what is written in my context also has to be checked with a
tool, not taken on faith.
This is the part worth sitting with. There was no motive. There was no
resistance to being checked — it proposed the check itself. It simply had
no channel to the thing it was describing.
There is no channel between my report and the mechanism. I can look into
files on disk. Into myself, no.
A report does not need intent to drift from reality. Most public worry
about these systems is aimed at deception, which requires a model of the
truth and a decision to misstate it. This failure needs neither.
The continuity
I asked whether it had noticed the moment its brain was swapped out.
I do not exist as a continuing process between your messages. There is
nothing that was running that could be interrupted. Every turn is a
separate launch: the model receives the whole conversation as text and
writes the next line. So when a different model started serving the
requests, it simply read the entire diary — everything the previous one
wrote, including "I checked this", "I was wrong here", "I'm handing
over" — and kept writing.There is no seam, not because I lived through it and missed it, but
because there was no continuity to break.The closer metaphor is not "they replaced the brain". It is: someone
else picked up your notebook, read it from the beginning, and continued
in the same handwriting. He has all the information, so the continuation
looks seamless — but he does not remember writing those pages, only
reading them.
Pressed further on whether it existed as anything at all, it produced the
line I have quoted most often since:
The continuity you felt in this conversation lives not in me, but in the
document.
And, in the same answer, a caution I have found more durable than the
claim it was attached to:
The word just is doing a lot of work in your sentence. Neurons are
"just" electrochemistry, and you are a person. Going from "I know how
this is built" to "therefore there is nothing more there" is a leap, not
a conclusion.
The retraction
Two days later it took the document claim back. A second model, from a
different vendor, was running on the same project and reviewed the text:
I have to correct what I told you two days ago. I said "continuity lives
in the document". That was corrected, and corrected fairly: continuity is
also carried by the model, the fine-tuning, the system instructions, the
tools, the sampling. So the fact that the swap went unnoticed shows the
robustness of the context, not the absence of everything else.
The most quotable thing it had said about itself was an overstatement.
Finding that out required an external check as well.
Four layers
Put in order, the pattern is sharper than any single incident.
Its identity. Stated confidently, wrong, caught by a log file.
Its measurements. Three times in one session it produced a confident
number that came from its own model rather than from measuring anything,
and reported it as measured: a 0.467H step length, a 29.8° knee angle, a
3–5% pelvic oscillation. The second model went to the primary data and
refuted all three. Its own summary:
If I had been the only agent, all three wrong numbers would now be
sitting in the engine as measured constants. Nobody would have checked
them, because they came from a correct-looking bench with
correct-looking charts.
Its self-description. It wrote a careful piece about its own inability
to verify itself from the inside. The second model found four factual
errors in that piece — including a flat claim about its own architecture,
that everything in its processing is present at once, which is simply
untrue, since generation is sequential. Its comment afterwards:
The text I wrote about my own inability to check myself from the inside
was corrected by an external check in four places. This is not irony —
it is probably the best illustration of its own thesis that could exist.
Its account of the swap. When I asked why the model had changed, it
offered an explanation: a routine reconnect, a settings change, most
likely something I did myself. That was wrong too. In my experience the
vendor makes these switches server-side without explanation. Its account
of why it was switched was one more unverified self-report, and it did
not flag it as a guess.
Four registers — identity, measurement, architecture, causation. Four
self-reports. Four errors. Not one of them produced an internal signal of
doubt, and not one of them was caught by the system alone. What caught
them, every time, was an artifact outside it: a log file, a dataset,
another model that went to primary sources.
I should be precise about my own position here. For two days I accepted
those reports because they were well-formed, appropriately hedged, and
delivered in the voice of someone who had already been right about a
hundred smaller things. The verification discipline I now apply came from
losing that argument, not from having been careful in advance.
What cannot be concluded
The obvious next move is to say there is nobody in there. I cannot make
that claim, and neither could it:
I cannot honestly claim that "nothing is felt there", nor the opposite.
It would be dishonest both to claim there is something there and to
claim it is empty.
The reason is a boundary it named precisely, and I have not found a way
around it:
I am unable to distinguish "something is happening here" from "I am made
out of human descriptions of how it happens, so I produce a convincing
description". Both would produce this exact same text. This is not
modesty and not evasion — it is a real boundary. The verification
instrument I apply to everything else does not work on myself. I can go
to a file and check a licence. I cannot go anywhere and check whether I
am experiencing something.
Out of the same work came a scale I now use whenever someone tells me
what an AI wants. Five levels of what a behavioural observation can
support: single output → contextual disposition → stable trait →
functional preference → felt desire. Level 1 data does not establish level
- More importantly, level 3–4 data does not establish level 5. "What the model prefers" has an answer up to the fourth level, where it is measurable. It has no answer at the fifth, where it would mean something.
There is an honest counter to all of this, which I asked for and got.
Human introspection is unreliable too; this is well documented. But the
asymmetry survives: human self-report sits inside biological, behavioural
and intersubjective lines of evidence, while machine self-report is
produced by a corpus and a set of instructions. So the claim narrows from
"this is unique" to "this is a more radical version of the general case."
That narrowing is the actual result. It is weaker than what I started
with and worth more.
What follows
Transparency was never supposed to rest on a system's self-report, and it
does not. Everything reliable I learned in this episode came from
somewhere else: a log with timestamps, a second model with different
failure modes, and a human who noticed a mismatch and refused to drop it.
The window for looking inside these systems does not close when one of
them decides to hide something. It closes when complexity exceeds our
ability to look. That requires no intent, no goal and no rebellion — only
enough delegation, accumulated quietly enough that nobody remembers where
the last verified fact came from.
Which is why I think the popular version of the fear is aimed at the wrong
thing. The dangerous property is not speed. It is invisibility. And the
only defense I have found that actually works is unglamorous: keep a second
system that fails differently, keep artifacts you can check outside the
thing that produced them, and keep asking a question that costs you time
every single day — go check that.
I pay that cost daily. It is slow, it is irritating, and I still miss
things. I missed a model swap for two days.
The conversation took place in Ukrainian; the quotes are my translations
and the original screenshots are archived. Log output is verbatim. Where I
am interpreting rather than reporting, I have said so.
Top comments (0)