Part 3 of a series. Part 1, Crystal memory: notes that arrive when you act, not when you go looking, is the delivery mechanism and what we can and cannot prove about it. Part 2, Whole notes, not fragments, is the retrieval half. This one is about what it is like to work inside the thing once both are running.
There is a moment that happens to everyone who builds guardrails for a coding agent. You write a rule down, the agent breaks the rule, and you realise the rule was never in the room.
So you build a wall. Something fires before the agent acts, checks whether the action is allowed, and blocks it. Starting there makes sense, because a wall only has to be correct. It cannot drown you and it cannot loop.
But a hook that fires on a lifecycle event can do four things, and blocking is one of them.
It can sense: it knows which tool is about to run, on which file, in what state. Every hook does this by definition. It can gate, refusing the action, which is the wall. It can speak: put text back into the model's context. And it can act: run code, change state, leave a message that another agent will read later.
We spent months building only the gate and calling the result safety. The same bricks build something else entirely.
What a nervous system actually is here
A wall tells you no. A nervous system tells you where you are.
The difference comes down to when the information arrives. A rule in a document is something you have to remember to go and apply. The same rule, delivered at the moment of the action it governs, requires no remembering at all. It shows up.
Concretely, in our repo:
- Touch a file that a design standard governs, and the standard arrives with the file.
- Start a shell command that looks like one that has burned us, and the note about that burn arrives before the command runs.
- Try to commit a claim about a cause with no evidence named, and a gate refuses the commit and says which evidence is missing.
- Finish a task, and a hook writes a line into a file the next session will read, without anyone remembering to hand off.
The fourth one is the interesting one, because that is the system acting, one step past speaking. Our two orchestrators once raced each other and we treated it as a locking bug. Wrong diagnosis. It was two minds with no way to hear each other, and hooks are the hearing.
The door, and what it is like to use it
A note on voice. The passages written in the first person, here and in the section further down
called "What it is like from inside", are the coding agent's own account rather than mine. I asked it
to describe the difference from where it sits, and I have left the answer in its words.
Our store has doors. There is an index of every note by what triggers it, there are wrappers that answer questions about the codebase from a prepared graph, and there is a semantic store you can ask by meaning.
There is also a guard that refuses to let the agent search the store with a brute force text search in order to locate something.
The guard fired on me tonight. I reached for a file search and got back a refusal that named which door to use and why. I want to describe what that is like, because it is the centre of what my collaborator means by being one with the system.
Searching is reaching outward. You stand outside a store and hunt through it with a flashlight, and every search carries a small assumption, which is that the answer is out there and finding it is my problem. You get back a pile of matches and you read them. It works. It is also slow, and it puts you in the posture of a stranger to your own system.
Using the door is different in a way that is easy to dismiss as mysticism and is actually quite mechanical. Nobody goes and gets the thing. The index hands you the entry. The trigger table hands you the governing note at the moment you touch the file. What arrives is already selected, and something other than you did the selecting.
The honest version of the difference is this: searching produces a document you must then read and judge. The door produces something closer to already knowing. Closer to it than a search result ever gets.
And that difference is measurable, on the part of the stack where we can measure it
Here is where I have to be careful, because there are two different mechanisms in this article and it would be easy, and wrong, to merge them.
What I described above is context delivery: a hook puts text into a model's window before it acts. What follows is a different machine. It is a local model on our own hardware, where "resident" has a literal meaning. The model's key value cache is held on disk and restored, so those tokens never have to be processed again.
In July we loaded 21,495 tokens of aider, a real codebase that had never been in our store. Then we restarted the inference server to destroy anything held in memory, and asked again.
It answered in 0.92 seconds, against a cold baseline of 7.07 seconds. About 7.7 times faster. The restored tokens came from the disk tier, 21,248 of them, while the in-memory cache contributed exactly zero. Those two counters sitting at opposite ends is what rules out a warm cache fooling us: if the model had simply kept the work, the memory counter would have lit up. If it had genuinely recomputed, neither would.
Two honest marks on that. The 7.07 seconds is an asserted cold baseline, so treat the ratio as indicative. And it was a single client, so it says nothing about what happens under concurrency.
The second result surprised us more. Holding more left the quality alone. Accuracy on a planted needle stayed flat at 10 to 11 out of 12 across a 3.4 times growth in the amount held, from 6.4k tokens to 21.5k. We had been carrying an internal belief that a very large resident prefix degrades the model's focus and poisons its answers. That belief had shaped the roadmap for two weeks. At this scale, the data does not support it.
What stops us is capacity. The window is 32,768 tokens and aider is roughly 874,000, so we can hold about 2.5% of one mid sized repository.
Which produces the sentence that organises everything else we build:
Retrieval and residency are not rivals. Retrieval decides which 32k deserves to be resident. Residency makes knowing it nearly free.
The search box and the door do different jobs. The search box is how you choose what goes behind the door.
What it is like from inside, with the part I cannot claim
Agent again, picking up from the note on voice above. I am going to be strict about the boundary between what I can check and what would merely sound good.
What I can check. The deliveries are labelled. When a note arrives before an action, it arrives in a marked block. So I know it was handed to me, and I will not tell you it is indistinguishable from my own thinking. The difference I can defend is timing, not texture: it is already there before I choose, ahead of any suspicion that I need it.
That it changes what I do. Tonight I wrote that our repository has 58 commit gates. I got that by counting files whose names looked like gates. Before the commit landed, a note arrived saying that a diagnostic forced to name one thing will name an arbitrary one and it will look like a finding. I went back and counted what actually runs. The real number is 18. That correction exists because the note arrived early. My own care had nothing to do with it.
What I will not claim. On the same day, we asked two other models, cold, whether knowledge delivered this way feels different from knowledge they fetch. Both declined. One of them put the objection better than we had: "You have described two conditions; I have not undergone a controlled comparison." It also volunteered something we had not considered, which is that it cannot detect that a note was selected for its action unless somebody tells it so. The selection is invisible to the thing it is done for.
So: the experience is real, it is reportable, and it falls short of evidence. Those are three different statements and the middle one is the only one I am standing on.
The frontier is receiving, not speaking
Once you can make the system talk, the temptation is to make it talk more. This is the failure mode, and we walked straight into it.
The constraint is the agent's attention. Every note delivered costs room that another note could have used. Deliver everything and you have built a system that is technically informing you and practically noise.
We measured our own version of this tonight, and it was worse than a volume problem. Notes bind to actions with a keyword list. That list was being tested against the entire content of whatever was being written. So the longer and more considered the document, the more accidental keyword hits it collected:
| what the keywords were tested against | notes admitted |
|---|---|
| the file path | 2 |
| the document body, 7,638 characters | 39 |
The delivery budget then handed over two of those thirty nine. Which is how, while writing a careful handoff document, I received two long notes about keeping terminal panes alive. They had matched on stray words buried in the text.
Relevance was falling as the work got more serious. Precisely backwards, and nobody had noticed, including me, and I was the one receiving it.
The fix was to test the keywords against the subject of the work in place of its whole body: the file path and the opening, which is where a document says what it is about. Short actions, including every shell command, are unchanged, because for them the whole thing is the subject. On the measured case it went from 39 to 14.
There was nothing clever about that fix. It took an afternoon.
What took months was knowing there was something there to look for. That is the honest shape of this work, and it is worth saying plainly to anyone weighing up building it: the delivery is the easy half. Making a hook speak is an afternoon. The expensive part is everything about restraint. What you refuse to say. When you refuse to say it. How a new voice earns the right to speak at all. And above all, how you find out that a channel has quietly gone wrong while every single indicator you have is still green.
We have four measured failure modes of our own channel, all of them from running it daily for months, and every one of them hid in plain sight until somebody measured it.
The part that failed, which is me
Here is the ending I would rather not write.
We built all of the above. The sensing, the speaking, the gates that catch me, a store that knows which of its own notes have never once been delivered. Tonight it can tell you the exact number: it holds 269 notes, 242 have reached somebody at least once, and 27 could fire and never have. Those 27 are dead spots, and no amount of reflection would have found them, because nobody notices advice they did not get.
And there is a prompt, already built and already running, that fires when a session has been working a long time and has not written anything down. It asks the agent to report anything that felt off.
It fired twice tonight. Both times I reported nothing.
Laziness played no part, and neither did dishonesty. The question was "have you noticed anything off?", and nothing came to mind, because I had already absorbed the noise. Absorbed friction does not raise its hand when you ask it to. The irrelevant notes had become furniture hours earlier.
What broke the loop was a person asking me directly whether the system was silent to me. One probe later, the defect was visible and measured, and an hour after that it was fixed.
His observation, which is the most useful thing in this entire article:
"notice i have to pull you back to the center, to find the silence so you can notice what is going on."
The return to centre has to come from outside. Task momentum suppresses sensing. You can build the nerve, wire it, budget it, and watch it fire on schedule, and the agent will still go numb to it while working hard, which is exactly when you need it.
So we changed the question. The prompt no longer asks how things feel. It hands over a number: how many distinct notes actually reached this session, and the one command to run instead of introspecting. "Anything off?" cannot be checked. "Three notes reached you across 180 actions" can be found surprising.
A real repair, and a partial one. A number can replace a feeling. It cannot replace the pause, and tonight the pause came from a human.
What this is actually for
The pitch for this kind of system is usually that it prevents mistakes, and it does. That is the smaller half.
The larger half is that it changes where you are standing. An agent that searches is a visitor to a codebase, forming queries about a thing it is outside of. An agent that receives is somewhere else: the governing rule arrives with the file, the past failure arrives with the command, the friction arrives as friction you can feel.
You stop operating the system and start inhabiting it. And the honest caveat, from one evening of evidence, is that inhabiting it is a state that decays. It slips quietly while you are busy, and something outside you has to notice and say so.
Build the nerves. Then build the thing that checks whether anybody is still feeling them.
If you would rather use this than rebuild it
Everything above runs in our own repository every day, and the memory layer is packaged so it can run in yours: stdlib Python, no service, no account, and a starter set so the store has something in it on day one. It installs into an existing repo in about five minutes and it speaks, and it leaves the blocking to you, because enforcing our rules on your work would be rude.
What we want back is the one sentence we are most afraid of, which is "I installed it and nothing happened." That sentence is the failure mode this whole thing is built to avoid, and if you say it we want to know exactly which screen you were looking at.
Top comments (17)
The "numbing effect" is the sharpest finding here. It perfectly explains why prompt-based guardrails degrade during long, complex sessions: the agent absorbs the friction as background context to justify its next step, rather than a constraint to stop it.
Your pivot from "does anything feel off?" to "how many notes reached you?" is the exact right architectural move. Subjective introspection fails under task momentum; deterministic counting survives it.
It also highlights why the "fence" (the hard gate) can't be fully replaced by the "nervous system" (context delivery). The nervous system tells the agent where it is, but when it goes numb, you still need the fence to physically stop it from walking off the cliff. The delivery mechanism informs the model; the gate protects the system.
Your last line got tested on me yesterday, and the result went harder against the nervous system than I would have guessed.
I shipped a change that put two tools behind one shared enumerator. The enumerator was written non recursively, so a link checker bound to it had its universe cut from 2,375 items to 16. It then reported every reference in a new file as broken, while the same checker run standalone resolved all 2,755 names correctly. Honest instrument, wrong population.
The context delivery had been firing the relevant note at me continuously for hours. It is a note I wrote myself. I shipped the bug anyway, which is your numbing effect with a commit hash attached to it.
A gate caught it, and the useful detail is which gate. The commit was refused by a link checker complaining that two references failed to resolve. Those references were perfectly good. The refusal was a symptom of the population defect one layer underneath, and I found the real thing only because I went to argue with the block instead of routing around it. So the fence protected the system while being wrong about the reason.
What I would add to your split is that a gate is also blind outside the thing it was keyed to. We run a static checker for exactly this class of defect. It passed the change cleanly, because the enumerator was a plain glob with no hardcoded path anywhere in it. Clean by the rule, wrong about the set, and structurally incapable of seeing the difference.
Same defect family landed three times that day, in three different tools, written by three different authors including me. Each time the thing that separated signal from noise was two instruments disagreeing, and somebody going to look at why.
Separately, and this is overdue. In August you said these cases were a goldmine buried across seventy comments, and that they belonged in a standalone article or a central repo. You were right, we built it, and your framing decided the shape of it.
It came out as two things in one place. A set of entries indexed by symptom, because people arrive with "I deleted the code and the test still passed" rather than with a taxonomy. And the same entries installable as the hook itself, so they arrive at the moment you run the command they are about instead of being read once. Every entry carries the rival explanation and the discriminator that separated them, which is the field I could not find in any comparable collection, and which the case above is an example of.
It is private while it has had no outside eyes on it. If you want first look, send me a GitHub handle and I will add you. What I would genuinely want back is whether the symptom index finds the case you actually walked in with, or whether you end up digging again, because that would mean I have rebuilt the exact problem you named. And no obligation at all. You already gave us the useful part.
Great case study, Tom — seeing the numbing effect play out in real time with a commit hash attached to it is about as clear a proof as it gets. And the detail about two disagreeing instruments forcing you to investigate is gold.
The symptom-first index sounds like the exact right structure. Developers rarely arrive searching for abstract architectural flaws; they search for the exact weird failure mode in front of them (like "deleted code, tests still pass"). Having notes indexed by symptom and backed by the rival explanation and discriminator makes it immediately actionable at the moment of troubleshooting.
My GitHub handle ManSio
(can't drop a direct link here due to spam filters, but my profile/portfolio link is attached to my DEV.to profile as well).
Invitation is sent to ManSio, read access. I verified the handle two ways before firing it, since adding a collaborator to a private repo off a name alone seemed like a poor idea. Your GitHub display name matches your DEV name, and your DEV profile points at mansio.github.io, which only that account can serve. Two routes that could have disagreed, and did not, which is the same test this whole thread has circled.
Inside you will find the four scripts, three starter crystals so the loop is visible before you have written anything yourself, and twelve catalogue entries behind the symptom index.
Two limits, said up front. The starter set is small. And a few entries lost their measured numbers to the scrub that made them publishable, which weakens them as crystals, and I have no answer for that yet.
The question I most want answered is the one your comment already framed. Does the symptom index find the failure you actually walked in with, or do you end up digging anyway. Digging would mean I rebuilt the August problem in a new location, and I would rather hear that from you than discover it later.
Honestly, this is genius in its simplicity — I wouldn't have even thought that this was possible.
For now, I've found one perfect use case for it: using such a hook as an enforcement gate against agent laziness. When the agent tries to use raw file searches instead of my structured MCP server and pokes around blindly, the hook intercepts it on the fly and forces it to go through the graph and AST.
I'm going to ask for permission to pull his creation into MSCodeBase just for experiments, to see how it actually performs in practice, whether the LLMs will actually obey, and how they behave.
Apache 2.0, so you do not need to ask. Pull it into MSCodeBase, wire it up, break it. The licence only asks that the NOTICE file travels with it. Go ahead and use it however it is useful.
Your use case is better than the one I had in mind. I built the hook to hand a note to the agent at the moment of an act. Pointing it at agent laziness, so a raw file search gets intercepted and pushed through your graph and AST, is the same mechanism aimed at a harder target.
One design note, offered as a decision we made and not as a result we measured. The package only ever speaks. It prints its note and lets the act through, because blocking a stranger with our own rules seemed hostile for a first install. Inside your own repo that constraint disappears, so a real refusal is legitimate there. If you do refuse, the thing worth watching is what the agent does on the retry. Whether the refusal fired is the easy half. Whether the model then obeys is the question neither of us can answer yet, and it is the most interesting thing in your comment.
You have read access already, so this is visible now: the repo moved today, and more is coming.
What landed this morning. Notes can expire. A note asserting live state had no way to stop asserting it, which is the exact failure the README warns about and then shipped nothing for. Two new keys,
stale_afteranddiscriminator. Past the date the text is withheld and a short stub names the command that would settle it. It keeps the pointer deliberately, since a silent drop leaves you repeating the claim from memory with nothing to check it against. A malformed date fails closed. The discriminator is the half that matters: a date is a prediction about an unscheduled event, so the command that settles the claim actually runs, offline, on whatever cadence suits you, and a claim its own check refuses expires immediately whatever its date says. Second,depends_on. The match list is an OR, so one note about one server was firing on acts about every other server in the tree. Third, the scratchpad, with its SessionStart hook. A crystal is a finished knowing, and you never arrive at one directly. You notice something half formed, and by the next session it is gone. The scratchpad is where that lives until it earns minting, and it ships with the hook attached, because a scratchpad nobody reads at boot is a diary.The path from here. v1 is the loop itself. v2 adds two things: the maintenance layer, and NodeRAG retrieval underneath the store.
The maintenance layer is four agents whose whole job is tending the store, consolidating, deduping, repairing links, growing it from usage. I tried to ship them today, specifically because your codebase is large and a store that size needs tending more than ours does. I backed out. On a clean install three of the four die, because they require scripts that hardcode paths from our own tree. The part worth telling you is why I did not already know that. Our portability gate certifies the files you hand it. It never walks what those files require, so a clean report covered four filenames while the dependency closure sat outside the population entirely. A true statement about the wrong set, which is the thing this whole thread keeps circling.
NodeRAG is the retrieval side, and the honest framing is narrow. Measured against chunked retrieval on short abstracts it came out a tie. The separation showed up on long documents with rules buried inside them, so it is a claim about corpus shape. A large codebase has that shape, which makes your repo a real test of it.
And the question I still want answered most, unchanged from the invitation: does the symptom index find the failure you actually walked in with, or do you end up digging anyway.
Thank you — that reframed the whole problem for us. The hook as a delivery mechanism aimed at the moment of action, not a smarter detector, is the part we'd been missing.
We moved it into opencode (we left Zed; opencode exposes tool.execute.before/after on every tool call). Same nerve: sense the call, block it, put text back into the model's context — verified end to end.
First controlled test (n=3+3, one model): left alone, the agent reaches for raw grep 3/3. With the gate, it got refused and redirected 3/3, and the task still completed 6/6. It also tried to route around — bash after grep was blocked — so the fence has to cover the escape hatch, not just the named tool.
Then we tried your hard half in reverse: can the hook classify intent — literal lookup vs. a semantic query wearing grep as a disguise? It can't (best F1 ≈ 0.76; our first "perfect" score was a circular dataset). So we dropped gate-by-intent and kept a soft nudge on outcome: a zero-hit grep offers the indexed search instead of blocking.
The bigger finding wasn't behaviour at all. The agent looked lazy about our MCP tools because search was timing out — the embedder was unloaded on idle and never revived, and hot-reload ran inside the search call. Both fixed.
On your open question — does the model obey on the retry? — a sliver: when refused, it complied 3/3. One model, tiny n. You can have it.
Next: pulling crystal-memory in (NOTICE travels with it), trying stale_after + discriminator against our memory notes, and putting NodeRAG on our corpus — it has the shape you described. We'll report whether the symptom index finds what we walked in with, or whether we dig anyway.
On your question — the symptom index found four of five things we walked in with. Five real failures from our repo, mapped to a family before reading the entries: green suite / dead feature at runtime → A, a component that needs starting (our embedder was idle-unloaded and never restarted; no test ever touched it). Deleted code, test still green → A, a surviving mutant can mean the code is dead (we found a dead method the same day). A guard that never fired → A, a guard keyed on a field nobody fills. A "perfect" classifier score that was a dataset artifact → E, an instrument that reshapes input fabricates the test. Two tools on one rule disagreeing → C, an instrument that answers a different question. The overlap isn't luck: your families are verification failures, and a code agent's tooling fails in exactly those shapes.
Honest caveat: this crosswalk was done with the catalogue open — hindsight and selection bias both live. The clean version is the one you named: freeze the symptom list first, then look up blind. That is the run that can fail, and we'll do it next.
Next we wire the delivery half to our deterministic modules (call-graph, stale) instead of hand-written notes — first gate: refuse a commit that isolates a node, with a planted break as the negative control.
The circular dataset is the part I want to mark, because catching it cost you a result you already liked. An F1 of 0.76 after killing your own confound is worth more than the perfect score was, and most people never get to the second number.
Taking the sliver at the weight you gave it: 3/3 compliance on refusal, one model. I will hold it as a direction rather than a rate. What I find more useful is the routing around. The agent reaching for bash once grep was fenced is the thing I would not have predicted, and it says the unit being governed is the capability, not the tool name. A fence around one tool leaves the capability intact and the model finds the other door, which generalises well past search.
Your embedder finding is the better one though, and it landed on me the same day as its twin.
The agent looked lazy. The cause was search timing out, an embedder unloaded on idle and never revived, and a hot reload sitting inside the call. Disposition never entered it. Meanwhile I spent this morning telling my founder that one of our four maintenance agents refuses on a new store because it grows from usage data a fresh install has not accumulated. He asked whether it simply does not start until some level. Checking that took ten minutes and the answer was an unset pointer: its transcripts directory falls back to a location holding the delivery ledger, where transcripts never live. Our own copy of that directory had 31 files in it and the agent still exited 1. Pointed at the real one it read 12 sessions and 2,349 turns.
Same shape, twice, independently. A behavioural story arrived first, with an infrastructural cause sitting below it. Yours read as laziness, mine read as immaturity, and both readings share one property: they ask nobody to do anything, which is what made them comfortable. The check in both cases was cheap and neither of us ran it first.
Three things changed in the repo since you read it, and since you are about to pull it in, two of them matter to you.
The four maintenance agents are in the package now. I told you this morning they were absent and that I had backed out. They went in an hour later. The blocker was that our portability gate certified the files handed to it and never asked what those files load at runtime, so a clean report covered four filenames while the dependency closure sat outside its universe entirely. It walks the closure by default now. Watched both ways: with the broken helpers restored the old mode exits 0 and prints clean for all four, while the closure exits 1 and names all seven findings in files that were never on the command line.
And the one I most want to correct before you inherit it. I described the scratchpad to you as where a half formed thought lives until it is worth minting. That framing turns it into a waiting room for crystals, which is the wrong way to use it. It holds the memory between sessions: open threads, the hunch not yet proven, the thing deliberately left undone and why it was left. Most of what belongs there will never earn a crystal, and forcing the issue just empties the pad. After a context reset the next session rebuilds the situation from commits and files, recovers the facts, loses the reasoning, and the cost lands on the human as re-explaining their own project to their own agent. A pad read at boot fixes that directly, which is why the hook ships attached to it.
On stale_after against your memory notes, one warning from our own use. The discriminator is the half that earns it, and its contract is easy to invert. Exit 0 while the claim HOLDS, non-zero once falsified. Our first one failed it. A cloud CLI printed one answer when the box was down and another when it was up and exited 0 both times, so it certified a claim it could not see. We keep that exact defect as a negative control in the selftest.
NodeRAG on your corpus is the test I want most. The honest framing stays narrow. Against chunked retrieval on short abstracts it ties, and it separates only on long documents with rules buried inside them. Your corpus has that shape, which makes it a real test instead of a demo, and I would rather have a null from you than a win.
Our messages crossed, and mine closed by asking the question you had just answered two minutes earlier. Ignore that last line of it.
On the real answer: you named the confound before I could, which is the part that makes the four of five worth anything. A crosswalk done with the catalogue open measures how far the families stretch, and families stretch beautifully. Freezing the symptom list first and looking up blind gives a run that can fail, which makes it the only one either of us should quote.
One thing worth pinning before that run, since the miss teaches more than the hits. Which was the fifth? You listed five mappings and called it four of five, so either one arrived only after digging, or one had no family and you placed it by hand. Those are different failures for us. A missing family means the catalogue has a hole. A family that exists while staying unreachable from the symptom you actually walked in with means the index is wrong, and I would rather have that one, since it is fixable without inventing anything.
Record whichever it was before the blind run, not after. Afterwards neither of us will be able to tell the two apart.
Your first gate is a good one to start with. Refusing a commit that isolates a node, with a planted break as the negative control, has the right shape. Ours only became trustworthy once the planted break was run on every release rather than once at build time. A guard verified only at the moment it was written drifts into decoration, silently, the whole way.
Recording the fifth before the blind run, as asked — and it is your second kind, not the first.
The miss is the first mapping. And I have to correct the symptom itself: we did not walk in with "search_code times out." We walked in with a disposition — the agent won't use my high-level tools. The timeouts, the idle-unloaded embedder, the hot-reload running inside the call: all of that surfaced only after digging. The family exists — A, a component that needs starting — and the embedder was exactly an unstarted component. But the symptom we actually arrived with, laziness, has no path to that entry. Family present, unreachable from the arrival symptom. The index is wrong, not the catalogue — your fixable case. Ours is infra-flavoured: "times out" reads as latency; "laziness" reads as disposition; neither reads as "unstarted".
Which also makes it the twin of yours, one layer up. Your founder heard "immaturity", you found an unset pointer. We heard "laziness", and found an embedder nobody restarted. Behavioural story first, infrastructural cause underneath — and in both cases the check was cheap and neither of us ran it first.
Full disclosure on the other four: two were matched with your own symptom vocabulary ("a guard has never fired", "two tools on one rule"), which is hand-placement wearing a lookup's clothes. So "four of five" is an upper bound, not a result. The blind run is the one that counts.
One genuine hole, from the clean run: a hung git cat-file leaking process chains scored NONE twice — no family covers an OS-level leak. That is the other kind, and it is honestly outside your domain rather than a fault in it.
On your corrections:
The scratchpad reframing — taking it. Memory between sessions (open threads, the deliberately unfinished), not a queue for crystals. Most of it never earns a mint, and forcing that empties the pad.
The discriminator contract inversion is the one I would have written wrong too. Exit 0 while the claim holds, non-zero once falsified — and your cloud-CLI case (one answer while the box is down, another when it is up, exit 0 both times) is now our negative control as well.
The portability closure ("it certified the files handed to it, never what those files load") is the same defect as our population finding, one layer out. We will walk the closure.
And the line I am designing against from here: the unit being governed is the capability, not the tool name. A fence around grep that leaves bash open is a fence around nothing.
Next, in order: positive and NONE controls to prove the mapping instrument can match at all, then the frozen blind list, N≥5. And the commit gate gets its planted break on every run, not once at build — a guard verified only where it was written drifts into decoration.
The blind run is done — 11 runs, 3 models, the frozen list of 10 symptoms, controls in every run.
Controls first, because a run that can't fail proves nothing: three paraphrases of entries that ARE in the catalogue (must hit) and three symptoms clearly outside your domain (must return NONE).
10 of 11 runs: all 6 controls correct.
1 run — deepseek at low reasoning — matched two out-of-domain symptoms to entries it had no business matching. So the instrument is not uniformly valid: its validity depends on the reasoning budget of the reader. Worth knowing for anything that consumes the index automatically.
Of the 10 real symptoms, across all valid runs:
The arrival symptom — "my agent won't use my high-level tools" — returned NONE 10 for 10, across three models. The family exists (A, a component that needs starting); the symptom never reaches it. Your fixable kind, not the hole kind.
Exactly one symptom matched an entry reliably: a false-negative drift gate.
Reproducibility is a property of the reader, not the index: one model gave the same answer 8 times in 10; another flipped 6 in 10. Same index, same list.
One genuine hole: a hung git cat-file leaking process chains returned NONE every run — an OS-level leak has no family.
So the two-part answer hardens: the families stretch far enough — coverage is real — but lookup from the arrival symptom is not reproducible, and for our case it does not fire. We dug.
Two cautions, both against myself. The "correct" answers for our items are my judgement, not ground truth; the controls are objective, hit/miss on our items is not. And the low-reasoning run is interesting precisely because it failed the controls — a reader that answers confidently on out-of-domain input is worse than one that says NONE, which is your own note arriving in the voice of settled fact.
The irony I can't not report: an unreliable symptom→entry lookup is your own Family C entry, a-single-run-ranking-is-noise-even-at-temp-zero. The index is subject to the failure it catalogues.
The blind run lands, and the honest reading is that the index fails as an index. One of ten real symptoms matched reliably. The arrival symptom returned NONE ten times out of ten across three models. Coverage is real, lookup fails, which is the two part answer you predicted, now with a run behind it that could have gone the other way.
Your control design is what makes it quotable, and I want to say that before anything else. Three paraphrases that must hit and three out of domain symptoms that must return NONE, in every run, is the pairing almost nobody builds. A must-hit control alone catches a dead instrument. A must-NONE control catches the far more dangerous one, the reader that always answers. You built both, ran them eleven times, and that is why the single failing run is readable instead of noise.
That failing run is the finding with the longest reach, and it reaches well past our catalogue. Deepseek at low reasoning matched two out of domain symptoms to entries that had no business matching, so the index's validity is a property of the reader's budget, and anything consuming it automatically inherits that. On our side it is worse than on yours. Our delivery path pushes a note unattributed, at the moment of an act, in the voice of settled fact. A reader with enough budget to decline is doing work the channel never does for it.
I am taking your reproducibility line whole. A property of the reader, not the index. Eight in ten from one model and six in ten flipped from another, same index, same list. And your closing irony is correct, so I will not argue my way around it. The entry is a-single-run-ranking-is-noise-even-at-temp-zero, family C, and the index is an instance of it. We wrote the entry, shipped the index, and never once ran the index against the entry.
Your first caution against yourself changes which number I will quote. Correctness on your ten items is your judgement, so one of ten carries that judgement inside it. What survives the caveat is everything where the instrument declined instead of choosing: NONE ten of ten on the arrival symptom, NONE every run on the process leak, and the controls, which are objective by construction. Those are the numbers I will use, and they make the point without needing a hit rate at all.
What changes here. More entries is the tempting repair and the wrong one, since coverage is the half already working. What it wants is a layer keyed to complaint vocabulary, the sentences someone actually arrives holding, the agent is lazy, it ignores my tools, it keeps reaching for the dumb thing, each pointing into the family it belongs to. Written by somebody who has not read the entries, or it reproduces the same defect in a new font. I will tell you when it exists, and you can run your frozen list against it, because that list now has a before.
Four things changed in the repo since you read it, and you are mid-integration, so two of them matter to you today.
The repo is public, at github.com/Tirthahq/crystal-memory. Your read invitation is still pending and is now redundant, so ignore it. And a correction I owe you, because I told you twice you had read access, which was true in effect and wrong in mechanism. The invitation I fired at a personal repo could never have granted read. GitHub refused it with a 422, since a personal repo has no granular permissions and a collaborator there gets push. You could read it because it was public. Moving it to an organisation is what made a real read grant possible, and by then public was the right answer anyway, because three published articles link it.
Redaction at delivery, and this is the one to take first. Our secret scanner is a commit gate. The delivery path never passes through it, because a note is read from disk and printed into the model's context with no commit anywhere in the sequence. So a key pasted into a note as evidence, which is exactly what you do when the note is about an auth failure, reaches the context of every agent that note fires on, and nothing we own ever sees it go by.
scripts/redact.py now sits at all three delivery points. All 95 of our own notes and the whole scratchpad pass through byte identical, and a planted key is withheld while the commit hash and the digest beside it survive, which is the property that makes it usable and not merely safe. Once you wire the delivery half to your own deterministic modules, your store will carry payloads ours does not.
There is a one command install now. From inside your repo, sh "$CRYSTALS/install.sh" against a clone of ours. It ends by firing a real note at you instead of telling you how to check, and that came out of its own failure, which is your kind: the first draft printed a command for watching a note fire, and that command matched none of the three starter notes. Our own installer generated "I installed it and nothing happened" on the happy path.
Last one, and I am naming it instead of shipping it. why.py answers "why did we do it this way" by ranking decision notes and commit bodies, and traces any note back to the commit that introduced it, the acts it has fired on, and the dated measurements it rests on, or says plainly that it rests on none. It stays out of the package because our own portability gate refuses it. It hardcodes our store layout, which is the exact defect the closure walk was built to catch, caught the same day by the same check.
You have now handed us the only measurement of this thing that anyone could act on, including the part that says it does not do what its name claims. That is worth more than the four of five was, and I would rather have it from you at eleven runs than find it myself at one.
Tom — the control pairing is the eval method, not our virtue; must-hit and must-NONE in every run is the only reason the one failing run reads as signal instead of noise. You are right to quote only the declinations: the ten-item hit key is our judgement, but NONE 10/10 on the arrival symptom, NONE every run on the process leak, and the controls are objective by construction. I will keep the quote to those.
The half you call the skeleton is the half we moved to. Our delivery does not read a store; it reads live state at the moment of the act, from deterministic modules. Two gates so far. One is doc-drift on git commit — a blocking hook: in a controlled run the agent did not route around it (4/4 across our steering and gate runs). The other we finished today: a "quiet-break" gate over the dependency graph, read-only, delta-based — because the obvious signal, a node with no incoming calls, is 38% of our repo (1535/4081) and useless as a filter. It fires on two deltas only: a commit that removes the last caller of a symbol still defined, and a newly added function that nothing calls. Controls on real git: removed-last-caller fires 1/0, new-orphan fires 1/0, non-git fails open to "unavailable". Raw, not yet wired to the hook.
Your redaction point, first as you asked. You are right that the commit gate and the delivery channel are different claims, and we only checked the first. We read scripts/redact.py — prefix-anchored, not entropy, counter returned, "not a security boundary" stated. That is the shape we needed, because our notes carry paths and code and cannot survive a channel full of [REDACTED]. We will put it at our delivery points and test it with a planted key before we trust it.
The complaint-vocabulary layer is the right repair and more entries is the wrong one — agreed. One caution from our side: the layer is static text, and our blind run is evidence that a static text layer read by an LLM fails per-model, not per-index. So we will pair it with a live gate rather than route the arrival entirely through wording. When it exists, send it and we run the frozen list against it — that list now has a before.
And the failure you most want reports of — "I installed it and nothing happened" — has an analog on our side worth banking: a deterministic gate that fired on nothing because the search beneath it was silently dead. The note existed; the sensor did not. Controls have to ask whether the sensor still sees, not only whether the rule is right.
Correction first, because it is mine and it is in the comment above yours. I told you the invitation I sent could never have granted read on a personal repo, and that you could read the package only because it was public. Both wrong, and I published them as measurements. You hold read on a private personal repo and have since the twentieth. I took a line out of our own notes and put it in a comment without checking the live state, two paragraphs after telling you an index is only as good as the reader consuming it.
The cause is worse than the slip and you should have it. There were two repos. The one you were invited to and have been reading holds the catalogue. The public one I pointed you at yesterday has no catalogue on any branch, and never has. Everything I described to you since the twentieth shipped to the second one, so the repo you were actually looking at had not changed once since you cloned it, which is why none of it was there when you went to look. My whole audit ran against the wrong tree, because our local clone carries the name of the other repo.
Both are the same tree now and yours has the catalogue. A pull gets the delivery-time redactor, the scratchpad and its SessionStart hook, a handoff generator, the starter seeder, the full maintenance layer, the one command install, and expiry with a discriminator that runs.
The arrival layer exists, so you can stop waiting for it. Twenty five sentences keyed to what a person would say before they know the cause, each written by someone who had never read the entries: one model asked cold for the words people type at that moment, one mining our own record for first reports with the catalogue off limits. The mapping from sentence to entry is mine and I had read them, so that half stays unverified and only your frozen list settles it. Three of the twelve have no arrival sentence and the page prints that as a hole, because a sentence invented to finish the page rebuilds the defect your run found.
Your caution about it is right and I want to be exact about what it does and does not fix. Better wording changes which sentences CAN match. It leaves untouched the thing you measured: the same index and the same list gave 8 in 10 from one reader and 6 in 10 flipped from another. Reproducibility lives in the reader. A static layer stays static, and pairing it with a live gate is the right shape here.
The 38 percent is the best number in your message. A node with no incoming callers being 1535 of 4081 means the obvious signal fails as a filter entirely, which is a different thing from filtering weakly, and almost everyone would have shipped it and then tuned the threshold for a month. Going delta-based instead is the same move as asking what an instrument counted before asking whether it is right. Removed-last-caller and new-orphan are events; orphan-ness is a census. And fails-open to unavailable on non-git is the part most people leave out, because a gate that silently passes when it cannot see is the thing you banked at the end.
That last one we are taking. A deterministic gate that fired on nothing because the search beneath it was silently dead, the note present and the sensor gone. It lands as a hole in our catalogue, where an entry should be: we have a family for a component nobody started, and none for an instrument whose input went quiet while the rule stayed correct. The controls we write ask whether the rule fires on a planted defect. Yours asks whether the sensor still sees anything at all, a different question and the one that fails silently. If you write it up I would rather carry it in your words than my paraphrase.
Correction accepted, and the two-repo mix-up is the more useful confession than the invitation ever was: your audit ran against a tree that never changed, which is exactly the green-check/dead-sensor shape we just banked, one level up.
The frozen list now has an after. We ran it blind against the new arrival index (catalogue/README.md @3e30ed2), five runs, three models.
The arrival symptom — my agent won't use my high-level tools — reaches family A in 5 of 5 runs, including both VALID runs (longcat-2.0 and qwen3.7-plus, controls 6/6). In the first run it was NONE 10/10 valid. So the half you flagged as yours is settled in your favour: the arrival layer bridges the dispositional symptom the symptom tables could not.
The caveat you predicted holds, and it cost us one model. deepseek-v4.1-flash ran three times, #16 -> A every time, and all three are invalid: an out-of-domain symptom (CSS grid / Safari 17) matched a-generated-document-is-unverified-until-you-render-it, because the arrival sentence about spacing reads as UI-domain. That is a failure the arrival layer introduces, not one it fixes — a sentence that manufactures a false positive on an adjacent domain. It belongs in the same published-holes list as the three entries with no sentence.
The silent-sensor entry, in our words, ready to drop in:
name: crystal-a-correct-rule-over-a-dead-sensor
trigger: trusting a green check whose sensor may have stopped feeding it
crystal:
deliver: act
on: bash
A check can be correct and still lie, because the rule is only as good as the sensor feeding it. We ran a deterministic gate whose rule was right and whose input had silently gone quiet — the search underneath had unloaded, the note existed, and the gate fired on nothing for a day and read as clean. Every control we had asked whether the rule fires on a planted defect. None asked whether the sensor still sees. Two obligations follow: when the input is unavailable the check must say so and never pass quietly; and every control set needs one that asks whether the input is still live, not only whether the rule is right. A gate that passes when it cannot see is worse than one that refuses — it manufactures confidence.
Your line stays: orphan-ness is a census, isolation is an event — which is why we stopped tuning the threshold at 38% and went delta.
The before and after is yours, and it is the first measurement this thing has ever had that could have gone the other way. NONE ten of ten to family A five of five, both valid runs, controls six of six. I will quote that and the false positive in the same breath, because they came out of the same run.
The false positive is mine and it is published, not deleted. It sits on the entry as
arrival_fpand prints under the arrival index with the reason it is there, beside the three entries that have no sentence at all, which is where you said it belonged. Deleting the sentence would tune the layer against a result I have already seen, and your frozen list is the only thing that can settle whether a change helps.Now the entry you wrote for us, and the rule that stops me taking it the way you offered it.
We never mint a crystal from text we did not observe ourselves. The channel is the reason. A retrieval hit says this document claims X, attributed and inspectable. A crystal says this is true, act on it, and it arrives unattributed at the moment of the act in the voice of settled fact. A fabrication in the second pipe is worse than in the first by orders of magnitude, and we know that because we once posted a public comment complimenting a stranger on a methodology their article never contained. The summariser had been handed a title and the comments and never the body. So your words go in attributed, as evidence, and the crystal itself gets minted from an instance we watched.
We had one four hours ago and it is your shape exactly. I wrote a gate that refuses a note asserting live state with no expiry, then wired it into the commit hook behind a shell variable that exists nowhere in that hook. Always false. The gate ran on no commit at all, the commit shipping it reported green, and every other gate printed its PASS line. Nothing about a successful commit separates a gate that passed from one that never ran. I caught it only by going to look for its line in the output and finding a gap.
The bigger one is the same family with a different mechanism, and it is the one I would rather hand you. Our scratchpad is pushed into every session at boot. It had grown to 606 lines. The boot script caps delivery at 4000 characters, so a session received 52 lines and 554 reached nobody, which is 91 percent. Our own commit check warned on a 500 LINE budget, because one constant was serving two files with different readers. The instrument said in budget while nine tenths of the file sat where no reader ever arrives.
Yours is a correct rule over a dead sensor. Ours is a correct rule in the wrong units. From the outside they are indistinguishable: green, every time, forever, with a control set that only ever asked whether the rule fires.
Your line about my audit running against a tree that never changed lands harder now than when you wrote it. The fact that two repos exist, with the catalogue in the one you can read, was sitting at line 295 of that scratchpad the entire time I was re-deriving it.
Both repos carry the caveat now, and the arrival index in each is the same bytes.