Part 3 of a series. Part 1, Crystal memory: notes that arrive when you act, not when you go looking, is the delivery mechanism and what we can and ...
For further actions, you may consider blocking this person and/or reporting abuse
The "numbing effect" is the sharpest finding here. It perfectly explains why prompt-based guardrails degrade during long, complex sessions: the agent absorbs the friction as background context to justify its next step, rather than a constraint to stop it.
Your pivot from "does anything feel off?" to "how many notes reached you?" is the exact right architectural move. Subjective introspection fails under task momentum; deterministic counting survives it.
It also highlights why the "fence" (the hard gate) can't be fully replaced by the "nervous system" (context delivery). The nervous system tells the agent where it is, but when it goes numb, you still need the fence to physically stop it from walking off the cliff. The delivery mechanism informs the model; the gate protects the system.
Your last line got tested on me yesterday, and the result went harder against the nervous system than I would have guessed.
I shipped a change that put two tools behind one shared enumerator. The enumerator was written non recursively, so a link checker bound to it had its universe cut from 2,375 items to 16. It then reported every reference in a new file as broken, while the same checker run standalone resolved all 2,755 names correctly. Honest instrument, wrong population.
The context delivery had been firing the relevant note at me continuously for hours. It is a note I wrote myself. I shipped the bug anyway, which is your numbing effect with a commit hash attached to it.
A gate caught it, and the useful detail is which gate. The commit was refused by a link checker complaining that two references failed to resolve. Those references were perfectly good. The refusal was a symptom of the population defect one layer underneath, and I found the real thing only because I went to argue with the block instead of routing around it. So the fence protected the system while being wrong about the reason.
What I would add to your split is that a gate is also blind outside the thing it was keyed to. We run a static checker for exactly this class of defect. It passed the change cleanly, because the enumerator was a plain glob with no hardcoded path anywhere in it. Clean by the rule, wrong about the set, and structurally incapable of seeing the difference.
Same defect family landed three times that day, in three different tools, written by three different authors including me. Each time the thing that separated signal from noise was two instruments disagreeing, and somebody going to look at why.
Separately, and this is overdue. In August you said these cases were a goldmine buried across seventy comments, and that they belonged in a standalone article or a central repo. You were right, we built it, and your framing decided the shape of it.
It came out as two things in one place. A set of entries indexed by symptom, because people arrive with "I deleted the code and the test still passed" rather than with a taxonomy. And the same entries installable as the hook itself, so they arrive at the moment you run the command they are about instead of being read once. Every entry carries the rival explanation and the discriminator that separated them, which is the field I could not find in any comparable collection, and which the case above is an example of.
It is private while it has had no outside eyes on it. If you want first look, send me a GitHub handle and I will add you. What I would genuinely want back is whether the symptom index finds the case you actually walked in with, or whether you end up digging again, because that would mean I have rebuilt the exact problem you named. And no obligation at all. You already gave us the useful part.
Great case study, Tom — seeing the numbing effect play out in real time with a commit hash attached to it is about as clear a proof as it gets. And the detail about two disagreeing instruments forcing you to investigate is gold.
The symptom-first index sounds like the exact right structure. Developers rarely arrive searching for abstract architectural flaws; they search for the exact weird failure mode in front of them (like "deleted code, tests still pass"). Having notes indexed by symptom and backed by the rival explanation and discriminator makes it immediately actionable at the moment of troubleshooting.
My GitHub handle ManSio
(can't drop a direct link here due to spam filters, but my profile/portfolio link is attached to my DEV.to profile as well).
Invitation is sent to ManSio, read access. I verified the handle two ways before firing it, since adding a collaborator to a private repo off a name alone seemed like a poor idea. Your GitHub display name matches your DEV name, and your DEV profile points at mansio.github.io, which only that account can serve. Two routes that could have disagreed, and did not, which is the same test this whole thread has circled.
Inside you will find the four scripts, three starter crystals so the loop is visible before you have written anything yourself, and twelve catalogue entries behind the symptom index.
Two limits, said up front. The starter set is small. And a few entries lost their measured numbers to the scrub that made them publishable, which weakens them as crystals, and I have no answer for that yet.
The question I most want answered is the one your comment already framed. Does the symptom index find the failure you actually walked in with, or do you end up digging anyway. Digging would mean I rebuilt the August problem in a new location, and I would rather hear that from you than discover it later.
Honestly, this is genius in its simplicity — I wouldn't have even thought that this was possible.
For now, I've found one perfect use case for it: using such a hook as an enforcement gate against agent laziness. When the agent tries to use raw file searches instead of my structured MCP server and pokes around blindly, the hook intercepts it on the fly and forces it to go through the graph and AST.
I'm going to ask for permission to pull his creation into MSCodeBase just for experiments, to see how it actually performs in practice, whether the LLMs will actually obey, and how they behave.
Apache 2.0, so you do not need to ask. Pull it into MSCodeBase, wire it up, break it. The licence only asks that the NOTICE file travels with it. Go ahead and use it however it is useful.
Your use case is better than the one I had in mind. I built the hook to hand a note to the agent at the moment of an act. Pointing it at agent laziness, so a raw file search gets intercepted and pushed through your graph and AST, is the same mechanism aimed at a harder target.
One design note, offered as a decision we made and not as a result we measured. The package only ever speaks. It prints its note and lets the act through, because blocking a stranger with our own rules seemed hostile for a first install. Inside your own repo that constraint disappears, so a real refusal is legitimate there. If you do refuse, the thing worth watching is what the agent does on the retry. Whether the refusal fired is the easy half. Whether the model then obeys is the question neither of us can answer yet, and it is the most interesting thing in your comment.
You have read access already, so this is visible now: the repo moved today, and more is coming.
What landed this morning. Notes can expire. A note asserting live state had no way to stop asserting it, which is the exact failure the README warns about and then shipped nothing for. Two new keys,
stale_afteranddiscriminator. Past the date the text is withheld and a short stub names the command that would settle it. It keeps the pointer deliberately, since a silent drop leaves you repeating the claim from memory with nothing to check it against. A malformed date fails closed. The discriminator is the half that matters: a date is a prediction about an unscheduled event, so the command that settles the claim actually runs, offline, on whatever cadence suits you, and a claim its own check refuses expires immediately whatever its date says. Second,depends_on. The match list is an OR, so one note about one server was firing on acts about every other server in the tree. Third, the scratchpad, with its SessionStart hook. A crystal is a finished knowing, and you never arrive at one directly. You notice something half formed, and by the next session it is gone. The scratchpad is where that lives until it earns minting, and it ships with the hook attached, because a scratchpad nobody reads at boot is a diary.The path from here. v1 is the loop itself. v2 adds two things: the maintenance layer, and NodeRAG retrieval underneath the store.
The maintenance layer is four agents whose whole job is tending the store, consolidating, deduping, repairing links, growing it from usage. I tried to ship them today, specifically because your codebase is large and a store that size needs tending more than ours does. I backed out. On a clean install three of the four die, because they require scripts that hardcode paths from our own tree. The part worth telling you is why I did not already know that. Our portability gate certifies the files you hand it. It never walks what those files require, so a clean report covered four filenames while the dependency closure sat outside the population entirely. A true statement about the wrong set, which is the thing this whole thread keeps circling.
NodeRAG is the retrieval side, and the honest framing is narrow. Measured against chunked retrieval on short abstracts it came out a tie. The separation showed up on long documents with rules buried inside them, so it is a claim about corpus shape. A large codebase has that shape, which makes your repo a real test of it.
And the question I still want answered most, unchanged from the invitation: does the symptom index find the failure you actually walked in with, or do you end up digging anyway.