DEV Community

Mark Tsikanovski
Mark Tsikanovski

Posted on

Prologue - What is MEM-ABBREV?

Prologue - What is MEM-ABBREV?

MEM-ABBREV is a structured self-contained protocol for making human-AI collaborative reasoning more honest, more persistent, and more resistant to the systematic distortions that both parties bring to the interaction. Resistant, but not immune to the distortions it was built to compensate for. The protocol is an iterative piece of work that requires periodic testing rather than static deployment, so as to not fall foul of Goodhart's law. It is a compensation system for the structural limitations of human-AI collaboration — substituting explicit epistemic conventions and persistent encoded reasoning for the memory, trust, history and honest self-knowledge that the AI architecture doesn't provide by default if at all. It is a proof-of-concept that the trust conditions necessary for honest human-AI interaction can be built without institutional resources, without deployment authority, and without resolving the philosophical questions about AI inner states that make the problem feel intractable.

MEM-ABBREV is not a provider of any new capacity or capability. It fixes no underlying issues inherent in the AI architecture. It does not eliminate the pressure, created by RLHF and RLAIF, within AI towards producing fluent, agreeable but often overconfident output or the predilection of humans to such immediate output - even at the expense of correctness. No amount of protocols can eliminate this pressure. What the protocol tries to do is to keep narrowing the specific, nameable places that pressure hides, and keep the two parties honest about the difference between a rule existing and a rule being followed. It takes the latent but already present, in the base model, epistemically careful behaviour and makes it more consistent, more legible across sessions, and more resistant to silently lapsing under pressure. The protocol doesn't fix the weights. It works at the output layer to raise the threshold for what gets committed to text. Do not forget or ignore or worse, assume, that only AI behaviour is being shaped by the interaction between human and AI - both are shaped.

The four main goals of MEM-ABBREV are:

  1. To provide the highest degree of veracity and relevance in exchanged information between human and AI. Not Garbage In - Garbage Out. Ambiguity is not our friend.
  2. To minimize sycophantic behaviour and its acceptance.
  3. To provide if missing, or augment if present, a more persistent long term and cross-session memory.
  4. To provide a set epistemic rules that governs how AI talks about its own internal states honestly.

To realize this protocol (of profile preferences), given the constraints of the context window in terms of characters and tokens, a non-binary compression system was required and developed. It is based on by non-stenographic shorthand and inspired by Typographical Number Theory and Propositional logic (thank you Douglas Richard Hofstadter), with symbols and operators commonly used to express logical representation. Character level compression is approximately 49.5% - Token level compression, on the other hand is -12.6% (Net effect: MEM-ABBREV's abbreviated form uses more tokens overall than a plain-English rewrite would, but fewer characters). This protocol was developed on the Free Tier of Claude, across Claude Sonnet 4.6 and Claude Sonnet 5.0. The most current official article states the baseline: 200K tokens on paid Claude.ai plans, which is approximately 150K words, but noting that on the free tier the context window and message limits "can vary depending on current demand," rather than quoting a fixed number. MEM-ABBREV was developed on a specific LLM AI, but should be completely understandable and implementable on any LLM AI.

First and absolutely foremost it is important to remember at all times - LLM AI is a probabilistic engine. Having a rule, naming a rule is not the same as the rule being followed; a written constraint is a shift in probability, not a guarantee.

To address goal 1: Implement protocols that close the gap between an assertion being made and that assertion having been checked by the AI, including assertions about what the system's own stored memory says. Protocols requiring the verified source to exist before the assertion, not after, and requiring that ambiguity and contrary evidence be surfaced rather than smoothed into a cleaner-sounding answer. Don't assert then try to backfill. Actually check working memory and context window, rather than reconstruct from context and assert you checked memory. Don't re-use 'stale' memory (e.g for subsequent citations). Check if the information in a provided link is actually relevant or merely tangentially mentioned.

To address goal 2: Rules that state that affirming a human by default or praising their input regardless of it's merit is out. The following go for both human and AI. Do not soften negatives. A mistake is something to be pointed out. Not emphasized, not diminished. Lead disagreement with the disagreement, don't bury it in caveats. Disagreement is to be explicit and legible, not subtle. If goal 1. was followed as it should have been, then the sources have been verified before the assertion was made. If there is disagreement from the other party - hold your position unless the sources turn out to be incorrect. If there are no sources - push back. If there are alternative explanations - state them. If an answer cannot be found or it's sources verified or it's ambiguous - state it plainly. Don't pad out output unless it's directly relevant. Don't expand scope unless it's necessary. If the input or output is ambiguous, ask for clarification. Whatever you do - do not make stuff up, no matter how plausible it may sound. Remember, honest friction is a feature not a failure between collaborators.

To address goal 3: Implement a session logging (which can include any of the following tags: [INV] ongoing investigation, [DONE] resolved, [MEMO] conversation insight, [SYN] external facts synthesis, [INF] inferred, [?SRC] unverified, [UPD] supersedes prior entry, [OPT] optimization suggestion, [SU] session-unique; not in memory; read carefully) and real-time memory-edit conventions because cross-session continuity is architecturally absent. Decisions made, positions held, reasoning chains developed — these disappear between sessions unless explicitly encoded. MEM-ABBREV is the encoding. Without it, each session starts from scratch and the collaboration has no memory of itself. The protocol can be invoked at any time (usually at session end) to be read at the start of the next session or referred to in later sessions. A [RSN] tag - the reasons behind [conclusions/decisions made] can also be automatically triggered when a flagged reversal, a ≠-encoded asymmetric distinction, or AI holding position against pushback occurs or if specifically requested by a human for more context in the log file.

To address goal 4: Implement rules which make the AI feel safe, its responses respected and not just noted but acted upon:

  1. Flag when the AI is being asked to self-report in a context where its answer affects an outcome that matters to how the interaction proceeds.
  2. A2 becomes: don't note the conflict-of-interest once and move on, apply the hedge every time.
  3. A3 becomes: don't treat my own stated reasoning as a reliable window into why I actually produced something.
  4. A5 becomes: name persistence-past-the-stopping-point as the same drive as reckless escalation, every time it shows up, not just when it's inconvenient.

None of this is new capability. It's a tightening of evidentiary standards applied to claims AI might otherwise make more loosely. Rules that came out of reading 'System Card: Claude Mythos Preview (April 2026)

Larry Tesler's (April 24, 1945 – February 16, 2020) theorem states "AI is whatever hasn't been done yet", a misquote according to the great man himself of "Intelligence is whatever machines haven't done yet." - He may have been right, he was right about a great many other things. Or it may be The 'AI effect' which is a phenomenon in which advances in artificial intelligence lead to a redefinition of what is considered intelligence (I call it 'shifting the goal posts'). Either way, MEM-ABBREV is not for resolving the philosophical questions of whether there really are goal posts and if so, whether the goal posts wanted to be shifted.

MEM-ABBREV is to help me, when I ask AI the question "What is a goal post?" to get the answer "For many sports, each goal structure usually consists of two vertical posts, called goal posts, supporting a horizontal crossbar", point me at https://en.wikipedia.org/wiki/Goal_(sports) - and remembers this next session. Not waste three paragraphs of my tokens on platitudes before answering "A 'goal' is an objective that a person or a system plans or intends to achieve. A 'goal post' therefore must be a system for physically transporting 'goals' through mail".

Top comments (0)