DEV Community

Cover image for Engram Corruption: What Happens When a Skill Container Doesn't Own Its Payload
owly
owly

Posted on

Engram Corruption: What Happens When a Skill Container Doesn't Own Its Payload

The setup

In a modular AI framework like LivinGrimoire, behavior comes from small, swappable units called skills. A Brain holds them in lobes, and a skill's input() runs on every think cycle.

On top of plain skills sits a second tier I'll call AH skills: skills that manage other skills. Examples:

  • AHNight / AHDay add and remove sets of skills depending on the time of day.
  • AHDebuff / AHBuff unplug or plug in skill sets on a timer or a spoken command.
  • AHHibernate wipes the brain to save compute, then restores it, in the spirit of Soulkiller from Cyberpunk.

The skills these managers control are their payload. The trouble starts when two managers, or a manager and some other piece of machinery, both think they control the same payload skill.

The engram

The original AHHibernate works through an engram: a snapshot of every skill in every lobe.

self.engram = BrainEngram(self.brain)   # snapshot everything
self.engram.remove_skill(self)          # except me
self.algPartsFusion(4, APHibernate(self, self.brain))  # wipe the brain
Enter fullscreen mode Exit fullscreen mode

On wake, APImprintEngram replays the snapshot back into the brain with plain add_skill calls.

This is a global operation. It captures skills it has no relationship with, including payload skills that belong to other managers. That is the root of everything below.

Failure mode 1: duplication

brain.add_skill calls skill.manifest(), the skill's life-cycle hook. A container like AHNight uses manifest() to load its payload.

Now trace a wake-up:

  1. The engram holds [AHNight, NightSkillA, NightSkillB].
  2. It restores AHNight first. manifest() runs and adds NightSkillA and NightSkillB.
  3. The engram keeps going and restores NightSkillA and NightSkillB again, from its own list.

Result: each payload skill is in the brain twice, and a skill that's in the brain twice responds twice. Your companion says everything double, or an action fires double.

Nothing errors. The lobe's add_regular_skill has no duplicate check. The only symptom is behavior that's slightly wrong.

Order matters, which makes it worse. If the payload happens to be restored before its container, the container's manifest() then adds them a second time instead. Either way you get two copies, and which skills double depends on list order.

Failure mode 2: stale restore

Duplication is at least loud. This one is quiet.

Time-of-day containers change their payload on a schedule. Suppose a nag-cooldown or hibernation starts at 17:59, snapshotting the day payload, and sunset hits at 18:00 in the middle of it.

  • During the cooldown, the container does its job and swaps day skills for night skills.
  • At wake, the engram restores its old snapshot, so the day skills come back at night.

The snapshot was true when taken and false by the time it was applied. Both containers now believe they're in the right state, since each updated its own flags correctly, so nothing corrects it until the next flip, up to about 12 hours later.

The cause is that two owners wrote to the same skills. The container decided what should be loaded, and the hibernator overwrote that decision with a stale copy.

Failure mode 3: flags that lie

Containers track their own state, like engaged = True. When a restore re-adds a container, manifest() runs again and recomputes the flag from the current time. The payload, restored from an older snapshot, may no longer match what the flag claims. The skill and its manager disagree about reality, and if not self.engaged guards then skip the repair that would have fixed it.

Failure mode 4: leaks and orphans

Global wipes also hit skills nobody asked them to touch. A hardware output skill like the console printer gets wiped along with everything else, which means announcements made during hibernation have nowhere to go. A skill removed by one manager can be quietly resurrected by another.

Things that look like fixes but aren't

I tried several patches in this ecosystem, and each one had a cost.

  • Existence checks everywhere (brain.contains_skill before every add): stops duplicates, and is still worth having as a safety net. It does nothing about stale restores, because the stale skill genuinely isn't present at that moment.
  • Marking containers "defcon" so hibernators skip them: protects the container but not its payload, so payload snapshots still go stale.
  • Branding payload with an ownership tag: works, but needs every container to participate, and the hibernator ends up knowing about other managers' conventions.
  • Per-tick re-verification: self-heals, at the price of constant checking for a rare event.

All of these treat a symptom of shared ownership instead of removing it.

The fix: single ownership

The rule that eliminates the whole class of bugs:

A skill is controlled by exactly one manager. A manager only ever touches skills that were handed to it.

Here is the shape of a hibernating container built on that rule:

class AHHibernate(Skill):
    def __init__(self, brain, hibernation_minutes=30, standby_minutes=2):
        super().__init__()
        self.brain = brain
        self.skills: list[Skill] = []          # the payload: mine, and only mine
        self.hibernating = False
        self.tg = TimeGate(hibernation_minutes)
        self.standby = AXStandBy(standby_minutes)

    def add_skill(self, *skills):
        for skill in skills:
            if skill is not self and skill not in self.skills:
                self.skills.append(skill)
        return self

    def manifest(self):                         # adopt my payload
        for skill in self.skills:
            if not self.brain.contains_skill(skill):
                self.brain.add_skill(skill)

    def ghost(self):                            # payload leaves with me
        for skill in self.skills:
            self.brain.remove_skill(skill)

    def input(self, ear, skin, eye):
        if not self.hibernating:
            if ear == "hibernate" or self.standby.standBy(ear):
                self.hibernating = True
                self.tg.openForPauseMinutes()
                self.algPartsFusion(4, APSkillsRemover(self.brain, *self.skills))
        elif self.tg.isClosed() or ear == "wake up":
            self.hibernating = False
            self.algPartsFusion(4, APAddUniqueSkills(self.brain, *self.skills))
Enter fullscreen mode Exit fullscreen mode

What changed conceptually:

Old engram model Single-ownership model
Snapshots the whole brain Holds a list it was given
Restores from a stale copy Restores its own live list
Can capture other managers' payload Never sees skills it doesn't own
Needs defcon/brand exemptions Needs nothing, the boundary is structural
Duplication depends on restore order Idempotent restore via APAddUniqueSkills

The container is the source of truth for its payload, and nothing else writes to it. A day/night swapper and a hibernator can coexist because they never share a skill, so there's no snapshot to go stale and no second manifest() to double-add.

The two supporting pieces

1. An existence check on the brain, so adds are idempotent:

# Brain
def contains_skill(self, skill: Skill) -> bool:
    match skill.get_skill_lobe():
        case 1: return self.logicLobe.exist(skill)
        case 2: return self.hardwareLobe.exist(skill)
        case 3: return self.ear.exist(skill)
        case 4: return self.skin.exist(skill)
        case 5: return self.eye.exist(skill)
    return False
Enter fullscreen mode Exit fullscreen mode

Routing by the skill's own lobe attribute matters. Containers can hold skills from any lobe and of either type, and a check against just the logic lobe's regular list would miss them.

2. An idempotent adder algorithm part, so restoring never doubles anything:

class APAddUniqueSkills(AlgPart):
    def __init__(self, brain, *skills_to_add):
        super().__init__()
        self.brain = brain
        self.skills_to_add = list(skills_to_add)
        self.done = False

    def action(self, ear, skin, eye):
        for skill in self.skills_to_add:
            if not self.brain.contains_skill(skill):
                self.brain.add_skill(skill)
        self.done = True
        return ""

    def completed(self):
        return self.done
Enter fullscreen mode Exit fullscreen mode

Both pieces are cheap, and they turn "restore" from a destructive replay into a safe, repeatable operation.

A timing gotcha worth knowing

Skills must not be added or removed from a lobe while it's mid-think. Lobe.remove_logical_skill silently returns if _isThinking is set, and input() runs inside the think loop. That's why containers queue their changes through algorithm parts (APSkillsRemover, APAddUniqueSkills), which run after the loop finishes, rather than mutating the lobe directly from input(). If you call brain.remove_skill straight from input(), it does nothing and gives you no error.

The rules to take away

  1. One owner per skill. If a skill is in a container's payload, nothing else adds, removes, snapshots or restores it.
  2. Never snapshot what you don't own. A global engram is a liability the moment a second manager exists.
  3. Make adds idempotent. Check existence before adding, so order and repetition stop mattering.
  4. Defer structural changes to algorithm parts that run outside the think loop.
  5. Anything shared is a bug waiting for a schedule. A day/night flip, a timer, or a standby trigger will eventually land at the wrong moment, and shared ownership turns that moment into a corrupted brain.

Modular systems get their power from composability, and composability only holds up when each module's boundary is real. Handing a manager an explicit payload list keeps that boundary real. The bug it prevents does not show up in a quick test, only after a weekend of uptime when a sunset lands inside a cooldown.

Built on LivinGrimoire, a modular AI skill framework.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to