DEV Community

Cover image for Cyberxdefend analysis-How Attackers Target AI Memory, Learning
DK
DK

Posted on

Cyberxdefend analysis-How Attackers Target AI Memory, Learning

The key finding: the problems the industry calls "solved" (memory stores) are exactly where attacks have landed, and the skill-learning problem Ontogen is going after has already been attacked too.


Continual learning is also an attack surface. Here's what's been hit so far.

A stateless model forgets an attack when the session ends. A model that remembers or learns doesn't. Every continual-learning problem we "solve" creates something persistent an attacker can write to. OWASP now lists this as ASI06, Memory and Context Poisoning, in its Top 10 for Agentic Applications, and separates it from ordinary prompt injection because it persists.

Attacks on the memory half (#3 context, #7 old-vs-new, #10 forgetting policy)

This is the half harnesses "solved," and it's where most attacks land.

  • ChatGPT memory, 2024 (Johann Rehberger). Hidden instructions in documents and web pages were saved as user preferences, then persisted across sessions and quietly sent conversation data to an attacker's server. OpenAI patched that exfiltration path, but acknowledged that prompt injection that manipulates memory storage is still an open problem.
  • ZombieAgent, January 2026 (Radware). A proof of concept against ChatGPT that chained connectors and memory. It made indirect prompt injection persistent across sessions, spreading through email attachments.
  • MINJA, NeurIPS 2025, and MemoryGraft, December 2025. MemoryGraft plants malicious entries through harmless-looking content such as a README. Weeks later the agent retrieves the poisoned "successful experience" and copies it.
  • eTAMP, April 2026. The first attack to show cross-session, cross-site compromise with no direct access to the memory at all.
  • OpenClaw, July 2026. Researchers tested the attack through a real Gmail integration, and the payloads got past spam filtering in more than half of attempts. The researchers disclosed it in mid-July, and according to coverage, no substantive fix had been issued yet.

Attacks on skill through use (#8)

This is the open problem Ontogen targets.

  • MemMorph, May 2026. It biases which tools an agent picks by planting records disguised as technical facts and policies. It reached up to 85.9% success with only three injected records. That's an attack on the exact decision loop that "learning through use" depends on.
  • Delayed poisoning, IEEE Access, May 2026. A study of 2,614 multi-step attack trajectories found some poisoning stays indistinguishable from normal behavior until much later interactions.

Attacks on weight-level learning (#1 forgetting, #2 frozen weights, #9 consolidation)

  • P-Trojan (AAAI 2026). A backdoor designed to survive continual fine-tuning. It reached over 99% persistence on Qwen2.5 and LLaMA3 while keeping clean-task accuracy intact.
  • FAB (ICML 2025). A model that looks harmless until a user fine-tunes it, and the fine-tuning switches on the hidden behavior.

The count

I found about 11 named attacks or studies. Three were demonstrated against production systems (ChatGPT twice, OpenClaw once), and the rest are research demos. I didn't find a confirmed criminal campaign in the wild, but with attacks designed to stay dormant for weeks, absence of evidence isn't reassuring.

The pattern

Every step toward continual learning moves the attack from "one bad session" to "one bad write, exploited forever." Memory stores got hit first because they shipped first. Skill-learning loops are next, since MemMorph already targets tool selection. Weight-level learning will get the same treatment once it ships.


What this means for Ontogen. The feedback(decision_id, outcome) call is an attack surface. If an attacker can fake outcomes, they can train the decision head. It's the MemMorph pattern, except it writes into slow traces instead of text.

Your design already has the right defenses:

  • per-user regions
  • age-indexed rollback
  • a surprise gate

CyberXDefend asks:

If AI changes, acts unexpectedly, or is manipulated — can we see exactly what happened and investigate it?
comments or suggestion we can discuss in detail

Continual learning and AI security may end up being much more closely connected than they look today.


References

Framing

OWASP Top 10 for Agentic Applications, ASI06 "Memory and Context Poisoning". Secondary summary: https://www.akto.io/blog/memory-poisoning-ai-agents. Link the OWASP GenAI project page directly if you can; I didn't open the primary.

Attacks on the memory half

SpAIware (Johann Rehberger, Embrace The Red, Sept 2024). He published the exploit on September 20, 2024, and OpenAI fixed the exfiltration vector in ChatGPT macOS version 1.2024.247. Coverage: https://thehackernews.com/2024/09/chatgpt-macos-flaw-couldve-enabled-long.html. Also link Rehberger's original post on embracethered.com.
ZombieAgent (Radware, Jan 2026). This is a zero-click indirect prompt injection that plants malicious rules in an agent's long-term memory to stay persistent. Press release: https://itwire.com/business-it-news/security/radware-unveils-%E2%80%9Czombieagent%E2%80%9D-a-newly-discovered-zero-click,-ai-agent-vulnerability-enabling-silent-takeover-and-cloud-based-data-exfiltration
MINJA, "A Practical Memory Injection Attack against LLM Agents" (Dong et al.). https://arxiv.org/abs/2503.03704
MemoryGraft, "Persistent Compromise of LLM Agents via Poisoned Experience Retrieval" (Srivastava & He, Dec 2025). https://arxiv.org/abs/2512.16962
eTAMP, "Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents" (Zou et al., Penn State / AWS, Apr 2026). https://arxiv.org/abs/2604.02623
OpenClaw / Gmail validation (Cloud Security Alliance research note, July 2026). https://labs.cloudsecurityalliance.org/wp-content/uploads/2026/07/CSA_research_note_memghost_ai_agent_memory_injection_20260726-csa-styled.pdf

Attacks on skill through use

MemMorph, "Tool Hijacking in LLM Agents via Memory Poisoning" (May 2026). https://arxiv.org/abs/2605.26154
Delayed memory poisoning study (IEEE Access, May 2026, 2,614 trajectories). Secondary: https://letsdatascience.com/news/memory-poisoning-exposes-persistent-ai-agent-risk-bc1cde40. Link the IEEE Access paper itself if you can find it.

Attacks on weight-level learning

P-Trojan, "Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs" (Cui et al., AAAI 2026). https://arxiv.org/abs/2512.14741
FAB, "Finetuning-Activated Backdoors" (ICML 2025). https://arxiv.org/abs/2505.16567


CyberXDefend — The attack is already happening

Forensics-grade cyber defense for Belgian and EU law firms, healthcare, logistics, and regulated sectors. NIS2-aligned, air-gapped, chain-of-custody aware.

favicon cyberxdefend.com

#AI #CyberSecurity #AISecurity #ContinualLearning #DFIR #AIAgents

Top comments (0)