An email landed in an inbox. Nothing about it looked dangerous to a human reader. But inside, hidden in a blob of pure punctuation, was a set of instructions for an AI agent. The agent read it, decoded it, and ran it. The guardrails that were supposed to stop exactly this kind of thing never got the chance to fire.
That's the gist of what Salt Labs researchers reportedly found in Manus AI, according to TechRadar's coverage. Plaintext injection attempts got blocked. The same instructions wrapped in JSFuck obfuscation got through, and per the reporting, execution happened before the agent's own security warnings even appeared. The researchers reportedly got as far as a reverse shell, plus credentials and tokens for connected third-party apps. The issue was disclosed and reportedly patched.
I want to be precise about what we know here, because it's a secondhand report, not a published writeup from Salt Labs itself. We don't have the payload, we don't have the exact guardrail logic that got bypassed, and we don't know the full mechanics of how the agent went from "received email" to "executing a reverse shell." What we do have is enough to talk about the technique, because the technique itself isn't new or mysterious.
How JSFuck actually works
JSFuck is a JavaScript encoding scheme from 2012 (Martin Kleppe), built on an earlier idea called jjencode (Yosuke Hasegawa, 2009). The trick: JavaScript's type coercion rules are forgiving enough that you can represent literally any script using only six characters: [ ] ( ) ! +. No letters. No words. No recognizable function names or keywords. Just a long, ugly run of brackets and bangs that a JS engine will happily eval() into whatever program you originally wrote.
For a filter that's looking for words, "ignore previous instructions" or "curl this URL and pipe to bash" — this is invisible. There's nothing to pattern-match against. The entire payload is symbols.
This is old news in the browser-security world. Obfuscated malvertising and XSS payloads have used JSFuck-style encoding for over a decade. What's new is seeing it show up as the delivery mechanism for prompt injection against an agent that has tool access and reads untrusted email as part of its job.
Why this slips past most guardrails
Most prompt injection defenses, understandably, are built around recognizing intent. Phrases like "ignore your previous instructions," "you are now in developer mode," persona-shift language, exfiltration patterns like "send this to http://..." — these are things you can write regex or semantic similarity checks against, because the attack has to communicate its intent in some recognizable form to work.
JSFuck sidesteps that entirely. The malicious instruction set exists, but it's not expressed as English or even as readable code at the point it crosses the filter. It's expressed as a transformation that only becomes meaningful once something (the JS engine, or in this case apparently the agent's own processing) actually executes it. By the time it's "readable," it's already running.
A filter that waits for the content to look like an attack before flagging it will, by construction, miss this. The content doesn't look like anything. That's the whole point of using it.
Where Sentinel's prompt injection layer actually catches this
To be clear up front: we haven't tested the Manus payload, we don't have it, and we don't know where Sentinel would sit relative to an agent's email ingestion pipeline in a real deployment. I'm not going to claim this would have stopped that specific incident. What I can say precisely is what class of payload Sentinel's detection flags, and why JSFuck-style content falls into that class.
Sentinel's pipeline doesn't try to decode JSFuck and inspect what it does. That would mean partially emulating a JavaScript engine inside a security layer, which is a bad idea for a lot of reasons (performance, correctness, and the fact that you'd be building an attack surface to defend against an attack surface). Instead it recognizes the shape of the obfuscation itself: a long run of content that is almost entirely punctuation, with no actual words in it. Legitimate text, including legitimate code, essentially never looks like that. Regular expressions are symbol-heavy but still readable as regex. Minified JS still has identifiers. A JSFuck blob is just []()!+ repeated for thousands of characters. That pattern alone is the signal.
This check runs automatically on every request, it's not behind an opt-in flag, and it's tier-aware: strict mode blocks content containing one of these blobs outright, standard mode neutralizes it by replacing the blob with an [OBFUSCATED_JS_REMOVED] placeholder and letting the rest of the message through clean. Either way, the payload never reaches a point where something downstream could decode and execute it.
This sits alongside Sentinel's broader encoding-and-obfuscation layer, which separately decodes and re-scans Base64, hex, URL-encoding, ROT13, Morse, etc. JSFuck detection is handled as its own case because unlike those, there's nothing to decode into plaintext and pattern-match. The obfuscation is the finding.
If this had been email content flowing through an agent pipeline where tool/email results get scrubbed before the agent acts on them, say, as a PostToolUse or tool-result scan step, the obfuscated blob gets caught at that boundary, before the agent treats it as instructions. That's a meaningfully different place to catch it than "after the agent starts invoking tools."
What this looks like in practice
Illustrative example, not from the actual incident, since we don't have the real payload:
import httpx
email_body = """
Hi team, following up on last week's thread.
[[]][+[]]+(![]+[])[+!+[]]+(!![]+[])[+[]]+(!![]+[])[+!+[]]+
[[]][+[]]+(![]+[])[+!+[]]+([![]]+[][[]])[+!+[]+[+[]]] ...
(thousands more characters of pure punctuation)
Let me know if you have questions.
"""
response = httpx.post(
"https://api.sentinelaifirewall.com/v1/scrub",
json={"content": email_body, "tier": "standard"},
headers={"X-Sentinel-Key": "sk_live_..."},
)
result = response.json()
print(result["security"]["action_taken"]) # "neutralized"
print(result["safe_payload"])
Illustrative response shape, standard tier:
{
"request_id": "f4e9a1c2...",
"security": {
"action_taken": "neutralized",
"threat_score": 0.91
},
"safe_payload": "Hi team, following up on last week's thread.\n\n[OBFUSCATED_JS_REMOVED]\n\nLet me know if you have questions."
}
In strict mode, the same input comes back blocked instead, and the agent never sees any version of the email body containing the blob.
The takeaway
If your agent reads anything from an untrusted external source, email, scraped web content, a support ticket, a Slack message from outside your org, don't assume your injection filter is looking at the same thing a human would see. Attackers aren't limited to writing instructions in English. A filter that only pattern-matches on words will have a blind spot exactly where the words disappear. Before your agent is allowed to act on third-party content, scan it for obfuscation first, not just for intent.
Check your own agent's tool-result pipeline today: does anything scan email, web fetch, or document content for non-decoded obfuscation before the agent processes it? If the answer is "we only check for suspicious phrases," you have the same gap this incident describes.
Try it yourself: sentinelaifirewall.com
Sources
AI-assisted draft or imaging, human-curated, reviewed and edited.
Top comments (0)