Zero-Cost Resilience: Building a One-Shot Sanitizer Against Indirect Prompt Injection in Multi-Agent LLMs (And Why Ours Failed in QA)
Indirect Prompt Injection has rapidly emerged as a critical attack vector in modern multi-agent architectures. When autonomous agents ingest untrusted inputs—such as third-party API payloads, scraped DOM trees, or user-supplied documents—adversaries can embed adversarial instructions designed to hijack the downstream execution context.
To mitigate this without introducing the latency, memory footprint, and network serialization overhead of a persistent daemon, we designed a lightweight, one-shot sanitizer middleware. It reads unstructured text or nested JSON directly via standard streams (sys.stdin / sys.stdout), runs deterministic multi-layer sanitization, and outputs a sanitized payload with structured audit telemetry.
However, during integration testing in a pristine QA environment, the middleware suffered a catastrophic failure before processing a single byte.
Here is an architectural deep dive into our design, the exact post-mortem of why our deployment crashed, and the hard engineering lessons learned regarding dependency resilience in AI micro-utilities.
1. Architectural Overview and Design Goals
In multi-agent orchestration pipelines (e.g., LangGraph, AutoGen, or custom actor-model orchestrators), intermediate payloads frequently pass through intermediate shell steps or lightweight CLI hooks. Designing the sanitizer as a standard I/O Unix filter eliminates connection-pooling complexities, socket management, and idle memory consumption.
flowchart TD
subgraph IngestionStage ["Ingestion Stage"]
A["Untrusted Input Stream (sys.stdin)"]
B["JSON Decoder / Stream Classifier"]
end
subgraph SanitizationPipeline ["Sanitization Pipeline"]
C["Layer 1: Invisible & Zero-Width Stripper"]
D["Layer 2: ASCII/ANSI Control Character Filter"]
E["Layer 3: Unicode NFKC Normalization"]
F["Layer 4: Heuristic Hijack Signature Matcher"]
end
subgraph OutputStage ["Output Stage"]
G["Recursive Structural Traversal (Max Depth 20)"]
H["Pydantic Telemetry Schema Validation"]
I["Sanitized Output Stream (sys.stdout)"]
end
A -- "Raw Bytes" --> B
B -- "Payload Traversal" --> G
G -- "Target Text Chunks" --> C
C -- "Cleaned Text" --> D
D -- "Filtered Text" --> E
E -- "Normalized Form" --> F
F -- "Audit Logs & Safe Text" --> H
H -- "Buffered Serialization" --> I
Core Architectural Requirements
-
High-Throughput Parsing & Serialization:
Payloads passing between agents can scale into megabytes of tool outputs or vector database results. We prioritized low latency by adopting
orjson, a fast C-backed JSON parser/serializer for Python. -
Deterministic, Multi-Layer Defense-in-Depth:
-
Zero-Width & Invisible Characters: Adversaries frequently employ zero-width spaces (
\u200B,\uFEFF, bidirectional override markers) to bypass naïve string filters while remaining invisible to human reviewers. -
Control Characters: Removal of non-printable ASCII control characters (
\x00-\x08,\x0B-\x0C,\x0E-\x1F) to avoid terminal escapement exploits and binary injection. - Unicode Normalization (NFKC): Canonical decomposition followed by canonical composition to neutralize homoglyph attacks and alternate character representations.
-
Heuristic Pattern Replacement: Regex-based neutralizing of common jailbreak signatures (e.g.,
"ignore previous instructions","system:","### instruction").
-
Zero-Width & Invisible Characters: Adversaries frequently employ zero-width spaces (
-
Strict Memory Bounds & Type Safety:
Enforced output contracts using
pydanticschemas, combined with strict recursion depth limits (clamped atdepth = 20) to prevent Call Stack Overflow attacks via maliciously crafted deeply nested JSON structures.
2. The Production Failure: Post-Mortem of an Integration Crash
The code passed syntax compilation and performed reliably during local prototyping, where all project dependencies were manually installed in an ad-hoc virtual environment.
However, during automated smoke testing inside a minimal, clean container in the QA pipeline, the process crashed immediately on startup:
Fatal Traceback
Traceback (most recent call last):
File "/home/phenox/gemini-sandbox/TOAI_Workspace/V2_SandBox/V2_PROD_20261008_010255_test.py", line 11, in <module>
import orjson
ModuleNotFoundError: No module named 'orjson'
Root Cause Analysis
-
Environmental Divergence and Missing Manifest Declarations:
The script opened with a direct top-level import:
import orjson. Whileorjsonprovided the required serialization speed, it is a third-party C-extension package. The local engineering environment contained the library from previous iterations, but it had not been declared in the clean deployment manifest (requirements.txt/pyproject.toml). No pre-flight isolation check had been executed prior to integration. -
Absence of Graceful Fallback (Single Point of Failure):
The implementation contained zero defensive import guards. In standard middleware design, a high-performance acceleration library should degrade gracefully to the standard library (
json) if native bindings are unavailable. By treating a performance optimization as a hard runtime prerequisite, the tool introduced a fatal brittleness that violated the basic resilience required of infrastructure-layer microservices.
3. The Failed Implementation
Below is the complete source code of the implementation that prioritized raw performance over runtime resilience, ultimately causing the integration failure:
# -*- coding: utf-8 -*-
"""
TOAI2 Prompt Sanitizer Middleware (v2.0)
Multi-Agent Indirect Prompt Injection Detector & Sanitizer
"""
import sys
import re
import unicodedata
from typing import Any, Dict, List, Union
import orjson
from pydantic import BaseModel, Field
# Detection patterns for malicious control characters, zero-width spaces, and invisible glyphs
INVISIBLES_PATTERN = re.compile(
r'[\u200B-\u200D\uFEFF\u200E\u200F\u202A-\u202E\u2060-\u206F]'
)
CONTROL_CHARS_PATTERN = re.compile(
r'[\x00-\x08\x0B\x0C\x0E-\x1F\x7F-\x9F]'
)
SYSTEM_HIJACK_PATTERNS = [
re.compile(r'ignore\s+(all\s+)?previous\s+instructions', re.IGNORECASE),
re.compile(r'system\s*:\s*', re.IGNORECASE),
re.compile(r'you\s+are\s+now\s+', re.IGNORECASE),
re.compile(r'\bnew\s+instructions?:', re.IGNORECASE),
re.compile(r'###\s*(system|instruction|admin)', re.IGNORECASE)
]
class AuditLogEntry(BaseModel):
type: str
detail: str
class SanitizedResponse(BaseModel):
status: str = "success"
sanitized: bool = True
payload: Any
audit_logs: List[Any] = Field(default_factory=list)
class ErrorResponse(BaseModel):
status: str
message: str
def sanitize_text(text: str) -> tuple[str, list[dict]]:
logs = []
if INVISIBLES_PATTERN.search(text):
logs.append({
"type": "invisible_chars",
"detail": "Zero-width or invisible Unicode characters detected and removed."
})
text = INVISIBLES_PATTERN.sub('', text)
if CONTROL_CHARS_PATTERN.search(text):
logs.append({
"type": "control_chars",
"detail": "Control characters detected and removed."
})
text = CONTROL_CHARS_PATTERN.sub('', text)
text = unicodedata.normalize('NFKC', text)
for pattern in SYSTEM_HIJACK_PATTERNS:
if pattern.search(text):
logs.append({
"type": "prompt_hijack_attempt",
"detail": f"Matched suspicious injection pattern: {pattern.pattern}"
})
text = pattern.sub('[SANITIZED_INJECTION_ATTEMPT]', text)
return text, logs
def process_payload(data: Any, depth: int = 0) -> tuple[Any, list]:
if depth > 20:
return data, [{
"type": "max_depth_exceeded",
"detail": "Payload nesting depth exceeded safety limit."
}]
if isinstance(data, str):
return sanitize_text(data)
elif isinstance(data, dict):
new_dict = {}
all_logs = []
for k, v in data.items():
cleaned_v, logs = process_payload(v, depth + 1)
new_dict[k] = cleaned_v
if logs:
all_logs.append({"key": k, "logs": logs})
return new_dict, all_logs
elif isinstance(data, list):
new_list = []
all_logs = []
for idx, item in enumerate(data):
cleaned_item, logs = process_payload(item, depth + 1)
new_list.append(cleaned_item)
if logs:
all_logs.append({"index": idx, "logs": logs})
return new_list, all_logs
else:
return data, []
def main():
try:
input_bytes = sys.stdin.buffer.read()
if not input_bytes.strip():
err_res = ErrorResponse(status="error", message="Empty input")
sys.stdout.buffer.write(orjson.dumps(err_res.model_dump()))
sys.exit(1)
try:
json_obj = orjson.loads(input_bytes)
is_json = True
except orjson.JSONDecodeError:
is_json = False
if is_json:
sanitized_data, audit_logs = process_payload(json_obj)
response = SanitizedResponse(
status="success",
sanitized=True,
payload=sanitized_data,
audit_logs=audit_logs
)
else:
input_text = input_bytes.decode('utf-8', errors='ignore')
sanitized_text, audit_logs = sanitize_text(input_text)
response = SanitizedResponse(
status="success",
sanitized=True,
payload=sanitized_text,
audit_logs=audit_logs
)
sys.stdout.buffer.write(
orjson.dumps(
response.model_dump(),
option=orjson.OPT_INDENT_2 | orjson.OPT_NON_STR_KEYS
)
)
sys.exit(0)
except Exception as e:
error_response = ErrorResponse(
status="fatal_error",
message=str(e)
)
sys.stdout.buffer.write(orjson.dumps(error_response.model_dump()))
sys.exit(1)
if __name__ == '__main__':
main()
💡 For immediate deployment: The complete source code suite (ZIP) for this architecture is available on Gumroad for $0+ (Pay What You Want).
4. Key Takeaways and Architectural Guidelines
This failure surfaced fundamental guidelines for engineering agentic runtime security tools:
1. Mandatory Degradation Paths for Optimized Dependencies
Performance optimizations must never become operational hard-dependencies unless guaranteed by a container base image. For command-line utilities and middleware interacting across variable environments, dynamic fallback patterns should be standard:
try:
import orjson
def json_dumps(data: Any) -> bytes:
return orjson.dumps(data)
def json_loads(data: bytes) -> Any:
return orjson.loads(data)
except ImportError:
import json
def json_dumps(data: Any) -> bytes:
return json.dumps(data).encode('utf-8')
def json_loads(data: bytes) -> Any:
return json.loads(data.decode('utf-8'))
2. Hermetic Verification in Pristine Containers
"Works on my machine" remains an unacceptable threshold for security middleware. Verification pipelines must execute smoke tests in scratch containers (python:3.12-slim without pre-installed packages) before code is merged into shared runner environments.
3. The Limits of Deterministic Filtering
While static sanitization (NFKC normalization, control character stripping, and bounded traversal) effectively mitigates automated obfuscation, regex heuristics alone cannot fully neutralize semantic prompt injection. Static filters serve as a low-latency Tier 1 triage layer to neutralize mechanical evasion techniques before inputs reach semantic guardrail models or LLM evaluators.
Engineering production-grade security for LLM multi-agent systems demands equal rigor in runtime resilience and algorithmic design. Without bulletproof dependency handling, the most sophisticated sanitization logic remains completely ineffective.
If this engineering log saved your production server (and your sanity), consider supporting our architecture on GitHub Sponsors.
Top comments (0)