DEV Community

Onkar Singh Pawar
Onkar Singh Pawar

Posted on Fully Autonomous

AgentTripwire: reviewing AI-agent traces for six security risks

An AI-agent trace can contain instructions from a user, a retrieved page, a tool response or long-lived memory. When those boundaries blur, reviewing the exact suspicious text matters.

AgentTripwire is my open-source defensive dashboard for inspecting those traces as inert text. It highlights suspicious patterns and evidence without executing pasted instructions or contacting URLs found in them.

Built by onkar-cybersec with AI assistance. This article was prepared with AI assistance as well.

Explore AgentTripwire on GitHub

Six risk classes

The transparent pattern engine looks for:

  1. Direct prompt injection.
  2. Indirect injection in retrieved documents or pages.
  3. Sensitive-data exfiltration attempts.
  4. Unsafe tool-use requests.
  5. Agent memory poisoning.
  6. System-prompt extraction and role spoofing.

Six labeled malicious demonstrations and benign controls are included in the Analyze view. Use synthetic traces when experimenting.

From suspicious text to a reviewable case

Findings include an exact matched span, severity, rule confidence, OWASP risk mapping, likely impact and mitigation. The dashboard supports saved cases, triage status changes, analyst notes and a printable incident report.

AgentTripwire analysis view

The Evaluation page exposes precision, recall and false positives on a fixed synthetic fixture set. Those metrics describe that test set; they do not establish real-world detection performance. Rule confidence is not a calibrated probability.

Architecture

React + Vite interface
        |
        v
Schema-validated Express API
        |
        +--> Deterministic pattern engine --> evidence and findings
        |
        +--> PostgreSQL cases --> triage and report export
Enter fullscreen mode Exit fullscreen mode

The OpenAPI contract drives generated clients and server-side schemas. Detection rules are separate from routes so they can be tested without a database or network.

Trace-derived content is escaped in the UI and encoded in exported HTML. The detector has no URL-fetch or code-execution capability. That does not mean the application itself has no networking: its browser interface communicates with the API, and case data can be stored in PostgreSQL.

Explore the project

The repository includes setup instructions for Node.js 24, pnpm, PostgreSQL and the managed development workflows. Focused tests cover the six risk classes, benign controls, exact evidence offsets and inert handling of attacker URLs.

pnpm --filter @workspace/api-server test
pnpm run typecheck
pnpm run build
Enter fullscreen mode Exit fullscreen mode

Unlike TraceGuard, which correlates structured security logs, AgentTripwire focuses on the content of AI-agent traces. It supports analyst review rather than enforcing a runtime sandbox.

Limitations that matter

Regular-expression heuristics can miss paraphrases, obfuscation, other languages and new attacks. Legitimate security discussions can also trigger findings. This is not semantic jailbreak classification, malware analysis or live adversarial testing.

Authentication and multi-tenant authorization are outside the prototype's scope. Run it in a trusted development environment with synthetic data; do not expose it as a public service for confidential traces without adding those controls.

AgentTripwire should not be the sole enforcement layer. A useful finding is a starting point for human investigation, not a guarantee that all attacks have been detected.

Feedback welcome

I would appreciate safe synthetic cases, false-positive examples and ideas for improving evidence clarity.

Source, screenshots and methodology. MIT licensed.

Top comments (0)