DEV Community

Cover image for Treat AI Security Findings as Untrusted: Building an OS-Enforced Evidence Pipeline
zimei07
zimei07

Posted on

Treat AI Security Findings as Untrusted: Building an OS-Enforced Evidence Pipeline

Nothing drains an on-call engineer's sanity faster than a 3:00 AM P1 alert triggered by an LLM hallucinating an exploit chain across three decoupled microservices. You scramble out of bed, parse frantic Slack threads, and trace code execution paths only to discover the agent invented a phantom boundary breach that physical network topologies render impossible. The real crisis in modern engineering workflows isn't generating candidate vulnerabilities—any foundation model can spit out dozens of plausible code smells in thirty seconds. The crisis is distinguishing a verified, reproducible boundary collapse from an agent's persuasive fiction without granting untrusted repository scripts ambient access to your developer credentials, SSH keys, or cloud infrastructure.

When our team evaluated cloudflare/security-audit-skill for our automated review pipeline, we were looking for an architecture that solved this exact failure mode. Too many off-the-shelf security agents fall victim to confirmation loops: one agent proposes an exploit narrative, a second summarizes it, and the resulting Markdown report presents an unvalidated hypothesis as ground truth.

A production-grade audit pipeline cannot operate on trust. It must enforce distinct trust domains between discovery, execution, and verification.

The Failure Mode: Phantom Exploits and Hallucinated Breaches

Toy audit workflows collapse discovery, testing, and reporting into a single interactive execution loop. One model instance scans the repository, flags an unvalidated sink, optionally spins up a local script, and drafts an executive summary stamped with a CVSS score. This naive pipeline breaks down across three fundamental operational axes:

  1. Confirmation Bias in Context: The discovering agent naturally defends its original premise, interpreting ambiguous logs or partial execution failures as positive exploit confirmations.
  2. Ambient Credential Hijacking: When agents execute repository-controlled build scripts, npm postinstall hooks, or unit tests to "verify" an exploit, they run with the host machine's ambient authority—inheriting SSH agent sockets, AWS metadata service endpoints, and local secrets.
  3. Unstructured Markdown Amnesia: Standard Markdown reports fail to preserve provenance. Downstream triage teams cannot determine whether a finding stems from a verified input-output trace or speculative static pattern matching.

Integrating security-audit-skill into our internal harness shifted our operational baseline. Instead of relying on a single agent run, the workflow partitions discovery into distinct lifecycle phases: reconnaissance, coverage-driven hunting, candidate validation, structured schema generation, and adversarial record verification. Crucially, validation runs on a fresh, isolated verifier whose explicit objective is to disprove the candidate exploit rather than confirm it.

In our pipeline, the primary deliverable is never the human-readable REPORT.md. The real contract is the structured, machine-verifiable finding record. The skill enforces three unambiguous lifecycle states:

  • confirmed: Supported by a full, deterministic source trace and an observed, reproducible boundary failure.
  • needs_validation: Pinned to an explicit, unresolved dependency or unreachable external state, explicitly omitting arbitrary severity ratings.
  • rejected: Formally disproved by the adversarial verification harness.

Treating suspicious sinks as unverified leads rather than immediate vulnerabilities eliminates the alarm fatigue that paralyzes security teams.

Enforce Execution Boundaries at the OS Layer

Prompt engineering cannot replace an operating system boundary. Instructing an agent to "never run malicious code or access production credentials" is wishful thinking, not defense in depth. When an audit agent exercises target code, test fixtures, or build dependencies, that execution must occur within an OS-enforced, network-isolated sandbox.

Target repositories must never reach internal network interfaces, Docker daemons, or host filesystem hierarchies. The sandbox must sanitize environment variables, enforce strict CPU and memory budgets, drop dangerous syscalls, and pin filesystem writes to a clean temporary workspace.

Below is the production-grade runner policy we use to sandbox local agent-driven audits before any dynamic validation script touches target source code:

# audit-runner-policy.yaml
run:
  working_directory: /workspace/target
  writable_paths:
    - /workspace/audit-output
    - /tmp/audit
  read_only_paths:
    - /workspace/target
  network:
    enabled: false
    loopback_only: false
  environment:
    clear: true
    allow:
      - HOME=/tmp/audit-home
      - PATH=/usr/local/bin:/usr/bin:/bin
      - LANG=C.UTF-8
      - NODE_OPTIONS=--no-deprecation --max-old-space-size=3072
      - PYTHONDONTWRITEBYTECODE=1
    deny_prefix:
      - AWS_
      - GITHUB_TOKEN
      - NPM_TOKEN
      - GCP_
      - DOCKER_
  resources:
    cpu_seconds: 900
    memory_mb: 4096
    process_limit: 128
    file_size_mb: 256
    disk_quota_mb: 2048
  mounts:
    deny:
      - /home
      - /root
      - /run/secrets
      - /var/run/docker.sock
      - /proc/sys
      - /sys
      - ~/.ssh
      - ~/.aws
      - ~/.config/gcloud
  syscalls:
    deny:
      - socket(AF_INET, *)
      - socket(AF_INET6, *)
      - connect
      - ptrace
      - mount
Enter fullscreen mode Exit fullscreen mode

Whether backed by gVisor, Firecracker microVMs, or hardened seccomp-bpf profiles, the core rule remains absolute: target code must operate inside an airtight execution jail. An unprofiled Docker container will not save you when an upstream dependency runs a hostile build hook designed to probe container escapes.

To bootstrap the skill into our CLI workflow, we install it directly into our agent runtime:

npx skills add https://github.com/cloudflare/security-audit-skill \
  --skill security-audit
Enter fullscreen mode Exit fullscreen mode

We then invoke the auditing agent, enforcing an explicit, mounted output destination completely isolated from the target source tree:

security audit this codebase; write artifacts to /workspace/audit-output
Enter fullscreen mode Exit fullscreen mode

The resulting /workspace/audit-output directory serves as our durable evidence vault. It preserves the coverage ledger, raw findings.json manifests, schema assertions, and downstream diffs. If an alert cannot point to concrete artifacts in this ledger, it does not get escalated.

Coverage Ledgers: Stop Relying on Prose Checklists

Traditional audit reports rely on fuzzy, unfalsifiable assurances: "Audited auth middleware; reviewed SQL queries; verified dependency graph." In an automated CI/CD pipeline, this narrative prose is useless. As soon as a pull request merges, past manual audit notes turn into stale documentation debt.

The skill replaces subjective summaries with an auditable coverage ledger. The ledger maps explicit AST nodes, functions, and file paths to their verification status, executing schema validation on every update. When a team modifies internal token-refresh routines or introduces new API routes, subsequent agent sweeps pinpoint unexamined execution paths instead of re-auditing pristine code or declaring the repository secure by omission.

This precision introduces clear engineering trade-offs:

  • Compute Overhead: Running independent hunter-verifier pairs across an immutable schema demands significantly more LLM token volume and compute cycles than a single-pass summary.
  • Network Isolation Constraints: Strict network-disabled sandboxes prevent dynamic verification of workloads dependent on external third-party APIs.

Never disable the network boundary simply to force a proof-of-concept to run. If an exploit path depends on third-party webhook callbacks or live SaaS endpoints, document the exact missing dependency as needs_validation. You can then graduate that isolated finding to a dedicated staging harness configured with synthetic credentials, explicit egress filtering, and mock connection-pool exhaustion fixtures. Most agent-generated exploit proofs completely overlook distributed system failure modes—such as connection-pool deadlocks, 429 backoff storms, or cache race conditions. Testing those dynamics requires deliberate staging infrastructure, not a compromised local sandbox.

Sanitizing Model Egress and Upstream Data Paths

Securing your local sandbox solves only half of the threat model. When security agents audit proprietary enterprise codebases, your intellectual property leaves the developer environment through model API calls. Stack traces, database schemas, internal hostnames, and proprietary business logic are continuously dispatched over the wire.

Treat model routing as a critical attack surface. Our architecture terminates all outbound agent prompts through zero-data-retention (ZDR) gateways like B-Lost. This ensures proprietary source code inspected during automated audits cannot be retained for upstream model training, cached without encryption, or exposed via cross-tenant data leakage. Securing an audit pipeline requires protecting both sides of the boundary: OS-level sandboxing for the target runtime, and cryptographically verified, privacy-preserving transport for the agent's inferences.

At the end of the day, an agent's confidence score is completely irrelevant. A security finding is confirmed if, and only if, an automated evidence pipeline demonstrates an observable, bounded failure under controlled execution constraints. Everything else is just an unverified lead.

Where does your team draw the line when automating security reviews? Are you running dynamic exploit verifiers inside microVMs, or relying strictly on static rule engines and manual triage? Drop your pipeline architecture or production battle scars in the comments below.

Disclosure: Compute infrastructure and multi-model benchmark relays for this writeup are sponsored by b-lost.com — an enterprise AI gateway offering 0.8x official pricing, native prompt caching, and zero user-data retention. All benchmark metrics reflect independent reproducible testing.

Top comments (0)