Decepticon is an autonomous red-team agent built on LangChain and LangGraph. It chains reconnaissance, exploitation, and privilege escalation tools without human intervention. The project hit GitHub Trending at #13 for Python with 5,825 stars and 1,100 forks, making it a real-world case study of agentic orchestration in a domain where tool execution must be both autonomous and auditable.
The red-team use case exposes hard problems. How do you let an agent run nmap, parse results, decide on next steps, and recover when a tool fails, all while maintaining an audit trail? Decepticon's architecture reveals how LangGraph state machines handle multi-phase adversarial workflows where tool boundaries, failure recovery, and context preservation matter.
Why Offensive Security Needs Orchestration
Traditional penetration testing is a linear, human-driven process. An operator runs nmap, reads the output, picks an exploit, runs Metasploit, checks for success, and repeats. Each step depends on the previous one. The operator holds state in their head.
Autonomous agents flip this model. The agent must:
- Execute tools in sequence without losing context.
- Parse tool output and decide the next action.
- Handle failures, partial results, and dead-end attack paths.
- Maintain an audit trail for compliance and post-mortem analysis.
Decepticon uses LangGraph to model this as a state machine. Each node represents a phase (reconnaissance, exploitation, post-exploitation). Edges define transitions based on tool output. The graph holds state across phases, so the agent does not duplicate work or lose context when a tool fails.
LangGraph State Machine Architecture
LangGraph is a stateful orchestration layer on top of LangChain. It models workflows as directed graphs where nodes are functions and edges are conditional transitions. State flows through the graph, accumulating context as the agent progresses.
Decepticon's graph has three primary phases:
- Reconnaissance: Run nmap, parse open ports, identify services.
- Exploitation: Match services to exploits, execute Metasploit modules, check for shells.
- Post-Exploitation: Escalate privileges, pivot to other hosts, exfiltrate data.
Each phase is a node. The agent transitions to the next phase only if the current phase succeeds. If a tool fails, the graph can backtrack, retry with different parameters, or escalate to a human operator.
State Schema
The state object carries information between nodes. For Decepticon, this includes:
- Target metadata: IP address, hostname, OS fingerprint.
- Tool results: nmap output, open ports, service versions.
- Exploit history: Which exploits were attempted, which succeeded.
- Session handles: Active shells, credentials, pivot points.
LangGraph persists this state in memory or a database. If the agent crashes, it can resume from the last checkpoint without re-running reconnaissance.
Conditional Edges
Edges in LangGraph can be conditional. After reconnaissance, the agent checks if any exploitable services were found. If yes, it transitions to exploitation. If no, it either retries with different nmap flags or exits.
This is where orchestration gets interesting. The agent does not blindly execute a fixed sequence. It adapts based on tool output. If Metasploit fails, the graph can route to a fallback exploit or a manual review node.
Tool Boundary Design
Offensive security tools are dangerous. An agent that runs arbitrary commands can cause unintended lateral movement, data loss, or compliance violations. Decepticon needs tool boundaries that let the agent operate autonomously while preventing runaway execution.
Sandboxed Tool Execution
Each tool runs in a sandboxed subprocess. The agent passes input via stdin and reads output from stdout. The subprocess has a timeout. If the tool hangs, the agent kills it and logs the failure.
This prevents a single tool from blocking the entire workflow. It also isolates tool crashes. If nmap segfaults, the agent logs the error and moves to the next phase.
Allowlisted Commands
The agent does not execute arbitrary shell commands. It has an allowlist of approved tools (nmap, Metasploit, custom scripts). Each tool has a schema that defines valid arguments.
For example, nmap can only scan a predefined IP range. The agent cannot run nmap -A 0.0.0.0/0 and scan the entire internet. The schema enforces constraints at runtime.
Output Parsing and Validation
Tool output is untrusted. The agent parses it with a strict schema. If nmap returns malformed XML, the agent logs the error and retries. If Metasploit reports a shell but the session is not reachable, the agent marks the exploit as failed.
This prevents the agent from making decisions based on garbage data. It also creates an audit trail. Every tool invocation, input, and output is logged.
Failure Recovery and Backtracking
Red-team workflows are full of dead ends. An exploit might fail because the target is patched. A shell might die because the network is unstable. The agent needs a recovery strategy.
Retry with Exponential Backoff
If a tool fails, the agent retries with exponential backoff. The first retry happens after 1 second, the second after 2 seconds, the third after 4 seconds. After three retries, the agent gives up and transitions to a fallback node.
This handles transient failures (network glitches, rate limits) without wasting time on permanent failures (patched targets, wrong exploits).
Fallback Exploits
If the primary exploit fails, the agent tries a fallback. For example, if a remote code execution exploit fails, the agent tries a privilege escalation exploit. The graph defines fallback edges for each exploit.
This is where LangGraph shines. The agent does not need a hardcoded decision tree. The graph defines the fallback logic declaratively. If the primary edge fails, the graph follows the fallback edge.
Human-in-the-Loop Escalation
Some failures require human judgment. If all exploits fail, the agent escalates to a human operator. The graph pauses, logs the state, and sends a notification. The operator reviews the logs, decides on a next step, and resumes the graph.
This hybrid model balances autonomy and safety. The agent handles routine tasks. Humans handle edge cases.
Observability and Audit Trails
Offensive security is a regulated activity. Every action must be logged for compliance and post-mortem analysis. Decepticon's observability stack includes:
- Structured logs: Every tool invocation, input, and output is logged in JSON.
- State snapshots: The graph saves state after each phase. If the agent crashes, it can resume from the last snapshot.
- Metrics: Tool execution time, success rate, failure modes.
The logs are queryable. An operator can search for all nmap scans against a specific IP, or all Metasploit exploits that failed with a specific error.
Trace Propagation
LangGraph supports trace propagation. Each tool invocation gets a trace ID. The trace ID flows through the graph, linking all related actions. This makes it easy to reconstruct the attack path.
For example, if the agent escalates privileges on a target, the trace shows:
- nmap scan that discovered the target.
- Metasploit exploit that gained initial access.
- Privilege escalation script that gained root.
The trace is a directed acyclic graph (DAG) of tool invocations. It is the audit trail.
Deployment Shape
Decepticon runs as a long-lived service. The agent polls a queue for new targets, executes the workflow, and writes results to a database. The deployment has three components:
| Component | Role | Technology |
|---|---|---|
| Agent runtime | Executes LangGraph workflows, manages state | Python, LangChain, LangGraph |
| Tool sandbox | Runs offensive security tools in isolated subprocesses | Docker, systemd-nspawn |
| Observability backend | Stores logs, traces, and metrics | Elasticsearch, Prometheus, Grafana |
The agent runtime is stateless. It reads state from the database, executes the workflow, and writes state back. This makes it easy to scale horizontally. Multiple agent instances can process targets in parallel.
The tool sandbox is ephemeral. Each tool invocation gets a fresh container. After the tool exits, the container is destroyed. This prevents state leakage between invocations.
Likely Failure Modes
Autonomous red-team agents have sharp edges. Here are the failure modes to watch for:
Runaway Execution
If the agent's allowlist is too permissive, it can execute unintended commands. For example, if the agent can run arbitrary Python scripts, it can exfiltrate data or pivot to unintended hosts.
Mitigation: Strict allowlists, schema validation, and human-in-the-loop escalation for high-risk actions.
State Corruption
If the agent crashes mid-workflow, the state might be inconsistent. For example, the agent might log a successful exploit but fail to save the session handle.
Mitigation: Atomic state updates, checkpointing, and idempotent tool execution.
Tool Failures
Offensive security tools are brittle. They crash, hang, or return malformed output. If the agent does not handle these failures gracefully, the workflow stalls.
Mitigation: Timeouts, retries, fallback exploits, and structured error logging.
Compliance Violations
If the agent scans or exploits unintended targets, it can violate compliance policies or legal boundaries.
Mitigation: IP allowlists, pre-flight validation, and audit trails.
Code Snippet: LangGraph Workflow
Here is a simplified LangGraph workflow for Decepticon. It shows how the agent transitions between reconnaissance, exploitation, and post-exploitation phases.
from langgraph.graph import StateGraph, END
from typing import TypedDict
class AgentState(TypedDict):
target_ip: str
open_ports: list[int]
exploit_results: dict
session_handle: str | None
def reconnaissance(state: AgentState) -> AgentState:
# Run nmap, parse open ports
ports = run_nmap(state["target_ip"])
return {**state, "open_ports": ports}
def exploitation(state: AgentState) -> AgentState:
# Match ports to exploits, run Metasploit
results = run_metasploit(state["open_ports"])
return {**state, "exploit_results": results}
def post_exploitation(state: AgentState) -> AgentState:
# Escalate privileges, pivot
session = escalate_privileges(state["exploit_results"])
return {**state, "session_handle": session}
def should_exploit(state: AgentState) -> str:
return "exploitation" if state["open_ports"] else END
def should_post_exploit(state: AgentState) -> str:
return "post_exploitation" if state["exploit_results"].get("success") else END
workflow = StateGraph(AgentState)
workflow.add_node("reconnaissance", reconnaissance)
workflow.add_node("exploitation", exploitation)
workflow.add_node("post_exploitation", post_exploitation)
workflow.set_entry_point("reconnaissance")
workflow.add_conditional_edges("reconnaissance", should_exploit)
workflow.add_conditional_edges("exploitation", should_post_exploit)
workflow.add_edge("post_exploitation", END)
app = workflow.compile()
This workflow runs reconnaissance, checks if any exploitable ports were found, runs exploitation if yes, and escalates privileges if the exploit succeeded. The conditional edges handle failures gracefully.
Technical Verdict
Use Decepticon when:
- You need autonomous penetration testing for continuous security validation.
- You have a mature observability stack and can audit every tool invocation.
- You can enforce strict tool boundaries and IP allowlists.
- You have human operators on standby for edge cases.
Avoid Decepticon when:
- You cannot tolerate false positives or unintended lateral movement.
- You lack the infrastructure to sandbox tool execution.
- You need real-time human oversight for every action.
- Your compliance regime prohibits autonomous offensive security tools.
Decepticon is not a drop-in replacement for human penetration testers. It is an orchestration layer that automates routine tasks and escalates edge cases to humans. The LangGraph architecture is sound, but the deployment requires careful boundary enforcement, observability, and failure recovery.
Source Links
- Primary source: Decepticon on GitHub
- Documentation: docs.decepticon.red
- Live app: app.decepticon.red
Top comments (1)
tr.ee/dev-to