Decepticon Autonomous Red Team Agent
Plans and executes complete red-team attack chains within an authorized engagement scope.
The evidence shows separation between management and operational networks, on-demand specialist workloads, read-only contents permission in CI, and an engagement process requiring written authorization, RoE, ConOps, deconfliction, and an OPPLAN. This supports moderate least-privilege and external-effects controls. Deductions apply because the agent can still perform reconnaissance, exploitation, privilege escalation, lateral movement, and C2, while LangGraph controls the sandbox through the Docker socket; no per-action confirmation mechanism or RoE enforcement implementation is shown. The architecture identifies LiteLLM, PostgreSQL, Neo4j, sandbox, and HTTP service flows, but hosted-service retention, telemetry, and third-party transfers are unspecified. Credential configuration and security-reporting scope are documented, but storage, redaction, log filtering, and rotation are not. Dependency controls include locked installs, commit-pinned Actions, a checksum-pinned MITRE bundle, overrides, and release image digests, but no vulnerability-scan results or critical-dependency remediation evidence. No rollback, compensating-action, or post-attack recovery mechanism is shown, so rollback scores zero. The organization name, license, contributor route, and security contact provide attribution, although publisher identity remains unverified.
The README, workspace configuration, CI, and supplied regression tests are broadly consistent about package boundaries, network architecture, knowledge-graph contracts, and testing workflow. The deduction is that many product capabilities are supported only by documentation rather than corresponding implementation files. Dependency availability is addressed through Docker prerequisites, locked uv synchronization, provider fallback, platform matrices, and digest publication; however, operation depends on numerous services, containers, and remote installation endpoints, with no offline or degraded-mode evidence. Tests contain specific assertion messages and the security policy gives a fallback reporting channel, but user-facing runtime diagnostics, retries, and recovery messages are not demonstrated.
The material clearly addresses self-hosted, hosted, SDK, research-integration, custom-orchestration, development, and CI scenarios, with macOS, Linux, Windows, WSL2, and many model providers, earning full marks for audience coverage and environment fit. SDK-versus-runtime boundaries and the core package's prohibition on heavyweight runtime dependencies are also clear, but tool permissions and unsupported scenarios are not comprehensively enumerated. Specialists and the dashboard are triggered on demand through ops_start, /web, or orchestration and are described as engagement-bound; full marks are withheld because trigger implementation, target validation, and mis-trigger safeguards are not supplied.
The README has strong information architecture, a useful documentation index, platform prerequisites, quick installation, SDK guidance, and contribution entry points. Knowledge-graph enums, relationship names, and technology-key formats are protected by explicit regression tests, justifying full naming-stability credit. Numerous examples and documentation links are present, but no substantive FAQ is included. Authorization requirements, SDK service dependencies, and some platform-test boundaries are disclosed, while a centralized and comprehensive limitations section is absent. The complete Apache-2.0 text matches the license metadata. SECURITY.md identifies 1.0.x as supported, but no changelog or release-history evidence is supplied, so versioning scores weakly. A maintenance email, private reporting path, response targets, and contribution guide exist, but named responsibility or governance is unclear and publisher identity is unknown.
The repository describes usable CLI, Web, and SDK surfaces plus engagement documents, knowledge-graph records, and defense briefs, suggesting actionable outputs; deductions reflect the absence of actual output samples or independently established quality. Persistent interactive sessions, attack-chain orchestration, and a proposed defensive loop offer plausible value beyond a basic scanner, and a 102/104 benchmark result is reported; however, it is project-published and the supplied evidence cannot validate methodology or real-world generalization. On-demand containers, eco/max/test profiles, and provider fallback show cost awareness, but resource use, model expense, deployment burden, and measured benefit are not quantified.
Major claims are linked internally to named documentation, architecture material, CI configuration, regression tests, and a claimed per-challenge benchmark index with traces. Deductions apply because most referenced documents and trace records are not included, leaving many marketing claims unverified from the supplied files. README, pyproject, CI, and tests cross-support package boundaries, platform claims, and knowledge-graph contracts, but all evidence is repository-controlled and not independent. The README explicitly labels the Offensive Vaccine as planned and distinguishes the SDK from required runtime services, showing some fact-versus-future separation; claims such as professional-grade operation, hardened isolation, and realistic attack capability remain only partially demonstrated by the provided source.
- This is an autonomous offensive tool capable of exploitation, privilege escalation, lateral movement, and C2. Use it only in an isolated environment with written authorization defining targets, timing, permitted techniques, and stop conditions.
- Docker-socket control and dual-homed Neo4j create high-impact trust boundaries. Independently review container privileges, network routes, secret mounts, and sandbox-escape exposure before deployment.
- Do not treat the reported 102/104 benchmark or the hardened-sandbox description as independent validation. This assessment did not execute the software and was not given the referenced per-challenge results, traces, or full implementation.
- The curl-to-shell and PowerShell irm-to-iex installation paths execute remote content directly. High-assurance deployments should pin a release, inspect the installer first, and verify artifacts using published digests or signatures.
- The supplied material does not demonstrate per-action approval, rollback, or recovery for offensive operations. Add deployment-level human gates, mandatory scope validation, a kill switch, audit logging, and a recovery plan.
What does this agent do, and when should you use it?
Decepticon is an autonomous agent for professional red-team work, spanning reconnaissance, exploitation, privilege escalation, lateral movement, and C2 rather than stopping at scanning and reporting. It organizes 16 specialist agents by kill-chain phase, orchestrates them with LangGraph, and gives each objective a fresh context window. Before sending network traffic, it produces an RoE, ConOps, Deconfliction Plan, and MITRE ATT&CK-mapped OPPLAN, then operates within those rules. Commands execute in a Kali Linux sandbox on a separate operational network, while persistent tmux sessions and prompt detection support interactive tools; the management plane includes LiteLLM, PostgreSQL, Neo4j, Skillogy, and LangGraph. Teams can use a terminal CLI, an on-demand web dashboard, or a Python client SDK, but agent execution still depends on model-proxy and sandbox runtime services.
The workflow starts with decepticon onboard, which collects the model provider, API key, and model profile. For an engagement, Decepticon prepares an RoE, ConOps, Deconfliction Plan, and MITRE ATT&CK-mapped OPPLAN before the orchestrator assigns objectives to reconnaissance, exploitation, post-exploitation, vulnerability-research, and domain specialists for AD, cloud, smart contracts, reversing, and analysis. Agents run commands inside the Kali Linux sandbox and control interactive programs such as msfconsole, sliver-client, and evil-winrm through persistent tmux sessions. The orchestrator can dynamically launch workloads such as BloodHound CE, Sliver C2, and Ghidra MCP using calls including ops_start("ad"), while Neo4j stores attack-chain findings. Running decepticon starts the core management plane and opens the terminal CLI; /web starts the dashboard on demand. In library mode, the decepticon package exposes agent factories, middleware, tools, skills, declarative PluginBundle plugins, and a safety gate, while DECEPTICON_LLM__PROXY_URL and SANDBOX_URL route model calls and command execution to runtime services.
- An enterprise red team with written authorization wants an agent to pursue a multi-stage internal assessment under a formal OPPLAN.
- A penetration-testing team needs continuous control of interactive tools such as
msfconsole,sliver-client, orevil-winrm, not just one-shot scanner commands. - A security research group wants to evaluate autonomous attack chains inside an isolated Kali Linux operational network and persist findings in Neo4j.
- AD, cloud, smart-contract, or reverse-engineering specialists need workloads such as BloodHound CE, Sliver C2, or Ghidra MCP to start only when an objective requires them.
- A product or research team wants to reuse agent factories, middleware, tools, and skills from PyPI while supplying its own model-proxy and sandbox services.
- An evaluation team wants to inspect the published XBOW validation results together with the per-challenge index, attack-class matrix, and LangSmith traces.
What are this agent's strengths and limitations?
- It explicitly covers reconnaissance through exploitation, privilege escalation, lateral movement, and C2, with the ability to adapt an OPPLAN as paths open.
- RoE, ConOps, a Deconfliction Plan, and a MITRE ATT&CK-mapped OPPLAN are generated before network activity, providing defined operating boundaries.
- Persistent tmux sessions and automatic prompt detection address interactive offensive tools directly.
- The management and sandbox networks are separated, while specialist workloads are spawned only when needed.
- The credential-aware fallback system supports Anthropic, OpenAI, Gemini, MiniMax, DeepSeek, xAI, Mistral, OpenRouter, Nvidia NIM, and local Ollama.
- Adopters can choose among a self-hosted CLI, an on-demand dashboard, a hosted cloud app, and an extensible Python SDK.
- Self-hosting involves Docker Compose plus LiteLLM, databases, a knowledge graph, LangGraph, and a Kali sandbox, creating a substantial operational footprint.
- The PyPI package is only a client SDK; agents cannot run without model-proxy and sandbox runtime services.
- Model execution requires provider credentials, subscription OAuth, or local Ollama, so cost, rate limits, and capability depend on the selected provider.
- Autonomous exploitation, lateral movement, and C2 are high-risk activities that require written authorization and strict scope controls.
- The material reports a 102/104 XBOW benchmark result but provides no evidence here about false-positive rates, stability, or long-running performance on enterprise networks.
- Specialized features introduce additional resource and deployment requirements through BloodHound CE, Sliver C2, or Ghidra MCP workloads.
How do you install or deploy this agent?
Prerequisites are Docker and Docker Compose v2, plus credentials for at least one supported model provider or a supported subscription OAuth service. On macOS, Linux, or WSL2, run:
curl -fsSL https://decepticon.red/install | bash
decepticon onboard
decepticonFor native Windows PowerShell, run:
irm https://decepticon.red/install.ps1 | iex
decepticon onboard
decepticonThe documented platforms are macOS on Apple Silicon or Intel, Linux on amd64 or arm64, and Windows on amd64 or arm64 through native PowerShell or Ubuntu/Kali WSL2. For client-library use, install pip install decepticon, or pip install "decepticon[neo4j]" for the knowledge-graph attack-chain tools. The SDK is not a self-contained runtime: deploy the Docker stack or configure DECEPTICON_LLM__PROXY_URL and SANDBOX_URL to point at equivalent services.
How do you use this agent?
Run decepticon onboard first and use the interactive wizard to select a provider, enter its API key, and choose a model profile. Then run decepticon; it starts the core LiteLLM, PostgreSQL, Neo4j, Skillogy, LangGraph, and sandbox services and opens the terminal CLI. Enter /web inside the CLI when a web interface is needed. Specialist services including BloodHound CE, Sliver C2, and Ghidra MCP are launched on demand by the orchestrator through ops_start(...). Available model profiles are the per-agent eco default, max with every agent on the HIGH tier, and test with every agent on LOW. Use the system only against systems and networks for which the owner has provided explicit written authorization.
How does this agent compare with similar options?
The project names Strix, PentestGPT, MAPTA, Cyber-AutoAgent, and commercial XBOW in a dedicated comparison document. Its stated points of distinction are full attack-chain execution, control of genuinely interactive shells, formal engagement artifacts before action, and isolated command execution; the supplied material does not include the competitors' itemized results.