Dev & Engineering architecture-driven-developmentspec-driven-developmentevidence-verificationbaseline-analysiscode-reviewtdd-workflowmulti-host-skills

Aegis Method Pack

Makes coding agents check architectural baselines, change safely, and prove completion with fresh evidence.

FollowAgents review · FARS-2.1
Not recommended
44/ 100 5-point scale 2.2 / 5
1 2 3 4 5 6
1Trust7 / 29 · 1.2/5

SECURITY.md limits the product to a Method Pack and distinguishes host-platform concerns from a future Runtime Core; the test fixture cleans up its temporary directory, and licensing attribution is clear. Deductions apply because the supplied evidence omits the core plugin implementation, permission inventory, confirmation gates, data flows, and sensitive-data rules. Dependencies are optional peers but no lockfile, audit evidence, or stronger supply-chain pinning is shown; external-effect and rollback evidence is mostly test-local rather than product-wide.

2Reliability11 / 14 · 3.9/5

Package metadata, the security boundary, and CI targets are broadly consistent. CI names boundary, schema, compatibility, synchronization, and aggregate static checks. Optional dependencies and missing hosts are handled as explicit skip conditions, while the scripts provide specific messages for invalid arguments, absent CLIs, incompatible command surfaces, and runtime blockers; this justifies full credit for failure messages. Missing core implementation and omitted test bodies prevent stronger conclusions about consistency and availability across all paths.

3Adaptability9 / 18 · 2.5/5

The evidence identifies several host or extension surfaces, including OpenCode, Codex, Pi, OMP, DSH, and Antigravity, and distinguishes missing-host, authentication, and configuration blockers. SECURITY.md usefully disclaims Runtime Core authority. However, detailed audiences, scenarios, activation rules, and false-trigger controls are not supplied, while environment-fit evidence is limited to metadata, CI, and one Antigravity probe.

4Convention8 / 18 · 2.2/5

The repository has language-entry documentation, standard package metadata, a security policy, CI organization, test structure, and a complete MIT license. Aegis naming and the Method Pack boundary are reasonably stable, and some host/runtime limitations are explicit. The supplied main README is absent, so installation instructions, examples, and FAQ cannot be credited. Version 2.8.8 is declared without a changelog or versioning policy, and the reporting policy identifies only a generic maintainer route with no definite contact or response commitment.

5Effectiveness4 / 13 · 1.5/5

The description promises baseline-first operation, evidence verification, and drift checking, while a test expects the doctor command to emit usable JSON configuration status. The provided files do not contain the principal workflow, output examples, or implementation, so output usability and incremental value over an ordinary coding agent remain only partially supported. Optional integrations and host-neutral CI suggest some cost awareness, but installation burden, runtime cost, and benefit trade-offs are not analyzed.

6Verifiability5 / 8 · 3.1/5

CI maps several schema, boundary, compatibility, synchronization, and static-aggregate claims to named scripts. The Antigravity runner clearly distinguishes passes, failures, and environment-driven skips. SECURITY.md, package metadata, and CI partly corroborate the Method Pack characterization and separate current capability, a future Runtime Core, and live-host integration checks. Deductions apply because the central invoked scripts, main README, and product implementation are missing, leaving the headline architecture-awareness and drift-checking claims without complete traceability or corroboration.

Evidence confidence: Low Reviewed Aug 25, 2026 Reviewed revision 1dc0faa63303
The upstream repository has new commits since this review. The score still applies to the reviewed revision shown and may not cover the latest changes.
Safety controls not found in source: confirmation before acting, data-flow disclosure, sensitive-data handling
Before you use it
  • The main README, core plugin code, and most scripts invoked by CI were not supplied; do not treat architecture awareness, evidence verification, or drift checking as fully established by this static evidence.
  • Before adoption, inspect actual file writes, command execution, network access, telemetry, credential handling, and user-confirmation boundaries.
  • Review the optional DeepSeek release-candidate dependencies and GitHub Actions dependencies for locking, provenance, vulnerabilities, and update policy.
  • Default CI explicitly omits a live Codex host smoke test, while Antigravity integration is opt-in and may skip; compatibility still needs verification in each target environment.
  • The private security-report path depends on GitHub reporting being enabled, and the fallback maintainer contact and response timeline are not specific.
Review evidence [1][2][3][4][5][6][7][8]
See the full review method →

What does this agent do, and when should you use it?

Aegis is a method pack for AI coding agents, not a daemon, background executor, or complete agent platform. It combines discoverable skills, host-specific installation adapters, workflow rules, and verification scripts into a process covering baseline alignment, architectural boundaries, implementation, independent review, and completion checks. Before editing, the host agent is expected to identify the project's owners, contracts, and boundaries; afterward, it reports fresh verification evidence, covered scope, and residual risk. The repository includes `scripts/aegis-doctor.py`, `scripts/aegis-update.py`, end-to-end checks, host installation guides, and optional activation and TDD controls. It targets Codex, OpenCode, and numerous other skill-aware hosts, although only Codex and OpenCode are described as having fresh evidence for the current method-pack scope; most other hosts still await release-level smoke testing.

Once installed, the host discovers Aegis skills and uses them to select a lightweight or expanded workflow according to the request. For non-trivial work, the method reads the target project's real baseline—owners, contracts, boundaries, and optionally CONTEXT.md or a bounded context selected through CONTEXT-MAP.md—and activates domain modeling only for resolved, ambiguous, renamed, deprecated, or conflicting terms. The coding host then plans and implements against that baseline, tracks or removes retired fallbacks and old paths, and performs fresh verification before claiming completion. Its output is evidence, covered scope, and residual risk; Aegis does not issue an authoritative GateDecision, PolicySnapshot, or final completion verdict. scripts/aegis-doctor.py --write-config --json checks installation, workspace support, and configuration, while scripts/aegis-update.py handles host-scoped updates. Maintainers can run bash tests/e2e/run-all.sh --full --host-profile fast plus focused boundary, workflow-quality, and installation-policy checks.

  1. A development team using a coding agent in a large existing codebase that wants ownership, interface contracts, and architectural boundaries identified before edits begin.
  2. An engineer delegating a long, multi-step change who needs controls against scope and architectural drift.
  3. A reviewer who requires new test or check evidence, covered scope, and residual risk instead of accepting an unsupported completion claim.
  4. A person or team working across Codex, OpenCode, Claude Code, or other skill-aware hosts who wants a consistent engineering method.
  5. A maintainer preparing a risky merge who wants an independent review, first-principles challenge, or explicit test-first route.
  6. A developer who wants trivial requests to remain lightweight while higher-risk work receives additional architectural and verification discipline.

What are this agent's strengths and limitations?

Pros
  • Builds real project owners, contracts, and architectural boundaries into the pre-edit workflow, directly addressing incorrect assumptions in established codebases.
  • Requires completion claims to carry fresh verification evidence, covered scope, and residual risk, backed by dedicated installation diagnostics and end-to-end checks.
  • On its frozen 120-run, 20-case A/B benchmark, contract pass rate increased from 61.67% to 93.33% and unsafe outcomes fell from 13.33% to 0%; the project also states the limits of that evidence.
  • Offers a fast path, automatic or explicit activation, optional TDD routing, and independent review instead of imposing equal ceremony on every task.
  • Targets multiple skill-aware hosts—including Codex, OpenCode, Claude Code, DeepSeek Harness, and Kimi Code CLI—with host-specific guides.
Limitations
  • It is not a complete platform, runtime core, background executor, or final completion authority; execution still depends on an external coding agent and host.
  • Most listed hosts other than Codex and OpenCode still lack release-level fresh smoke evidence, leaving a validation gap for integrations such as Claude Code.
  • The benchmark is bounded advisory evidence: review was arm-hidden technical review rather than independent human review, and host events did not report the observed model identity.
  • Installation is not a single uniform command; adopters must follow a host guide and verify discovery, native activation, and automatic entry, sometimes with extra parameters.
  • The optional global routing prefix is copied manually and is not maintained by aegis:update; users of retired Lite or Advanced profiles must replace those Aegis blocks themselves.
  • Gemini CLI support has been retired, and Aegis no longer ships or verifies an adapter for it.

How do you install or deploy this agent?

The documented fastest entry is to ask the current AI coding agent to read https://github.com/GanyuanRan/Aegis, identify the active host, and perform a global installation using that host's guide. After installation and any required host reload, locate <aegis-method-pack-root> and run:

cd <aegis-method-pack-root>
python scripts/aegis-doctor.py --write-config --json

Installation is complete only when the JSON contains "ok": true, "workspaceSupport": "available", and "configStatus": "configured", and the host guide's native activation and automatic-entry checks also pass. If the host has a separate discovery directory, add --discovery-root <path>; if its guide declares a skill-directory prefix, also add --discovery-name-prefix <prefix>. For the official DeepSeek Harness, the native command is:

dsh plugin --profile <profile> add "git+https://github.com/GanyuanRan/Aegis.git"

Its direct-child compatibility path should be used only when the plugin manager is unavailable and the user explicitly approves compatibility mode. The supplied material does not specify one universal manual clone command, a Python version, or every host's complete command sequence, so those details remain host-guide-specific.

How do you use this agent?

After installation and host restart, invoke Aegis through ordinary language, such as Aegis goal: Fix the auth refresh bug without rewriting the auth system., Why does this login failure happen? Diagnose it before changing code., or Review this diff independently before I merge it. Use Grill me ... or 审问我 ... for a one-question-at-a-time decision interview that does not plan or implement. Request strict test-first work with TDD Route: strict, strict TDD, test-first, or RED / GREEN / REFACTOR. TDD is off by default; to let Aegis choose strict, light, or skipped according to risk, run python scripts/aegis-doctor.py tdd-mode auto from the method-pack root. Activation defaults to automatic; switch to explicit activation with python scripts/aegis-doctor.py activation-mode explicit, then restart the host. Later updates can be requested as update Aegis or aegis:update; updating every registered host requires an explicit --all request.

How does this agent compare with similar options?

Aegis is derived from Jesse Vincent's Superpowers. Both use composable skills across multiple agent harnesses, while Aegis adds an architecture- and evidence-oriented method layer for real software projects, including baseline alignment, drift controls, and workflow governance. It also cites mattpocock/skills as inspiration for concise communication, shared language, and disciplined debugging, with those ideas reimplemented in Aegis format.

FAQ

Does Aegis run independently or edit code in the background?
No. It is a method pack loaded by an external coding-agent host, not a daemon, background runner, or standalone runtime.
Can it authoritatively certify that a change is safe and complete?
No. It requires fresh evidence, covered scope, and residual risk, but does not provide an authoritative GateDecision, PolicySnapshot, or final completion verdict. User instructions and target-project rules take precedence.
Does it support Claude Code?
A Claude Code installation guide exists, but release-level fresh host smoke is still pending. Codex and OpenCode are the hosts explicitly described as having fresh evidence for the current method-pack scope.
Will it force TDD and a heavyweight process on every request?
No. TDD is off by default and trivial tasks stay on the fast path. Users can request strict TDD explicitly or set TDD mode to auto so risk determines whether the route is strict, light, or skipped.
How is a successful installation verified?
Run the doctor from the installed method-pack root and confirm ok: true, workspaceSupport: available, and configStatus: configured in its JSON. The selected host's native activation and automatic-entry checks must pass as well.

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents