Athena
A local workspace that carries your memory, retrieval, and decision rules across AI models.
Per-dimension scores and reasoning
The evidence shows local Markdown storage, read-only CI permissions, private-path leak checks, a tested dangerous-command gate, a public-repository privacy warning, and an explicit mitigation for the DiskCache pickle risk. Deductions apply because the complete runtime permission model, per-action confirmation flow, and exact data flows when Google, Anthropic, or Supabase SDKs are enabled are not shown. Dependency auditing is non-blocking. Git provides a rollback path, but recovery for effects such as archival moves is not fully documented. Authorship and licensing are clear, while publisher identity remains unverified rather than suspect.
The README, package metadata, CI, and tests are broadly consistent about the product, version, and gating behavior, with installation troubleshooting and an explanatory veto message. Deductions apply because the test configuration skips several private-only modules absent from the public distribution, type checking and dependency auditing do not gate CI, and the coverage floor is only 15%. Public-distribution completeness and general failure-message quality therefore receive only partial support.
Developer audiences, lightweight/full/deep modes, problem classes, uncertainty postures, and varied scenarios are documented thoroughly. The source also clearly excludes ordinary web-chat interfaces and requires an IDE capable of reading local files, justifying full marks for audience definition and capability boundaries. Trigger precision is deducted because only a small classification sample covers false positives and negatives. Environment fit is deducted because cross-platform support and Windows Unicode guidance are documented but not backed by a platform matrix or equivalent CI evidence.
The README has clear navigation, quickstart material, workflow explanations, concepts, and a linked documentation structure. Virtual-environment setup, optional installs, workspace requirements, and common installation mistakes are covered well, and the MIT text matches package metadata. Deductions apply because the FAQ and many examples are only linked rather than included, while limitations are scattered among notices. The README points to docs/CHANGELOG.md but package metadata points to a root CHANGELOG.md, creating a small update-path inconsistency. A maintainer and vulnerability channel are named, but the contact details, governance, and succession responsibility are not specific.
The material presents actionable boot/end workflows, structured decision postures, ruin-path vetoes, and modes scaled to task complexity, supporting practical output usability. Portable local memory and model-independent governance provide plausible marginal value over ordinary chat assistants. Deductions apply because claims about compounding benefits, cross-model superiority, context-window percentages, and capability after hundreds of sessions are primarily promotional assertions. Cost discussion covers installation time and token budgets but not the systematic costs of curation, model usage, maintenance, or a growing file corpus.
Some claims trace to configuration, CI, tests, the security policy, and named protocol or case-study references; version, license, dependencies, gate behavior, and archival behavior also align across files. Deductions apply because most linked targets, benchmarks, the changelog, and the security-patch implementation are absent from the supplied evidence, preventing static verification of many quantitative and comparative claims. The README usefully separates deterministic, assumption-sensitive, and stochastic judgments and acknowledges curation needs, but promotional projections are not consistently distinguished from demonstrated behavior.
- The public test suite automatically skips tests that require private-repository modules; a green CI result should not be treated as verification of every advertised public capability.
- Several cloud-service SDKs are dependencies, but the supplied material does not fully describe activation conditions, transmitted fields, retention, or opt-out behavior. Review implementation and configuration before storing health, financial, journal, or similar sensitive data.
- pip-audit and mypy are non-blocking, most dependencies have lower bounds only, and the implementation of the stated CVE mitigation was not supplied for this static review.
- Do not treat session-count milestones, context percentages, compounding effects, or comparative claims in the README as performance guarantees established by the provided source.
- For personal memory, use a private repository and verify remotes, commit history, IDE-provider data policies, and recovery procedures for archived files.
What does this agent do, and when should you use it?
Athena is a local-first personal knowledge management and decision-support system, not a standalone model or web chatbot. Its repository combines an athena.yaml agent manifest, a Python SDK, boot and shutdown workflows, hybrid RAG, semantic search, an MCP server, protocols, skills, and governance controls. Users open the entire Athena-Public directory in an AI-enabled editor and use chat commands such as /start, /ultrastart, and /end to load or consolidate context. The chosen model produces the responses, while memory, configuration, session material, and version history remain in inspectable files that the user can edit and manage with Git. The design is model-agnostic, but governance is not yet equally portable: the code-enforced meta-awareness hook is currently specific to Claude Code. It suits people willing to maintain a file-based knowledge workspace and a regular curation loop; it does not suit users seeking a plug-in that runs directly inside ChatGPT.com, Claude.ai, or Gemini web.
The user opens Athena-Public as the root of an AI editor workspace. /start loads roughly 10K tokens of identity, memory, and active state, while /ultrastart loads roughly 20K tokens for deeper work; lightweight conversations can omit the full boot. During a session, Athena retrieves context from Markdown files, session logs, and its memory store using Hybrid RAG described as chunk retrieval, RRF fusion, and cross-encoder reranking; the documented stack also names gemini-embedding-001, Supabase, and pgvector. athena.yaml declares the model, tools, skills, hooks, and governance settings, while src/athena/ implements configuration, permissions, security, search, memory synchronization, boot orchestration, CLI functions, and mcp_server.py. More than 89 slash commands plus protocol and skill files route work into research, risk analysis, planning, writing, project management, and other procedures. At the end, /end or /ultraend extracts decisions, patterns, and lessons from the session into version-controlled memory. The CLI also exposes athena init --ide claude, antigravity, cursor, gemini, vscode, kilocode, and zoocode for supported editor setups.
- A user who regularly switches between models can keep personal context in local Markdown and reuse it from supported environments such as Claude Code, Cursor, and Gemini CLI.
- A researcher working across many sessions can retrieve older sources through keyword, semantic, and vector search, then preserve later syntheses in a searchable workspace.
- Someone managing health, career, financial, and client commitments can keep those projects in one context store so later discussions account for earlier constraints and decisions.
- A user facing consequential or potentially ruinous choices can structure the analysis with protocols, risk checks, capability levels, and the No Irreversible Ruin rule while retaining authority over non-ruinous choices.
- A writer can accumulate personal writing samples and session history so later model outputs can retrieve evidence of the writer's established voice.
- A developer building an inspectable agent workspace can extend the Python SDK, athena.yaml manifest, YAML tool definitions, lifecycle hooks, and MCP server.
What are this agent's strengths and limitations?
- Memory is stored as local Markdown that can be inspected, edited, moved, and audited or rolled back through git log, git diff, and git blame.
- The memory layer is separated from the reasoning model, with documented support for Claude Code, Cursor, Gemini CLI, Antigravity, VS Code with Copilot, Kilo Code, and Zoo Code.
- The repository goes beyond prompts: it includes a Python SDK, athena.yaml, hybrid retrieval, an MCP server, lifecycle hooks, permissions, security modules, protocols, and skills.
- Its validation table distinguishes shipped functions from N=1 experience and governance claims that have not been adversarially tested.
- Athena is MIT-licensed and free, with a lightweight installation path and an optional full installation for vector search and reranking.
- It requires an AI editor that can read local files and explicitly does not work through ChatGPT.com, Claude.ai, or Gemini web.
- Its value depends on repeatedly completing the /start–/end loop and curating memory; the project warns that unpruned archives decay and confidently retrieved stale memories can be worse than no memory.
- The workspace contains hundreds of Markdown files, scripts, protocols, and skills, creating more setup and maintenance overhead than custom instructions or a simple project chat.
- The anti-sycophancy meta-awareness gate is a Claude Code hook; in Cursor, Antigravity, Gemini CLI, and other environments, that protection and parts of governance remain agent-discretion behavior.
- Compounding personalization is supported mainly by the author's 1,900-plus-session N=1 experience, not a controlled multi-user study, and the governance gates lack a published adversarial test suite.
- The full search stack names Gemini Embeddings, Supabase, and pgvector, which may add network-service, configuration, and migration burdens even though the basic quickstart requires no database setup.
How do you install or deploy this agent?
Athena documents support for macOS, Windows, and Linux. Run:
git clone https://github.com/winstonkoh87/Athena-Public.git
cd Athena-Public
python3 -m venv .venv
source .venv/bin/activatepip install -e .
On Windows, activate with:
.venv\Scripts\activateFor vector search and reranking, install the full extra instead:
pip install -e ".[full]"Do not install the unrelated athena-cli package from PyPI. Open Athena-Public as the workspace root in Claude Code, Antigravity, Cursor, Gemini CLI, VS Code with Copilot, Kilo Code, or Zoo Code. Where desired, initialize an editor with athena init --ide claude, athena init --ide antigravity, athena init --ide cursor, athena init --ide gemini, athena init --ide vscode, athena init --ide kilocode, or athena init --ide zoocode. The basic quickstart says no API key or database setup is required; the selected AI editor or model service may still require its own account or subscription.
How do you use this agent?
Enter /start in the AI editor's chat panel, not in the terminal. On the first session, enter /tutorial for the roughly 20-minute guided tour and profile-building flow. Then discuss a problem, research material, plan a project, or draft content; Athena loads relevant workspace memory, protocols, and skills as context for the selected model. For quick general conversation, chat without a full boot; for complex multi-domain or architectural work, use /ultrastart. Finish an ordinary session with /end, or a deep session with /ultraend, so decisions, patterns, and lessons can be extracted into memory. If the workspace will contain health records, finances, journals, or other private material, do not rely on a public GitHub fork: copy the files into a new private repository or change the cloned repository's remote to a private one.
How does this agent compare with similar options?
Compared with ChatGPT Projects, Gemini Gems, and provider-managed memory, Athena keeps searchable, version-controlled memory on the user's disk and reuses it across models and editors; the tradeoff is requiring an AI editor with local filesystem access. Compared with Claude Code alone, Athena is a memory and governance workspace layered between the editor and model and is not limited to coding, although Claude Code currently receives its strongest code-enforced anti-sycophancy hook. Compared with managed agents such as Manus or Lindy, Athena trades always-on cloud convenience for local ownership and portability. Compared with Custom Instructions, it can load roughly 2K–20K tokens of structured context and retain session history, but demands a substantially more involved file and shutdown workflow.