Keep the Why
The why layer of repo-native project memory: preserve your codebase's reasoning as Git-versioned Markdown, so coding agents and humans never re-propose what was already rejected.
- Source repo
- oliver-zehentleitner/keep-the-why
- Stars
- ★ 163
- Last updated
- today
- License
- MIT
- Primary language
- Python
- FA score
- 75/100 · Good
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code
- Cost
- Free, no paid service needed
- Setup effort
- Low · running in minutes
- You'll need
- Typical use
- A team running coding agents on a long-lived codebase wants to stop re-litigating settled architecture questions in every fresh session
- Not a fit if
- Teams wanting session memory, transcripts or task tracking — this does none of those
- One-off scripts with no long-lived maintenance burden to protect
- Workflows unwilling to commit context/ Markdown into the repository
- Source review
- 75/100 · Good
What does this agent do, and when should you use it?
Keep the Why is a SKILL.md-format agent skill that preserves the reasoning a codebase can't explain on its own: architecture decisions, rejected alternatives, workarounds, incident learnings, and operational constraints. Everything is plain Markdown in the repository's context/ directory, committed with the code — Git provides the storage, history and distribution, with no database, no daemon, and no account. The skill works in four modes: continuous capture during development, retrospective recovery on existing repos, knowledge-transfer interviews, and maintenance of existing entries. Companion tools built in this repository include keep-the-why-lint, a structural validator on PyPI and the GitHub Marketplace, and keep-the-why-dashboard, a read-only viewer. It supports 70+ agent tools (Claude Code, Codex CLI, Gemini CLI, Cursor, and more) and backs its core claim with a controlled experiment: with a context/ record, zero of ten fresh agent sessions re-proposed an already-rejected simplification; without one, seven of ten did. MIT-licensed, one-command install, and after a one-time setup it keeps working with the project automatically.
Once installed, the agent notices rationale worth keeping during normal development and writes it to topic files in context/, located via context/index.md, using fields like a UUIDv4 Id, Status (active/superseded/…), Evidence (confirmed/inferred/unknown), rejected alternatives and reasons. The flow: install the skill package via npx skills add or gh skill install, tell your agent "initialize Keep the Why in this project", complete the one-time setup wizard (creating a .keep-the-why marker file at the project root, configuring capture mode, CI linting, and how the skill loads in future sessions), after which the agent reads that marker at each session start and keeps capturing. keep-the-why-lint validates required fields, value sets and index consistency in CI; keep-the-why-dashboard serves a local read-only view (or a static export) over context/, the config, lint findings and Git history, with a topic-reference graph and queues of what needs a person.
- A team running coding agents on a long-lived codebase wants to stop re-litigating settled architecture questions in every fresh session
- A new maintainer or a fresh agent inheriting a legacy project needs to know whether odd code is a Chesterton's Fence or safe to delete
- Before a veteran maintainer leaves, the interview mode analyzes the codebase and asks targeted questions — or just listens — to capture tacit knowledge
- An agent abandons a proposed simplification after discovering a real constraint (no commit, no PR); the skill still records the reason so the next person doesn't rediscover it
- A team wants mechanically checkable structure for context/ in CI via the keep-the-why-lint GitHub Action
How do you install or deploy this agent?
Recommended — skills CLI (needs Node.js; npx ships with it):
bash
npx skills add https://github.com/oliver-zehentleitner/keep-the-why/tree/latest/skills/keep-the-whyPrompts for one of 70+ agents and project/personal scope. Or via GitHub CLI (v2.90.0+):
bash
gh skill install oliver-zehentleitner/keep-the-why keep-the-why@latestAs a Claude Code plugin:
bash
claude plugin marketplace add oliver-zehentleitner/keep-the-why --sparse .claude-plugin skills
claude plugin install keep-the-why@keep-the-whyIn CI (GitHub Actions):
yaml
uses: oliver-zehentleitner/keep-the-why@lint-latestOr locally:
bash
pip install keep-the-why-lintHow do you use this agent?
After installing, start a new session and tell your agent "initialize Keep the Why in this project" — on a brand-new project, setup only runs from such a request. The wizard is one list; answering "defaults" gives the fully integrated setup: proactive capture, a project-carried start path, and lint installed from PyPI. Then work as usual — the agent writes to context/ when rationale surfaces; entries live one file per topic, found through context/index.md. Retrospective mode reconstructs what it can from git history, issues and code on an existing repo; interview mode asks targeted questions or listens to free narration before knowledge departs. The dashboard shows entry statuses and what still needs a human.
What are this agent's strengths and limitations?
- No database, no daemon, no account: context/ is Markdown in the repo, so a PR shows the reasoning diff beside the code diff
- The core claim is measured: across 20 fresh sessions, 10/10 avoided the rejected simplification with a record vs. 7/10 repeating it without one
- Open Agent Skills format spanning 70+ tools with multiple install paths (skills CLI, gh, Claude/Codex/Copilot plugins)
- Complete in-repo ecosystem: structural validator (keep-the-why-lint, with a GitHub Marketplace action) and read-only dashboard (keep-the-why-dashboard)
- Retrospective recovery on existing repos is limited — history, issues and code give back only part of the why, and the docs admit filling the rest takes real effort
- Whether captured rationale is true stays a human judgment; the linter only checks structure, and the project ships no opinion on keeping docs honest over time
- A skill cannot load itself — session-start loading depends on hooks or entry-point sections set up by the wizard, and unlisted agent tools are unmeasured
- Requires committing and reviewing context/ files in the repo — a workflow change for teams unwilling to carry docs in their diffs
How does this agent compare with similar options?
Architecture Decision Records remain the right tool for major, discrete architectural decisions; Keep the Why's topic files handle the larger, messier volume of smaller rationale — complementary, not competing. AGENTS.md is treated as the lean agent entry point rather than a rival. Unlike session memory (e.g. Claude Code's auto memory), it preserves project-level why, not a transcript of agent activity. The maintainer publishes a dated comparison against Claude Code Auto Memory, MemoryCustodian and AgentsRoom (September 2026, on the blog).
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Keep the Why This agent | 75 · Good | Agent plugin / skillFree | ★ 163 | today | Python | Codex · Claude Code |
| Softaworks Agent Skills | 51 · Major gaps | Agent plugin / skillFree + model costs | ★ 2.5k | 6mo ago | Python | Claude Code |
| Agent Skills for Context Engineering | 49 · Major gaps | Agent plugin / skillFree + model costs | ★ 18k | 17d ago | Python | Codex · Claude Code |
| Finding-Unknowns Skills | 76 · Good | Agent plugin / skillFree | ★ 339 | today | Python | Claude Code · Claude.ai |
How does FollowAgents rate this agent?
Why each dimension lost points
Privilege evidence is solid: all CI workflows declare read-only contents: read tokens and pin actions by SHA; dedicated PathConfinement tests prove the linter refuses to read outside the project tree (../, absolute paths, symlink escapes all trigger E009), and dashboard tests show symlinked and external directories are never read — least_privilege earns 3. Write confirmation is supported via config (capture-confirmation, confirm-when-unsure) but the confirmation flow itself is not shown in the provided files: 2. Data flow (plain Markdown, no DB/daemon) and sensitive-data handling (anonymize tests, script-tag escaping, hidden-character detection) are test-backed but cover only the dashboard side: 2 each. Dependencies: black is pinned in CI, but the dashboard tests install jsdom@24 without a full pin and no lockfile is provided: 2. External effects are limited to in-repo file writes and optional PyPI linter installation with an ask-first option: 2. Rollback rests on Git itself plus superseded/pinned-version mechanisms — reasonable but indirect: 2. Source attribution is thorough (git attribution, .mailmap merging, Evidence/Source fields, author-statistics tests): 3.
test_version.py hard-tests that the version prefix equals SUPPORTED_SCHEMA and that all schema gates are covered — self_consistency earns 3. The core linter claims stdlib-only and frontend tests claim no dependencies, but no complete dependency manifest appears in the provided files: 2. Failure output has structured codes (E001–E301) and a GITHUB_ACTIONS annotation mode; the tests claim 'never a traceback and never a guess', but that is a self-description in test comments, not independent evidence: 2.
Audiences (maintainers, teams, legacy projects) and four modes are covered, with honest limits on retrospective recovery ('history gives back only part'), but the scenario description is largely narrative: 2. Boundaries are stated in several places ('has not been measured, not failed'; honest eval failure analysis) — 2 rather than 3 because the boundary statements mostly point to off-repo pages not verifiable here. Trigger precision is deliberately designed: activation only on conversation match, setup only from an explicit request, hook/entry-point loading statistics (10/10, 3/3) recorded but self-reported: 2. Environment fit is the strongest: 70+ agents, a per-agent directory table, shared-path fallbacks, and Windows path-parsing platform differences all test-covered: 3.
Information architecture is excellent: a 'Where this fits' table separates README/docs/CHANGELOG/context responsibilities, and the lettered index skeleton has full lint rules (E205/E206): 3. Install notes cover skills CLI, gh, asm, plugin marketplaces, and manual-clone fallback with pitfalls explained: 3. Naming stability has a fixed directory-name requirement and pinned-version/pinned-path validation tests, but legacy-format migration appears only in tests: 2. The examples/ directory and per-mode walkthroughs are referenced but not present in the provided files: 2. Known limitations are stated (eval caveats, retrospective limits) but mostly off-repo: 2. The MIT license file is complete and matches metadata: 3. Versioning: CHANGELOG is linked and the four-segment scheme is test-locked, but the CHANGELOG itself is not in the provided files: 2. Maintenance: named maintainer, 48-hour security response commitment, one-issue-per-problem policy — but a documented solo commitment: 2.
Output is plain Markdown plus index plus a self-contained dashboard export (script-tag escaping tested); usability is good but unverified by execution: 2. Marginal value is backed by the 20-session controlled experiment (10/10 no longer re-proposing the rejected change); the transcripts are claimed to live in-repo but were not provided, so it is a self-reported experiment: 2. The cost-benefit argument (reasoning as a byproduct, token savings) is plausible but likewise rests on self-reported measurement: 2.
The core claim points to full transcripts and grades under experiments/rejected-change/, so traceability is designed in, but those files were not part of this review: 2. Cross-source corroboration lists multiple registries and security-scan badges, none independently verified here; the agent matrix is off-repo: 2. Fact/inference separation is a first-class schema requirement (Evidence: confirmed/inferred/unknown; E111 forces explanations for contradicted claims) with tests proving it — but as a self-assessment instrument it says more about the product's design than about the project's own claims: 2.
- This is a static source review with no executed tests or evals; all behavioral claims (10/10 hook loading, the 20-session experiment) are self-reported by the repository — confidence is low.
- Key evidence (eval numbers, install details, security notes) lives behind off-repo links to keepthewhy.com and blog posts; verify them at the pinned revision before adopting.
- The linter reads Markdown in context/; path confinement is test-protected, but context/ content is agent-generated and could still carry injected instructions — review context/ PR changes manually.
- One install path pulls via npx skills, a supply-chain trust decision; for production prefer an exact release tag or manually copying skills/keep-the-why/.
- Publisher identity is unverified by the FollowAgents curated registry; the maintainer is an individual — assess single-maintainer risk yourself.