Mnemo Cortex
A local-first cognitive coprocessor for AI agents: active memory, semantic recall, and cross-agent sharing with no cloud.
- Source repo
- GuyMannDude/mnemo-cortex
- Stars
- ★ 155
- Last updated
- 2d ago
- License
- MIT
- Primary language
- Python
- FA score
- 75/100 · Good
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code · Claude.aiChatGPT (Partial support)
- Cost
- Free, no paid service needed
- Setup effort
- Medium · a few setup steps
- You'll need
- Typical use
- A solo developer running multiple agents (Claude Code, Claude Desktop, Codex CLI, Hermes) who wants Agent A to recall what Agent B learned
- Not a fit if
- Teams wanting a zero-deployment hosted cloud memory
- Users needing OpenAI cloud connectors to reach a self-hosted MCP directly
- Users unwilling to maintain a local server process
- Source review
- 75/100 · Good
What does this agent do, and when should you use it?
Mnemo Cortex is a self-hosted cognitive coprocessor that gives AI agents persistent, local, cross-agent memory. It watches agent session files from the outside (JSONL for OpenClaw and Claude Code), ingests every message into a local SQLite database, and applies rolling LLM-backed compaction via a local Ollama model (default qwen2.5:32b-instruct) for roughly 80% token reduction. The system runs as an API server (default http://localhost:50001) and connects to clients — Claude Desktop, LM Studio, AnythingLLM, OpenClaw, Agent Zero, Hermes — through MCP bridges; ChatGPT is served through a hardened two-route REST gate instead of MCP. Core components include tiered smart notes versus session logs, a structured facts store with a confidence ladder and authority tiers, nightly cross-agent 'dreaming' synthesis, The Librarian workspace document index (SQLite FTS5), the encrypted USB-courier Cortex Stick, the Disco-Bus agent messaging system, and a tamper-evident Memory Ledger. According to the README, it has run since March 2026 as the production memory of a five-agent fleet, holding about 10,000 memories and ~7,000 verified facts across 84 shipped versions.
After mnemo-cortex start, the API server listens on port 50001; mnemo-cortex watch --backfill starts a session watcher that tails JSONL session files and auto-ingests messages every 2 seconds (no hooks or agent changes needed). Saves are tiered at write time: Tier 1 smart notes are classified by the reasoning model into one of eight categories (topology, current_state, doctrine, incident, identity, relationship, decision), while Tier 2 holds raw session logs excluded from default recall. Recall runs semantic vector search and a BM25 keyword lane (SQLite FTS5) in parallel; since v4.22 it also matches misspelled proper names via consonant-skeleton keys. Structured facts are stored as (entity, attribute, value) triples with an automatic verified→high_probability→false confidence ladder and four MCP tools: mnemo_fact_save, mnemo_fact_get, mnemo_fact_query, and mnemo_fact_demote. The nightly dreaming pipeline map-reduces each agent's recent memories into themes, then merges them cross-agent into a shared brief. mnemo-cortex health verifies the server, database, compaction model, and per-agent recall; mnemo-cortex stick sync carries deltas between two machines via an optionally AES-256-encrypted USB stick.
- A solo developer running multiple agents (Claude Code, Claude Desktop, Codex CLI, Hermes) who wants Agent A to recall what Agent B learned
- A user working across two desks who syncs agent memory via the encrypted Cortex Stick USB courier — no cloud, VPN, or account
- Local-LLM users (LM Studio, Open WebUI, Ollama, AnythingLLM) who want persistent semantic memory at zero API cost
- A ChatGPT Plus subscriber who wants a Custom GPT with memory they own, accessed through a private gate rather than an exposed server
- Users who need exact facts (names, settings, identifiers) rather than fuzzy semantics — served by the structured facts store and the BM25 lexical lane
- Multi-agent teams wanting overnight cross-agent dream synthesis plus the Lane Protocol session workflow
How do you install or deploy this agent?
Prerequisites: Python 3.11+, optionally Ollama (recommended for local reasoning and embeddings), Node.js 18+ only when running an MCP-bridge integration. Install from source:
bash
git clone https://github.com/GuyMannDude/mnemo-cortex.git
cd mnemo-cortex
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activatepip install -e .
This registers the mnemo-cortex CLI (alias mnemo). For non-interactive install (LLM agents or CI):
bash
./robot-install.shThe script emits a single JSON object on stdout (ok, steps, etc.) with progress on stderr; sandbox runs are supported via MNEMO_INSTALL_DRY_RUN=1 and path overrides. A dedicated macOS guide exists (docs/install-macos.md); Windows is natively supported as of v4.4.1.
How do you use this agent?
Five-step flow:
bash
mnemo-cortex init # Interactive wizard: reasoning/embedding models, bind address, port; writes ~/.agentb/agentb.yaml
mnemo-cortex start # Detached start; server listens on http://localhost:50001
mnemo-cortex health # 14-point health check; use mnemo-cortex doctor for deeper diagnosticsEnable auto-capture:
bash
mnemo-cortex watch --backfillOr set export MNEMO_AUTO_CAPTURE=true in your shell profile so the watcher starts with the server. Then connect a client: Claude Code via integrations/claude-code/; Claude Desktop via a drag-and-drop .mcpb bundle; LM Studio, AnythingLLM, OpenClaw, Agent Zero via a one-line MCP config; ChatGPT via Custom GPT Actions against the two-route gate (Plus tier or higher). Direct API example:
bash
curl -s -X POST http://localhost:50001/context \
-H "Content-Type: application/" \
-d '{"prompt": "what happened today", "agent_id": "YOUR-AGENT-ID", "max_results": 5}'Recommended companions: a forked mnemo-plan brain repo pointed at via the BRAIN_DIR env var, and the THE-LANE-PROTOCOL.md operating ritual.
What are this agent's strengths and limitations?
- Only memory system claiming cross-agent overnight dreaming — map-reduce compaction of each agent's memory into a shared brief so all agents wake up knowing what the others did
- Fully local-first: no telemetry, no cloud dependency, single-file SQLite storage, and $0 compaction/embedding via local Ollama models
- Multi-lane recall: semantic vectors + BM25 keyword lane (SQLite FTS5) + phonetic name matching; author's test shows 11/11 exact hits versus 8/11 with vectors alone
- Structured facts with a confidence ladder, authority tiers (probe/declared/open), tamper-evident Memory Ledger, and a demote endpoint for wrong memories
- Rich operational tooling: per-platform install guides, robot-install non-interactive setup, health/doctor diagnostics, Disco-Bus agent bus, The Librarian document index (~107K files, ~17s full rebuild)
- Requires a continuously running local server process (default 127.0.0.1:50001) that you must keep alive via systemd/launchd/Task Scheduler yourself
- Depends on local Ollama models for compaction and embeddings (qwen2.5:32b-instruct, nomic-embed-text); embedding model-name drift is a documented failure mode causing empty recall
- ChatGPT cannot use standard MCP — a custom hardened gate is required, and OpenAI plan tiers restrict functionality (read-only custom MCP connectors on Plus)
- Authority tiers are currently an audit trail plus honor system; enforcement that actually blocks agent writes is not yet implemented
- Personal/small-team project by a self-described non-developer; the documented timeline (2026 dates) and long-term maintenance cannot be independently verified
How does this agent compare with similar options?
The README explicitly compares against OpenClaw's native Active Memory plugin (2026.4.10): that plugin is single-agent, per-workspace scratchpad storage, while Mnemo is a centralized SQLite + embeddings memory bus that survives resets and machine moves — the two stack rather than compete. It also self-positions against Mem0, Zep, and Letta, which store memory per agent, whereas Mnemo dreams across all of them (author's claim). For ChatGPT, instead of exposing a standard MCP server over HTTPS, it ships a hardened two-route gate — a deliberate design contrast with typical self-hosted MCP servers.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Mnemo Cortex This agent | 75 · Good | Self-hosted serviceFree | ★ 155 | 2d ago | Python | Codex · Claude Code · Claude.ai |
| Compartment | 76 · Good | CLIFree | ★ 582 | 6d ago | Python | Codex · Claude Code |
| Remnic Agent Memory | 85 · Good | CLIFree + model costs | ★ 210 | 9d ago | TypeScript | ChatGPT · Codex · Claude Code · OpenAI API |
| Ori Mnemos | 80 · Good | CLIFree | ★ 326 | 8d ago | TypeScript | Claude Code · OpenAI API · Claude API |
How does FollowAgents rate this agent?
Why each dimension lost points
Evidence shows local-first binding (localhost/private net), a hardened bearer-auth tenant-pinned rate-limited ChatGPT gate with audit logging, Librarian index exclusion of keys/.env/credentials, AES-256-SIV stick encryption with key never on the stick, and secret redaction at ingest. Deductions: authority tiers (4.23) are self-described as an 'audit trail plus honor system' — any agent holding the server token can accept/lock on the user's behalf; fact conflicts push outbound notifications to a Discord webhook; agents can demote each other's records; the loopback guard and gate hardening are code-level claims whose behavior a static review cannot confirm. The proactive exclusion of malicious fastapi 0.136.3 is a plus.
CI covers pytest (with drift-guard and security regression tests), wheel smoke, and Docker boot checks; the classifier falls back to regex tagging so saves are never blocked; stick sync plans before moving bytes, writes the manifest last, and refuses loudly on torn state. Deductions: comments admit a flaky suite hang (pytest-timeout mitigates rather than removes it), and version drift shipped in v4.9.1/4.9.2/nearly 4.9.4 before CI — process stability is newly constructed, not long-proven.
Excellent audience coverage: per-client install docs for Claude Code/Desktop, ChatGPT, LM Studio, AnythingLLM, Agent Zero, Hermes, Ollama, Open WebUI, etc., plus shared/isolated/hybrid deployment modes. Deductions: the ChatGPT route is plan-gated and not MCP; Python 3.11+ floor; local-LLM classification quality directly determines tier accuracy and is not quantified across environments.
Strong information architecture: robot.info JSON manifest, llms.txt, explicit INSTALL.md vs robot.install disambiguation, 84-version CHANGELOG, complete MIT LICENSE, SECURITY.md with private reporting. Known limitations are unusually candid (seed anti-memory warning, authority-tier enforcement gap, unreviewed Fable pass). Deductions: single-maintainer with a self-stated multi-day response time; mnemo/mnemo-cortex dual entry points and near-identical filenames are a confusion risk; the 'production since March 2026' timeline cannot be statically verified.
Output usability is good: sub-millisecond structured-fact lookup, tiered recall, BM25/phonetic hybrid ranking, each with an off switch and CLI usage. High marginal value: cross-agent overnight synthesis, USB sneakernet, and a tamper-evident ledger are an uncommon combination for local memory systems. Deductions: cost side includes an extra LLM classification call per save, nightly dreaming Ollama spend, and 84 releases in 5 months implies an upgrade burden; all performance figures (11/11, 107K files in 17s) come from the author's own deployment and are unverifiable statically.
Claim traceability is moderate: core mechanisms (tiers, lexical lane, seed warning) ship with design rationale and failure narratives, and pyproject carries checkable dependency annotations. Deductions: ~10,000 memories, ~7,000 facts, the 'only system that does cross-agent synthesis' exclusivity claim, and comparisons to Mem0/Zep/Letta are assertions without in-repo data or tests; promotional narrative (73-year-old maker, Fable pass) is interleaved with facts, so the fact/inference boundary must be peeled by the reader.
- Authority-tier accept/lock can be performed by any agent holding the server token; the author admits enforcement is not yet implemented — do not treat it as a real permission boundary in multi-agent setups.
- Fact-conflict notifications make outbound calls to a Discord webhook; confirm that path fits your network and privacy policy before deploying.
- No runtime behavior was verified in this static review; the gate, loopback guard, and stick encryption are code-level claims — audit before sensitive deployments.
- Read the example-file warning before scheduling seed-facts: a stale seed file continuously clobbers your corrections.
- Single-maintainer project with multi-day stated security response time; do not depend on it as a high-availability or strongly-assured component.