Mnemo Cortex

A local-first cognitive coprocessor for AI agents: active memory, semantic recall, and cross-agent sharing with no cloud.

Stars
★ 155
Last updated
2d ago
License
MIT
Primary language
Python

At a glance

How it runs
Self-hosted serviceMCP serverCLI
Works with
Universal · cross-platformCodex · Claude Code · Claude.aiChatGPT (Partial support)
Cost
Free, no paid service needed
Setup effort
Medium · a few setup steps
You'll need
Python 3.11+SQLite FTS5Ollama (recommended)Node.js 18+ (only for MCP-bridge integrations)Shell / CLINetwork accessLocal filesystemMCP Server
Typical use
A solo developer running multiple agents (Claude Code, Claude Desktop, Codex CLI, Hermes) who wants Agent A to recall what Agent B learned
Not a fit if
  • Teams wanting a zero-deployment hosted cloud memory
  • Users needing OpenAI cloud connectors to reach a self-hosted MCP directly
  • Users unwilling to maintain a local server process
Source review
75/100 · Good

What does this agent do, and when should you use it?

Mnemo Cortex is a self-hosted cognitive coprocessor that gives AI agents persistent, local, cross-agent memory. It watches agent session files from the outside (JSONL for OpenClaw and Claude Code), ingests every message into a local SQLite database, and applies rolling LLM-backed compaction via a local Ollama model (default qwen2.5:32b-instruct) for roughly 80% token reduction. The system runs as an API server (default http://localhost:50001) and connects to clients — Claude Desktop, LM Studio, AnythingLLM, OpenClaw, Agent Zero, Hermes — through MCP bridges; ChatGPT is served through a hardened two-route REST gate instead of MCP. Core components include tiered smart notes versus session logs, a structured facts store with a confidence ladder and authority tiers, nightly cross-agent 'dreaming' synthesis, The Librarian workspace document index (SQLite FTS5), the encrypted USB-courier Cortex Stick, the Disco-Bus agent messaging system, and a tamper-evident Memory Ledger. According to the README, it has run since March 2026 as the production memory of a five-agent fleet, holding about 10,000 memories and ~7,000 verified facts across 84 shipped versions.

After mnemo-cortex start, the API server listens on port 50001; mnemo-cortex watch --backfill starts a session watcher that tails JSONL session files and auto-ingests messages every 2 seconds (no hooks or agent changes needed). Saves are tiered at write time: Tier 1 smart notes are classified by the reasoning model into one of eight categories (topology, current_state, doctrine, incident, identity, relationship, decision), while Tier 2 holds raw session logs excluded from default recall. Recall runs semantic vector search and a BM25 keyword lane (SQLite FTS5) in parallel; since v4.22 it also matches misspelled proper names via consonant-skeleton keys. Structured facts are stored as (entity, attribute, value) triples with an automatic verified→high_probability→false confidence ladder and four MCP tools: mnemo_fact_save, mnemo_fact_get, mnemo_fact_query, and mnemo_fact_demote. The nightly dreaming pipeline map-reduces each agent's recent memories into themes, then merges them cross-agent into a shared brief. mnemo-cortex health verifies the server, database, compaction model, and per-agent recall; mnemo-cortex stick sync carries deltas between two machines via an optionally AES-256-encrypted USB stick.

  1. A solo developer running multiple agents (Claude Code, Claude Desktop, Codex CLI, Hermes) who wants Agent A to recall what Agent B learned
  2. A user working across two desks who syncs agent memory via the encrypted Cortex Stick USB courier — no cloud, VPN, or account
  3. Local-LLM users (LM Studio, Open WebUI, Ollama, AnythingLLM) who want persistent semantic memory at zero API cost
  4. A ChatGPT Plus subscriber who wants a Custom GPT with memory they own, accessed through a private gate rather than an exposed server
  5. Users who need exact facts (names, settings, identifiers) rather than fuzzy semantics — served by the structured facts store and the BM25 lexical lane
  6. Multi-agent teams wanting overnight cross-agent dream synthesis plus the Lane Protocol session workflow

How do you install or deploy this agent?

Prerequisites: Python 3.11+, optionally Ollama (recommended for local reasoning and embeddings), Node.js 18+ only when running an MCP-bridge integration. Install from source:

bash

git clone https://github.com/GuyMannDude/mnemo-cortex.git
cd mnemo-cortex
python -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate

pip install -e .

This registers the mnemo-cortex CLI (alias mnemo). For non-interactive install (LLM agents or CI):

bash

./robot-install.sh

The script emits a single JSON object on stdout (ok, steps, etc.) with progress on stderr; sandbox runs are supported via MNEMO_INSTALL_DRY_RUN=1 and path overrides. A dedicated macOS guide exists (docs/install-macos.md); Windows is natively supported as of v4.4.1.

How do you use this agent?

Five-step flow:

bash

mnemo-cortex init      # Interactive wizard: reasoning/embedding models, bind address, port; writes ~/.agentb/agentb.yaml
mnemo-cortex start     # Detached start; server listens on http://localhost:50001
mnemo-cortex health    # 14-point health check; use mnemo-cortex doctor for deeper diagnostics

Enable auto-capture:

bash

mnemo-cortex watch --backfill

Or set export MNEMO_AUTO_CAPTURE=true in your shell profile so the watcher starts with the server. Then connect a client: Claude Code via integrations/claude-code/; Claude Desktop via a drag-and-drop .mcpb bundle; LM Studio, AnythingLLM, OpenClaw, Agent Zero via a one-line MCP config; ChatGPT via Custom GPT Actions against the two-route gate (Plus tier or higher). Direct API example:

bash

curl -s -X POST http://localhost:50001/context \
-H "Content-Type: application/" \
-d '{"prompt": "what happened today", "agent_id": "YOUR-AGENT-ID", "max_results": 5}'

Recommended companions: a forked mnemo-plan brain repo pointed at via the BRAIN_DIR env var, and the THE-LANE-PROTOCOL.md operating ritual.

What are this agent's strengths and limitations?

Pros
  • Only memory system claiming cross-agent overnight dreaming — map-reduce compaction of each agent's memory into a shared brief so all agents wake up knowing what the others did
  • Fully local-first: no telemetry, no cloud dependency, single-file SQLite storage, and $0 compaction/embedding via local Ollama models
  • Multi-lane recall: semantic vectors + BM25 keyword lane (SQLite FTS5) + phonetic name matching; author's test shows 11/11 exact hits versus 8/11 with vectors alone
  • Structured facts with a confidence ladder, authority tiers (probe/declared/open), tamper-evident Memory Ledger, and a demote endpoint for wrong memories
  • Rich operational tooling: per-platform install guides, robot-install non-interactive setup, health/doctor diagnostics, Disco-Bus agent bus, The Librarian document index (~107K files, ~17s full rebuild)
Limitations
  • Requires a continuously running local server process (default 127.0.0.1:50001) that you must keep alive via systemd/launchd/Task Scheduler yourself
  • Depends on local Ollama models for compaction and embeddings (qwen2.5:32b-instruct, nomic-embed-text); embedding model-name drift is a documented failure mode causing empty recall
  • ChatGPT cannot use standard MCP — a custom hardened gate is required, and OpenAI plan tiers restrict functionality (read-only custom MCP connectors on Plus)
  • Authority tiers are currently an audit trail plus honor system; enforcement that actually blocks agent writes is not yet implemented
  • Personal/small-team project by a self-described non-developer; the documented timeline (2026 dates) and long-term maintenance cannot be independently verified

How does this agent compare with similar options?

The README explicitly compares against OpenClaw's native Active Memory plugin (2026.4.10): that plugin is single-agent, per-workspace scratchpad storage, while Mnemo is a centralized SQLite + embeddings memory bus that survives resets and machine moves — the two stack rather than compete. It also self-positions against Mem0, Zep, and Letta, which store memory per agent, whereas Mnemo dreams across all of them (author's claim). For ChatGPT, instead of exposing a standard MCP server over HTTPS, it ships a hardened two-route gate — a deliberate design contrast with typical self-hosted MCP servers.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Mnemo Cortex This agent 75 · Good Self-hosted serviceFree ★ 155 2d ago Python Codex · Claude Code · Claude.ai
Compartment 76 · Good CLIFree ★ 582 6d ago Python Codex · Claude Code
Remnic Agent Memory 85 · Good CLIFree + model costs ★ 210 9d ago TypeScript ChatGPT · Codex · Claude Code · OpenAI API
Ori Mnemos 80 · Good CLIFree ★ 326 8d ago TypeScript Claude Code · OpenAI API · Claude API

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Good
75/ 100 5-point scale 3.8 / 5
Trust 21/29
Reliability 9/14
Adaptability 14/18
Convention 16/18
Effectiveness 10/13
Verifiability 5/8
Why each dimension lost points
Trust21 / 29 · 3.6/5

Evidence shows local-first binding (localhost/private net), a hardened bearer-auth tenant-pinned rate-limited ChatGPT gate with audit logging, Librarian index exclusion of keys/.env/credentials, AES-256-SIV stick encryption with key never on the stick, and secret redaction at ingest. Deductions: authority tiers (4.23) are self-described as an 'audit trail plus honor system' — any agent holding the server token can accept/lock on the user's behalf; fact conflicts push outbound notifications to a Discord webhook; agents can demote each other's records; the loopback guard and gate hardening are code-level claims whose behavior a static review cannot confirm. The proactive exclusion of malicious fastapi 0.136.3 is a plus.

Reliability9 / 14 · 3.2/5

CI covers pytest (with drift-guard and security regression tests), wheel smoke, and Docker boot checks; the classifier falls back to regex tagging so saves are never blocked; stick sync plans before moving bytes, writes the manifest last, and refuses loudly on torn state. Deductions: comments admit a flaky suite hang (pytest-timeout mitigates rather than removes it), and version drift shipped in v4.9.1/4.9.2/nearly 4.9.4 before CI — process stability is newly constructed, not long-proven.

Adaptability14 / 18 · 3.9/5

Excellent audience coverage: per-client install docs for Claude Code/Desktop, ChatGPT, LM Studio, AnythingLLM, Agent Zero, Hermes, Ollama, Open WebUI, etc., plus shared/isolated/hybrid deployment modes. Deductions: the ChatGPT route is plan-gated and not MCP; Python 3.11+ floor; local-LLM classification quality directly determines tier accuracy and is not quantified across environments.

Convention16 / 18 · 4.4/5

Strong information architecture: robot.info JSON manifest, llms.txt, explicit INSTALL.md vs robot.install disambiguation, 84-version CHANGELOG, complete MIT LICENSE, SECURITY.md with private reporting. Known limitations are unusually candid (seed anti-memory warning, authority-tier enforcement gap, unreviewed Fable pass). Deductions: single-maintainer with a self-stated multi-day response time; mnemo/mnemo-cortex dual entry points and near-identical filenames are a confusion risk; the 'production since March 2026' timeline cannot be statically verified.

Effectiveness10 / 13 · 3.8/5

Output usability is good: sub-millisecond structured-fact lookup, tiered recall, BM25/phonetic hybrid ranking, each with an off switch and CLI usage. High marginal value: cross-agent overnight synthesis, USB sneakernet, and a tamper-evident ledger are an uncommon combination for local memory systems. Deductions: cost side includes an extra LLM classification call per save, nightly dreaming Ollama spend, and 84 releases in 5 months implies an upgrade burden; all performance figures (11/11, 107K files in 17s) come from the author's own deployment and are unverifiable statically.

Verifiability5 / 8 · 3.1/5

Claim traceability is moderate: core mechanisms (tiers, lexical lane, seed warning) ship with design rationale and failure narratives, and pyproject carries checkable dependency annotations. Deductions: ~10,000 memories, ~7,000 facts, the 'only system that does cross-agent synthesis' exclusivity claim, and comparisons to Mem0/Zep/Letta are assertions without in-repo data or tests; promotional narrative (73-year-old maker, Fable pass) is interleaved with facts, so the fact/inference boundary must be peeled by the reader.

Risks and how to mitigate them
  • Authority-tier accept/lock can be performed by any agent holding the server token; the author admits enforcement is not yet implemented — do not treat it as a real permission boundary in multi-agent setups.
  • Fact-conflict notifications make outbound calls to a Discord webhook; confirm that path fits your network and privacy policy before deploying.
  • No runtime behavior was verified in this static review; the gate, loopback guard, and stick encryption are code-level claims — audit before sensitive deployments.
  • Read the example-file warning before scheduling seed-facts: a stale seed file continuously clobbers your corrections.
  • Single-maintainer project with multi-day stated security response time; do not depend on it as a high-availability or strongly-assured component.
Evidence confidence: Low Reviewed Sep 28, 2026 Reviewed revision d7fb0294d360
See the full review method →

FAQ

Does Mnemo Cortex cost anything to use?
The software is MIT-licensed and free. Running compaction and embeddings on local Ollama models costs $0. If you configure OpenAI/Google/Anthropic/OpenRouter as your reasoning or embedding provider, you supply the API keys and pay their usage. The project is donation-funded overall.
Can ChatGPT connect directly?
Not via standard local MCP (ChatGPT has no stdio transport). The project ships a bearer-authenticated, tenant-pinned, rate-limited, audit-logged two-route gate that a Custom GPT calls through Actions. Plus tier or higher is required, and full custom MCP save+recall in the main ChatGPT app needs Business/Enterprise. The Mnemo server itself never faces the internet.
Why does recall return 'No chunks'?
The most common cause is an embedding model name that no longer matches the provider's current model (Ollama: nomic-embed-text; Google: gemini_embedding_001; text-embedding-004 was shut down in January 2026). Fix the embedding setting in your config.
Is it production-proven?
Per the README, it has been the daily production memory of a five-agent fleet (CC, Opie, Dreamer, Rocky, Cody on Claude Code, Claude Desktop, Hermes, and Codex CLI) since March 2026, with ~10,000 memories, ~7,000 verified facts, and 84 versions in five months. These figures are author-reported and not independently verifiable.
Do agents see each other's memories? What about privacy?
Deployment is configurable: shared (cross-agent search and dreaming for the whole fleet), isolated (per-agent or per-customer stores with zero bleed), or hybrid (shared internal, isolated customer-facing). All data stays in a local SQLite database you control, with no telemetry or third-party services, and you can read or delete it anytime.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents