PRO-LONG
Durable programmatic memory for coding agents: one local append-only log of session events, searched by the agent with code instead of being stuffed into the prompt.
- Source repo
- alexisfox7/PRO-LONG
- Stars
- ★ 458
- Last updated
- 1mo ago
- License
- MIT
- Primary language
- Python
- FA score
- 57/100 · Major gaps
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code
- Cost
- Free, no paid service needed
- Setup effort
- Low · running in minutes
- You'll need
- Typical use
- Engineers running large refactors in Claude Code or Codex that span compaction or multi-day sessions and need earlier decisions preserved
- Not a fit if
- Teams whose policy forbids retaining transcripts that may contain secrets
- Developers using coding CLIs outside the four supported ones who won't adapt hooks
- Buyers expecting the 97.4% ARC-AGI-3 result to be a direct benchmark of the coding-tool MVP
- Source review
- 57/100 · Major gaps
What does this agent do, and when should you use it?
PRO-LONG (GitHub: alexisfox7/PRO-LONG, MIT license) gives coding CLIs like Codex, Claude Code, OpenCode, and pi programmable long-term memory. Long tasks outlive context windows: after compaction or a fresh session, agents can lose earlier decisions and repeat failed work. PRO-LONG uses project-local adapters to append prompts, tool activity, assistant handoffs, and session boundaries to .prolong/log.l, and ships an agent skill that teaches the agent to search the log with ordinary tools like rg, jq, or Python. It never injects the full log into prompts, adds no wrapper, server, or database, and excludes log reads from recording so memory cannot recursively copy itself. It originated in long-horizon agent research, reaching 97.4% best@2 with Fable 5 on the full ARC-AGI-3 public game set—an 18-point average improvement over matched baselines—though the authors note this does not yet directly benchmark the coding-tool MVP.
After prolong init detects supported clients on PATH, it installs: .prolong/runtime.mjs (a dependency-free event writer), .prolong/log.l (the append-only log), .prolong/install. (install manifest), .agents/skills/prolong/SKILL.md (the retrieval skill), a managed AGENTS.md pointer, a .prolong/ entry in .gitignore, and client-specific integrations (.codex/hooks. for Codex, .claude/settings. lifecycle hooks for Claude Code, .opencode/plugins/prolong.ts for OpenCode, .pi/extensions/prolong.ts for pi). The adapter appends lifecycle events to the log; when prior work matters, the skill teaches the agent to search it with rg, jq, or Python and read only relevant entries. The full CLI is three commands: prolong init (optionally --client codex,claude-code,opencode,pi), prolong status (with -- for automation), and prolong uninstall (preserves the log; --purge deletes history too).
- Engineers running large refactors in Claude Code or Codex that span compaction or multi-day sessions and need earlier decisions preserved
- Debugging long tasks where the agent should retrieve earlier failed tool calls and results instead of repeating them
- Developers who start frequent fresh sessions in one project and want relevant history recovered each time
- OpenCode or pi users wanting a local, serverless, database-free memory layer
- Privacy-conscious teams requiring that memory stays inside the project and is gitignored by default
How do you install or deploy this agent?
Requires Node.js/npm and a supported coding client already on PATH. Clone and build:
bash
git clone https://github.com/alexisfox7/PRO-LONG.git && cd PRO-LONG
npm install && npm run build && npm linkThen initialize inside your project:
bash
cd /path/to/your/project
prolong initTo select clients explicitly:
bash
prolong init --client codex,claude-code,opencode,piHow do you use this agent?
After initialization, use your coding agent normally—PRO-LONG records in the background and the agent retrieves relevant history via the prolong skill when tasks span compaction or sessions. Check status or uninstall with:
bash
prolong status
prolong status --
prolong uninstall
prolong uninstall --purgeNote: hooks and extensions run with your user permissions; review the generated integration and accept the client's project or hook trust prompt before use.
What are this agent's strengths and limitations?
- The log is never injected into every prompt, and log reads are excluded from recording, preventing context bloat and recursive self-copying
- Zero-dependency local runtime.mjs; no server, no database, log stays in the project and is gitignored by default
- One command integrates four mainstream coding CLIs (Codex, Claude Code, OpenCode, pi) without changing your workflow
- Backed by ARC-AGI-3 research (97.4% best@2, +18 points average) with reproduction code and a paper
- The log can contain prompts, tool inputs/results, and any secrets that appeared in them—sensitive projects must check local policy
- The 97.4% figure comes from the research harness with Fable 5; the authors state it is not yet a direct benchmark of the coding-tool MVP
- Only four clients are supported out of the box; other CLIs require hand-written hooks/plugins
- Memory value depends on the agent following SKILL.md and actively searching—no retrieval means no benefit
How does this agent compare with similar options?
The README names no specific competitors, but its positioning implies a contrast with approaches that inject full transcripts/memory into prompts or rely on external memory servers/databases: PRO-LONG instead uses a local append-only log searched with code, with no wrapper or service.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| PRO-LONG This agent | 57 · Major gaps | CLIFree | ★ 458 | 1mo ago | Python | Codex · Claude Code |
| MonoCode | 56 · Major gaps | Desktop appFree | ★ 1.6k | 1d ago | TypeScript | Codex · Claude Code |
| Coder Eval | 86 · Good | CLIFree + model costs | ★ 141 | today | Python | Codex · Claude Code · OpenAI API · Claude API |
| Emulo | 78 · Good | Agent plugin / skillFree + model costs | ★ 292 | 3d ago | HTML | Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
Trust: README states hooks run with user permissions, the runtime is dependency-free, the log stays project-local and gitignored, privacy notes warn the log may contain secrets, and --purge exists; uninstall preserves existing user config. Deductions: src/ code is not among the evidence, so runtime.mjs behavior, hook injection content, and log-writing logic rest on README assertions only — not full marks. Source attribution relies on a CITATION.cff that is not provided and the publisher is unverified: 1.
Reliability: tests/lifecycle.test.ts has evident defects — assertions reference undefined claudeConfig and pluginPath and use helpers inconsistent with the imports — so the suite likely fails to compile, undercutting self-consistency: 1. The tests do demonstrate idempotent init, safe failure on invalid hook JSON without partial writes, status --, and uninstall warnings: 2. Node>=20 with zero runtime dependencies keeps availability risk low: 2.
Adaptability: the coding-agent audience is clear, four clients are supported with explicit selection, and the README honestly separates ARC-AGI-3 research results from the coding-tool MVP, bounding claims. Deduction: trigger precision — when and how the agent searches — is one sentence with no SKILL.md or source shown: 1.
Convention: MIT license file is present and consistent with package.: 3. README structure, install notes, stable naming (pro-long/prolong), and candid limitations earn 2 each. Deductions: no CHANGELOG, only a README 'Updates' line dated August 2026 (a suspicious future date) at version 0.1.0: 1; no usage examples or FAQ, notably none for the core log-search workflow: 1; maintenance responsibility is just an issues URL with an unverified publisher and no governance: 1.
Effectiveness: the output (.prolong/log.l) is searchable with standard rg/jq/Python and status supports --: 2. The marginal-value argument for compaction-induced context loss is plausible but the MVP itself has no benchmark data: 2. Deduction: cost-benefit is unverifiable — the README admits the 97.4% result does not directly benchmark the coding tool, and real overhead/gain is unevidenced: 1.
Verifiability: paper and reproduction-code links give a traceability path, but the arXiv ID 2607.20064 maps to July 2026, outside what can be corroborated, and the research code and CITATION.cff are not included: 1. No independent corroboration; 97.4% and 18pp are self-reported: 1. Credit: the README explicitly states the research results are 'not yet a direct benchmark of the coding-tool MVP', a good fact/inference separation: 2.
- The test file has evident code defects (undefined variables, incomplete imports); under static review the suite cannot be assumed to run — do not infer functional correctness from it.
- Core source (src/, runtime.mjs, SKILL.md) is absent from this evidence set; all safety and behavior claims come from the README. Review generated hooks and plugin files yourself before use.
- The arXiv ID and the August 2026 update date are unverifiable; 97.4% and 18pp are self-reported, and the authors concede they do not directly benchmark the coding-tool MVP.
- The log records prompts and tool results and may contain secrets; do not enable transcript retention where local policy forbids it, and run prolong uninstall --purge when needed.
- Version 0.1.0, no CHANGELOG, unverified publisher: suitable for evaluation trials, not production reliance.