Agent Memory Benchmark (AMB)
Reproducibly evaluate agent memory systems on accuracy, speed, and token cost.
- Source repo
- vectorize-io/agent-memory-benchmark
- Stars
- ★ 144
- Last updated
- today
- Primary language
- Python
- FA score
- 39/100 · Major gaps
At a glance
- How it runs
- Works with
- Portable with changes
- Cost
- Free software; you pay for model usage
- Setup effort
- Medium · a few setup steps
- You'll need
- Typical use
- Memory-system developers comparing providers on answer accuracy, speed, and token cost over the same dataset.
- Not a fit if
- Teams that need to run evaluations without an API key
- Teams evaluating MemBench without access to its local data directory
- Source review
- 39/100 · Major gaps 2 safety controls not found
What does this agent do, and when should you use it?
Agent Memory Benchmark (AMB) is a benchmark suite for agent memory systems, publishing its datasets, prompts, scoring logic, and results. Its CLI, invoked with `uv run amb`, provides commands to inspect datasets and providers, run evaluations, view dataset statistics, and browse results. In the standard flow, dataset documents are ingested into a memory provider, the provider retrieves context for each query, Gemini generates an answer, and a second Gemini call scores it against gold answers. AMB tracks retrieval and generation time separately and records ingestion time; results are saved under `outputs/{dataset}/{memory}/{mode}/{domain}.json` and can be browsed with `uv run amb view`. PrecisionMemBench uses `--mode retrieval` to score whether returned belief-ID sets satisfy per-case assertions, without an answer LLM or judge call.
AMB loads documents from a selected dataset into a chosen memory provider, asks that provider to retrieve relevant context for each query, and, in the standard mode, calls Gemini to generate an answer and Gemini again to score it against gold answers. It records retrieval, generation, and ingestion time and saves evaluation results, including accuracy, speed, and token-cost measurements. The --oracle option ingests only gold documents to evaluate generation quality in isolation; dataset-stats reports dataset statistics. For PrecisionMemBench, --mode retrieval checks the returned belief-ID set against each case's required and forbidden assertions directly.
- Memory-system developers comparing providers on answer accuracy, speed, and token cost over the same dataset.
- Researchers using
--oracleto isolate generation quality and distinguish generation issues from retrieval-context issues. - Agent teams evaluating how preferences accumulated across interactions affect multi-step decisions with datasets such as PersonaMem.
- Memory-provider authors checking whether retrieval returns specified beliefs with PrecisionMemBench, where noisy results count as failures.
- Evaluation leads reproducing and auditing benchmark results using the published harness, prompts, and model configuration.
How do you install or deploy this agent?
Requires Python ≥ 3.11. The README does not specify an installation command or package installation process; it gives these setup steps for the Gemini API key:
cp .env.example .envSet the key in .env:
GEMINI_API_KEY=...To run MemBench, also set MEMBENCH_DATA_PATH to your local data directory.
How do you use this agent?
List available datasets, memory providers, and modes:
uv run amb providersRun a benchmark:
uv run amb run --dataset personamem --domain 32k --memory bm25Limit the number of queries for a quick run:
uv run amb run --dataset personamem --domain 32k --memory bm25 --query-limit 20Browse results:
uv run amb viewRun PrecisionMemBench in retrieval mode:
uv run amb run --dataset precisionmembench --split single-turn --memory hindsight-cloud --mode retrievalWhat are this agent's strengths and limitations?
- Publishes the evaluation harness, judge prompts, answer-generation prompts, and exact model configuration for reproducibility and review.
- Tracks retrieval speed, generation speed, ingestion time, and token cost alongside accuracy.
- The
--oraclemode can separate generation quality from memory retrieval performance. - PrecisionMemBench explicitly asserts which beliefs must and must not be returned for each case; noisy results affect the score.
- Standard evaluations require
GEMINI_API_KEYand use Gemini for answer generation and judging. - Requires Python 3.11 or newer.
- MemBench requires an existing local data directory configured through
MEMBENCH_DATA_PATH. - The README shows
uv run ambcommands but does not provide installation steps or a dependency list.
How does this agent compare with similar options?
The README describes LoComo and LongMemEval as datasets designed for chatbot use and the 32k-context era, noting that dumping all context can score competitively with million-token windows. AMB adds datasets for agent tasks, including memory across tool calls, knowledge built from document research, and preferences applied to multi-step decisions.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Agent Memory Benchmark (AMB) This agent | 39 · Major gaps | CLIFree + model costs | ★ 144 | today | Python | — |
| ReLE Chinese LLM Benchmark & Defect Library | 14 · Major gaps | Web appFreemium | ★ 6.5k | 9d ago | — | — |
| Dolt: Git for Data | 55 · Major gaps | CLIFree | ★ 25k | today | Go | — |
| MobileGym | 50 · Major gaps | CLIFree + model costs | ★ 805 | today | Python | — |
How does FollowAgents rate this agent?
Why each dimension lost points
The README describes document ingestion, retrieval, generation, judging, result storage, and the Gemini key environment variable, so data flow is partly visible. Least privilege, sensitive data handling, and external effects score 1–2. The supplied material shows no confirmation step, permission boundaries, key protection details, or rollback process; user_confirmation and rollback score 0, and related safety criteria receive no higher marks. Dependencies have minimum versions and one compatibility note but no security review evidence, so dependency_security scores 1. The README names the upstream PrecisionMemBench project but gives no maintainer or source governance details. Publisher identity is unknown; source_attribution scores 1.
README usage generally matches the command entry point, package name, and Python requirement in pyproject.toml, so self_consistency scores 2. Dependencies have minimum versions and a limited compatibility explanation, but no lockfile or availability guarantees, so dependency_availability scores 1. The files do not describe failure messages or error handling, so failure_messages scores 0.
The README describes conversation memory, document research, and cross-tool tasks, and gives examples for datasets, memory providers, modes, and query limits; audience_and_scenarios scores 2. It distinguishes ordinary answer evaluation from PrecisionMemBench retrieval mode, but does not systematically define capability boundaries or triggering rules, so capability_boundaries and trigger_precision score 1 each. Python, environment variables, the MemBench data path, and command line steps are documented, so environment_fit scores 2.
The README is organized into overview, problem, evaluation flow, setup, usage, results, and requirements, and includes command examples. Setup and naming guidance are clear enough for 2 in install_notes, naming_stability, and examples_and_faq. It gives only limited discussion of benchmark shape and cost rather than a complete limitations list, so known_limitations scores 1. The supplied files contain no license, changelog, or maintenance owner evidence, so license, versioning_changelog, and maintenance_responsibility score 0. Information architecture is assessable only from the README, so it scores 1.
The tool offers dataset and provider listings, statistics, benchmark runs, result files, and browser viewing, making outputs reasonably usable; output_usability scores 2. The README gives a concrete motivation about benchmark discrimination with million-token context windows and identifies accuracy, speed, and token cost as measures, so marginal_value scores 2. It emphasizes cost measurement but supplies no concrete estimate of run costs or resource needs, so cost_benefit scores 1.
The README names the evaluation flow, modes, result path, and materials it says are needed for reproduction, and identifies the PrecisionMemBench upstream project; this makes some claims traceable, so claim_traceability scores 2. The available evidence is limited to README and project configuration. The configuration corroborates the entry point, dependencies, and Python requirement but cannot independently establish results, dataset contents, or reproducibility claims, so cross_source_corroboration scores 1. The README separates the evaluation steps from PrecisionMemBench retrieval scoring and presents performance statements as project motivation; however, the claims cannot be independently checked from the supplied material, so fact_inference_separation scores 2 rather than full marks.
- Not found in source: confirmation before actingTurn on (or add) a confirmation step before it acts, and try it in a sandbox or test environment before real data.
- Not found in source: rollback or recovery pathBack up first, or work on a git branch or snapshot, so its changes can be undone.
- The supplied files do not show a license, changelog, maintenance owner, dependency lock, security review, or error handling guidance; do not infer that these are covered.
- The README’s claims about benchmark advantages and reproducibility are not corroborated by results or independent sources in the supplied material.
FAQ
Which API key is needed to run evaluations?
GEMINI_API_KEY; Gemini is used to generate and judge answers.Do all datasets require answer generation and judging calls?
--mode retrieval skips the answer LLM and judge, scoring whether belief-ID sets satisfy the assertions.Can I run MemBench without preparing data?
MEMBENCH_DATA_PATH to it.Where are results saved?
outputs/{dataset}/{memory}/{mode}/{domain}.json and can be browsed with uv run amb view.