Agent Memory Benchmark (AMB)

Reproducibly evaluate agent memory systems on accuracy, speed, and token cost.

Stars
★ 144
Last updated
today
Primary language
Python

At a glance

How it runs
CLIWeb app
Works with
Portable with changes
Cost
Free software; you pay for model usage
Setup effort
Medium · a few setup steps
You'll need
Python ≥ 3.11Gemini API keyMEMBENCH_DATA_PATH (for MemBench)Shell / CLINetwork accessLocal filesystem
Typical use
Memory-system developers comparing providers on answer accuracy, speed, and token cost over the same dataset.
Not a fit if
  • Teams that need to run evaluations without an API key
  • Teams evaluating MemBench without access to its local data directory

What does this agent do, and when should you use it?

Agent Memory Benchmark (AMB) is a benchmark suite for agent memory systems, publishing its datasets, prompts, scoring logic, and results. Its CLI, invoked with `uv run amb`, provides commands to inspect datasets and providers, run evaluations, view dataset statistics, and browse results. In the standard flow, dataset documents are ingested into a memory provider, the provider retrieves context for each query, Gemini generates an answer, and a second Gemini call scores it against gold answers. AMB tracks retrieval and generation time separately and records ingestion time; results are saved under `outputs/{dataset}/{memory}/{mode}/{domain}.json` and can be browsed with `uv run amb view`. PrecisionMemBench uses `--mode retrieval` to score whether returned belief-ID sets satisfy per-case assertions, without an answer LLM or judge call.

AMB loads documents from a selected dataset into a chosen memory provider, asks that provider to retrieve relevant context for each query, and, in the standard mode, calls Gemini to generate an answer and Gemini again to score it against gold answers. It records retrieval, generation, and ingestion time and saves evaluation results, including accuracy, speed, and token-cost measurements. The --oracle option ingests only gold documents to evaluate generation quality in isolation; dataset-stats reports dataset statistics. For PrecisionMemBench, --mode retrieval checks the returned belief-ID set against each case's required and forbidden assertions directly.

  1. Memory-system developers comparing providers on answer accuracy, speed, and token cost over the same dataset.
  2. Researchers using --oracle to isolate generation quality and distinguish generation issues from retrieval-context issues.
  3. Agent teams evaluating how preferences accumulated across interactions affect multi-step decisions with datasets such as PersonaMem.
  4. Memory-provider authors checking whether retrieval returns specified beliefs with PrecisionMemBench, where noisy results count as failures.
  5. Evaluation leads reproducing and auditing benchmark results using the published harness, prompts, and model configuration.

How do you install or deploy this agent?

Requires Python ≥ 3.11. The README does not specify an installation command or package installation process; it gives these setup steps for the Gemini API key:

cp .env.example .env

Set the key in .env:

GEMINI_API_KEY=...

To run MemBench, also set MEMBENCH_DATA_PATH to your local data directory.

How do you use this agent?

List available datasets, memory providers, and modes:

uv run amb providers

Run a benchmark:

uv run amb run --dataset personamem --domain 32k --memory bm25

Limit the number of queries for a quick run:

uv run amb run --dataset personamem --domain 32k --memory bm25 --query-limit 20

Browse results:

uv run amb view

Run PrecisionMemBench in retrieval mode:

uv run amb run --dataset precisionmembench --split single-turn --memory hindsight-cloud --mode retrieval

What are this agent's strengths and limitations?

Pros
  • Publishes the evaluation harness, judge prompts, answer-generation prompts, and exact model configuration for reproducibility and review.
  • Tracks retrieval speed, generation speed, ingestion time, and token cost alongside accuracy.
  • The --oracle mode can separate generation quality from memory retrieval performance.
  • PrecisionMemBench explicitly asserts which beliefs must and must not be returned for each case; noisy results affect the score.
Limitations
  • Standard evaluations require GEMINI_API_KEY and use Gemini for answer generation and judging.
  • Requires Python 3.11 or newer.
  • MemBench requires an existing local data directory configured through MEMBENCH_DATA_PATH.
  • The README shows uv run amb commands but does not provide installation steps or a dependency list.

How does this agent compare with similar options?

The README describes LoComo and LongMemEval as datasets designed for chatbot use and the 32k-context era, noting that dumping all context can score competitively with million-token windows. AMB adds datasets for agent tasks, including memory across tool calls, knowledge built from document research, and preferences applied to multi-step decisions.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Agent Memory Benchmark (AMB) This agent 39 · Major gaps CLIFree + model costs ★ 144 today Python —
ReLE Chinese LLM Benchmark & Defect Library 14 · Major gaps Web appFreemium ★ 6.5k 9d ago — —
Dolt: Git for Data 55 · Major gaps CLIFree ★ 25k today Go —
MobileGym 50 · Major gaps CLIFree + model costs ★ 805 today Python —

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Major gaps
39/ 100 5-point scale 2.0 / 5
Trust 8/29
Reliability 5/14
Adaptability 9/18
Convention 6/18
Effectiveness 7/13
Verifiability 4/8
Why each dimension lost points
Trust8 / 29 · 1.4/5

The README describes document ingestion, retrieval, generation, judging, result storage, and the Gemini key environment variable, so data flow is partly visible. Least privilege, sensitive data handling, and external effects score 1–2. The supplied material shows no confirmation step, permission boundaries, key protection details, or rollback process; user_confirmation and rollback score 0, and related safety criteria receive no higher marks. Dependencies have minimum versions and one compatibility note but no security review evidence, so dependency_security scores 1. The README names the upstream PrecisionMemBench project but gives no maintainer or source governance details. Publisher identity is unknown; source_attribution scores 1.

Reliability5 / 14 · 1.8/5

README usage generally matches the command entry point, package name, and Python requirement in pyproject.toml, so self_consistency scores 2. Dependencies have minimum versions and a limited compatibility explanation, but no lockfile or availability guarantees, so dependency_availability scores 1. The files do not describe failure messages or error handling, so failure_messages scores 0.

Adaptability9 / 18 · 2.5/5

The README describes conversation memory, document research, and cross-tool tasks, and gives examples for datasets, memory providers, modes, and query limits; audience_and_scenarios scores 2. It distinguishes ordinary answer evaluation from PrecisionMemBench retrieval mode, but does not systematically define capability boundaries or triggering rules, so capability_boundaries and trigger_precision score 1 each. Python, environment variables, the MemBench data path, and command line steps are documented, so environment_fit scores 2.

Convention6 / 18 · 1.7/5

The README is organized into overview, problem, evaluation flow, setup, usage, results, and requirements, and includes command examples. Setup and naming guidance are clear enough for 2 in install_notes, naming_stability, and examples_and_faq. It gives only limited discussion of benchmark shape and cost rather than a complete limitations list, so known_limitations scores 1. The supplied files contain no license, changelog, or maintenance owner evidence, so license, versioning_changelog, and maintenance_responsibility score 0. Information architecture is assessable only from the README, so it scores 1.

Effectiveness7 / 13 · 2.7/5

The tool offers dataset and provider listings, statistics, benchmark runs, result files, and browser viewing, making outputs reasonably usable; output_usability scores 2. The README gives a concrete motivation about benchmark discrimination with million-token context windows and identifies accuracy, speed, and token cost as measures, so marginal_value scores 2. It emphasizes cost measurement but supplies no concrete estimate of run costs or resource needs, so cost_benefit scores 1.

Verifiability4 / 8 · 2.5/5

The README names the evaluation flow, modes, result path, and materials it says are needed for reproduction, and identifies the PrecisionMemBench upstream project; this makes some claims traceable, so claim_traceability scores 2. The available evidence is limited to README and project configuration. The configuration corroborates the entry point, dependencies, and Python requirement but cannot independently establish results, dataset contents, or reproducibility claims, so cross_source_corroboration scores 1. The README separates the evaluation steps from PrecisionMemBench retrieval scoring and presents performance statements as project motivation; however, the claims cannot be independently checked from the supplied material, so fact_inference_separation scores 2 rather than full marks.

Risks and how to mitigate them
  • Not found in source: confirmation before actingTurn on (or add) a confirmation step before it acts, and try it in a sandbox or test environment before real data.
  • Not found in source: rollback or recovery pathBack up first, or work on a git branch or snapshot, so its changes can be undone.
  • The supplied files do not show a license, changelog, maintenance owner, dependency lock, security review, or error handling guidance; do not infer that these are covered.
  • The README’s claims about benchmark advantages and reproducibility are not corroborated by results or independent sources in the supplied material.
Evidence confidence: Low Reviewed Oct 09, 2026 Reviewed revision d97d4960dd9a
Review evidence README.mdpyproject.toml
See the full review method →

FAQ

Which API key is needed to run evaluations?
The standard flow requires GEMINI_API_KEY; Gemini is used to generate and judge answers.
Do all datasets require answer generation and judging calls?
No. PrecisionMemBench with --mode retrieval skips the answer LLM and judge, scoring whether belief-ID sets satisfy the assertions.
Can I run MemBench without preparing data?
No. You need a local data directory and must set MEMBENCH_DATA_PATH to it.
Where are results saved?
Results are written to outputs/{dataset}/{memory}/{mode}/{domain}.json and can be browsed with uv run amb view.
View on GitHub ↗ Install ↓

Related agents