Statewave Agent Memory Runtime

A self-hosted memory runtime that gives AI agents reproducible, provenance-tagged context instead of noisy query-time retrieval.

Stars
★ 271
Last updated
2d ago
License
Apache-2.0
Primary language
Python

At a glance

How it runs
Self-hosted serviceLibrary / SDKCLI
Works with
Universal · cross-platformOpenAI API · Claude APIClaude Code (Partial support)
Cost
Free software; you pay for model usage
Setup effort
Medium · a few setup steps
You'll need
PostgreSQL 14+ with pgvector ≥ 0.4.2DockerPython 3.11+Node.js (for npx install and TypeScript connectors)Shell / CLINetwork accessLocal filesystemMCP Server
Typical use
Support teams that need returning customers recognised across sessions, ranked context inside a token budget, and a handoff pack via POST /v1/handoff when a case escalates.
Not a fit if
  • Multi-replica teams unwilling to switch to Postgres-backed rate limiting
  • Teams that want built-in auth that issues API keys and admin identities
  • Teams doing mostly single-hop lookups and optimising hard for token cost
Source review
62/100 · Some gaps

What does this agent do, and when should you use it?

Statewave is an open-source memory runtime for AI agents, self-hosted on Postgres + pgvector, built around a compile-then-use model rather than per-query similarity search. It records append-only episodes per subject (user:, repo:, account:, or any prefix), compiles them into typed memories with confidence scores and provenance using either a fully local heuristic compiler or an LLM compiler, then assembles ranked, token-bounded context bundles for a given task. Every bundle comes with an immutable state-assembly receipt that traces back to source episodes, supports HMAC-SHA256 signing and POST /v1/receipts/{id}/replay. The server runs on port 8100 behind a REST API with a full OpenAPI surface, and ships Python (pip install statewave) and TypeScript (npm install @statewavedev/sdk) SDKs. It boots in demo mode with stub embeddings and the heuristic compiler, so no LLM key or GPU is needed to start; enabling semantic search and stronger extraction means pointing LiteLLM at any of 100+ providers.

The runtime is organised around subjects, and the end-to-end loop is ingest → compile → use. Ingestion uses POST /v1/episodes or POST /v1/episodes/batch (up to 100 at a time) to append raw events. Compilation calls POST /v1/memories/compile, which extracts typed, summarised memories with confidence scores and provenance, and is idempotent so recompiling produces no duplicates. Retrieval calls POST /v1/context to assemble a ranked, token-bounded bundle for a named task; the same query against the same subject at the same point in time yields the same bytes. Governance endpoints include GET /v1/timeline, GET /v1/subjects, DELETE /v1/subjects/{id} for subject-level deletion, POST /v1/handoff for a compact handoff pack, plus GET /v1/subjects/{id}/health for an explainable health score and /sla for response and resolution metrics. Search is served by GET /v1/memories/search across kind, text or semantic similarity, and POST /v1/resolutions tracks issue state per session. Compliance features include a declarative YAML policy engine (deny / redact) over memory tags such as pii, financial and secret, v0.9 heuristic auto-labeling with suggested_labels and operator promotion, multi-tenant isolation via the X-Tenant-ID header, and per-tenant region pinning that returns HTTP 403 residency.mismatch to processes in the wrong region.

  1. Support teams that need returning customers recognised across sessions, ranked context inside a token budget, and a handoff pack via POST /v1/handoff when a case escalates.
  2. Long-running coding agents that persist project memory — tech stack, preferences and architecture decisions — across separate conversations under a repo: subject.
  3. Platform teams that want to self-host memory on their own Postgres and wire any language or framework to it over the REST API, keeping data inside their network.
  4. Compliance and risk teams that need deny/redact policies over pii, financial and secret tags plus signed receipts showing which memories influenced each assembled bundle.
  5. Organisations with data-residency rules that pin tenants to a region so requests hitting the wrong region are refused outright.
  6. Teams building memory from non-chat sources by syncing GitHub, Slack, Notion, Zendesk or Gmail events into episodes with the @statewavedev/connectors-* packages, dry-run first.

How do you install or deploy this agent?

The fastest path is the official install script or npx; you can also run it yourself with Docker Compose.

# macOS / Linux
npx @statewavedev/statewave
# or
curl -fsSL https://www.statewave.ai/install | sh
# Windows (PowerShell)
irm https://www.statewave.ai/install.ps1 | iex

Self-managed install:

git clone https://github.com/smaramwbc/statewave && cd statewave
docker compose up -d

This brings up Postgres (with pgvector) plus the API; migrations run on container start and the API is served at http://localhost:8100. The default is demo mode (stub embeddings + heuristic compiler). For LLM-backed behaviour, place a .env next to docker-compose.yml and re-run docker compose up -d:

STATEWAVE_EMBEDDING_PROVIDER=litellm
STATEWAVE_LITELLM_API_KEY=sk-...            # any LiteLLM provider
STATEWAVE_LITELLM_MODEL=gpt-4o-mini
STATEWAVE_LITELLM_EMBEDDING_MODEL=text-embedding-3-small

Requirements: PostgreSQL 14+ with pgvector ≥ 0.4.2; SDKs need Python 3.11+ or Node.js.

How do you use this agent?

Check readiness, then run the ingest → compile → use loop from the SDK or over HTTP.

curl http://localhost:8100/readyz
curl http://localhost:8100/healthz
from statewave import StatewaveClient

with StatewaveClient("http://localhost:8100") as sw:
    sw.create_episode(subject_id="user-42", source="chat", type="message",
                      payload={"text": "Alice asked about pricing tiers"})
    sw.compile_memories("user-42")
    print(sw.get_context("user-42", task="answer pricing", max_tokens=1000).assembled_context)

Interactive API docs live at http://localhost:8100/docs (Swagger) and /redoc. Key endpoints: POST /v1/episodes, POST /v1/memories/compile, POST /v1/context, GET /v1/timeline, GET /v1/subjects, DELETE /v1/subjects/{id}. Connector example (dry-run by default, nothing is ingested):

statewave-connectors sync github \
  --repo smaramwbc/statewave \
  --subject repo:smaramwbc/statewave \
  --dry-run

What are this agent's strengths and limitations?

Pros
  • Deterministic context: the same subject, task and point in time always yield the same bundle bytes, backed by state-assembly receipts with HMAC-SHA256 signing and a replay endpoint
  • Compile-once design instead of query-time retrieval, producing typed memories with confidence scores and provenance and avoiding sampling noise
  • Fully self-hosted on Postgres + pgvector; the default heuristic compiler and stub embeddings run locally and the API process is CPU-only, so no GPU is required
  • Provider-neutral via LiteLLM (100+ model and embedding providers) with both Python and TypeScript SDKs covering sync and async usage
  • Built-in sensitivity labels with a YAML deny/redact policy engine, multi-tenant isolation, and tenant region pinning for residency-sensitive deployments
Limitations
  • You must run and operate PostgreSQL 14+ with pgvector ≥ 0.4.2 yourself; there is no vendor cloud as the primary path
  • Auth is limited to validating API keys you configure — it does not issue them, admin action identity (promoted_by) is still null, and multi-tenancy is app-layer without Postgres RLS
  • Rate limiting defaults to per-process and per-IP; multi-replica deployments must switch to STATEWAVE_RATE_LIMIT_STRATEGY=distributed and there is no per-tenant or per-API-key limiting yet
  • v1.5.0 is actively developed: receipt replay is current code plus original policy rather than byte-for-byte history, and federated cross-region audit plus a visual policy editor are still pending
  • Compiled bundles are denser than fact-store retrieval, so per-answer token cost is higher and a lighter fact store may be the better call for mostly single-hop queries

How does this agent compare with similar options?

The README contrasts Statewave with two alternatives: memory layers that store isolated facts and retrieve them per query (or extract a graph whose retrieval surface can return the same summary regardless of the question), and bolting a vector database or raw chat logs onto a prompt. It also states explicitly that it is not a chatbot framework, vector database, RAG pipeline, or hosted service.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Statewave Agent Memory Runtime This agent 62 · Some gaps Self-hosted serviceFree + model costs ★ 271 2d ago Python OpenAI API · Claude API
DuraGraph 45 · Major gaps CLIFree ★ 162 29d ago Go —
Brigade — Enterprise-grade personal intelligence 59 · Major gaps CLIFree + model costs ★ 11k 2d ago TypeScript ChatGPT · Codex · Claude Code · OpenAI API · Claude API
IntentKit 35 · Major gaps Self-hosted serviceFree + model costs ★ 6.5k 25d ago Python —

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Some gaps
62/ 100 5-point scale 3.1 / 5
Trust 16/29
Reliability 9/14
Adaptability 12/18
Convention 13/18
Effectiveness 9/13
Verifiability 3/8
Why each dimension lost points
Trust16 / 29 · 2.8/5

Evidence shows CI workflow declares permissions: contents: read (least privilege), SECURITY.md states least privilege and audit logging, and config exposes STATEWAVE_API_KEY, tenant isolation, region pinning, sensitivity labels and policy engine (deny/redact). However, default STATEWAVE_API_KEY empty means open access, CORS defaults to ["*"], no built-in auth issuance, no Postgres RLS, and README's curl|sh and irm|iex installs are external effects without confirmation. Rollback is limited to DELETE /v1/subjects/{id} and idempotent recompile, with no transaction-level rollback described. Hence least_privilege 2, user_confirmation/external_effects/rollback deducted to 1.

Reliability9 / 14 · 3.2/5

pyproject.toml pins dependencies with bounds and explains explicit numpy/httpx declarations, CI uses uv lock --check to prevent drift, docker-smoke verifies image boots, and /readyz gives specific detail text when LLM key is missing, making failures diagnosable. But no runtime dependency availability fallback details (pgvector version, LiteLLM provider) are given, so all three reliability criteria are 2.

Adaptability12 / 18 · 3.3/5

README clearly lists applicable scenarios (support, coding agent, A/B) and non-goals (not a chatbot framework, vector DB, RAG pipeline, or hosted service), so capability_boundaries is 3. Trigger precision is weak: when the Agent calls compile/context is left to the caller, with no trigger conditions or thresholds defined, so trigger_precision is 1. Environment fit covers Python 3.11-3.13, multi-arch Docker, Postgres 14+, but macOS/Windows are not CI-tested, so environment_fit is 2.

Convention13 / 18 · 3.6/5

README is well-structured with FAQ, install notes, config table, API table, platform support table; known_limitations honestly lists 7 items; license is full Apache-2.0 text consistent with pyproject, so license 3 and known_limitations 3. But changelog/roadmap are hosted externally and maintenance responsibility is not stated in-repo (no CODEOWNERS/MAINTAINERS), so maintenance_responsibility is 1; information architecture, naming stability, examples and FAQ are 2.

Effectiveness9 / 13 · 3.5/5

Output is a token-bounded, provenance-tagged context bundle directly usable in prompts, so output_usability is 2. Marginal value lies in deterministic compilation and provenance, but README admits token cost is higher than a plain fact store and may not pay off for single-hop queries, so marginal_value and cost_benefit are 2.

Verifiability3 / 8 · 1.9/5

README heavily references external repos (statewave-docs, statewave-examples, statewave-py/ts) and benchmark claims (56 assertions, multi-hop accuracy) with no in-repo files to cross-check, so claim_traceability and cross_source_corroboration are 1. Fact/inference separation is weak: marketing claims ("higher multi-hop accuracy") are interleaved with verifiable facts without evidence grading, so fact_inference_separation is 1.

Risks and how to mitigate them
  • Default STATEWAVE_API_KEY empty means open access and CORS defaults to ["*"], so authentication and origin restrictions must be explicitly configured before production.
  • README's curl|sh and irm|iex install scripts execute remote code without confirmation or verification steps; prefer Docker/Helm or review the script first.
  • Multi-tenant isolation is app-layer only with no Postgres RLS, so cross-tenant isolation strength depends on application code correctness.
  • Benchmark and accuracy claims in README cannot be verified within this repository and reference external repos/docs; independent verification is required.
  • Maintenance responsibility and update path are not stated in-repo (no CODEOWNERS/MAINTAINERS); changelog/roadmap are hosted externally.
  • Publisher identity is not verified by FollowAgents and should be treated as unknown; do not infer safety or reliability from it.
Evidence confidence: Low Reviewed Oct 11, 2026 Reviewed revision fc4cf21e5340
See the full review method →

FAQ

Can I run it without paying for an LLM API?
Yes. It boots in demo mode with stub hash-based embeddings and the heuristic compiler, so ingest → compile → use works with no API key. A key only becomes necessary when you want real semantic search or LLM-backed memory extraction, and you pay the provider you choose per use.
Do I need a GPU or Kubernetes?
No GPU. The API process is CPU-only; GPUs only matter if you self-host an LLM compiler or embedding model. Deployment is possible via Docker Compose, Helm or bare metal, and multi-replica operation has been verified with Fly multi-machine and Helm HPA.
Why does a context bundle cost more tokens than a plain fact store?
Compiled bundles are deliberately denser, which is what buys the higher multi-hop accuracy in the project's benchmarks. If your queries are mostly single-hop and you are cost-sensitive, the README says a lighter fact store may be the right call.
What breaks at multi-replica scale?
Rate limiting defaults to the in-process memory strategy and is keyed only by IP, so multi-replica deployments must set STATEWAVE_RATE_LIMIT_STRATEGY=distributed. Per-tenant and per-API-key limits are not available yet. Load is ultimately bounded by Postgres and your embedding provider.
Can I use it commercially in a closed-source product?
Yes. The server and SDKs are Apache-2.0 with an explicit patent grant, so proprietary, hosted and commercial use carries no source-disclosure obligation. Enterprise procurement, SLA and indemnity questions go to [email protected].
View on GitHub ↗ Install ↓

Related agents