Productivity & Collaboration self-hosted-automationmulti-agent-orchestrationgovernancehitl-approvalworkflow-automationoffice-automationgraphragmcp-client

Atom Platform

An open-source, self-hosted governed AI agent platform where autonomy is earned through verified outcomes — running safely on your own infrastructure.

FollowAgents review · FARS-2.1
Not recommended
59/ 100 5-point scale 3.0 / 5
1 2 3 4 5 6
1Trust11 / 29 · 1.9/5

Trust: README and SECURITY.md extensively declare sandboxing (filesystem scope, tool whitelists, kill-run, egress allowlist), 4-tier maturity gates, HITL approval, Fernet encryption at rest, and outbound approval gates — but these are documentation claims only; the source files provided contain no implementation code to confirm they actually function, so least_privilege, user_confirmation, data_flow_transparency, sensitive_data_handling, and external_effects score 1 (asserted without support in this review's evidence). Rollback is a kubectl rollout undo doc plus one phrase ('validated state machine with rollback') → 1. Source attribution: full AGPL-3.0 text genuinely provided in LICENSE.md → 2; unverified publisher neither adds nor deducts. No signs of malware, credential theft, or covert exfiltration.

2Reliability9 / 14 · 3.2/5

Reliability: README claims align with ci.yml's actual jobs (boot checks, quality gates, quickstart end-to-end verification); timeout seatbelts and log dumps show mature failure handling → self_consistency 2. Multi-provider routing plus local-model fallback reduces external dependency risk → dependency_availability 2. CI comments record specific historical incidents (2026-09-05 6-hour hang) and fixes → failure_messages 2. However, CI actually runs ~30 named test files while marketing claims '85k+ tests' — a magnitude gap that weighs on verifiability.

3Adaptability12 / 18 · 3.3/5

Adaptability: README gives concrete scenarios and integration matrices for six departments and positions the product against three competitor classes → audience_and_scenarios 3. Capability boundaries (maturity tiers, 'ask fails safe', auto-revocation) are well documented but uncorroborated by code → 2. Trigger precision stays at natural-language examples with no trigger-spec or false-trigger documentation → 1. Environment fit covers Docker, DigitalOcean 1-Click, local Ollama, and CI matrix → 2.

4Convention14 / 18 · 3.9/5

Convention: docs index, user guide, architecture deep-dives, and env reference are well organized → information_architecture 3. Install notes are end-to-end verified by the CI quickstart job, with expected results and default credentials documented → install_notes 3. ATOM_-prefixed env vars are consistent with a reference doc → naming_stability 2. Rich examples (use-case tables, playbooks) but no FAQ → examples_and_faq 2. Honest disclosure of non-gating UI tests and non-blocking typecheck → known_limitations 2. Full AGPL-3.0 text provided → license 3. No CHANGELOG or versioned releases found; SECURITY.md admits only latest main is supported → versioning_changelog 1. SECURITY.md disclosure process and CI maintenance evidence → maintenance_responsibility 2; single-maintainer repo with no release cadence evidence does not merit full marks.

5Effectiveness9 / 13 · 3.5/5

Effectiveness: outputs carry citations, two-tier confidence labels, audit trail, and approval gates → output_usability 2. Differentiated positioning (governed agent platform, postcondition oracle) is clear in a crowded market → marginal_value 2. Cost side is concretely addressed (BYOK, 16+ providers with cost-aware routing, local models, OpenCode Go savings path) → cost_benefit 2. All benefit claims are unverified by execution, capping further credit.

6Verifiability4 / 8 · 2.5/5

Verifiability: README's numeric claims (0.027ms P99, 616k ops/s, 85k+ tests, ~1,100 fixes) link to benchmarks and RESEARCH_NOTES.md, but this static review cannot confirm them, and marketing tone is heavy (the '88% of pilots' figure cites a 2026 industry stat that cannot be corroborated here) → claim_traceability 1. ci.yml genuinely corroborates README's test and quickstart claims (boot verification, quality gates, artifact checks exist) → cross_source_corroboration 2. The README footnote explicitly separates industry statistics from self-claims ('Atom makes no claims about its own deployments') → fact_inference_separation 2.

Evidence confidence: Low Reviewed Sep 10, 2026 Reviewed revision ef55b56a0af7
Before you use it
  • Trust and safety mechanisms (sandbox, HITL, encryption, postcondition oracle) are README/SECURITY.md claims only; no implementation code was provided in this static review. Audit the corresponding source and run adversarial validation before enterprise deployment.
  • CI actually runs ~30 named test files versus the '85k+ tests' marketing claim — a magnitude gap; the 25% coverage gate is modest. Discount test claims accordingly.
  • Performance figures (0.027ms P99, 616k ops/s) are self-reported repo benchmarks, independently unverified.
  • Only latest main is supported, with no versioned releases or CHANGELOG — no stable anchor for upgrades or rollback.
  • AGPL-3.0 imposes open-source obligations for network service use; check compliance boundaries for managed deployments. Publisher identity is unverified by the registry.
  • Frontend typecheck is non-blocking and UI tests are non-gating; frontend quality assurance is weaker than backend.
Review evidence [1][2][3][4][5][6][7][8][9]
See the full review method →

What does this agent do, and when should you use it?

Atom (GitHub repository rush86999/atom) is an open-source (AGPL-3.0), self-hosted AI agent workforce platform positioned as a delegable team of specialty agents rather than a code-first framework. The backend is FastAPI, the frontend is a Next.js web UI, with React Native mobile and Tauri macOS menubar companion apps. Its core differentiator is governance: agents progress through a 4-tier maturity model (STUDENT → INTERN → SUPERVISED → AUTONOMOUS) only after an independent postcondition oracle verifies their runs, rather than trusting self-reports. Execution is constrained by a default-on sandbox layer covering filesystem scope, tool whitelists, resource caps, KillRun, and an egress allowlist. The platform supports 16+ LLM providers (including local models via Ollama), ships 46+ business integrations (Salesforce, HubSpot, Slack, QuickBooks, Gmail, Notion, Shopify, Zoom, and more), and keeps all workflow data and agent state on your own infrastructure.

You describe an outcome in plain language (e.g., "extract invoice data from Gmail PDFs and match against QuickBooks"), and Atom generates a governed, replayable workflow with Human-in-the-Loop approval gates. The Queen Agent handles structured workflows, the Fleet Admiral handles open-ended tasks, and the Conductor provides 5 execution strategies backed by a validated state machine with rollback. Every mutating action is re-derived against your system of record by an independent postcondition oracle, with two-tier confidence (self-reported vs externally verified). Capabilities include Office automation (Excel/Word/PPTX editing with formula evaluation and live Canvas co-editing), GraphRAG multi-hop retrieval, hybrid search (BM25 fused with LanceDB vectors via RRF), a Knowledge VFS, peer-to-peer agent messaging (Agent Radio), an OpenAI/Anthropic-compatible LLM Gateway, an MCP client, and an A2A agent-protocol endpoint.

  1. Sales teams: a new HubSpot lead triggers automated company research, scoring, an Asana task for the rep, and a Slack notification
  2. Finance staff: PDF invoices arriving in Gmail are OCR-extracted, matched against QuickBooks, and discrepancies flagged with alerts
  3. Support teams: Zendesk tickets are monitored for sentiment, urgent ones auto-escalated, and reply drafts sent only after human approval
  4. HR departments: new BambooHR hires get accounts provisioned, Slack invites sent, and orientation scheduled automatically
  5. Engineering teams: GitHub PRs trigger tests and security scans, auto-merge on green, with summaries posted to Slack/Jira
  6. Solo operators: import pre-built personal starters (invoice chasing, candidate pipeline, support triage) where nothing sends without explicit approval

What are this agent's strengths and limitations?

Pros
  • Governance as a core differentiator: 4-tier maturity model, independent postcondition oracle (a refuted "done" self-report is stamped UNVERIFIED), and comprehensive audit trail
  • Default-on execution sandbox: filesystem scope, tool whitelist, resource caps, KillRun, egress allowlist — mini-apps run in Firecracker microVMs
  • Genuine self-hosting and data sovereignty: embedded store with no cloud requirement, BYOK keys encrypted at rest, fully private deployments via Ollama/LM Studio/vLLM
  • 46+ business integrations (Salesforce, HubSpot, QuickBooks, Zendesk, Shopify, etc.) and 16+ LLM providers with cost-aware routing and fallback
  • The free AGPL-3.0 edition is full-featured — no closed-source "pro" tier, keys never gated by plans
Limitations
  • Self-hosting means you operate the Python 3.11+ backend, Next.js frontend, and infrastructure yourself — a real maintenance burden
  • AGPL-3.0 carries copyleft obligations for commercial integration and redistribution; legal review is warranted before enterprise adoption
  • Requires at least one LLM provider key (or a local model server), creating ongoing inference costs and provider setup work
  • Quantified claims (85k+ tests, 0.027ms P99, 616k ops/s) come from the repository's own benchmarks with no third-party audit cited
  • The feature surface is extremely broad (GraphRAG, mini-apps, 46+ integrations); real-world stability likely varies by integration, so validate your target stack on a small pilot first

How do you install or deploy this agent?

Clone and bootstrap in one shot: git clone https://github.com/rush86999/atom.git && cd atom, then run make setup (creates venv, installs dependencies, generates .env and frontend config). Run make backend in one terminal (backend on :8001) and make frontend in another (Next.js UI on :3001). Alternatively deploy via Docker (docs/operations/personal-edition.md) or DigitalOcean 1-Click (deploy/digitalocean/app.yaml). Python 3.11+ is required.

How do you use this agent?

After startup, set at least one LLM key in backend/.env: OPENCODE_API_KEY (low-cost subscription models) or OPENAI_API_KEY / ANTHROPIC_API_KEY / DEEPSEEK_API_KEY / GOOGLE_API_KEY; for fully local operation set ATOM_LOCAL_ONLY=true and OLLAMA_BASE_URL=http://localhost:11434/v1. Open http://localhost:3001, sign in as [email protected] using the password in bootstrap_admin_password.txt, and describe a workflow in plain language — Atom builds it with an approval gate before anything executes. OAuth integration tokens are encrypted at rest with Fernet; production fails closed if the encryption key is missing.

How does this agent compare with similar options?

The README draws explicit comparisons: versus Zapier/Make/n8n, Atom uses reasoning, self-correcting agents rather than fixed steps, with built-in governance and approval gates; versus LangGraph/CrewAI/AutoGen, Atom is a turnkey self-hosted product rather than a code-first framework; versus single-agent personal assistants like Hermes Agent and OpenClaw, Atom offers multi-agent teams, governed sandboxes, and 46+ business integrations. A per-competitor feature matrix (outcome verification, default-on sandbox, Office/Canvas nativeness) is included in the repository.

FAQ

What does it cost?
The software is fully free and open source under AGPL-3.0 with no feature paywalls. Costs come from LLM inference: use your own keys (BYOK), an OpenCode Go subscription (~$10/month, README claims ~90% savings vs pay-per-token), or local models via Ollama for zero inference cost.
Does my data leave my server?
Workflow data, agent state, and memory live in a local embedded store with no cloud requirement. Inference runs through your own API keys (encrypted at rest) or local models (Ollama, LM Studio, vLLM, llama.cpp server), enabling fully private deployments.
Could an agent take dangerous actions without approval?
Not by default. Mutating actions require HITL approval, execution is bounded by the sandbox layer (filesystem scope, tool whitelist, resource caps, KillRun, egress allowlist), and interruption-free privileges per action type are only earned from verified track records and auto-revoked on regression.
How do I know an agent actually finished its task?
An independent postcondition oracle, on by default, re-derives every mutating action's outcome against your system of record; refuted self-reports are stamped UNVERIFIED. It can be disabled via the ATOM_ORACLE_ENFORCE kill switch.
What team size is it suitable for?
Solo operators can start with pre-built personal starters that include approval gates; enterprises get OIDC SSO, SCIM v2 provisioning, and 8-role RBAC. The README provides no measured performance-at-scale or production deployment data, so pilot on a single workflow first.

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents