Comma Personal Agent
Keep long-running work moving across devices without managing the right chat session.
- Source repo
- AFK-surf/Comma
- Stars
- ★ 232
- Last updated
- 7d ago
- License
- AGPL-3.0
- Primary language
- Elixir
- FA score
- 65/100 · Some gaps
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code
- Cost
- Free tier plus a paid hosted plan
- Setup effort
- Medium · a few setup steps
- You'll need
- Typical use
- People with unfinished work across devices can use persistent Tasks to carry context forward without choosing a particular chat session.
- Not a fit if
- People or teams that need fully offline operation
- Users unwilling to configure Docker and a model provider
- Source review
- 65/100 · Some gaps
What does this agent do, and when should you use it?
Comma is an open-source personal agent from AFK Inc. that organizes work, memory, and context around persistent Tasks. People can use Comma Web or the Mac app, or run a self-hosted service; larger requests become Tasks handled by independent agent loops that plan, execute, review, and verify. The backend is primarily built with Elixir/OTP, critical system components use Lean, and TLA+ is used to model-check core distributed protocols. Salix can connect a Mac so Comma can read files; with Allow operations enabled, it can also change files, run commands, and hand work to Codex, Claude Code, Pi, or Kimi. The self-hosted Compose stack includes the Web, Admin, API, database, cache, and object storage services, while real model calls require a provider account.
A user can make a request in chat: quick questions receive chat answers, while larger work becomes a Task tracked on a board with Backlog, In progress, Needs Review, and Done. Comma assigns each Task to an independent agent loop that plans, executes, checks, and verifies work; Tasks wait in Needs Review when the user must decide. Users can schedule briefings with Routines or use Loops to watch an inbox, repository, or feed and wake the agent when attention is needed. After connecting a Mac, Salix can read local files; with Allow operations enabled, Comma can also change files, run commands, and hand work to Codex, Claude Code, Pi, or Kimi. The Mac app can record a call, save it to Drive, and start a summary Task.
- People with unfinished work across devices can use persistent Tasks to carry context forward without choosing a particular chat session.
- Users with larger goals can let agent loops break work into ongoing planning, execution, and review.
- People who need to watch an inbox, repository, or feed can use Loops to wake the agent when something needs attention.
- Users who want an assistant to inspect files or run commands on their Mac can connect Salix and enable Allow operations as needed.
- Developers already using Codex, Claude Code, Pi, or Kimi can hand work to those tools through Comma.
- Individuals or teams who want to run their own service can deploy the Web, Admin, and backend stack with Docker Compose.
How do you install or deploy this agent?
A self-hosted instance requires Docker Engine with Compose v2.24 or later; the first build downloads toolchains including Lean, Elixir, Go, and Node. Copy the sample configuration and set COMMA_OWNER_EMAIL to the owner email; you can provide model-provider settings in the configuration or set up a private model in Settings after login. Then start the stack:
cp .env.example .envSet COMMA_OWNER_EMAIL in .env and, if desired, the model settings, then run:
docker compose up -d --buildOpen http://localhost:8080 to request a login code and read it in the local inbox at http://localhost:8025. Real model calls require a model-provider account; self-hosting does not require a Comma account or private repository token.
How do you use this agent?
After startup, open Comma Web:
http://localhost:8080Request a login code with the owner email; the local inbox used by the default stack is:
http://localhost:8025After logging in, submit requests, track Tasks, configure a model in Settings, or connect a Mac in device settings. The self-hosted model configuration supports COMMA_LLM_BASE_URL, COMMA_LLM_PROTOCOL, COMMA_LLM_MODEL, COMMA_LLM_API_KEY, COMMA_LLM_CONTEXT_TOKENS, and COMMA_LLM_MAX_TOKENS. To let Comma operate on the local computer, connect Salix and enable Allow operations.
What are this agent's strengths and limitations?
- Tasks persist across sessions and devices, and independent agent loops keep working on them.
- The complete single-node stack can be self-hosted with Docker Compose without a Comma account or private repository token.
- Users can configure a model provider and hand work to Codex, Claude Code, Pi, or Kimi.
- Elixir/OTP, formal properties in Lean, and TLA+ protocol models provide explicit design evidence for failure handling and runtime behavior.
- Self-hosting requires building multiple services and downloading Lean, Elixir, Go, and Node toolchains, so setup involves several steps.
- Real model calls require a provider account and may incur that provider's charges; the default stack does not include a mock model.
- Local file and command operations require connecting Salix and enabling Allow operations.
- The runtimes needed for meetings and Agent VMM hosts are not included in this repository; group cloud computers also require Cloudflare configuration.
How does this agent compare with similar options?
The README describes Codex, Claude Code, and other agents as tools Comma can invoke or register as Workers/Sub-agents. It also positions Comma as an open-source alternative to Muse, Dots, Instinct, and Town.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Comma Personal Agent This agent | 65 · Some gaps | Self-hosted serviceFreemium | ★ 232 | 7d ago | Elixir | Codex · Claude Code |
| Rakazo AI Teammates | 63 · Some gaps | Self-hosted serviceFree + model costs | ★ 3.6k | today | TypeScript | — |
| GAIA — Personal AI Assistant | 73 · Some gaps | Hosted serviceFreemium | ★ 308 | today | Python | — |
| Commonly | 44 · Major gaps | Self-hosted serviceFree + model costs | ★ 1.4k | today | TypeScript | Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
The README describes separation of device identity from cloud work, an opt-in computer operations setting, Needs Review decisions, and no automatic rerun of interrupted side effects. The Admin test excerpts also show confirmation fields and checks for sensitive-field projections. Least privilege scores 1 because the product is described as always on and able to read files, run commands, and operate applications, while the supplied material does not establish default permissions or per-action authorization scope. Data-flow transparency and sensitive-data handling score 2: the docs explain model-provider charges, device/cloud responsibilities, and that the config volume holds an encryption key, but do not provide a complete lifecycle or retention account. Dependency security scores 1: package versions are pinned, but no vulnerability audit or mitigation evidence is supplied. External effects and rollback score 2: review gates, an operations toggle, and backup/restore guidance are documented, but recovery paths for every mutating operation are not shown. Source attribution scores 1: AFK Inc. and security/contribution contacts are named, but the prompt says publisher identity is unverified and the supplied material does not establish a verifiable maintainer identity.
Scores 2: the README's descriptions of task state, process restarts, persistence of accepted messages, and interrupted side effects are mutually consistent, and it points to Lean/TLA+ verification material; test excerpts provide partial support for Admin API contracts. Dependency availability scores 2: self-hosting docs list Docker Compose, public toolchains, and an external model account, but the supplied material does not include a complete dependency inventory or alternatives. Failure messages score 2: docs describe some failure semantics, and tests assert specific errors for missing build configuration and direct API requests; user-facing error coverage is not shown.
Scores 2: the product targets continuous personal tasks across computers, phones, cloud environments, and cooperating agents, but the supplied setup chiefly covers single-node Docker and Mac devices. Capability boundaries score 2: docs say meetings and Agent VMM require runtimes outside the repository and group cloud computers need Cloudflare configuration; they also distinguish drafts from autonomous work. Per-integration permission boundaries are not fully described. Trigger precision scores 2: Loops wake the agent only when a watched target needs attention, and Routines can brief at a chosen time; trigger configuration and false-wakeup controls are not shown. Environment fit scores 2: local self-hosting, development, Electron, and public deployment are documented, though deployment involves multiple services and an external model provider.
Scores 3 for clear repository layout, documentation index, architecture entry point, and install guidance covering requirements, configuration, startup, access, public deployment, upgrades, and backups. Naming stability scores 2: Tasks, Loops, and Needs Review are explained, but the supplied evidence is insufficient to assess consistency across code and docs. Examples and FAQ score 2: commands, configuration variables, and a development sign-in example are provided, but no FAQ is shown. Known limitations score 2: external runtime and cloud configuration requirements are stated, though coverage is limited. License scores 3: README and LICENSE identify AGPL-3.0-only and note that third-party components retain their own licenses. Versioning/changelog scores 1 because no release policy or changelog is included in the supplied material. Maintenance responsibility scores 2: AFK Inc., a security contact, and a CLA are named, but publisher identity and ongoing maintenance are not verified.
Scores 2: the task board, persistent execution, review decisions, scheduled briefings, and device operations describe usable work paths, but the supplied material contains no independent run results establishing task quality. Marginal value scores 2: cross-session task state, device collaboration, and agent orchestration offer a distinct product proposition, while some capabilities depend on external accounts or runtimes. Cost-benefit scores 2: docs disclose model-provider charges and say self-hosting needs no Comma credits or Stripe, but do not quantify resource costs or ongoing operational burden.
Scores 2: the README points to specific repository paths for formal verification, deployment, development, and architecture, and test excerpts show selected API and boundary assertions; the linked verification docs and full implementation are not part of the supplied material. Cross-source corroboration scores 2: README claims about confirmation, secret handling, and API boundaries have partial counterparts in test excerpts, but those excerpts are limited and cannot establish product-wide behavior. Fact/inference separation scores 2: the README explicitly says formal verification does not prove model judgment or task success, appropriately limiting that claim; other performance and product claims remain primarily project assertions.
- The product can run continuously, read files, and execute commands. Before enabling computer operations, review the actual permission scope, external-service data handling, and available recovery procedures.
- Self-hosting requires a model provider and several persistent volumes. Back them up together as documented, and preserve the config volume containing encryption and authentication secrets.