Octomind

A multi-model, MCP-native coding-agent runtime for terminals and unattended automation.

Source repo
Muvon/octomind
Stars
★ 147
Last updated
1d ago
License
Apache-2.0
Primary language
Rust

At a glance

How it runs
CLISelf-hosted service
Works with
Universal · cross-platformOpenAI API · Claude API
Cost
Free tier plus a paid hosted plan
Setup effort
Low · running in minutes
You'll need
macOS, Linux, or WindowsOctomind Cloud login or a supported provider API keyRust 1.95+ for Cargo installation or source buildsShell / CLINetwork accessLocal filesystemMCP Server
Typical use
A developer runs developer:general in a local repository to inspect code, execute tests, and make edits.
Not a fit if
  • Users who require a graphical desktop interface
  • Teams unwilling to let an agent run commands or access project files
  • Environments requiring every tool dependency inside one binary
Source review
70/100 · Some gaps

What does this agent do, and when should you use it?

Octomind is an open-source, CLI-first AI agent client and runtime distributed as a single Rust binary. Its session engine is exposed through an interactive terminal, stdin pipelines, a long-running daemon, a WebSocket server, and an ACP sub-agent interface. Models use MCP tools to inspect and edit files, execute shell commands, search code, and delegate work to other agents. Roles, model choices, tool permissions, workflows, budgets, and guardrails are configured in TOML rather than framework code. Users can obtain model access through Octomind Cloud or bring credentials for providers such as OpenRouter, OpenAI, Anthropic, DeepSeek, and Ollama. The principal deployment boundary is a locally or self-hosted process, although external models and MCP tools may introduce network, credential, and dependency requirements.

A user starts a role or tap specialist with octomind run; the runtime resolves its model settings, system prompt, MCP servers, and allowed tools. With the octofs integration, an agent can call view, text_editor, batch_edit, extract_lines, shell, and workdir to inspect a repository, run commands, and change files. It can also enable servers with mcp, delegate to specialists with tap, register dynamic workers through agent, and handle delayed or event-driven work through schedule and monitor. Requests can arrive interactively or over stdin, while responses can be plain text or JSONL; supported models can additionally follow a supplied JSON Schema. Repository-level guards can deny matching calls before execution, hooks can turn failed checks into corrective feedback, and validators can report post-turn failures back to the agent. Named sessions persist across runs, and adaptive compaction reduces accumulated context while retaining critical task knowledge.

  1. A developer runs developer:general in a local repository to inspect code, execute tests, and make edits.
  2. A CI maintainer pipes a review or analysis task into Octomind and captures plain-text or JSONL output.
  3. A platform team runs a named daemon or WebSocket server for monitoring jobs, internal dashboards, or editor integrations.
  4. A multi-agent system invokes octomind acp to use a configured specialist as an ACP sub-agent.
  5. A team governing unattended automation uses guards, hooks, and validators to block risky calls and feed test failures back automatically.
  6. A cost-conscious engineering group assigns different providers to research and review stages while tracking request and session spending.

How do you install or deploy this agent?

The installation script supports macOS, Linux, and Windows; Windows users run it under Git Bash or MSYS2. The default destination is ~/.local/bin/.

curl -fsSL https://raw.githubusercontent.com/muvon/octomind/master/install.sh | bash

Sign in for model access through the default octohub:auto configuration:

octomind login

Alternatively, skip Cloud login, supply a supported provider key, and replace the model in [model], [supervisor.model], and [compression.model]. For example:

export OPENROUTER_API_KEY="your_key"
[model]
name = "openrouter:google/gemini-2.5-flash"

[supervisor.model]
name = "openrouter:google/gemini-2.5-flash"

[compression.model]
name = "openrouter:google/gemini-2.5-flash"

Cargo installation is also documented and requires Rust 1.95 or newer:

cargo install octomind

Verify the binary, generate the default configuration, and start a session:

octomind --version
octomind config
octomind run

How do you use this agent?

Start the general development specialist supplied by the default tap. Its first use may install tools and request credentials:

octomind run developer:general

For scripts or CI, provide the request through stdin. The positional argument to octomind run is a role or tap tag, not a message:

echo "Explain the auth module" | octomind run developer:general --format plain

Select JSONL for machine-readable output:

echo "List TODO items" | octomind run developer:general --format jsonl

For a long-running session, start a named daemon and send subsequent requests from another terminal. The daemon stays attached to its original terminal:

echo "watch the build" | octomind run --name watcher --daemon --format jsonl
octomind send --name watcher "run the test suite"

Expose a specialist as an ACP sub-agent:

octomind acp developer:general

Named sessions can be resumed later:

octomind run --name my-feature
octomind run --resume my-feature

What are this agent's strengths and limitations?

Pros
  • One session engine supports interactive CLI use, stdin automation, daemon operation, WebSocket integrations, and ACP orchestration.
  • It supports stdio and Streamable HTTP MCP servers and can activate capabilities or delegate to tap specialists during a session.
  • Deterministic pre-call guards, post-result hooks, and post-turn validators provide concrete controls for unattended execution.
  • Model selection is configurable per role and workflow step across multiple providers, with mid-session switching and cache-aware cost accounting.
  • Adaptive compaction, persistent named sessions, and an out-of-band supervisor are designed for multi-hour or multi-day work.
Limitations
  • Although the core is one binary, octofs, octobrain, and tap-provided MCP tools can bring separate dependencies and credentials.
  • The generated configuration defaults to octohub:auto; BYOK users must update the main, supervisor, and compression model tables and inspect role overrides.
  • Spending thresholds inspect already-recorded costs, so an individual provider call can exceed the configured amount; they are not hard prepaid caps.
  • Schema-constrained output depends on model support, and --format at an interactive terminal errors unless stdin is piped or daemon mode is enabled.
  • Anonymous telemetry is enabled by default. It can be disabled, and the documented payload excludes code, prompts, paths, tool arguments, and environment values.

How does this agent compare with similar options?

The repository reports July 30, 2026 octobench results on 25 tasks taken from merged pull requests and scored by held-out tests. Octomind with glm-5.2 solved 24 tasks, Claude Code with claude-opus-5 solved 23, Codex with gpt-5.6-sol solved 21, and OpenCode with glm-5.2 solved 19. These figures evaluate complete setups: the Octomind run used a staged tap and a binary override with an unfinished-handback pre-gate, so the difference cannot be attributed to a single runtime feature or model. The README also identifies them as published historical results rather than measurements of the current checkout.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Octomind This agent 70 · Some gaps CLIFreemium ★ 147 1d ago Rust OpenAI API · Claude API
Selectools 65 · Some gaps Library / SDKFree + model costs ★ 11 2mo ago Python OpenAI API · Claude API
DSPy-Go 49 · Major gaps CLIFree + model costs ★ 197 12d ago Go OpenAI API · Claude API
Neuron AI — Agentic Framework for PHP 45 · Major gaps Library / SDKFree + model costs ★ 2.1k 2d ago PHP OpenAI API · Claude API

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Some gaps
70/ 100 5-point scale 3.5 / 5
Trust 16/29
Reliability 9/14
Adaptability 16/18
Convention 14/18
Effectiveness 10/13
Verifiability 5/8
Why each dimension lost points
Trust16 / 29 · 2.8/5

The evidence shows role- and capability-scoped tool exposure, an optional sandbox, pre-call guards, post-result hooks, post-turn validators, and request/session spending thresholds. The README also discloses consequential capabilities such as file writes, shell execution, remote model traffic, cloud login, remote taps, and automatic dependency installation. Deductions apply because sandboxing is not shown as the default, the design deliberately substitutes scripted policy for per-action confirmation, spending limits are disabled by default and can be exceeded before stopping, and credential persistence is mentioned without storage location, encryption, redaction, or deletion details. A weekly workflow updates dependencies, tests them, and runs cargo audit, while Cargo.toml documents avoidance of known advisories; however, the CI imports a reusable workflow from a mutable master branch and several Actions are pinned only to major tags. Saved sessions provide limited recovery, but no atomic file rollback or undo facility is demonstrated. Apache licensing, a company author, and contact metadata establish attribution, although publisher identity remains unverified and the provenance chain for taps and auto-installed tools is not fully documented.

Reliability9 / 14 · 3.2/5

The README, Cargo manifest, CI, and fixtures present a broadly coherent multi-surface MCP agent. CI covers Linux, macOS, Windows, stable/beta/nightly, and musl builds; fixtures exercise MCP timeouts and protocol errors, while the ACP smoke script checks initialization, sessions, cancellation, and tool use. Deductions apply because no product implementation or executed results are supplied, the README banner shows v0.50.1 while Cargo declares v0.55.2, and operation depends on external models, Octomind Cloud, taps, Hugging Face assets, ONNX libraries, and a shared master-branch CI workflow. Scripts contain useful missing-binary, timeout, and failure diagnostics and describe graceful MCP initialization failure, but direct evidence for user-facing authentication, networking, quota, and partial-failure messages is limited.

Adaptability16 / 18 · 4.4/5

The material thoroughly covers interactive terminals, stdin/CI, daemons, WebSocket, ACP, SSH, tmux, specialist roles, BYOK, and multiple model providers, earning full credit for audiences, scenarios, and environment fit. Capabilities are layered through roles, skills, capability bundles, and MCP servers. Semantic activation, deterministic rules, abstention on near ties, deactivation, and compression interactions are described precisely, supporting full trigger-precision credit. Capability boundaries are deducted because roles may dynamically enable servers, install dependencies, register agents, and delegate work, with effective authority controlled by external taps and configuration; the supplied files do not prove a uniform least-privilege ceiling over every dynamic extension.

Convention14 / 18 · 3.9/5

The README has a table of contents, quick start, mode comparison, configuration samples, command examples, architecture links, and documentation routes, giving it strong information architecture and examples. The complete Apache-2.0 license agrees with Cargo metadata. Installation covers a curl pipeline, cargo install, and source builds and notes the Rust requirement plus first-run credential/tool setup, but checksum, signature, and version-pinned curl installation guidance are absent. Naming is generally clear and ambiguous behavior such as run arguments, default roles, and daemon attachment is explained; deductions reflect the stale banner version and provider availability varying with octolib. Limitations are candid but dispersed rather than consolidated. Semantic versioning is present, but no changelog, compatibility policy, or migration guide is supplied. The company, email, repository, and contribution route identify responsibility at an organizational level, but named maintainers, response commitments, and a security-reporting path are absent, and publisher identity remains unknown.

Effectiveness10 / 13 · 3.8/5

Plain text, JSONL, schema-constrained output, the interactive UI, ACP, and WebSocket modes make results usable by people and automation, supporting full output-usability credit. Roles, dynamic capabilities, guardrails, compaction, spending controls, and multiple surfaces offer plausible marginal value beyond a basic chat client, but much of that evidence is README-level assertion without implementation source. The benchmark identifies task count, cost, duration, date, and explicitly says it is not a measurement of the current checkout; because its repository and artifacts are not supplied, it cannot establish current effectiveness. Cost tracking, cache accounting, thresholds, and per-step model choice support cost management, but thresholds default off, can be crossed before enforcement, and subscription, model, auto-installation, and maintenance costs are not comprehensively quantified.

Verifiability5 / 8 · 3.1/5

The README maps many claims to named documentation sections, commands, configuration keys, and a specific benchmark revision. It also carefully distinguishes model tool calls from shell commands, historical benchmark results from the current checkout, and continuation thresholds from prepaid limits. Cargo metadata, CI, dependency automation, MCP fixtures, and the ACP smoke script independently corroborate some version, platform, protocol, timeout, and audit claims. Deductions apply because implementation source, linked documentation, Cargo.lock, benchmark artifacts, and actual CI outputs are absent, leaving major security, compaction, cost, and performance claims traceable only to the README. Marketing conclusions such as “no surprise tools” and the asserted governance benefits are not directly proven by the supplied source.

Risks and how to mitigate them
  • The agent can read and write files, execute shells, contact remote models and MCP servers, auto-install tap dependencies, and delegate to sub-agents. Before using it outside a trusted repository, explicitly enable sandboxing and review every role, tap, capability, and guardrail configuration.
  • Per-action user confirmation is not the primary safety control. Spending thresholds are disabled by default and may trigger only after a provider call has exceeded them; configure deny rules, budgets, and network scope before unattended operation.
  • Login and specialist setup may request and persist credentials, but the supplied evidence does not explain encryption, file permissions, redaction, or removal. Avoid entering high-privilege keys on unaudited shared hosts or in environments with exposed logs.
  • The quick installer uses remote curl-to-shell, and CI imports a mutable master-branch workflow. Pin revisions, verify downloaded artifacts, and audit taps plus their auto-installed dependencies before production adoption.
  • The published benchmark explicitly does not measure the reviewed checkout, and its artifacts are absent from the evidence. Do not treat the 24/25 result, cost, or duration as independently verified for this revision.
Evidence confidence: Low Reviewed Oct 05, 2026 Reviewed revision ecda35115d91
See the full review method →

FAQ

Is an Octomind Cloud subscription required?
No. Cloud login supplies model access for the default octohub:auto setup, but users can instead configure supported providers such as OpenRouter, OpenAI, Anthropic, DeepSeek, or Ollama.
What local resources can the agent access?
Access depends on the active role and configuration. With octofs, it can view, search, and edit files, execute shell commands, and manage the working directory; executable scripts in .agents/tools/ can also become project-local MCP tools.
How can a team restrict dangerous operations?
A .agents/guardrails.toml file can define pre-call guards, post-result hooks, and post-turn validators. Rules may match capability names, argument patterns, call history, roles, and result text.
Are the spending thresholds guaranteed hard limits?
No. Both thresholds are disabled by default and operate on recorded costs, so one provider request can take spending beyond the configured amount.
Does it support fully unattended workflows?
Yes. It accepts stdin with plain or JSONL output and also offers daemon, WebSocket, and ACP modes. A session spending threshold stops piped, ACP, or WebSocket work, while an interactive terminal can ask whether to continue.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents