Ordewell
Multi-agent task orchestration for coding agents: turn one goal into an editable plan of tasks, run them in parallel on isolated runners, models and branches, and verify the results.
- Source repo
- ordewell/ordewell
- Stars
- ★ 177
- Last updated
- today
- License
- Apache-2.0
- Primary language
- TypeScript
- FA score
- 67/100 · Some gaps
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code · OpenAI API · Claude API
- Cost
- Free software; you pay for model usage
- Setup effort
- Low · running in minutes
- You'll need
- Typical use
- A solo developer adding a cross-cutting feature (e.g. rate limiting a public API) to a large existing repo who wants to read and change the task plan before anything runs.
- Not a fit if
- Developers without Claude Code, Codex or OpenCode installed
- Projects not managed with git
- Windows users unwilling to use WSL for the terminal UI
- Source review
- 67/100 · Some gaps
What does this agent do, and when should you use it?
Ordewell is an open-source (Apache-2.0) task orchestration tool for coding agents that turns a single goal into an ordered, editable plan of tasks. Its planner is read-only: it researches your repository, asks about anything you left vague, and returns a dependency graph of tasks, each naming the coding agent, model and thinking effort it will use. At execution time each task starts a fresh agent session inside its own git worktree branch, with independent tasks running three at a time by default. A task counts as done only when a unique completion marker appears in the runner's output, not on a model's opinion of its own work; passing tasks land on an integration branch in plan order, and you decide when to merge. It supports Claude Code, Codex and OpenCode out of the box, freely mixable within one plan, and ships four entry points: a terminal UI, a VS Code extension, a CLI, and a local API.
Ordewell's workflow has five stages. 1) Plan: run ordewell plan --goal "..."; the planner explores your workspace without modifying it, asks clarifying questions, and produces an ordered list of tasks with dependencies. 2) Edit: use commands like ordewell task-runner, task-model and task-deps to change any task's runner, model or dependencies without another round trip to the model. 3) Run: ordewell run starts a fresh coding agent session per task in its own worktree, fed the results of its dependencies; independent tasks run concurrently (three at a time by default). 4) Verify: a task passes only when its unique completion marker appears in the agent's output, with the exit code kept as supporting evidence; passing tasks merge onto an integration branch in plan order, with a chance to self-resolve conflicts before waiting for you. 5) Hand off: ordewell handoff review shows the diff, ordewell handoff merge lands or discards it. Folders containing several repositories are treated as one workspace, with a worktree of every repository per task. Pick planner and runners with /planner and /runners; the planner can be a coding agent you already pay for, or any of 25 API providers.
- A solo developer adding a cross-cutting feature (e.g. rate limiting a public API) to a large existing repo who wants to read and change the task plan before anything runs.
- Teams using Claude Code, Codex and OpenCode together who want a security refactor on a stronger model and a README update on a cheaper one, mixed in one plan.
- Engineers who want independent subtasks executed in parallel without agents interfering with each other, relying on git worktree isolation and ordered merges to an integration branch.
- Cautious reviewers who distrust model self-assessment and want completion decided by an evidence marker plus exit code.
- Automation-minded users who prefer scripts: the CLI subcommands and local API fit existing pipelines, with the terminal UI for interactive work.
How do you install or deploy this agent?
Install the CLI globally:
bash
npm install -g ordewellFor VS Code, install Ordewell from the Marketplace, or run:
bash
code --install-extension ordewell.ordewellRequirements: Node.js 20 or newer, at least one of Claude Code, Codex or OpenCode, and git for task isolation. The terminal UI also needs tmux; on Windows, run it under WSL. The VS Code extension bundles its own core and needs nothing from npm.
How do you use this agent?
Run ordewell in your project to open the terminal UI and type a goal; on first run pick a planner and runners with /planner and /runners. The same workflow is available from the command line:
bash
export AI_PROVIDER=claude-code # plan with Claude Code, Codex or OpenCodeordewell plan --goal "Add rate limiting to the public API"
ordewell run
ordewell handoff review # read the diff
ordewell handoff merge # bring it onto your branchBetween plan and run, edit the plan:
bash
ordewell task-runner 2 opencode # move a task to another agent
ordewell task-model 3 sonnet # or just change its model
ordewell task-deps 3 1,2 # make it wait for tasks 1 and 2What are this agent's strengths and limitations?
- The plan is fully editable before execution: change any task's prompt, runner, model, effort or mode, add or remove tasks, and rewire dependencies without another model round trip.
- Per-task model assignment, with every assignment visible before a token is spent.
- Completion verdicts come from evidence — a unique marker in the runner's output plus the exit code — never a model's opinion of its own work.
- The planner is read-only and refuses commands that would change your repository; every task runs isolated on its own worktree branch, landing in plan order on one integration branch.
- No extra API key required: a coding agent you already pay for can act as planner, or any of 25 providers' keys.
- Task isolation depends on git worktrees, so non-git projects cannot use it.
- Requires an installed, paid coding agent — Claude Code, Codex or OpenCode — to run at all.
- The terminal UI needs tmux and is not supported natively on Windows; WSL is required.
- The Ordewell name and logos are excluded from the Apache-2.0 license (see NOTICE), limiting brand reuse.
- Default concurrency is three, and tasks that fail to self-resolve merge conflicts still require manual intervention.
How does this agent compare with similar options?
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Ordewell This agent | 67 · Some gaps | CLIFree + model costs | ★ 177 | today | TypeScript | Codex · Claude Code · OpenAI API · Claude API |
| Plannotator | 78 · Good | CLIFreemium | ★ 9k | today | TypeScript | Codex · Claude Code |
| Claudexor | 76 · Good | CLIFree + model costs | ★ 487 | 2d ago | TypeScript | Codex · Claude Code |
| senpi | 74 · Some gaps | CLIFree + model costs | ★ 453 | 1d ago | TypeScript | OpenAI API · Claude API |
How does FollowAgents rate this agent?
Why each dimension lost points
Least privilege: SECURITY.md defines a read-only planner envelope (ADR-0008, commandPolicy.ts) and candidly admits the denylist classifier is not an OS sandbox; deducted because core code is absent and the classifier cannot be statically verified. User confirmation: plan approval is the consent gate and merges are user-driven; partially deducted since some autonomy levels skip per-step confirmation. Data flow: worktree isolation, integration branch, completion markers are clearly described, but as documented claims only. Sensitive data: key handling and log/argv leakage are in-scope statements, not evidence. Dependency security: npm ci with lockfile exists, but no audit/Dependabot evidence — 1. External effects: running agents on the user's repo is the product; landing requires user action. Rollback: discard run and clean up worktrees. Source attribution: explicit acknowledgement of Matt Pocock's MIT skills plus NOTICE — full marks.
Self-consistency: README, SECURITY, and ADR references agree; version 0.5.5 matches CI/release flow; deducted for invisible internals. Dependency availability: Node 20+/tmux/WSL requirements documented and CI comments explain the lockfile cross-platform problem; deducted because the missing macOS/Windows CI matrix leaves platform availability unproven. Failure messages: merge conflicts name files and wait for the user; deducted because concrete error-handling code was not provided.
Audience and scenarios: developer-facing with four entry points (TUI/VS Code/CLI/API), clearly described — full marks. Capability boundaries: SECURITY.md is the standout — classifier is not a sandbox, local access is not trusted, hostile README content should not be assumed contained — full marks. Trigger precision: CLI surface is clear but parsing logic unseen. Environment fit: Windows requires WSL; cross-platform build gap is openly acknowledged.
Information architecture: monorepo packages, docs/adr/, CONTEXT.md vocabulary — full marks. Install notes: npm, Marketplace, and requirements complete — full marks. Naming stability: pre-1.0, latest-version-only support, no forward naming commitment. Examples and FAQ: concrete quick-start commands, but deep docs live on an external site. Known limitations: the candid SECURITY.md statements are excellent — full marks. License: full Apache-2.0 text with NOTICE — full marks. Versioning/changelog: version and tag verification exist but no CHANGELOG file is present — 1. Maintenance responsibility: security SLA (3/14 days) and defined release flow, but publisher is unverified and community signals unknowable.
Output usability: diff review and handoff merge/discard workflow are clear; deducted because static review cannot verify actual output quality. Marginal value: mixed-agent plans, editable plans, evidence-based verification are differentiators, but self-declared. Cost benefit: model assignments visible before spend is good design; no cost data evidence.
Claim traceability: ADR numbers and concrete file paths (commandPolicy.ts, verify-release-tag.mjs) are cited, but some cited files are not in the evidence set. Cross-source corroboration: README and SECURITY agree, but core source, classifier implementation, and ADR contents are absent — 1. Fact/inference separation: CI comments separate facts from design trade-offs; README carries some marketing tone.
- Static review only: the planner read-only envelope and command classifier (commandPolicy.ts) implementations are not in the evidence set; verify they actually prevent writes/out-of-workspace reads before use — audit the file and ADR-0008/ADR-0011 yourself.
- SECURITY.md itself states the classifier is not a sandbox: hostile repo content (README, commit messages) could influence the planner or sub-agents via prompt injection; run in an isolated environment/container.
- Some autonomy levels run tasks without per-step confirmation; set carefully and review the diff line-by-line before merging.
- Publisher is unverified; pre-1.0 with latest-version-only support, no backward-compatibility commitment, and no CHANGELOG — derive changes from tags/releases yourself.
- No CI coverage on macOS/Windows; non-Linux behaviour is unverified.
- Runner plugins are installed from URLs and describe how to spawn processes; only install plugin manifests from trusted sources.