Cezar — AI Coding Agent Orchestrator

One control center to run Claude Code, Codex, OpenCode and other AI coding agents in parallel — locally or 24/7 on your own server.

Stars
★ 261
Last updated
today
License
MIT
Primary language
TypeScript

At a glance

How it runs
CLIWeb appSelf-hosted service
Works with
Universal · cross-platformCodex · Claude Code
Cost
Free, no paid service needed
Setup effort
Low · running in minutes
You'll need
Node.js 20+at least one logged-in agent CLI (claude, codex, opencode, or pi)git (optional)GitHub CLI gh (optional)Shell / CLINetwork accessLocal filesystem
Typical use
A solo developer who wants several coding agents working on different tasks across repos simultaneously, each in its own worktree.
Not a fit if
  • Teams that want a chat-first agent inside Slack or IM rather than a web cockpit
  • Projects unwilling to let agents modify the repo via git worktrees
Source review
63/100 · Some gaps

What does this agent do, and when should you use it?

Cezar is the open-source orchestrator and ADE from the open-mercato/cezar repository, providing a single web cockpit for running multiple AI coding agents concurrently. It reuses your existing logged-in claude, codex, opencode or pi CLIs, so no API keys are required. Every task runs in its own git worktree, letting several agents work simultaneously while extra tasks wait in a queue. Workflows are short YAML files and skills are Markdown files, with agents mixable per step; live runs stream agent text, tool calls, tokens and cost. It has no database — all state is plain files under .ai/cezar/ — and can be started locally with npx cezar-run or deployed to an Ubuntu VPS via server-install for 24/7 operation, reviewable from a fully responsive phone UI.

Cezar takes task descriptions (typed, with attached files, or launched from a GitHub issue / Jira / Linear) and runs them in a fresh git worktree according to a YAML workflow (such as the built-in quick-task), executing agent steps plus shell checks. A typical workflow:

yaml

name: fix-and-verify

steps:

- id: implement
prompt: "{{task}}"
skill: project-conventions
runner: codex
- id: verify
command: "npm test"
onFail: { retry: implement, max: 2 }

It invokes your locally logged-in agent CLIs to execute steps, retrying automatically with the error fed back when a check fails. The cockpit streams every step, tool call, token and cost live. You can run the same task ×2 or ×3 and compare diffs, agents can delegate child tasks (up to four children in flight per parent, sharing the parent budget), and Automations launch tasks on schedules or from GitHub/tracker events. Everything — workflows, skills, task state — is stored as plain files in .ai/cezar/. Results are reviewed as diffs, with notes or draft PRs; nothing auto-merges.

  1. A solo developer who wants several coding agents working on different tasks across repos simultaneously, each in its own worktree.
  2. A remote/freelance developer deploying Cezar on a VPS so agents keep working 24/7 after the laptop closes, reviewing diffs from a phone.
  3. An open-source maintainer handing a GitHub issue to an agent in one click and reviewing the resulting draft PR.
  4. Teams on Jira or Linear launching workflows directly from tracker issues and configuring event-driven automations.
  5. A developer comparing approaches by running the same task ×2 or ×3 and keeping the best diff.
  6. A team tracking per-project token usage and reported cost through the Usage & cost dashboard.

How do you install or deploy this agent?

Prerequisites: Node 20+ and at least one logged-in agent CLI (Claude Code, Codex, OpenCode, or pi); git and gh are optional. No database or API keys needed.

bash

cd your-repo
npx cezar-run

This opens the cockpit at http://localhost:4321. To try it without logging in:

bash

CEZ_DRY_RUN=1 npx cezar-run

(Uses a built-in mock agent.) To set up a server:

bash

npx cezar-run server-install --platform ubuntu-vps

This sets up HTTPS, a login and a system service so the cockpit is reachable from anywhere, including your phone. A macOS + ngrok guide is also documented.

How do you use this agent?

Type a task in the cockpit, pick a workflow and hit Start, or run headless from the CLI:

bash

npx cezar-run run "add a -- flag to the export command"

Scaffold project configuration:

bash

npx cezar-run init

Try the nightly build:

bash

npx cezar-run@nightly

Custom workflows are YAML files in .ai/cezar/workflows/; skills are Markdown files in .ai/skills. Workflows can be built by drag-and-drop in the UI and are saved as YAML. Each step can target a different agent via the runner field, and command steps pass when they exit 0. You can also connect GitHub, Jira or Linear to launch tasks from issues, or schedule recurring work and event triggers in Automations.

What are this agent's strengths and limitations?

Pros
  • Reuses your existing agent CLI logins — no API keys — to drive Claude Code, Codex, OpenCode or Pi.
  • Per-task git worktrees give genuine multi-agent parallelism, plus parent/child task delegation (up to four children in flight).
  • Live streaming of agent text, tool calls, tokens and cost, with per-project cost comparison in Usage & cost.
  • Automations combine schedules with GitHub/tracker event triggers, and VPS deployment enables 24/7 operation.
  • No database: all state lives as plain files under .ai/cezar/, easy to inspect and back up.
Limitations
  • Requires Node 20+ and an already-logged-in agent CLI (claude, codex, opencode, or pi); without one, no real tasks can run.
  • Parents are limited to four children in flight, and children share the parent's budget, which may constrain heavy parallel use.
  • Agent usage bills against your own Claude/Codex/OpenCode subscriptions — the software is free but heavy use costs you tokens.
  • Complex behavior depends on hand-written or drag-built YAML workflows and Markdown skills, which imposes a learning curve.
  • A promoted cloud sandbox (openmercatocloud.com) exists, but its pricing is not stated in the source material reviewed.

How does this agent compare with similar options?

The README positions Cezar as an orchestration layer above the individual coding agent CLIs — Claude Code, Codex, OpenCode and Pi are the execution backends it invokes, while Cezar adds parallel dispatch, worktree isolation, workflows, automations and one cockpit, which none of those CLIs provide alone.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Cezar — AI Coding Agent Orchestrator This agent 63 · Some gaps CLIFree ★ 261 today TypeScript Codex · Claude Code
Kandev 74 · Some gaps CLIFree + model costs ★ 848 2d ago Go Codex · Claude Code
Solo Agent 48 · Major gaps Self-hosted serviceFree + model costs ★ 695 15d ago Go Codex · Claude Code
Emdash: Parallel AI Coding Agent Desktop 45 · Major gaps Desktop appFree + model costs ★ 5.8k 1d ago TypeScript ChatGPT · Codex · Claude Code

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Some gaps
63/ 100 5-point scale 3.2 / 5
Trust 18/29
Reliability 9/14
Adaptability 10/18
Convention 13/18
Effectiveness 9/13
Verifiability 4/8
Why each dimension lost points
Trust18 / 29 · 3.1/5

Evidence shows local-only binding to 127.0.0.1, git-worktree isolation, a four-child dispatch cap, 'Nothing auto-merges', reversible server-install/uninstall, and a threat model in SECURITY.md (DNS rebinding, CSRF, path traversal) — least_privilege, external_effects, rollback, user_confirmation score 2. Deductions: sensitive_data_handling only mentions tracker secrets inside vulnerability scope with no key-storage detail (1); no code evidence of per-action permission confirmation beyond human diff review for user_confirmation.

Reliability9 / 14 · 3.2/5

self_consistency: dry-run, runs., and provider-disable behavior described in README match e2e assertions (2). dependency_availability: Node 20+, agent CLIs, and optional gh are clearly stated (2). failure_messages: tests cover runtime auth failure with recovery events, disabled-provider errors, and unknown-platform exit codes (2). Deduction: core server source was not provided, so breadth of runtime error handling cannot be confirmed.

Adaptability10 / 18 · 2.8/5

audience_and_scenarios: local, VPS, mobile, GitHub/Jira/Linear triggers are covered (2). capability_boundaries: autonomous mode, child budgets, no auto-merge are documented (2). environment_fit: Node 20+, Ubuntu VPS and macOS+ngrok guides, CEZ_HOME isolation (2). Deduction: trigger_precision rests on one sentence ('Preview event filters before enabling them') with no filter semantics (1).

Convention13 / 18 · 3.6/5

install_notes are thorough (npx, headless, server-install, dry-run) — 3. license is full MIT text with named author — 3. information_architecture, naming_stability (the #851 incident with a dedicated guard test), examples_and_faq, and versioning_changelog (dist-tag system but no CHANGELOG file shown) score 2. Deductions: known_limitations only notes pre-1.0 in SECURITY.md, absent from README (1); maintenance_responsibility has a support table and SLAs but no successor/maintainer path (2).

Effectiveness9 / 13 · 3.5/5

output_usability: diff review, variant comparison, live streaming, usage/cost panel (2). marginal_value: parallel multi-agent orchestration with worktree isolation is a genuine increment over single-agent CLIs (2). cost_benefit: no database, plain-file state, token/cost tracking (2). Deduction: all utility claims are static declarations; no execution verification is possible in this review.

Verifiability4 / 8 · 2.5/5

claim_traceability: core claims (package contents, bin mapping, provider disable, server-install steps) are pointed at by e2e tests and CI (2). cross_source_corroboration: README, package., CI, SECURITY.md, and tests reinforce each other (2). Deduction: fact_inference_separation — the README mixes marketing language ('hundreds of agents', 24/7) with factual claims, and key documents like docs/reference.md were not supplied with the evidence, so they cannot be checked (1).

Risks and how to mitigate them
  • Publisher identity is unverified (open-mercato / Patryk Lewczuk); verify the supply chain independently before enterprise adoption.
  • The project is pre-1.0 with security fixes only for the latest minor line; pin an exact version for production use.
  • Sensitive-data handling (tracker tokens, agent logins) is described only as a threat model, with no storage/encryption implementation detail — review the source before deploying.
  • Autonomous mode runs to completion without confirmation, backed only by human diff review; disable or restrict it on critical repositories.
  • server-install configures nginx, HTTPS, and a system service via sudo — a high-impact system change; validate in an isolated environment first.
  • This is a static review with no executed tests; marketing claims in the README ('hundreds of agents', 24/7) are not independently verified.
Evidence confidence: Low Reviewed Sep 28, 2026 Reviewed revision 31b8598b7783
See the full review method →

FAQ

Do I need to pay or provide API keys?
Cezar itself is free and open source (MIT). It uses your own logged-in claude, codex, opencode or pi CLI sessions and explicitly requires no API keys; agent usage consumes your own subscription quota. A vendor cloud sandbox is promoted, but its pricing is not documented in the source.
Will agents merge code automatically?
No. The README states "Nothing merges on its own" — you review diffs, send notes back, or open draft PRs; merging is always a human decision.
What happens when a task fails?
When a shell check in a workflow fails, the agent sees the error and retries; retries are configurable, e.g. onFail: { retry: implement, max: 2 }.
Can it run fully offline?
Not guaranteed. It needs your logged-in coding agent CLIs to do real work, which typically require network access, and server deployment and GitHub/Jira/Linear integrations also depend on connectivity.
Where is my data stored?
There is no database. All workflows, skills and task state are plain files under .ai/cezar/ in your project directory.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents