Cezar — AI Coding Agent Orchestrator
One control center to run Claude Code, Codex, OpenCode and other AI coding agents in parallel — locally or 24/7 on your own server.
- Source repo
- open-mercato/cezar
- Stars
- ★ 261
- Last updated
- today
- License
- MIT
- Primary language
- TypeScript
- FA score
- 63/100 · Some gaps
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code
- Cost
- Free, no paid service needed
- Setup effort
- Low · running in minutes
- You'll need
- Typical use
- A solo developer who wants several coding agents working on different tasks across repos simultaneously, each in its own worktree.
- Not a fit if
- Teams that want a chat-first agent inside Slack or IM rather than a web cockpit
- Projects unwilling to let agents modify the repo via git worktrees
- Source review
- 63/100 · Some gaps
What does this agent do, and when should you use it?
Cezar is the open-source orchestrator and ADE from the open-mercato/cezar repository, providing a single web cockpit for running multiple AI coding agents concurrently. It reuses your existing logged-in claude, codex, opencode or pi CLIs, so no API keys are required. Every task runs in its own git worktree, letting several agents work simultaneously while extra tasks wait in a queue. Workflows are short YAML files and skills are Markdown files, with agents mixable per step; live runs stream agent text, tool calls, tokens and cost. It has no database — all state is plain files under .ai/cezar/ — and can be started locally with npx cezar-run or deployed to an Ubuntu VPS via server-install for 24/7 operation, reviewable from a fully responsive phone UI.
Cezar takes task descriptions (typed, with attached files, or launched from a GitHub issue / Jira / Linear) and runs them in a fresh git worktree according to a YAML workflow (such as the built-in quick-task), executing agent steps plus shell checks. A typical workflow:
yaml
name: fix-and-verifysteps:
- id: implement
prompt: "{{task}}"
skill: project-conventions
runner: codex
- id: verify
command: "npm test"
onFail: { retry: implement, max: 2 }It invokes your locally logged-in agent CLIs to execute steps, retrying automatically with the error fed back when a check fails. The cockpit streams every step, tool call, token and cost live. You can run the same task ×2 or ×3 and compare diffs, agents can delegate child tasks (up to four children in flight per parent, sharing the parent budget), and Automations launch tasks on schedules or from GitHub/tracker events. Everything — workflows, skills, task state — is stored as plain files in .ai/cezar/. Results are reviewed as diffs, with notes or draft PRs; nothing auto-merges.
- A solo developer who wants several coding agents working on different tasks across repos simultaneously, each in its own worktree.
- A remote/freelance developer deploying Cezar on a VPS so agents keep working 24/7 after the laptop closes, reviewing diffs from a phone.
- An open-source maintainer handing a GitHub issue to an agent in one click and reviewing the resulting draft PR.
- Teams on Jira or Linear launching workflows directly from tracker issues and configuring event-driven automations.
- A developer comparing approaches by running the same task ×2 or ×3 and keeping the best diff.
- A team tracking per-project token usage and reported cost through the Usage & cost dashboard.
How do you install or deploy this agent?
Prerequisites: Node 20+ and at least one logged-in agent CLI (Claude Code, Codex, OpenCode, or pi); git and gh are optional. No database or API keys needed.
bash
cd your-repo
npx cezar-runThis opens the cockpit at http://localhost:4321. To try it without logging in:
bash
CEZ_DRY_RUN=1 npx cezar-run(Uses a built-in mock agent.) To set up a server:
bash
npx cezar-run server-install --platform ubuntu-vpsThis sets up HTTPS, a login and a system service so the cockpit is reachable from anywhere, including your phone. A macOS + ngrok guide is also documented.
How do you use this agent?
Type a task in the cockpit, pick a workflow and hit Start, or run headless from the CLI:
bash
npx cezar-run run "add a -- flag to the export command"Scaffold project configuration:
bash
npx cezar-run initTry the nightly build:
bash
npx cezar-run@nightlyCustom workflows are YAML files in .ai/cezar/workflows/; skills are Markdown files in .ai/skills. Workflows can be built by drag-and-drop in the UI and are saved as YAML. Each step can target a different agent via the runner field, and command steps pass when they exit 0. You can also connect GitHub, Jira or Linear to launch tasks from issues, or schedule recurring work and event triggers in Automations.
What are this agent's strengths and limitations?
- Reuses your existing agent CLI logins — no API keys — to drive Claude Code, Codex, OpenCode or Pi.
- Per-task git worktrees give genuine multi-agent parallelism, plus parent/child task delegation (up to four children in flight).
- Live streaming of agent text, tool calls, tokens and cost, with per-project cost comparison in Usage & cost.
- Automations combine schedules with GitHub/tracker event triggers, and VPS deployment enables 24/7 operation.
- No database: all state lives as plain files under .ai/cezar/, easy to inspect and back up.
- Requires Node 20+ and an already-logged-in agent CLI (claude, codex, opencode, or pi); without one, no real tasks can run.
- Parents are limited to four children in flight, and children share the parent's budget, which may constrain heavy parallel use.
- Agent usage bills against your own Claude/Codex/OpenCode subscriptions — the software is free but heavy use costs you tokens.
- Complex behavior depends on hand-written or drag-built YAML workflows and Markdown skills, which imposes a learning curve.
- A promoted cloud sandbox (openmercatocloud.com) exists, but its pricing is not stated in the source material reviewed.
How does this agent compare with similar options?
The README positions Cezar as an orchestration layer above the individual coding agent CLIs — Claude Code, Codex, OpenCode and Pi are the execution backends it invokes, while Cezar adds parallel dispatch, worktree isolation, workflows, automations and one cockpit, which none of those CLIs provide alone.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Cezar — AI Coding Agent Orchestrator This agent | 63 · Some gaps | CLIFree | ★ 261 | today | TypeScript | Codex · Claude Code |
| Kandev | 74 · Some gaps | CLIFree + model costs | ★ 848 | 2d ago | Go | Codex · Claude Code |
| Solo Agent | 48 · Major gaps | Self-hosted serviceFree + model costs | ★ 695 | 15d ago | Go | Codex · Claude Code |
| Emdash: Parallel AI Coding Agent Desktop | 45 · Major gaps | Desktop appFree + model costs | ★ 5.8k | 1d ago | TypeScript | ChatGPT · Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
Evidence shows local-only binding to 127.0.0.1, git-worktree isolation, a four-child dispatch cap, 'Nothing auto-merges', reversible server-install/uninstall, and a threat model in SECURITY.md (DNS rebinding, CSRF, path traversal) — least_privilege, external_effects, rollback, user_confirmation score 2. Deductions: sensitive_data_handling only mentions tracker secrets inside vulnerability scope with no key-storage detail (1); no code evidence of per-action permission confirmation beyond human diff review for user_confirmation.
self_consistency: dry-run, runs., and provider-disable behavior described in README match e2e assertions (2). dependency_availability: Node 20+, agent CLIs, and optional gh are clearly stated (2). failure_messages: tests cover runtime auth failure with recovery events, disabled-provider errors, and unknown-platform exit codes (2). Deduction: core server source was not provided, so breadth of runtime error handling cannot be confirmed.
audience_and_scenarios: local, VPS, mobile, GitHub/Jira/Linear triggers are covered (2). capability_boundaries: autonomous mode, child budgets, no auto-merge are documented (2). environment_fit: Node 20+, Ubuntu VPS and macOS+ngrok guides, CEZ_HOME isolation (2). Deduction: trigger_precision rests on one sentence ('Preview event filters before enabling them') with no filter semantics (1).
install_notes are thorough (npx, headless, server-install, dry-run) — 3. license is full MIT text with named author — 3. information_architecture, naming_stability (the #851 incident with a dedicated guard test), examples_and_faq, and versioning_changelog (dist-tag system but no CHANGELOG file shown) score 2. Deductions: known_limitations only notes pre-1.0 in SECURITY.md, absent from README (1); maintenance_responsibility has a support table and SLAs but no successor/maintainer path (2).
output_usability: diff review, variant comparison, live streaming, usage/cost panel (2). marginal_value: parallel multi-agent orchestration with worktree isolation is a genuine increment over single-agent CLIs (2). cost_benefit: no database, plain-file state, token/cost tracking (2). Deduction: all utility claims are static declarations; no execution verification is possible in this review.
claim_traceability: core claims (package contents, bin mapping, provider disable, server-install steps) are pointed at by e2e tests and CI (2). cross_source_corroboration: README, package., CI, SECURITY.md, and tests reinforce each other (2). Deduction: fact_inference_separation — the README mixes marketing language ('hundreds of agents', 24/7) with factual claims, and key documents like docs/reference.md were not supplied with the evidence, so they cannot be checked (1).
- Publisher identity is unverified (open-mercato / Patryk Lewczuk); verify the supply chain independently before enterprise adoption.
- The project is pre-1.0 with security fixes only for the latest minor line; pin an exact version for production use.
- Sensitive-data handling (tracker tokens, agent logins) is described only as a threat model, with no storage/encryption implementation detail — review the source before deploying.
- Autonomous mode runs to completion without confirmation, backed only by human diff review; disable or restrict it on critical repositories.
- server-install configures nginx, HTTPS, and a system service via sudo — a high-impact system change; validate in an isolated environment first.
- This is a static review with no executed tests; marketing claims in the README ('hundreds of agents', 24/7) are not independently verified.