Codex Dynamic Workflows (Claude Code skill)
Claude authors dynamic-workflow scripts; a fleet of local Codex/GPT agents runs them and streams back a live execution map.
- Source repo
- scasella/claude-dynamic-workflows-codex
- Stars
- ★ 325
- Last updated
- 3mo ago
- License
- MIT
- Primary language
- JavaScript
- FA score
- 53/100 · Major gaps
At a glance
- Works with
- Platform-specificCodex · Claude Code
- You'll need
- Typical use
- A developer auditing every route under src/ for missing authorization, who wants parallel scanning plus an independent skeptic that refutes each finding unless it can point to file:line evidence.
- Main limitation
- Execution depends on the local codex CLI app-server; it is built and verified against codex 0.144.0, and other versions may require regenerating bindings.
- Source review
- 53/100 · Major gaps
What does this agent do, and when should you use it?
This is a Claude Code skill plus a standalone runner that re-hosts the Claude Code dynamic-workflow DSL (agent / parallel / pipeline / phase / budget) on a local Codex (GPT) backend. Claude compiles one or two rough sentences of intent into a workflow script, and the codex app-server executes the agents. It ships the runner (zero npm dependencies), a single-file HTML execution map, a terminal ASCII map, the summarize-run cost report, and a documented fleet supervision protocol. Beyond the native DSL it adds long-lived, steerable workers: agent.start, session.steer, agent.waitAny for races, session.cancel for losers, and --resume that re-attaches persisted threads warm. It is usable as the /codex-workflows skill in Claude Code, or entirely without Claude Code from the CLI.
After a manual /codex-workflows invocation, Claude preflights the Codex app-server, compiles the request into a concrete workflow script (agent / parallel / pipeline / phase / budget), writes it into your project, and runs it on Codex via thread/start plus turn/start, pinning every agent to the current frontier model (e.g. gpt-5.6-sol) and scaling thinking effort to layer width with --effort or --auto-effort. Progress is journaled to <project>/.workflow-journal/<name>.jsonl; view-run.js turns that journal into a self-contained HTML execution map (--watch patches the DOM live without reloading) and map-run.js renders the same run as a terminal ASCII map. summarize-run.js reads the journal plus sidecars to produce a token, timing, and reliability report broken down by phase, worker, and costliest agent. When a workflow reaches a human() gate, the run pauses and, under --gui, an answer card appears in the live page; unattended runs fall back to the gate's default on timeout. With --multi, Claude launches several concurrent variant workflows and supervises them itself using fleet status and fleet answer, killing dead ends and forking winners by copying the journal and resuming at zero tokens for already-completed turns.
- A developer auditing every route under src/ for missing authorization, who wants parallel scanning plus an independent skeptic that refutes each finding unless it can point to file:line evidence.
- An engineer who loads packages/core or a whole data room into one worker once, then asks a stream of follow-up questions without paying a cold agent to re-read the corpus each time.
- An on-call responder facing a 12x p99 latency spike who wants three hypotheses (N+1 query, pool exhaustion, cache stampede) raced in parallel, keeping only the first to land and steering it on its warm thread.
- A team migrating every call of legacyFetch() that still wants the plan shown and a human click before anything under payments/ is touched.
- Anyone who wants to GoalLint a vague /goal before an expensive run, or ClaimCheck a README/PR draft's claims against actual repo artifacts afterwards.
- A researcher attacking one stubborn problem from several angles at once while Claude supervises the fleet, answers gates, and keeps the total under a token ceiling.
How do you install or deploy this agent?
Install as a Claude Code plugin (recommended), or as a classic skills-dir clone:
/plugin marketplace add scasella/claude-dynamic-workflows-codex
/plugin install codex-workflows@codex-workflowsgit clone https://github.com/scasella/claude-dynamic-workflows-codex ~/.claude/skills/codex-workflowsPrerequisites: Node >= 18 (zero npm dependencies) and a logged-in codex CLI on your PATH:
codex loginVerify Codex is reachable at any time:
npx github:scasella/claude-dynamic-workflows-codex doctor # → state: readyHow do you use this agent?
In Claude Code, invoke the skill manually (it never auto-triggers) and describe the task in one or two rough sentences:
/codex-workflows Audit every route under src/ for missing auth checks
/codex-workflows Research the current state of on-device LLM inference and verify each claim, then watch it live
/codex-workflows --multi Find the cause of the checkout p99 regression — attack it from a few different angles at onceOr run a workflow directly without Claude Code:
node runner/bin/run-workflow.js examples/review.workflow.js --frontier --auto-effort \
--sandbox read-only --args '{"files":["runner/src/codexAgent.js"],"focus":"error handling"}'
node runner/bin/run-workflow.js examples/incident-demo/checkout-incident.workflow.js \
--frontier --auto-effort --sandbox read-only --guiInspect the bundled demo run offline, with no Codex and no tokens:
git clone https://github.com/scasella/claude-dynamic-workflows-codex
cd claude-dynamic-workflows-codex
node runner/bin/view-run.js examples/incident-demo --openView any past run, render a terminal map, or distill a cost report:
node runner/bin/view-run.js <project-dir> --open # add --watch for live
node runner/bin/map-run.js <project-dir> --watch
node runner/bin/summarize-run.js <project-dir> --markdown --out reports/summary.mdKey flags: --frontier, --auto-effort, --plan (dry run, no tokens), --budget N with --budget-meter total|output, --sandbox read-only|workspace-write, --tui / --gui / --monitor, --resume, --summary.
What are this agent's strengths and limitations?
- Adds sessionful workers the native one-shot DSL lacks: agent.start plus session.steer reuse one warm thread, and the repo's own benchmark measures ~69k tokens / ~6s per follow-up versus ~219k tokens / ~97s for a cold re-read.
- agent.waitAny wakes on the first worker to finish and session.cancel stops the losers, so you do not pay a parallel() barrier for the slowest strategy; losers are recorded as cancelled, not failed.
- Zero npm dependencies on Node >= 18, and the execution map is a self-contained offline HTML file you can share or open via file://, with Dark/Light themes, a Tree layout, and keyboard navigation.
- Every run is journaled to .workflow-journal as jsonl, and summarize-run reports tokens, time, cache hits, and risk flags by phase, worker, and costliest agent, while --resume replays completed turns for free.
- Fleet supervision is a documented file contract (references/fleet-protocol.md) with a supervise shim, so any long-running command can be wrapped in the same gates and status dashboard.
- Execution depends on the local codex CLI app-server; it is built and verified against codex 0.144.0, and other versions may require regenerating bindings.
- It is explicitly not the native Claude Code experience: no in-session background tasks, no /workflows progress UI, and no save-as-/command.
- Agents default to approvalPolicy "never" inside a workspace-write sandbox, meaning they read, write, and run shell commands without prompting unless you ask for read-only.
- Session resume depends on the persisted rollout; if it is gone or the codex version predates thread/resume, that worker re-runs live, and turn replay is positional, so editing session call order invalidates the replayed prefix.
- Budget accounting is per-process (--budget-meter selects total or the native output-token pool) and does not match native semantics 1:1, and the map models barrier/phase structure as an approximation for pipeline-shaped runs.
How does this agent compare with similar options?
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Codex Dynamic Workflows (Claude Code skill) This agent | 53 · Major gaps | — | ★ 325 | 3mo ago | JavaScript | Codex · Claude Code |
| Ordewell | 67 · Some gaps | CLIFree + model costs | ★ 199 | today | TypeScript | Codex · Claude Code · OpenAI API · Claude API |
| MCO | 73 · Some gaps | CLIFree + model costs | ★ 532 | 4d ago | Python | Codex · Claude Code |
| NTM (Named Tmux Manager) | 73 · Some gaps | CLIFree + model costs | ★ 456 | 4d ago | Go | Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
Evidence shows: default --sandbox workspace-write permits writes but a --sandbox read-only option exists; human() gates with safe timeout defaults, journaled answers, and --resume that does not re-ask support user confirmation and rollback; CI uses contents: read only and there are zero npm dependencies, lowering supply-chain risk. Deductions: the README does not explain how run data (code, prompts, results) flows to the local codex app-server or whether anything leaves the machine, so data-flow transparency is only partial; sensitive-data handling has no dedicated guidance (no PII/secret handling notes); publisher identity is unverified, so source attribution rests only on package.json and LICENSE claims — unknown, not suspicious.
Evidence shows: codex-session.test.js covers timeouts, transport death, interrupt races, and listener leaks; failure messages are specific (e.g. 'Timed out waiting', 'Transport is not connected'); completion never rejects, giving good self-consistency. Deductions: dependency availability is only asserted as Node >=18 plus an external codex CLI, with no pinned versions, offline fallback, or compatibility matrix — thin.
Evidence shows: the README targets Claude Code users with audit/research/migration/triage scenarios, states capability boundaries clearly (no model routing, one frontier model, read-only is safety not cost), and has precise triggering (manual /codex-workflows, never auto-triggered). Deductions: environment fit covers only Node 18+ and the codex CLI, with no notes on Windows, headless, or offline environments — thin.
Evidence shows: clear README structure, complete install notes (plugin and clone paths, doctor check), rich examples, and a full MIT license. Deductions: no CHANGELOG, version only 0.2.0 in package.json; naming stability is unaddressed (model names like gpt-5.6-sol shift externally); known limitations are scattered rather than a dedicated section; maintenance responsibility rests on an unverified author with an update path dependent on GitHub pushes.
Evidence shows: high output usability (inline execution map, viewer, summarize/compare reports, reproducible journal) and clear marginal value over a single agent (parallelism, race cancellation, warm sessionful threads, supervised fleets). Deductions: cost-benefit rests only on self-reported benchmarks (~3x cheaper, ~16x faster) with no independent verification, and defaults use a frontier model at high effort, so cost may be significant — thin.
Evidence shows: the README makes many concrete claims (token counts, timings, benchmarks, dogfood results) but mostly self-reported without raw data for independent checking; test files and CI provide partial cross-corroboration (protocol contract, offline suites). Deductions: claim traceability is weak (no per-claim citations or data files), cross-source corroboration is limited (in-repo self-attestation only), and fact/inference separation is insufficient (marketing prose mixed with verifiable facts).
- Default --sandbox workspace-write lets agents write files; use --sandbox read-only explicitly before running on untrusted repos or prompts.
- The README does not explain how run data (code, prompts, results) flows to the local codex app-server or whether anything leaves the machine; verify data flow before using on sensitive codebases.
- Depends on an external codex CLI and specific frontier model names (e.g. gpt-5.6-sol); version drift may change behavior, and versions are not pinned.
- Cost claims (~3x cheaper, ~16x faster) are self-reported benchmarks without independent verification; defaults use a frontier model at high effort, so real spend may be significant.
- Publisher identity is unverified; maintenance and update path rest on a single unverified author.