Dev & Engineering coding-agent-orchestrationgit-workflowtask-worktreesvalidation-evidencemerge-queueworkflow-runnerlocal-firstrelease-readiness

YYLO CLI

Orchestrate coding agents, repository tasks, validation, and protected merges with durable receipts.

FollowAgents review · FARS-2.1
Recommended
88/ 100 5-point scale 4.4 / 5
1 2 3 4 5 6
Per-dimension scores and reasoning
1Trust25 / 29 · 4.3/5

Least privilege, confirmation, data-flow transparency, and external-effect controls are thoroughly evidenced: observation is separated from mutation, publication authority is isolated, network and model-provider use are disclosed, merges use a fenced owner and expected-old-SHA CAS, and tests exercise refusal without authorization, dirty state, stale refs, and malformed state. Concrete rollback, backup, recovery-plan, and dirt-preservation behavior justifies full rollback credit. Sensitive-data handling only establishes externally configured provider credentials and token-free verification workflows; it does not document comprehensive secret storage, log redaction, or retention policy, so this is deducted to 2. npm ci, exact release identity, and compatibility boundaries support dependency security, but broad caret runtime ranges remain and no vulnerability scanning, provenance signing, or dependency-update policy is shown, so it scores 2. Repository, package, issue tracker, company author, and license attribution agree, but the publisher is unverified and fuller governance or individual stewardship is absent, so source attribution scores 2; unknown identity is not treated as suspicious.

2Reliability12 / 14 · 4.3/5

The README, package manifest, release workflows, and tests are consistent about command names, the Node floor, release channels, task boundaries, and release identity. Stable 0.2.2 versus the current 0.2.3-rc.1 manifest is expressly explained by the stable/next channel split, supporting full self-consistency. Prerequisites, separate Ledger/Benchmark installation, and compatibility versions are documented, but Pi, model providers, npx/Git fallback, and external registries remain dependencies whose availability the supplied files cannot ensure, so dependency availability scores 2. Failure handling is strongly evidenced: tests cover timeout, drift, missing authority, dirty state, detached HEAD, stale refs, duplicate configuration, and interrupted recovery while asserting actionable diagnostics, justifying full marks for failure messages.

3Adaptability18 / 18 · 5.0/5

The documentation addresses beginner developers, project operators, and maintainers across a model-free canary, interactive agents, bounded loops, YAML workflows, tasks, merges, releases, and relocation, so audience and scenarios receive full marks. Capability boundaries thoroughly separate orchestration from provider credentials, independent Ledger/Benchmark tools, controllers, task worktrees, and the integration owner, while identifying observational and mutating commands. Trigger conditions are precise through explicit authorization flags, exact task IDs, release channels, model selectors, and refusal rules. Node, npm, Git, shell completion, interactive-terminal requirements, source-checkout tooling, relocation, and local diagnostic commands provide thorough environment-fit guidance.

4Convention16 / 18 · 4.4/5

The README is comprehensively organized around quick start, ownership, beginner use, workflows, task/merge operations, recovery, release, delegated packages, development, and help. Installation includes prerequisites, stable and prerelease channels, exact pinning advice, and a success canary, justifying full information-architecture and install scores. The yylo/yy equivalence, ypl expansion, removed lifecycle command, and compatibility aliases are explicitly managed, supporting full naming stability. Numerous executable examples, placeholder explanations, nested-help directions, and disclosed operational constraints provide strong examples/FAQ and known-limitations coverage. MIT metadata matches the complete LICENSE, so license receives full marks. Although package.json says CHANGELOG.md is shipped, its contents are absent from the evidence; versions and channels alone do not establish an actual change history, so versioning/changelog scores 1. The company name, support address, repository, and issue route establish ordinary maintenance ownership, but no verified publisher identity, maintainer roster, response commitment, or governance policy is supplied, so maintenance responsibility scores 2.

5Effectiveness10 / 13 · 3.8/5

Structured status, named JSON schemas, truncation and cursor metadata, durable receipts, bounded logs, recovery packets, and explicit help make outputs usable by both people and automation, justifying full output-usability credit. Task worktrees, evidence capture, protected merging, workflow recovery, and release binding show credible value beyond directly invoking an agent, but no comparative baseline, adoption data, or independent outcome evidence is supplied, so marginal value scores 2. Iteration limits, bounded logs, risk-tiered review, and a model-free canary demonstrate cost controls, but token, time, storage, and operational costs are not measured, while Ledger, Benchmark, agent, and multi-worktree dependencies add complexity; cost-benefit therefore scores 2.

6Verifiability7 / 8 · 4.4/5

Important claims are tied to concrete commands, schema names, versions, file responsibilities, workflow checks, and test assertions. Release workflows verify tag, manifest, packed identity, and registry hash, supporting full claim traceability. Some README safety and recovery claims are corroborated by workflows and the two supplied test groups, but only a small subset of implementation and tests is present; many broad task, merge, and workflow claims remain README-only, so cross-source corroboration scores 2. The sources carefully distinguish source versions from published versions, observation from mutation, sync from push, and a model-free canary from provider-contacting commands, and they do not claim that this static review executed tests, justifying full fact/inference separation.

Evidence confidence: Low Reviewed Sep 11, 2026 Reviewed revision 83f05897c946
Before you use it
  • This is a low-confidence static review; the CLI, tests, release flows, and network operations were not executed.
  • Agent runs may send prompts and repository context to external model providers; the supplied material does not define comprehensive secret redaction, log retention, or data-deletion policy.
  • Runtime dependencies use several caret ranges, and the evidence does not include a lockfile, vulnerability scan, artifact signing, or dependency-update governance.
  • Ledger, Benchmark, Pi, model providers, and skill installation are separate dependencies; pin and review their versions and permissions before adoption.
  • Only part of the README's task, merge, and recovery guarantees is corroborated by the supplied tests; review the complete implementation and test suite before production use.
  • Publisher identity is unknown, not suspicious; organizations requiring supply-chain assurance should separately verify JUNO AI INC. and control of the npm package and repository.
Review evidence [1][2][3][4][5][6][7][8]
See the full review method →

What does this agent do, and when should you use it?

YYLO is a local-first command-line orchestrator exposed through equivalent `yylo` and `yy` launchers. It coordinates coding-agent runs and repeatable development workflows while separating command observation, validation evidence, typed tasks, task worktrees, merge ownership, and release readiness into explicit boundaries. Runs can produce bounded logs, terminal state, input-bound evidence, manifests, and durable receipts instead of relying on terminal scrollback. Agent executables, model access, and provider credentials remain external; the documented surface includes optional Pi usage, Codex and Claude Code agent paths, and OpenAI and Anthropic model selections. YYLO Ledger and YYLO Benchmark are separately installed products to which `yy ledger` and `yy benchmark` delegate. The system is a strong fit for teams that value controlled Git topology and auditable recovery, but its controller and queue model may be excessive for a simple one-shot prompt.

YYLO starts coding agents through yy pi, yy start, and the agent aliases exposed by the installed version, with -i available to bound iterations inside an invocation. yy loop and the YAML Workflow Runner execute ordered agent, test, and shell steps while retaining command identity, stdout, stderr, responses, session IDs, attempts, declared receipt hashes, and terminal manifests. yy watch exec|status|await captures bounded logs and machine-readable terminal truth for an ordinary local command, while yy evidence run|status|await creates content-addressed validation evidence tied to exact task inputs. The typed lifecycle—yy task start|run|status|checkpoint|preflight|finish—creates and manages a dedicated branch and worktree, then queues a clean committed task without merging it. yy merge status|plan|arbiter|drive|next|resolve handles delivery through a fenced target owner, expected-old-SHA compare-and-swap checks, and risk-based review or explicit recovery. Read-only commands such as yy info, yy where, yy doctor workspace, and yy integration status inspect registered topology; integration movement, pushing, and publishing remain separate explicit mutations.

  1. A development team that wants Pi, Codex, or Claude Code to implement and inspect increments while npm test runs after each pass.
  2. A repository operator who requires every feature to be developed in a dedicated worktree and admitted through preflight and a protected merge queue.
  3. An audit-sensitive project that needs durable records of test outcomes, agent responses, exact input hashes, attempts, and failure states.
  4. An operator recovering an interrupted workflow who wants to verify completed evidence and resume only the first invalid step.
  5. A maintainer who needs read-only diagnosis of controller, task-worktree, and integration-owner topology before authorizing synchronization or repair.
  6. A team converting a repeated multi-agent coding routine into reviewed YAML with explicit iteration and error-handling limits.

What are this agent's strengths and limitations?

Pros
  • Separates the metadata controller, task worktree, and integration owner so product edits, task records, and protected-target changes have distinct authorities.
  • Retains bounded command logs, terminal manifests, and content-addressed task evidence, reducing dependence on reconstructed terminal history.
  • Protects target mutation with one fenced owner and expected-old-SHA compare-and-swap while preserving conflicts and unrelated dirty content for recovery.
  • Supports both compact inline loops and reviewed YAML workflows, with separate bounds for outer workflow iterations and inner agent iterations.
  • Documents multi-provider Pi use alongside OpenAI Codex and Anthropic model shortcuts, while keeping credentials and model availability outside the orchestrator.
Limitations
  • Adoption requires Node.js 20.10 or newer, npm, Git, a separately installed agent CLI, provider credentials, and whatever usage costs or availability constraints that provider imposes.
  • The full lifecycle introduces controller registration, dedicated worktrees, clean commits, preflight, evidence, and queue concepts that are heavier than a basic one-command agent wrapper.
  • Ledger and Benchmark are not bundled and must be installed at compatible versions; the skill set also has a separate repository and installation lifecycle.
  • doctor workspace intentionally exits nonzero for actionable topology findings, so automation must distinguish diagnostic findings from an execution failure.
  • Synchronization, merge, push, and publication are independent authority boundaries, requiring operators to understand workspace roles and invoke each mutation explicitly.

How do you install or deploy this agent?

YYLO requires Node.js 20.10 or newer, npm, and Git. Install the stable channel and verify it:

npm install --global '@yylo/cli@latest'
yy --version

Run the first local canary without contacting a model provider:

mkdir yylo-demo
cd yylo-demo
git init
yy init --task "Document the onboarding path" --subagent pi
yy watch exec pwd

A successful run initializes .juno_task/ and emits a watch receipt containing "state":"COMPLETED", "exit_code":0, and nonzero log_bytes. The seven YYLO agent skills are not bundled; install them separately with yy skills install if needed. Before a model-backed run, install the selected coding-agent CLI and configure its provider credentials. For Pi, the documented installation is npm install --global '@mariozechner/pi-coding-agent'. Package installation, skill acquisition, and model calls generally require network access.

How do you use this agent?

Inspect the registered workspace and exact command inventory first:

yy info --json
yy doctor workspace
yy --help

Run a repository analysis without persisting a Pi session:

yy pi --no-session 'Summarize this repository and make no changes'

That invocation may contact the configured model provider. A bounded implementation-and-test loop can be run as:

yy loop -n 2 --step 'yy pi "Implement the next increment"' --step 'npm test'

For a typed manual task, run these commands from the metadata controller:

yy task start TASK_ID
# Edit, test, and commit in the task worktree returned by start
yy task preflight TASK_ID
yy task finish TASK_ID
yy merge arbiter run --through TASK_ID

The managed path is yy task run TASK_ID followed by the explicit mutation yy merge drive --through TASK_ID. Use yy task status TASK_ID, yy merge status, or yy evidence status TASK_ID for observation; status commands do not acquire merge, push, or release authority.

FAQ

Do I need model credentials to verify the installation?
No. The documented yy watch exec pwd canary runs locally without contacting a model provider. Model-backed commands such as yy pi require the corresponding agent, credentials, and an available model.
Will YYLO automatically merge, push, or publish changes?
No implicit authority is granted. task finish only queues a clean committed task; merge drive, integration push, and the maintainer publication script are separate explicit operations.
What happens when a repository has uncommitted changes?
Bootstrap does not commit or rearrange an existing or dirty repository. Integration sync refuses dirty, diverged, or ambiguous state, and merge recovery preserves conflicts and unrelated dirt instead of automatically resetting, stashing, rebasing, or squashing it.
Are Ledger, Benchmark, and YYLO skills included in the npm package?
No. Ledger and Benchmark are independently installed canonical packages that receive delegated commands. The versioned YYLO skills must also be acquired separately with yy skills install.
Must an interrupted workflow restart from the beginning?
Not necessarily. Workflow Runner provides recover-attempt RUN_DIRECTORY --dry-run and doctor RUN_DIRECTORY; recovery verifies unchanged successful evidence before resuming at the first invalid step.

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents