Coder Eval

Playwright for coding agents: run real agents against declarative YAML tasks in a sandbox, score them with weighted criteria, and gate CI on the result.

Stars
★ 141
Last updated
today
License
Apache-2.0
Primary language
Python

At a glance

How it runs
CLIAgent plugin / skill
Works with
Universal · cross-platformCodex · Claude Code · OpenAI API · Claude API
Cost
Free software; you pay for model usage
Setup effort
Medium · a few setup steps
You'll need
Python 3.13+uv 0.8+Claude Code CLINode.js 20Docker (only for container isolation driver)Shell / CLINetwork accessLocal filesystemMCP Server
Typical use
Benchmark authors who want to compare Claude Code, Codex, Gemini (Antigravity), OpenCode, and Pi on task suites from their own domain instead of a shared public dataset.
Not a fit if
  • Teams that want a ready-made fixed leaderboard instead of authoring their own tasks and scoring criteria
  • Users unwilling to install and configure a separate agent runtime such as Claude Code or OpenCode
Source review
86/100 · Good

What does this agent do, and when should you use it?

Coder Eval is UiPath's open-source, agent-agnostic framework for evaluating and benchmarking AI coding agents and their skills, published on PyPI as coder-eval and installed with pip install coder-eval or uv tool install coder-eval on Python 3.13+. An evaluation is a YAML task file: an initial_prompt, an agent config, a sandbox driver, and a list of success_criteria. At run time the framework starts a real agent — Claude Code, OpenAI Codex, Google Antigravity (Gemini), OpenCode, or Pi — inside an isolated sandbox, then scores the files and commands it actually produced. Scoring is continuous and weighted from 0.0 to 1.0, with criterion types ranging from file_exists and run_command to code similarity and LLM-graded rubrics, plus a skill_triggered activation check. It also ships an A/B experiment layer, per-tool-call token and cost telemetry, run.json/variant.json/task.json reports, a GitHub Actions composite action, and a Claude Code plugin marketplace front end. It never supplies model access: you bring the agent runtime and the model credentials.

You author a YAML task (task_id, description, initial_prompt, agent, sandbox, success_criteria), then use the coder-eval CLI: coder-eval plan tasks/hello_date.yaml validates it without spending tokens, coder-eval run tasks/hello_date.yaml starts the configured agent in a tempdir or Docker sandbox to complete the task, and coder-eval report runs/latest shows the scored result. agent.type is the only harness-specific line, so switching to codex, antigravity, or opencode (or overriding per run with -D agent.type=opencode) leaves tasks, criteria, scoring, telemetry, and reports unchanged. After the run, each success_criterion is evaluated against the artifacts and commands the sandbox produced, yielding a weighted score against pass/fail thresholds. The experiment layer fans the same tasks across different models, tools, or prompts side by side, while telemetry records every tool call, token count, and cost with real-time streaming. For CI, the UiPath/coder_eval composite action installs a pinned CLI version, runs the tasks, writes JUnit XML and run.md, and fails the step on any task failure. Inside Claude Code, the repo doubles as a plugin marketplace exposing six slash commands — /coder-eval:init, /coder-eval:check-skill, /coder-eval:task, /coder-eval:lint-tasks, /coder-eval:analyze, and /coder-eval:ci — which drive the same CLI.

  1. Benchmark authors who want to compare Claude Code, Codex, Gemini (Antigravity), OpenCode, and Pi on task suites from their own domain instead of a shared public dataset.
  2. Teams shipping Claude Code skills or plugins who need the skill_triggered criterion to confirm the skill actually fires in the agent their users run.
  3. Skill maintainers who wire a scheduled GitHub Actions job so skills are continuously re-validated and quietly stop triggering before users notice.
  4. Platform or infrastructure teams that want coding-agent quality to be a CI gate, failing the build when a regression appears.
  5. Researchers and product teams running A/B experiments — model vs. model, tool-on vs. tool-off, prompt vs. prompt — who also want per-tool token and cost data.
  6. Teams with existing task data that use Bring Your Own Dataset to fan one task across many rows and scale the suite.

How do you install or deploy this agent?

Prerequisites: Python 3.13+, uv 0.8+, and the runtime of at least one coding agent plus that agent's own model credentials. From a clone:

uv sync                      # install the framework (add --extra codex / --extra antigravity for those runtimes)
cp .env.example .env         # then set ANTHROPIC_API_KEY — or skip it: an existing `claude login` is picked up automatically

If you only want the CLI without cloning:

uv tool install coder-eval
uv tool install "coder-eval[codex,antigravity]"   # same, with agent extras
coder-eval --version                              # verify the install

To add it as a project dependency, use uv add coder-eval or pip install coder-eval. Two of the agents need their runtime installed separately:

brew install claude                     # Claude Code
npm install -g opencode-ai              # OpenCode

How do you use this agent?

Your first end-to-end run, using the default claude-code agent:

uv run coder-eval plan tasks/hello_date.yaml   # validate (no tokens spent)
uv run coder-eval run  tasks/hello_date.yaml   # run your first evaluation
uv run coder-eval report runs/latest           # view the result

A minimal task file, tasks/hello_world.yaml:

task_id: "hello_world"
description: "Create a Python script that prints Hello, World!"
initial_prompt: "Create hello.py that prints 'Hello, World!'"

agent:
  type: "claude-code"
  permission_mode: "acceptEdits"
  allowed_tools: ["Read", "Write", "Bash"]

sandbox:
  driver: "tempdir"
  python: {}

success_criteria:
  - type: "file_exists"
    path: "hello.py"
    description: "hello.py must be created"
  - type: "run_command"
    command: "python hello.py"
    timeout: 10
    description: "Script must execute successfully"

Install the Claude Code plugin front end for the six slash commands:

/plugin marketplace add UiPath/coder_eval
/plugin install coder-eval@coder-eval

Use it as a GitHub Actions gate (the default claude-code agent needs the claude CLI first):

- uses: actions/setup-node@v4
  with: { node-version: '20' }
- run: npm install -g @anthropic-ai/claude-code

- uses: UiPath/coder_eval@v0
  id: eval
  with:
    args: |
      tests/tasks/**/*.yaml
      --model
      claude-sonnet-5
    env: |
      ANTHROPIC_API_KEY=${{ secrets.ANTHROPIC_API_KEY }}

- if: always()
  run: cat "${{ steps.eval.outputs.run-md-path }}" >> "$GITHUB_STEP_SUMMARY"
- uses: mikepenz/action-junit-report@v5
  if: always()
  with:
    report_paths: ${{ steps.eval.outputs.junit-path }}

Do not run the action under pull_request_target with secrets exposed to untrusted fork PRs. To turn off usage telemetry, set TELEMETRY_ENABLED=false in your .env or environment.

What are this agent's strengths and limitations?

Pros
  • Tasks, criteria, scoring, telemetry, and reports stay identical across harnesses: agent.type is the only agent-specific line, and swapping to Codex, Antigravity, OpenCode, or Pi requires no task edits.
  • Scoring is continuous and weighted (0.0–1.0) with thresholds, and the skill_triggered criterion exists specifically to verify that a target skill actually fires.
  • Three documented delivery paths — the coder-eval CLI, the UiPath/coder_eval GitHub Action with JUnit XML and run.md outputs, and a Claude Code plugin marketplace — make it usable as a CI gate out of the box.
  • Full per-tool-call telemetry covering token counts and cost, plus an A/B experiment layer for comparing models, tool toggles, and prompts on the same tasks.
  • Sandboxing supports both a tempdir driver and a Docker container driver, and task dependencies can be pinned for reproducibility.
Limitations
  • Python 3.13+ only, which is a hard blocker for teams still on older interpreters.
  • No model access is provided: you must supply agent runtimes (Claude Code and OpenCode are separate CLIs) and your own model credentials, so real API spend is unavoidable.
  • Anonymous usage telemetry is on by default and must be explicitly disabled with TELEMETRY_ENABLED=false.
  • Tasks execute agent-generated code, and the docs state the tempdir driver is not a security boundary — untrusted tasks must use the container driver.
  • The GitHub Action exposes eight inputs but none of them is a coder-eval run flag, so flags and task globs must all go through args; GitHub silently ignores an input the referenced tag does not define, which can produce a run that measured something else and still exits 0.

How does this agent compare with similar options?

The docs compare Coder Eval against several named alternatives: versus fixed benchmarks like SWE-bench and SkillsBench, which score a canonical dataset, Coder Eval scores your tasks with continuous 0.0–1.0 weighted criteria (and can still wrap a fixed dataset via Bring Your Own Dataset); versus large-scale/RL harnesses such as Harbor, which target scale and RL rollouts, Coder Eval targets weighted, skill-aware suites gated in CI; versus model-output eval tools like OpenAI Evals, which grade model text, Coder Eval runs a full agent in a sandbox and scores the files and commands it produced; versus hand-rolled scripts, it ships reproducible sandboxes, weighted criteria, cost/token telemetry, A/B experiments, and CI-ready pass/fail gates.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
Coder Eval This agent 86 · Good CLIFree + model costs ★ 141 today Python Codex · Claude Code · OpenAI API · Claude API
Loop Engineering 69 · Some gaps CLIFree + model costs ★ 11k 1d ago TypeScript Codex · Claude Code
Helmor Local Agent Workbench 52 · Major gaps Desktop appFree + model costs ★ 1.3k 2mo ago TypeScript Codex · Claude Code
PRO-LONG 57 · Major gaps CLIFree ★ 458 1mo ago Python Codex · Claude Code

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Good
86/ 100 5-point scale 4.3 / 5
Trust 24/29
Reliability 12/14
Adaptability 15/18
Convention 16/18
Effectiveness 13/13
Verifiability 6/8
Why each dimension lost points
Trust24 / 29 · 4.1/5

Strong evidence: CI workflows default to zero permissions with per-job opt-in, persist-credentials: false, prompt-injection tool allowlists, env-only credential passthrough; dependencies carry CVE annotations, constraint pins, pip-audit/bandit/CodeQL. Deductions: telemetry is on by default (opt-out with a one-time notice, not opt-in); executing real agent code is an inherent external effect, with the tempdir driver explicitly documented as not a security boundary; no explicit rollback mechanism beyond pin-version advice; maintainer identity not independently verifiable (unverified publisher).

Reliability12 / 14 · 4.3/5

README, pyproject, and workflows are highly consistent (rationales for deliberately empty extras, version, entry points); missing optional dependencies fail at dispatch/start with documented clear hints instead of crashing. Deduction: failure-message quality is asserted in comments ('clear hint pointing back here'); the actual message text was not observable in this static review.

Adaptability15 / 18 · 4.2/5

Explicit audiences (benchmark authors, CLI/skill builders) and scenarios (A/B, CI gating, skill validation); harness switching via a single agent.type field; a clear 'Known limits & non-goals' section including the tempdir-is-not-a-boundary warning. Deductions: precision of skill_triggered detection cannot be verified statically; Python 3.13+-only and macOS-primary development narrow the environment fit (though declared).

Convention16 / 18 · 4.4/5

Complete docs index (tutorials, user guide, schemas, comparison), install notes covering clone/CLI/CI paths, explicit limitations, LICENSE+NOTICE, semantic-release versioning and CHANGELOG. Deductions: the GitHub Action is still @v0 ('@v1 once 1.0.0 ships'), so interface naming carries drift risk; pre-1.0 with no LTS, and maintenance commitment rests on the latest release plus an email channel with no visible cadence evidence.

Effectiveness13 / 13 · 5.0/5

Output usability well evidenced: JUnit XML, run.md, a documented report schema, evalboard; the differentiation vs SWE-bench/Harbor/OpenAI Evals is concrete; a no-token plan command, per-tool cost telemetry, and pin-to-release CI guidance show real cost/benefit awareness.

Verifiability6 / 8 · 3.8/5

README claims cross-check against pyproject extras, workflow configs, and SECURITY.md scope notes — good corroboration. Deductions: performance/comparison claims point to an external site (coder-eval.com) and are not reviewable within this repo; several technical judgments ('1.0.0 already proved…', glibc analysis) mix fact and author inference rather than independently checkable evidence.

Risks and how to mitigate them
  • Usage telemetry is on by default: set TELEMETRY_ENABLED=false in .env or the environment if you do not accept it.
  • The tempdir sandbox driver is not a security boundary: use the container (Docker) driver for untrusted tasks; evaluated agent code really executes.
  • The GitHub Action is still @v0: pin to an exact version, and never expose secrets under pull_request_target (the README itself warns of this).
  • The agent_judge criterion spawns a Claude Code agent with tool access using evaluator credentials — SECURITY.md flags it as a trust boundary; treat its inputs as untrusted.
  • Python 3.13+ only, and the publisher is unverified: verify publisher identity and supply chain separately before enterprise adoption.
Evidence confidence: Low Reviewed Sep 28, 2026 Reviewed revision 101bb5c85c5a
See the full review method →

FAQ

Does it ship models or a benchmark dataset?
No. Coder Eval never supplies model access, and it only bundles an agent runtime where an extra says so (codex, antigravity) — Claude Code and OpenCode are separate CLIs you install. It is also not a fixed leaderboard: you bring the tasks and the scoring, and the repo only ships example tasks.
What does a run cost?
The software is free under Apache-2.0; the cost is whatever your agent spends calling models. The framework records every tool call, token count, and cost via telemetry so you can track it, and coder-eval plan <task>.yaml validates a task without spending any tokens.
Is it safe to run tasks that execute generated code?
Tasks execute real agent-generated code. The documentation states that the tempdir driver is not a security boundary, so untrusted tasks must run under the Docker container driver, and in GitHub Actions you should avoid pull_request_target with secrets exposed to untrusted fork PRs.
How do I switch agents or models without rewriting tasks?
Change the agent.type field (claude-code, codex, antigravity, opencode, pi), or override it for a single run with -D agent.type=opencode. Tasks, criteria, scoring, telemetry, and reports stay the same.
How does it relate to Claude Code?
Claude Code is both one of the agents you can evaluate (the default agent.type) and the authoring front end: the repo is a Claude Code plugin marketplace whose six slash commands — init, check-skill, task, lint-tasks, analyze, ci — drive the same coder-eval CLI, which you must install separately for those commands to work.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents