Coder Eval
Playwright for coding agents: run real agents against declarative YAML tasks in a sandbox, score them with weighted criteria, and gate CI on the result.
- Source repo
- UiPath/coder_eval
- Stars
- ★ 141
- Last updated
- today
- License
- Apache-2.0
- Primary language
- Python
- FA score
- 86/100 · Good
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code · OpenAI API · Claude API
- Cost
- Free software; you pay for model usage
- Setup effort
- Medium · a few setup steps
- You'll need
- Typical use
- Benchmark authors who want to compare Claude Code, Codex, Gemini (Antigravity), OpenCode, and Pi on task suites from their own domain instead of a shared public dataset.
- Not a fit if
- Teams that want a ready-made fixed leaderboard instead of authoring their own tasks and scoring criteria
- Users unwilling to install and configure a separate agent runtime such as Claude Code or OpenCode
- Source review
- 86/100 · Good
What does this agent do, and when should you use it?
Coder Eval is UiPath's open-source, agent-agnostic framework for evaluating and benchmarking AI coding agents and their skills, published on PyPI as coder-eval and installed with pip install coder-eval or uv tool install coder-eval on Python 3.13+. An evaluation is a YAML task file: an initial_prompt, an agent config, a sandbox driver, and a list of success_criteria. At run time the framework starts a real agent — Claude Code, OpenAI Codex, Google Antigravity (Gemini), OpenCode, or Pi — inside an isolated sandbox, then scores the files and commands it actually produced. Scoring is continuous and weighted from 0.0 to 1.0, with criterion types ranging from file_exists and run_command to code similarity and LLM-graded rubrics, plus a skill_triggered activation check. It also ships an A/B experiment layer, per-tool-call token and cost telemetry, run.json/variant.json/task.json reports, a GitHub Actions composite action, and a Claude Code plugin marketplace front end. It never supplies model access: you bring the agent runtime and the model credentials.
You author a YAML task (task_id, description, initial_prompt, agent, sandbox, success_criteria), then use the coder-eval CLI: coder-eval plan tasks/hello_date.yaml validates it without spending tokens, coder-eval run tasks/hello_date.yaml starts the configured agent in a tempdir or Docker sandbox to complete the task, and coder-eval report runs/latest shows the scored result. agent.type is the only harness-specific line, so switching to codex, antigravity, or opencode (or overriding per run with -D agent.type=opencode) leaves tasks, criteria, scoring, telemetry, and reports unchanged. After the run, each success_criterion is evaluated against the artifacts and commands the sandbox produced, yielding a weighted score against pass/fail thresholds. The experiment layer fans the same tasks across different models, tools, or prompts side by side, while telemetry records every tool call, token count, and cost with real-time streaming. For CI, the UiPath/coder_eval composite action installs a pinned CLI version, runs the tasks, writes JUnit XML and run.md, and fails the step on any task failure. Inside Claude Code, the repo doubles as a plugin marketplace exposing six slash commands — /coder-eval:init, /coder-eval:check-skill, /coder-eval:task, /coder-eval:lint-tasks, /coder-eval:analyze, and /coder-eval:ci — which drive the same CLI.
- Benchmark authors who want to compare Claude Code, Codex, Gemini (Antigravity), OpenCode, and Pi on task suites from their own domain instead of a shared public dataset.
- Teams shipping Claude Code skills or plugins who need the skill_triggered criterion to confirm the skill actually fires in the agent their users run.
- Skill maintainers who wire a scheduled GitHub Actions job so skills are continuously re-validated and quietly stop triggering before users notice.
- Platform or infrastructure teams that want coding-agent quality to be a CI gate, failing the build when a regression appears.
- Researchers and product teams running A/B experiments — model vs. model, tool-on vs. tool-off, prompt vs. prompt — who also want per-tool token and cost data.
- Teams with existing task data that use Bring Your Own Dataset to fan one task across many rows and scale the suite.
How do you install or deploy this agent?
Prerequisites: Python 3.13+, uv 0.8+, and the runtime of at least one coding agent plus that agent's own model credentials. From a clone:
uv sync # install the framework (add --extra codex / --extra antigravity for those runtimes)
cp .env.example .env # then set ANTHROPIC_API_KEY — or skip it: an existing `claude login` is picked up automaticallyIf you only want the CLI without cloning:
uv tool install coder-eval
uv tool install "coder-eval[codex,antigravity]" # same, with agent extras
coder-eval --version # verify the installTo add it as a project dependency, use uv add coder-eval or pip install coder-eval. Two of the agents need their runtime installed separately:
brew install claude # Claude Code
npm install -g opencode-ai # OpenCodeHow do you use this agent?
Your first end-to-end run, using the default claude-code agent:
uv run coder-eval plan tasks/hello_date.yaml # validate (no tokens spent)
uv run coder-eval run tasks/hello_date.yaml # run your first evaluation
uv run coder-eval report runs/latest # view the resultA minimal task file, tasks/hello_world.yaml:
task_id: "hello_world"
description: "Create a Python script that prints Hello, World!"
initial_prompt: "Create hello.py that prints 'Hello, World!'"
agent:
type: "claude-code"
permission_mode: "acceptEdits"
allowed_tools: ["Read", "Write", "Bash"]
sandbox:
driver: "tempdir"
python: {}
success_criteria:
- type: "file_exists"
path: "hello.py"
description: "hello.py must be created"
- type: "run_command"
command: "python hello.py"
timeout: 10
description: "Script must execute successfully"Install the Claude Code plugin front end for the six slash commands:
/plugin marketplace add UiPath/coder_eval
/plugin install coder-eval@coder-evalUse it as a GitHub Actions gate (the default claude-code agent needs the claude CLI first):
- uses: actions/setup-node@v4
with: { node-version: '20' }
- run: npm install -g @anthropic-ai/claude-code
- uses: UiPath/coder_eval@v0
id: eval
with:
args: |
tests/tasks/**/*.yaml
--model
claude-sonnet-5
env: |
ANTHROPIC_API_KEY=${{ secrets.ANTHROPIC_API_KEY }}
- if: always()
run: cat "${{ steps.eval.outputs.run-md-path }}" >> "$GITHUB_STEP_SUMMARY"
- uses: mikepenz/action-junit-report@v5
if: always()
with:
report_paths: ${{ steps.eval.outputs.junit-path }}Do not run the action under pull_request_target with secrets exposed to untrusted fork PRs. To turn off usage telemetry, set TELEMETRY_ENABLED=false in your .env or environment.
What are this agent's strengths and limitations?
- Tasks, criteria, scoring, telemetry, and reports stay identical across harnesses: agent.type is the only agent-specific line, and swapping to Codex, Antigravity, OpenCode, or Pi requires no task edits.
- Scoring is continuous and weighted (0.0–1.0) with thresholds, and the skill_triggered criterion exists specifically to verify that a target skill actually fires.
- Three documented delivery paths — the coder-eval CLI, the UiPath/coder_eval GitHub Action with JUnit XML and run.md outputs, and a Claude Code plugin marketplace — make it usable as a CI gate out of the box.
- Full per-tool-call telemetry covering token counts and cost, plus an A/B experiment layer for comparing models, tool toggles, and prompts on the same tasks.
- Sandboxing supports both a tempdir driver and a Docker container driver, and task dependencies can be pinned for reproducibility.
- Python 3.13+ only, which is a hard blocker for teams still on older interpreters.
- No model access is provided: you must supply agent runtimes (Claude Code and OpenCode are separate CLIs) and your own model credentials, so real API spend is unavoidable.
- Anonymous usage telemetry is on by default and must be explicitly disabled with TELEMETRY_ENABLED=false.
- Tasks execute agent-generated code, and the docs state the tempdir driver is not a security boundary — untrusted tasks must use the container driver.
- The GitHub Action exposes eight inputs but none of them is a
coder-eval runflag, so flags and task globs must all go through args; GitHub silently ignores an input the referenced tag does not define, which can produce a run that measured something else and still exits 0.
How does this agent compare with similar options?
The docs compare Coder Eval against several named alternatives: versus fixed benchmarks like SWE-bench and SkillsBench, which score a canonical dataset, Coder Eval scores your tasks with continuous 0.0–1.0 weighted criteria (and can still wrap a fixed dataset via Bring Your Own Dataset); versus large-scale/RL harnesses such as Harbor, which target scale and RL rollouts, Coder Eval targets weighted, skill-aware suites gated in CI; versus model-output eval tools like OpenAI Evals, which grade model text, Coder Eval runs a full agent in a sandbox and scores the files and commands it produced; versus hand-rolled scripts, it ships reproducible sandboxes, weighted criteria, cost/token telemetry, A/B experiments, and CI-ready pass/fail gates.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Coder Eval This agent | 86 · Good | CLIFree + model costs | ★ 141 | today | Python | Codex · Claude Code · OpenAI API · Claude API |
| Loop Engineering | 69 · Some gaps | CLIFree + model costs | ★ 11k | 1d ago | TypeScript | Codex · Claude Code |
| Helmor Local Agent Workbench | 52 · Major gaps | Desktop appFree + model costs | ★ 1.3k | 2mo ago | TypeScript | Codex · Claude Code |
| PRO-LONG | 57 · Major gaps | CLIFree | ★ 458 | 1mo ago | Python | Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
Strong evidence: CI workflows default to zero permissions with per-job opt-in, persist-credentials: false, prompt-injection tool allowlists, env-only credential passthrough; dependencies carry CVE annotations, constraint pins, pip-audit/bandit/CodeQL. Deductions: telemetry is on by default (opt-out with a one-time notice, not opt-in); executing real agent code is an inherent external effect, with the tempdir driver explicitly documented as not a security boundary; no explicit rollback mechanism beyond pin-version advice; maintainer identity not independently verifiable (unverified publisher).
README, pyproject, and workflows are highly consistent (rationales for deliberately empty extras, version, entry points); missing optional dependencies fail at dispatch/start with documented clear hints instead of crashing. Deduction: failure-message quality is asserted in comments ('clear hint pointing back here'); the actual message text was not observable in this static review.
Explicit audiences (benchmark authors, CLI/skill builders) and scenarios (A/B, CI gating, skill validation); harness switching via a single agent.type field; a clear 'Known limits & non-goals' section including the tempdir-is-not-a-boundary warning. Deductions: precision of skill_triggered detection cannot be verified statically; Python 3.13+-only and macOS-primary development narrow the environment fit (though declared).
Complete docs index (tutorials, user guide, schemas, comparison), install notes covering clone/CLI/CI paths, explicit limitations, LICENSE+NOTICE, semantic-release versioning and CHANGELOG. Deductions: the GitHub Action is still @v0 ('@v1 once 1.0.0 ships'), so interface naming carries drift risk; pre-1.0 with no LTS, and maintenance commitment rests on the latest release plus an email channel with no visible cadence evidence.
Output usability well evidenced: JUnit XML, run.md, a documented report schema, evalboard; the differentiation vs SWE-bench/Harbor/OpenAI Evals is concrete; a no-token plan command, per-tool cost telemetry, and pin-to-release CI guidance show real cost/benefit awareness.
README claims cross-check against pyproject extras, workflow configs, and SECURITY.md scope notes — good corroboration. Deductions: performance/comparison claims point to an external site (coder-eval.com) and are not reviewable within this repo; several technical judgments ('1.0.0 already proved…', glibc analysis) mix fact and author inference rather than independently checkable evidence.
- Usage telemetry is on by default: set TELEMETRY_ENABLED=false in .env or the environment if you do not accept it.
- The tempdir sandbox driver is not a security boundary: use the container (Docker) driver for untrusted tasks; evaluated agent code really executes.
- The GitHub Action is still @v0: pin to an exact version, and never expose secrets under pull_request_target (the README itself warns of this).
- The agent_judge criterion spawns a Claude Code agent with tool access using evaluator credentials — SECURITY.md flags it as a trust boundary; treat its inputs as untrusted.
- Python 3.13+ only, and the publisher is unverified: verify publisher identity and supply chain separately before enterprise adoption.
FAQ
Does it ship models or a benchmark dataset?
What does a run cost?
coder-eval plan <task>.yaml validates a task without spending any tokens.Is it safe to run tasks that execute generated code?
How do I switch agents or models without rewriting tasks?
-D agent.type=opencode. Tasks, criteria, scoring, telemetry, and reports stay the same.