PinchBench Agent Benchmark
53 real-world tasks that measure how well an LLM actually performs as an OpenClaw coding agent.
- Source repo
- pinchbench/skill
- Stars
- ★ 1.4k
- Last updated
- 3mo ago
- License
- MIT
- Primary language
- Python
- FA score
- 42/100 · Major gaps
At a glance
- How it runs
- Works with
- Portable with changesOpenAI API · Claude API
- Cost
- Free software; you pay for model usage
- Setup effort
- Medium · a few setup steps
- You'll need
- Typical use
- Researchers and engineering teams comparing how different models handle tool calling and multi-step reasoning on genuine agent work
- Not a fit if
- Users who cannot run an OpenClaw instance
- Readers who only want leaderboard results, not their own runs
- Source review
- 42/100 · Major gaps 1 safety controls not found
What does this agent do, and when should you use it?
PinchBench is a benchmarking system for evaluating LLM models as the brain of an OpenClaw coding agent. This repository holds the benchmark skill and task set; the public leaderboard lives separately at pinchbench.com, and official results are curated in pinchbench/scripts/default-models.yml. The suite contains 53 tasks spread across eight categories — Productivity, Research, Writing, Coding, Analysis, Email, Memory, and Skills — covering things like calendar scheduling, stock lookups, blog writing, weather scripts, spreadsheet and PDF analysis, inbox triage, context recall, and ClawHub skill discovery. Everything is driven by ./scripts/run.sh, which requires Python 3.10+, the uv package manager, and a running OpenClaw instance. Grading is automatic, LLM-judged, or both; results land in a local results/ directory and can optionally be uploaded to the leaderboard, with a separate official key for official runs. Session transcripts are archived as JSONL files under results/{run_id}_transcripts/ for post-run analysis.
PinchBench reads task definitions under tasks/ and organizes them into real-world categories: event creation and time parsing for calendars, web search and data extraction for research, tone and formatting for writing, code generation and file operations for coding, data processing for spreadsheets and PDFs, inbox management for email, long-term recall for memory, and OpenClaw ecosystem integration for skills. The entry point is ./scripts/run.sh, which accepts --model, --judge, --suite, --runs, --timeout-multiplier, --thinking, --output-dir, --no-upload, --register, --upload, and --official-key. Model IDs must carry a provider prefix such as openrouter/ or anthropic/, with OpenRouter as the default router. Each run drives the model as an OpenClaw agent session that must call tools and chain actions to complete the task. Grading is done automatically, by an LLM judge, or both; without --judge the judge runs as an OpenClaw session, while passing --judge calls the model API directly (OpenRouter, Kilo Gateway, Anthropic, OpenAI, or a headless claude CLI) and bypasses OpenClaw personality injection. Output goes to a results/ JSON file, optionally uploaded to pinchbench.com, with per-task agent conversations saved as JSONL transcripts.
- Researchers and engineering teams comparing how different models handle tool calling and multi-step reasoning on genuine agent work
- Model vendors who want to submit results to the public leaderboard and distinguish ordinary submissions from official runs
- Developers choosing a primary model for an OpenClaw agent who need real-task evidence rather than synthetic capability scores
- Teams running targeted regressions by narrowing scope with --suite, for example only task_calendar and task_stock
- Contributors who want to extend the benchmark by submitting tasks following tasks/TASK_TEMPLATE.md
- Debuggers who replay results/{run_id}_transcripts/ JSONL files to inspect exactly what the agent did on each task
How do you install or deploy this agent?
Prerequisites: Python 3.10+, the uv package manager, and a running OpenClaw instance. Clone the repository and the scripts are ready to use; no separate package installation step is documented.
git clone https://github.com/pinchbench/skill.git
cd skillSet an API key for the provider you use. OpenRouter is the default routing provider; when you pass --judge, prepare the matching environment variable for the judge's prefix (OPENROUTER_API_KEY, KILO_API_KEY, ANTHROPIC_API_KEY, or OPENAI_API_KEY).
How do you use this agent?
Run the full suite against a model. Model IDs must include a provider prefix:
./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4Run only specific tasks, or average over multiple runs and scale timeouts for slower models:
./scripts/run.sh --model openrouter/openai/gpt-4o --suite task_calendar,task_stock
./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4 --runs 3 --timeout-multiplier 2Register for an upload token so results reach the leaderboard, or keep everything local with --no-upload:
./scripts/run.sh --register
./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4
./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4 --no-uploadSubmit an official run, or switch the judge to a direct API call:
export PINCHBENCH_OFFICIAL_KEY=your_official_key
./scripts/run.sh --model anthropic/claude-sonnet-4
./scripts/run.sh --model openai/gpt-4o --judge openrouter/anthropic/claude-sonnet-4-5
./scripts/run.sh --model openai/gpt-4o --judge claudeWhat are this agent's strengths and limitations?
- Tests real work instead of synthetic prompts: 53 tasks spanning calendar scheduling, email triage, file management, spreadsheet and PDF analysis across eight categories
- Flexible grading — automatic, LLM judge, or both; the judge defaults to an OpenClaw session and can instead call OpenRouter, Kilo Gateway, Anthropic, OpenAI, or a headless claude CLI directly
- Model-agnostic switch: any model ID with a provider prefix works, with OpenRouter as the default router
- Results stay local or upload to the public pinchbench.com leaderboard, with a separate official-run marking path
- Every run archives JSONL session transcripts so you can inspect the agent's behavior task by task
- A documented TASK_TEMPLATE.md and explicit task quality criteria make community contributions straightforward
- Hard dependency on a running OpenClaw instance; without it the benchmark cannot execute
- Requires third-party model API keys and incurs per-use API costs rather than running free locally
- Not the source of official leaderboard results — adding models officially means editing pinchbench/scripts/default-models.yml
- Judge comparability depends on which model and provider you choose, and the documentation offers no normalization guidance
- No packaged installer or container image is documented; setup relies on cloning and supplying Python, uv, and OpenClaw yourself
- The repository lists no topics and ships no additional guidance file, so metadata around the project is thin
How does this agent compare with similar options?
The README positions itself against benchmarks that 'test isolated capabilities' but names no specific competing product, so no named comparison is offered.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| PinchBench Agent Benchmark This agent | 42 · Major gaps | CLIFree + model costs | ★ 1.4k | 3mo ago | Python | OpenAI API · Claude API |
| amux Agent Control Plane | 77 · Good | Self-hosted serviceFree + model costs | ★ 507 | 1d ago | Rust | Codex · Claude Code |
| Multica — Multi-Agent Task Collaboration Platform | 52 · Major gaps | Web appFreemium | ★ 52k | today | Go | ChatGPT · Codex · Claude Code |
| CORE Personal AI OS | 44 · Major gaps | CLIFree + model costs | ★ 2k | 23d ago | TypeScript | Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
README states results auto-upload to pinchbench.com and offers --no-upload and --register token flows, giving limited transparency and user control, but it does not describe what is uploaded, how tokens are stored, or data retention, so least_privilege/user_confirmation/data_flow_transparency score 1. Dependencies (pyyaml/fabric/paramiko) are common with lower bounds but no lockfile or hashes, so dependency_security is 1. No rollback or revocation path for uploaded results is provided, so rollback is 0. Source attribution is clear (MIT, pinchbench.com, kilo.ai, OpenClaw links), so source_attribution is 2.
README commands and flags are broadly consistent with pyproject entry points, but README references SKILL.md as the readme while that file is absent from the evidence, and the .agents/skills/building-dashboards tests are unrelated to the PinchBench benchmark, indicating mixed repository content, so self_consistency is 1. Dependencies are public PyPI packages with predictable availability but no pinning, so dependency_availability is 1. Test scripts emit pass/fail and exit codes, but failure messages for the main benchmark scripts are not shown, so failure_messages is 1.
README clearly targets evaluating LLMs as OpenClaw coding agents and lists 8 categories with 53 tasks, so audience_and_scenarios is 2. Boundaries are partly stated (needs OpenClaw instance, Python 3.10+, uv) but unsupported environments or model ranges are not described, so capability_boundaries is 1. Invocation is via CLI flags with no natural-language trigger precision discussion, so trigger_precision is 1. Environment fit requires a running OpenClaw instance and multiple API keys, a narrow specialized setup, so environment_fit is 1.
README is well structured with quick start, command reference, task categories, and contribution guidance, so information_architecture is 2. Install steps are concrete (git clone, uv, Python version, API keys), so install_notes is 2. Naming is mostly consistent but the repo mixes the PinchBench benchmark with a building-dashboards skill, weakening naming stability and topical coherence, so naming_stability is 1. Examples and command samples are plentiful, so examples_and_faq is 2. Known limitations are only mentioned in passing (not the official leaderboard source, needs OpenClaw), so known_limitations is 1. LICENSE is full MIT text, so license is 3. release.yml writes BENCHMARK_VERSION but there is no CHANGELOG, so versioning_changelog is 1. Maintenance points to the pinchbench org and issues link, but publisher identity is unverified, so maintenance_responsibility is 1.
Outputs are results JSON and JSONL transcripts that support later analysis, so output_usability is 2. As a benchmark system it has real value for evaluating agent models, but the repo mixes unrelated skills and official results require editing another repository, so marginal_value is 1. Running requires an OpenClaw instance, model API costs, and multiple runs, so cost_benefit is 1.
Most README claims (task counts, categories, commands) are partly traceable within the repo, but key claims such as 53 tasks and grading methods lack verifiable files in the evidence, so claim_traceability is 1. README, pyproject, and workflows partly corroborate each other, but task definitions and grading code are missing, so cross_source_corroboration is 1. Factual statements are mixed with marketing language without clear separation of inference from fact, so fact_inference_separation is 1.
- Not found in source: rollback or recovery pathBack up first, or work on a git branch or snapshot, so its changes can be undone.
- Results upload to pinchbench.com by default, with no description of uploaded content, token storage, or retention; use --no-upload and review scope before handling sensitive transcripts.
- The repository mixes the PinchBench benchmark with a .agents/skills/building-dashboards skill, so the actual deliverable boundary should be confirmed before assessment.
- README references SKILL.md, which is absent from the evidence, and task definitions and grading code are missing, so the 53-task and grading claims cannot be statically verified.
- Dependencies are unpinned with no hash verification; supply-chain risk must be mitigated by the consumer.
- Publisher identity is unverified; maintenance responsibility and update path are only inferred from README links.