PinchBench Agent Benchmark

53 real-world tasks that measure how well an LLM actually performs as an OpenClaw coding agent.

Source repo
pinchbench/skill
Stars
★ 1.4k
Last updated
3mo ago
License
MIT
Primary language
Python

At a glance

How it runs
CLI
Works with
Portable with changesOpenAI API · Claude API
Cost
Free software; you pay for model usage
Setup effort
Medium · a few setup steps
You'll need
Python 3.10+uvRunning OpenClaw instanceOpenRouter API key (optional, default provider)KILO_API_KEY / ANTHROPIC_API_KEY / OPENAI_API_KEY (judge dependent)Shell / CLINetwork accessLocal filesystem
Typical use
Researchers and engineering teams comparing how different models handle tool calling and multi-step reasoning on genuine agent work
Not a fit if
  • Users who cannot run an OpenClaw instance
  • Readers who only want leaderboard results, not their own runs

What does this agent do, and when should you use it?

PinchBench is a benchmarking system for evaluating LLM models as the brain of an OpenClaw coding agent. This repository holds the benchmark skill and task set; the public leaderboard lives separately at pinchbench.com, and official results are curated in pinchbench/scripts/default-models.yml. The suite contains 53 tasks spread across eight categories — Productivity, Research, Writing, Coding, Analysis, Email, Memory, and Skills — covering things like calendar scheduling, stock lookups, blog writing, weather scripts, spreadsheet and PDF analysis, inbox triage, context recall, and ClawHub skill discovery. Everything is driven by ./scripts/run.sh, which requires Python 3.10+, the uv package manager, and a running OpenClaw instance. Grading is automatic, LLM-judged, or both; results land in a local results/ directory and can optionally be uploaded to the leaderboard, with a separate official key for official runs. Session transcripts are archived as JSONL files under results/{run_id}_transcripts/ for post-run analysis.

PinchBench reads task definitions under tasks/ and organizes them into real-world categories: event creation and time parsing for calendars, web search and data extraction for research, tone and formatting for writing, code generation and file operations for coding, data processing for spreadsheets and PDFs, inbox management for email, long-term recall for memory, and OpenClaw ecosystem integration for skills. The entry point is ./scripts/run.sh, which accepts --model, --judge, --suite, --runs, --timeout-multiplier, --thinking, --output-dir, --no-upload, --register, --upload, and --official-key. Model IDs must carry a provider prefix such as openrouter/ or anthropic/, with OpenRouter as the default router. Each run drives the model as an OpenClaw agent session that must call tools and chain actions to complete the task. Grading is done automatically, by an LLM judge, or both; without --judge the judge runs as an OpenClaw session, while passing --judge calls the model API directly (OpenRouter, Kilo Gateway, Anthropic, OpenAI, or a headless claude CLI) and bypasses OpenClaw personality injection. Output goes to a results/ JSON file, optionally uploaded to pinchbench.com, with per-task agent conversations saved as JSONL transcripts.

  1. Researchers and engineering teams comparing how different models handle tool calling and multi-step reasoning on genuine agent work
  2. Model vendors who want to submit results to the public leaderboard and distinguish ordinary submissions from official runs
  3. Developers choosing a primary model for an OpenClaw agent who need real-task evidence rather than synthetic capability scores
  4. Teams running targeted regressions by narrowing scope with --suite, for example only task_calendar and task_stock
  5. Contributors who want to extend the benchmark by submitting tasks following tasks/TASK_TEMPLATE.md
  6. Debuggers who replay results/{run_id}_transcripts/ JSONL files to inspect exactly what the agent did on each task

How do you install or deploy this agent?

Prerequisites: Python 3.10+, the uv package manager, and a running OpenClaw instance. Clone the repository and the scripts are ready to use; no separate package installation step is documented.

git clone https://github.com/pinchbench/skill.git
cd skill

Set an API key for the provider you use. OpenRouter is the default routing provider; when you pass --judge, prepare the matching environment variable for the judge's prefix (OPENROUTER_API_KEY, KILO_API_KEY, ANTHROPIC_API_KEY, or OPENAI_API_KEY).

How do you use this agent?

Run the full suite against a model. Model IDs must include a provider prefix:

./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4

Run only specific tasks, or average over multiple runs and scale timeouts for slower models:

./scripts/run.sh --model openrouter/openai/gpt-4o --suite task_calendar,task_stock
./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4 --runs 3 --timeout-multiplier 2

Register for an upload token so results reach the leaderboard, or keep everything local with --no-upload:

./scripts/run.sh --register
./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4
./scripts/run.sh --model openrouter/anthropic/claude-sonnet-4 --no-upload

Submit an official run, or switch the judge to a direct API call:

export PINCHBENCH_OFFICIAL_KEY=your_official_key
./scripts/run.sh --model anthropic/claude-sonnet-4
./scripts/run.sh --model openai/gpt-4o --judge openrouter/anthropic/claude-sonnet-4-5
./scripts/run.sh --model openai/gpt-4o --judge claude

What are this agent's strengths and limitations?

Pros
  • Tests real work instead of synthetic prompts: 53 tasks spanning calendar scheduling, email triage, file management, spreadsheet and PDF analysis across eight categories
  • Flexible grading — automatic, LLM judge, or both; the judge defaults to an OpenClaw session and can instead call OpenRouter, Kilo Gateway, Anthropic, OpenAI, or a headless claude CLI directly
  • Model-agnostic switch: any model ID with a provider prefix works, with OpenRouter as the default router
  • Results stay local or upload to the public pinchbench.com leaderboard, with a separate official-run marking path
  • Every run archives JSONL session transcripts so you can inspect the agent's behavior task by task
  • A documented TASK_TEMPLATE.md and explicit task quality criteria make community contributions straightforward
Limitations
  • Hard dependency on a running OpenClaw instance; without it the benchmark cannot execute
  • Requires third-party model API keys and incurs per-use API costs rather than running free locally
  • Not the source of official leaderboard results — adding models officially means editing pinchbench/scripts/default-models.yml
  • Judge comparability depends on which model and provider you choose, and the documentation offers no normalization guidance
  • No packaged installer or container image is documented; setup relies on cloning and supplying Python, uv, and OpenClaw yourself
  • The repository lists no topics and ships no additional guidance file, so metadata around the project is thin

How does this agent compare with similar options?

The README positions itself against benchmarks that 'test isolated capabilities' but names no specific competing product, so no named comparison is offered.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
PinchBench Agent Benchmark This agent 42 · Major gaps CLIFree + model costs ★ 1.4k 3mo ago Python OpenAI API · Claude API
amux Agent Control Plane 77 · Good Self-hosted serviceFree + model costs ★ 507 1d ago Rust Codex · Claude Code
Multica — Multi-Agent Task Collaboration Platform 52 · Major gaps Web appFreemium ★ 52k today Go ChatGPT · Codex · Claude Code
CORE Personal AI OS 44 · Major gaps CLIFree + model costs ★ 2k 23d ago TypeScript Codex · Claude Code

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Major gaps
42/ 100 5-point scale 2.1 / 5
Trust 10/29
Reliability 5/14
Adaptability 8/18
Convention 10/18
Effectiveness 6/13
Verifiability 3/8
Why each dimension lost points
Trust10 / 29 · 1.7/5

README states results auto-upload to pinchbench.com and offers --no-upload and --register token flows, giving limited transparency and user control, but it does not describe what is uploaded, how tokens are stored, or data retention, so least_privilege/user_confirmation/data_flow_transparency score 1. Dependencies (pyyaml/fabric/paramiko) are common with lower bounds but no lockfile or hashes, so dependency_security is 1. No rollback or revocation path for uploaded results is provided, so rollback is 0. Source attribution is clear (MIT, pinchbench.com, kilo.ai, OpenClaw links), so source_attribution is 2.

Reliability5 / 14 · 1.8/5

README commands and flags are broadly consistent with pyproject entry points, but README references SKILL.md as the readme while that file is absent from the evidence, and the .agents/skills/building-dashboards tests are unrelated to the PinchBench benchmark, indicating mixed repository content, so self_consistency is 1. Dependencies are public PyPI packages with predictable availability but no pinning, so dependency_availability is 1. Test scripts emit pass/fail and exit codes, but failure messages for the main benchmark scripts are not shown, so failure_messages is 1.

Adaptability8 / 18 · 2.2/5

README clearly targets evaluating LLMs as OpenClaw coding agents and lists 8 categories with 53 tasks, so audience_and_scenarios is 2. Boundaries are partly stated (needs OpenClaw instance, Python 3.10+, uv) but unsupported environments or model ranges are not described, so capability_boundaries is 1. Invocation is via CLI flags with no natural-language trigger precision discussion, so trigger_precision is 1. Environment fit requires a running OpenClaw instance and multiple API keys, a narrow specialized setup, so environment_fit is 1.

Convention10 / 18 · 2.8/5

README is well structured with quick start, command reference, task categories, and contribution guidance, so information_architecture is 2. Install steps are concrete (git clone, uv, Python version, API keys), so install_notes is 2. Naming is mostly consistent but the repo mixes the PinchBench benchmark with a building-dashboards skill, weakening naming stability and topical coherence, so naming_stability is 1. Examples and command samples are plentiful, so examples_and_faq is 2. Known limitations are only mentioned in passing (not the official leaderboard source, needs OpenClaw), so known_limitations is 1. LICENSE is full MIT text, so license is 3. release.yml writes BENCHMARK_VERSION but there is no CHANGELOG, so versioning_changelog is 1. Maintenance points to the pinchbench org and issues link, but publisher identity is unverified, so maintenance_responsibility is 1.

Effectiveness6 / 13 · 2.3/5

Outputs are results JSON and JSONL transcripts that support later analysis, so output_usability is 2. As a benchmark system it has real value for evaluating agent models, but the repo mixes unrelated skills and official results require editing another repository, so marginal_value is 1. Running requires an OpenClaw instance, model API costs, and multiple runs, so cost_benefit is 1.

Verifiability3 / 8 · 1.9/5

Most README claims (task counts, categories, commands) are partly traceable within the repo, but key claims such as 53 tasks and grading methods lack verifiable files in the evidence, so claim_traceability is 1. README, pyproject, and workflows partly corroborate each other, but task definitions and grading code are missing, so cross_source_corroboration is 1. Factual statements are mixed with marketing language without clear separation of inference from fact, so fact_inference_separation is 1.

Risks and how to mitigate them
  • Not found in source: rollback or recovery pathBack up first, or work on a git branch or snapshot, so its changes can be undone.
  • Results upload to pinchbench.com by default, with no description of uploaded content, token storage, or retention; use --no-upload and review scope before handling sensitive transcripts.
  • The repository mixes the PinchBench benchmark with a .agents/skills/building-dashboards skill, so the actual deliverable boundary should be confirmed before assessment.
  • README references SKILL.md, which is absent from the evidence, and task definitions and grading code are missing, so the 53-task and grading claims cannot be statically verified.
  • Dependencies are unpinned with no hash verification; supply-chain risk must be mitigated by the consumer.
  • Publisher identity is unverified; maintenance responsibility and update path are only inferred from README links.
Evidence confidence: Low Reviewed Sep 30, 2026 Reviewed revision 819384ae8304
See the full review method →

FAQ

What does running PinchBench cost?
The repository is MIT-licensed and free, but benchmarking calls third-party models, so you pay per-use API rates for the tested model and, if you pick one, for a separately invoked judge model.
Do I have to run it myself to see results?
No — the public leaderboard at pinchbench.com is browsable. Running locally is for validating your own model or a specific subset of tasks; the README notes this repository is not the source of official leaderboard results.
What is the difference between an ordinary and an official submission?
After ./scripts/run.sh --register gives you a token, results upload automatically as ordinary submissions. Adding --official-key or setting PINCHBENCH_OFFICIAL_KEY marks the run as official.
Can I keep results off the leaderboard?
Yes, pass --no-upload. The run still writes its results JSON and the JSONL session transcripts under results/{run_id}_transcripts/ to disk.
How do I add a new task?
Follow the format in tasks/TASK_TEMPLATE.md; good tasks are real-world, measurable, reproducible, and challenging, and are contributed through the repository's issue tracker or pull requests.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents