SWE-Bench Pro
Benchmark models and coding agents on long-horizon software engineering tasks.
- Source repo
- scaleapi/SWE-bench_Pro-os
- Stars
- ★ 534
- Last updated
- 10d ago
- License
- MIT
- Primary language
- Python
- FA score
- 55/100 · Major gaps
At a glance
- How it runs
- Works with
- Universal · cross-platform
- Cost
- Free, no paid service needed
- Setup effort
- Medium · a few setup steps
- You'll need
- Typical use
- Model research teams comparing how different models or coding harnesses perform on long-running, repository-level fixes.
- Not a fit if
- Developers seeking a ready-to-use coding assistant
- Teams unable to run containerized evaluations
- Users focused only on short, isolated coding problems
- Source review
- 55/100 · Major gaps 2 safety controls not found
What does this agent do, and when should you use it?
SWE-Bench Pro is a benchmark, dataset, and evaluation toolkit for long-horizon software engineering rather than an end-user coding assistant. Each task gives a model a codebase and an issue and asks it to produce a resolving patch. V2 contains 642 validated tasks across 11 repositories, the HARD-51 subset, and self-contained Harbor task directories with verifiers, reference solutions, and public container images. The repository also preserves the original v1 pipeline, including `swe_bench_pro_eval.py`, run scripts, Dockerfiles, and Docker Hub images. Users can load the data from Hugging Face, generate `.pred` files with their chosen harness or the SWE-agent submodule, consolidate them into JSON, and run containerized evaluations.
It reads tasks from the Hugging Face ScaleAI/SWE-bench_Pro dataset and supplies the codebase-and-issue workload for a selected model or harness to solve. The SWE-agent submodule provides one patch-generation scaffold, while other harnesses may also be used; generated patches are stored as per-instance .pred files. helper_code/gather_patches.py consolidates those files into a JSON array containing instance_id, patch, and prefix. For v1, swe_bench_pro_eval.py evaluates that JSON against task samples and run_scripts using container images and configurable parallel workers. V2 instead packages each case as a Harbor directory under v2/tasks/, together with its verifier, reference solution, and GHCR image. V2 also includes HARD-51 and locked-protocol tooling for an offline agent phase followed by re-grading in a fresh sandbox.
- Model research teams comparing how different models or coding harnesses perform on long-running, repository-level fixes.
- Agent developers who need executable, containerized patch verification instead of judging textual answers alone.
- Researchers measuring performance separately on the full V2 set, HARD-51, or the original v1 dataset.
- Leaderboard participants reproducing SWE-agent or mini-swe-agent runs and preparing consolidated prediction files.
- Evaluation-infrastructure engineers studying a locked protocol with offline execution and fresh-sandbox re-grading.
How do you install or deploy this agent?
Install the repository's declared Python dependencies:
pip install -r requirements.txtInstall Docker for reproducible containerized evaluation. The v1 workflow recommends configuring Modal:
modal setup # Follow the prompts to generate your tokenThen verify that ~/.modal.toml contains:
token_id = <token id>
token_secret = <token secret>
active = trueFor the v1 local Docker Beta path, no additional Modal setup is required; add --use_local_docker when running evaluations. V2 uses Harbor for evaluation. The README shows a harbor run invocation but does not provide a Harbor installation command.
How do you use this agent?
Load the default V2 test split, HARD-51, or the original v1 data:
from datasets import load_dataset
v2 = load_dataset('ScaleAI/SWE-bench_Pro', split='test') # V2, 642 tasks (default config)
hard = load_dataset('ScaleAI/SWE-bench_Pro', 'hard', split='test') # HARD-51
v1 = load_dataset('ScaleAI/SWE-bench_Pro', 'v1', split='test') # original 731 tasksGenerate .pred patch files with a harness of your choice. If using SWE-agent, first configure and run the scaffold following the instructions in the SWE-agent git submodule. Consolidate the predictions:
python helper_code/gather_patches.py \
--directory swe_bench_pro_results/sample1 \
--prefix sample1 \
--output sample1_patches.jsonRun the v1 evaluator:
python swe_bench_pro_eval.py \
--raw_sample_path=swe_bench_pro_full.csv \
--patch_path=<your_patches>.json \
--output_dir=<output_directory> \
--scripts_dir=run_scripts \
--num_workers=100 \
--dockerhub_username=jefzdaV2 tasks can be executed through Harbor. This example runs the reference patch with 50-way concurrency:
harbor run -p v2/tasks -e modal -n 50 -a oracleWhat are this agent's strengths and limitations?
- V2 provides 642 validated tasks drawn from 11 repositories, plus the more demanding HARD-51 subset.
- Every V2 task includes a verifier, reference solution, and public GHCR container image that does not require login to pull.
- The locked protocol separates offline agent execution from re-grading in a fresh sandbox, creating a defined evaluation boundary.
- Patch generation is harness-flexible, while documented SWE-agent and mini-swe-agent paths provide concrete starting points.
- Both v1 and V2 remain available, supporting reproduction of older results alongside adoption of the newer format.
- This is benchmark infrastructure, not a coding assistant that directly handles everyday development work.
- The full workflow spans Python dependencies, Docker, dataset downloads, patch collection, and evaluation commands.
- Modal is recommended for v1, while the local Docker alternative is explicitly labeled Beta.
- V1 and V2 use different task formats, image registries, and evaluation pipelines, adding migration overhead.
- The news section reports identified leaderboard issues under active remediation, so leaderboard consumers should account for that status.
How does this agent compare with similar options?
SWE-Bench Pro is inspired by SWE-Bench but explicitly targets more challenging, long-horizon software engineering tasks. The repository also presents SWE-agent and mini-swe-agent as patch-generation scaffolds; it reports comparable Sonnet 4.5 results for them, but supplies no broader head-to-head comparison.
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| SWE-Bench Pro This agent | 55 · Major gaps | CLIFree | ★ 534 | 10d ago | Python | — |
| SWE-bench | 45 · Major gaps | CLIFree + model costs | ★ 6k | 15d ago | Python | — |
| Multi-SWE-bench | 46 · Major gaps | CLIFree | ★ 362 | 9mo ago | Python | — |
| FrontierAgent | 74 · Some gaps | CLIFree + model costs | ★ 5k | today | Python | OpenAI API |
How does FollowAgents rate this agent?
Why each dimension lost points
The README describes the principal flows involving the dataset, container images, Modal, Docker, and result files. The test harness also provides limited recovery through traps, Git restoration, and temporary-file cleanup, while attributing the paper, upstream SWE-bench project, contributors, and dataset. Deductions apply because there is no least-privilege design or pre-action confirmation: scripts install packages, start Redis, modify and remove test files, and rely on network, container, and optional cloud effects. The Modal token location is documented, but credential protection, rotation, and log-redaction guidance are absent. Python packages use broad lower bounds, and task setup performs unlocked npm installs without hashes, vulnerability checks, or other supply-chain controls.
The V1/V2 distinction, patch collection format, evaluation workflow, and supplied scripts are broadly consistent. The parser emits structured test results, while the verifier retains raw logs and reports missing required tests. Deductions apply because operation depends on external datasets, registries, Docker, Modal, npm, and Redis with loose version constraints. In addition, run_script.sh converts timeouts or Mocha failures into an empty-test JSON payload, which can obscure the originating failure and surface only as missing tests. Unexecuted properties such as determinism and result correctness were not scored.
The material identifies users evaluating long-horizon software-engineering agents and supports arbitrary patch-generation harnesses, SWE-agent, local models, Modal, and local Docker, with separate V2, HARD-51, and legacy V1 paths. Commands, parameters, and selected-test behavior are reasonably precise. Deductions apply because local Docker is labeled Beta and the implementation assumes Linux-like Bash, Docker, and fixed container directory layouts; resource scaling, non-container operation, and alternative cloud backends receive little treatment.
The README is organized into overview, installation, images, usage, and reproduction sections, with runnable commands, a JSON example, release news, a known leaderboard issue, and V1/V2 migration guidance. The complete MIT text justifies full credit for licensing. Deductions apply because there is no dedicated FAQ, formal release record, or systematic changelog. Naming across SWE-Bench Pro, SWE-bench_Pro, V1/V2, and multiple image locations is understandable but not fully uniform. A copyright owner is named, but the unverified publisher identity and absence of a clear maintainer contact, support policy, or security-reporting route leave maintenance responsibility only partly established.
The framework turns agent-generated patches into a uniform JSON format and supplies containerized verification, per-test statuses, logs, and a reward artifact, offering clear marginal value for comparing software-engineering agents. Deductions apply because static evidence cannot establish actual result quality, and the recommended workflow may use 100 workers, model inference, cloud compute, large image downloads, and npm installation without quantified cost, duration, disk, memory, or reduced-resource guidance.
Important claims generally map to named directories, dataset configurations, image conventions, commands, and output schemas, and the README is internally corroborated by the parser, shell harness, and license. V1 versus V2 and public versus private leaderboards are explicitly separated. Deductions apply because claims such as 642 validated tasks and oracle success on 642/642 are primarily assertions in the supplied material without included aggregate evidence or auditable result artifacts. The linked paper, dataset, and leaderboards were not independently checked in this static review, and some reproduction language does not sharply separate author confirmation from independent validation.
- Not found in source: confirmation before actingTurn on (or add) a confirmation step before it acts, and try it in a sandbox or test environment before real data.
- Not found in source: dependency securityPin versions and run a dependency audit (npm audit, pip-audit) before installing; prefer running it in a container.
- Run the evaluation scripts inside an isolated container because they install dependencies, start services, rewrite tests, and remove matched paths.
- Pin and audit Python, npm, container-image, and submodule versions before substantial use; the current constraints do not establish reproducibility or supply-chain safety.
- Do not interpret an empty test result produced after a timeout as an ordinary test failure; inspect the retained stdout, stderr, and original process status.
- Modal configuration contains persistent token material; restrict file permissions, keep configuration and logs out of source control, and rotate credentials under organizational policy.
- Independently verify the 642/642, leaderboard reproduction, and validation claims through execution or attached result artifacts; this assessment did not run the code.
FAQ
Is this a coding agent I can use directly on my repository?
How do V2 and v1 differ?
swe_bench_pro_eval.py, run scripts, and Docker Hub images.Is Modal mandatory?
--use_local_docker. The documented V2 example runs Harbor with Modal.Do the container images require registry credentials?
docker pull. V1 uses prebuilt images from the jefzda/sweap-images Docker Hub repository.Can I evaluate my own model or harness?
vllm for local models, though adaptation depends on the selected harness.