OSWorld 2.1 Computer-Use Agent Benchmark

Run long-horizon desktop tasks on real Ubuntu VMs and score computer-use agents reproducibly.

Stars
★ 358
Last updated
12d ago
License
Apache-2.0
Primary language
Python

At a glance

How it runs
CLIFrameworkSelf-hosted service
Works with
Portable with changesOpenAI API · Claude API
Cost
Free software; you pay for model usage
Setup effort
High · needs real infrastructure
You'll need
Python >= 3.12uvDocker (KVM-capable host) or AWS accountHugging Face account with gated dataset accessimageio-ffmpeg>=0.6.0mocked website self-hosting (Task-Web/[email protected])self-hosted GitLab (for GitLab-backed tasks, needs GITLAB_URL and GITLAB_PRIVATE_TOKEN)Shell / CLINetwork accessLocal filesystem
Typical use
A research group comparing two VLMs or computer-use agents on identical tasks, using one pinned osworld-v2.1 release so numbers are directly comparable.
Not a fit if
  • Solo users who want a quick script, not VM clusters and self-hosted sites
  • Teams evaluating Windows or macOS desktops; only Docker Ubuntu and AWS images are provided

What does this agent do, and when should you use it?

OSWorld-V2 is xlang-ai's benchmark and evaluation environment for computer-use agents; the recommended release is osworld-v2.1. Host-side Python creates Ubuntu VMs through Docker or AWS, a DesktopEnv wraps each one, and the agent observes the screen as screenshots and acts with mouse and keyboard until an evaluator checks the resulting state and scores the run. Task classes are distributed through the gated xlangai/osworld_v2_tasks dataset, initial files and ground truth live in xlangai/osworld_v2_assets_gated, and OSWORLD_FILE_BASE_URL must point at the local asset snapshot. Heavier components are self-hosted: Task-Web/OSWorld-web provides mocked websites, and GitLab-backed tasks require your own GitLab plus a private token. The repo ships reference agents under mm_agents (Claude, GPT, Qwen), multi-environment runners, and manual_examine.py, so you can port your own agent and rerun the same tasks.

Evaluation is driven from the host. Runners such as scripts/python/run_multienv_claude.py and run_multienv_qwen.py read a benchmark release manifest, spin up N DesktopEnv instances against Docker or AWS, and hand each agent a screenshot observation (observation_type screenshot) plus an action interface it uses to drive the VM's browser and desktop apps. Task classes load from evaluation_examples/task_class, and their setup logic, ground-truth files and other assets are pulled via desktop_env.file_source.asset(...) from the local directory named by OSWORLD_FILE_BASE_URL. When a task ends, the evaluator compares files, UI state or mocked-website data and emits a score; screenshots and videos are written to result_dir and can be inspected in the hosted Trajectory Viewer. Helper scripts scripts/tools/download_osworld_v2_tasks.py and download_osworld_v2_assets.py fetch the artifacts pinned to a --benchmark-release, while manual_examine.py replays one example_id step by step for debugging. Default VM credentials are user / osworld-public-evaluation.

  1. A research group comparing two VLMs or computer-use agents on identical tasks, using one pinned osworld-v2.1 release so numbers are directly comparable.
  2. A model vendor running a large parallel regression sweep on AWS before shipping a new checkpoint, to catch GUI-operation regressions.
  3. An agent developer wiring their own agent into run_multienv_*.py, validating with a smoke run and a few examples before launching the full suite.
  4. An engineer debugging why a task fails, using manual_examine.py with a specific example_id to record screenshots and video and separate environment faults from model faults.
  5. A team that wants its score on the verified leaderboard, scheduling a run with the maintainers and sharing agent code plus trajectories under the public evaluation guide.

How do you install or deploy this agent?

You need Python >= 3.12; the project uses uv for dependency management. Clone the release tag and sync:

git clone --branch osworld-v2.1 https://github.com/xlang-ai/OSWorld-V2
cd OSWorld-V2
uv sync --frozen

Accept access to both gated datasets, log in to Hugging Face, then download task classes and assets:

uvx --from huggingface_hub hf auth login
uv run scripts/tools/download_osworld_v2_tasks.py --benchmark-release osworld-v2.1
uv run scripts/tools/download_osworld_v2_assets.py \
  --benchmark-release osworld-v2.1 \
  --target-dir cache/osworld_v2_assets \
  --clean
export OSWORLD_FILE_BASE_URL="$(pwd)/cache/osworld_v2_assets"

Provider prerequisites: Docker needs a Linux host with KVM support, AWS needs ports 3000 and 8000 open for the V2 task service in addition to the standard ports. Self-host Task-Web/[email protected] and export WEBSITE_HOST_SUFFIX; for GitLab tasks self-host GitLab and set GITLAB_URL and GITLAB_PRIVATE_TOKEN. Optional heavy agent or OCR stacks for v1 tasks install with uv sync --extra full.

How do you use this agent?

Fill in the environment variables at the top of the sample script (AWS credentials, subnet and security group, the model API key, OSWORLD_CLIENT_PASSWORD, WEBSITE_HOST_SUFFIX, OSWORLD_FILE_BASE_URL, and GitLab credentials if needed), then run:

bash scripts/bash/run_multienv_claude.sh

Qwen models behind an OpenAI-compatible endpoint use the dedicated runner:

export OPENAI_BASE_URL=http://127.0.0.1:8000/v1
export OPENAI_API_KEY=dummy

uv run python scripts/python/run_multienv_qwen.py \
  --smoke_only \
  --model your-served-model \
  --base_url "$OPENAI_BASE_URL"

Inspect one task by hand:

uv run python scripts/python/manual_examine.py \
  --headless \
  --provider_name aws \
  --observation_type screenshot \
  --result_dir ./results_human_examine \
  --test_config_base_dir evaluation_examples \
  --domain tasks \
  --eval_version v2 \
  --example_id 146 \
  --max_steps 3

For a custom agent, implement the agent interface and import it in run.py for single-threaded runs or in a scripts/python/run_multienv_xxx.py runner for parallel runs, following the Claude or GPT reference implementations.

What are this agent's strengths and limitations?

Pros
  • Tasks execute inside full Ubuntu VMs rather than a DOM or simplified API sandbox, which better reflects long-horizon real desktop work.
  • Benchmark releases pin code, task classes, assets, mocked websites and provider images together, so scores stay comparable across teams.
  • Task classes and complete asset snapshots ship through gated datasets, reducing benchmark leakage from agents searching the web for answers.
  • Two officially supported image paths (Docker and AWS) plus multi-environment runners support parallel, large-scale evaluation.
  • Reference Claude, GPT and Qwen agents, a manual examination script, and migration skills for porting OSWorld 1.0 agents are included.
Limitations
  • Provisioning is heavy: two gated dataset approvals, a local asset snapshot, self-hosted mocked websites, and a self-hosted GitLab instance with a private token for GitLab tasks.
  • Only Docker and AWS images are provided; VMware, Azure, GCP, Aliyun and Volcengine exist as code paths only and require migrating 1.0 images.
  • The main branch is explicitly development code; reproducible evaluation requires checking out a release tag and never mixing releases.
  • Evaluation consumes real cloud VMs and paid model API calls, and no cost ceiling is documented.
  • Some tasks need proxy configuration for website defenses or restricted networks; missing it silently lowers scores.

How does this agent compare with similar options?

The README names OSWorld 1.0 (xlang-ai/OSWorld) as the predecessor and publishes docs/MIGRATING_v2.1_FROM_OSWORLD_V1.md: agents, images and tasks must be migrated, and 2.1 introduces new dependencies such as imageio-ffmpeg>=0.6.0. No other competing benchmark is named in the source.

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
OSWorld 2.1 Computer-Use Agent Benchmark This agent 53 · Major gaps CLIFree + model costs ★ 358 12d ago Python OpenAI API · Claude API
CRAB Cross-Environment Benchmark 28 · Major gaps CLIFree + model costs ★ 428 4d ago Python OpenAI API
AppWorld Agent Benchmark Environment 56 · Major gaps CLIFree + model costs ★ 529 1mo ago Python OpenAI API · Claude API
SWE-Bench Pro 55 · Major gaps CLIFree ★ 541 17d ago Python —

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Major gaps
53/ 100 5-point scale 2.7 / 5
Trust 11/29
Reliability 8/14
Adaptability 10/18
Convention 13/18
Effectiveness 7/13
Verifiability 4/8
Why each dimension lost points
Trust11 / 29 · 1.9/5

least_privilege: README states some tasks require sudo inside the VM and default credentials are user/osworld-public-evaluation; this is a known benchmark design, but no least-privilege configuration or isolation guidance is provided, score 1. user_confirmation: the setup-osworld prompt asks before cloud spend, DNS, SSH, or secret steps, but this is an example prompt rather than an enforced mechanism, and evaluation itself consumes cloud resources, score 1. data_flow_transparency: README explains gated Hugging Face task/asset distribution and OSWORLD_FILE_BASE_URL, but does not systematically describe screenshots, trajectories, or model API data flows, score 1. sensitive_data_handling: sample scripts require AWS keys, Anthropic API key, GitLab private token, etc.; README only warns that a shared GitLab token is risky and offers no secret-management or redaction scheme, score 1. dependency_security: pyproject.toml has a very large dependency set, mostly unpinned, with only a few ~= or == constraints, and no lockfile or vulnerability-scan evidence, score 1. external_effects: evaluation creates cloud VMs, downloads images, and calls external model APIs; README gives cost/resource warnings but no automatic cleanup or blast-radius control, score 1. rollback: README offers version switching, backing up docker_vm_data, and --clean, but no evaluation-failure or state rollback mechanism, score 1. source_attribution: LICENSE carries XLANG NLP Lab copyright, and README provides paper citation, acknowledgements, and data-source links, score 2.

Reliability8 / 14 · 2.9/5

self_consistency: tests test_release_name_is_consistent_across_components and test_release_source_resolution verify release-manifest component tags, and README warns against mixing releases, score 2. dependency_availability: dependencies come from PyPI, gated Hugging Face datasets, Docker images, AWS AMI, and a local packages/osworld-ffmpeg path; some require manual approval or private credentials, so availability has external dependency risk, score 1. failure_messages: tests show ClaudeCodeExecutionError with error_kind and cc_ask returning structured outcome on error results, giving reasonably clear failure information, score 2.

Adaptability10 / 18 · 2.8/5

audience_and_scenarios: README targets researchers, evaluators, and agent developers, covering Docker/AWS scenarios, score 2. capability_boundaries: it clearly states only Docker and AWS images are provided, other providers require migration, and main is development code, score 2. trigger_precision: setup-osworld and migrate-osworld-agent skills are triggered by natural-language prompts with no precise trigger conditions or parameter validation, score 1. environment_fit: requires Python >=3.12, uv, KVM, cloud credentials, self-hosted websites, and GitLab; requirements are clear but heavy, score 2.

Convention13 / 18 · 3.6/5

information_architecture: README is well structured with updates, version matrix, setup, evaluation, FAQ, and citation, score 2. install_notes: provides uv sync --frozen, task/asset downloads, and provider setup steps, score 2. naming_stability: release tags such as osworld-v2.1 are consistent and tested, score 2. examples_and_faq: includes multi-env Claude, Qwen, and manual-examination examples plus FAQ, score 2. known_limitations: README mentions possible task failures from proxy/network issues and that main is development code, but lacks a systematic known-limitations list, score 1. license: full Apache-2.0 text is present and pyproject.toml declares it consistently, score 3. versioning_changelog: benchmark_releases manifests, version matrix, and update log are present, score 3. maintenance_responsibility: README lists current maintainer emails and a public evaluation process, but no long-term maintenance commitment or response-time expectation, score 2.

Effectiveness7 / 13 · 2.7/5

output_usability: evaluation outputs include result directories, a trajectory viewer, and manual-examination scripts, but require environment setup, score 2. marginal_value: as a long-horizon computer-use agent benchmark with gated tasks/assets and multi-provider support, it has clear value for the evaluation ecosystem, score 2. cost_benefit: running requires cloud VMs, model APIs, self-hosted websites, and GitLab, imposing high cost and operational burden, score 1.

Verifiability4 / 8 · 2.5/5

claim_traceability: README claims map to benchmark_releases manifests and test assertions, but some claims such as 'bug-fix release' are not elaborated, score 2. cross_source_corroboration: README, pyproject.toml, LICENSE, and tests corroborate version, license, and dependency-source claims, score 2. fact_inference_separation: README is mostly operational instructions and does not clearly separate verified facts from inference, and test coverage is limited, score 1.

Risks and how to mitigate them
  • Evaluation creates cloud VMs, downloads images, and calls external model APIs, which can incur significant cost; README warns but provides no automatic cleanup or cost cap.
  • Sample scripts require sensitive credentials such as AWS keys, model API keys, and GitLab private tokens; no secret-management or redaction scheme is provided.
  • The dependency set is large and mostly unpinned, creating supply-chain and reproducibility risk; use a lockfile and vulnerability scanning.
  • Task classes and complete assets are distributed through gated Hugging Face datasets requiring manual approval, which may be unavailable in offline or restricted networks.
  • Some tasks require sudo inside the VM and default credentials are public; isolate network and credentials if running in shared or production environments.
  • main is development code and README recommends release tags; mixing versions makes evaluation results incomparable.
Evidence confidence: Low Reviewed Oct 08, 2026 Reviewed revision acdd3493808e
See the full review method →

FAQ

Does it require paid cloud hosts, or can I run it locally?
Official images cover Docker and AWS. The Docker path needs a Linux host with KVM, so it can run on your own server; AWS suits large parallel evaluation or training. VMware, Azure, GCP, Aliyun and Volcengine are code paths only and require converting 1.0 images per the migration doc.
Can the evaluated agent just look up the answers online?
Task classes come from the gated xlangai/osworld_v2_tasks dataset and complete assets from xlangai/osworld_v2_assets_gated; the public xlangai/osworld_v2_assets repository only holds files browser-facing task state must reach anonymously. You download the snapshot locally and point OSWORLD_FILE_BASE_URL at it so tasks do not fall back to online asset resolution.
What will a run cost me?
The code is Apache-2.0 and free to use. Actual spend comes from the cloud VMs (AWS or your own Docker host) and the model API calls you evaluate; the repository states no pricing or quota. Cost scales with parallel environment count, task steps and model choice.
How do I get my agent onto the verified leaderboard?
Public evaluation requires scheduling a meeting with the maintainers so they run your agent code on their side. You upload your implementation for disclosure under the OSWorld framework (you may keep your model API private), or, as a trusted institution, share monitoring data and trajectories. Follow docs/PUBLIC_EVALUATION_GUIDELINE_v2.1.md.
Will my OSWorld 1.0 agent and experiments carry over?
Not directly. 2.1 expects code, task classes, assets, websites and images from the same release, and dependencies changed (for example imageio-ffmpeg>=0.6.0). docs/MIGRATING_v2.1_FROM_OSWORLD_V1.md and the migrate-osworld-agent skill cover dependency setup, task conversion, provider reuse, websites and GitLab, agent migration, and result comparability.
View on GitHub ↗ Install ↓

Compare agents like this one

The same FARS review applied across the shortlist this agent qualifies for.

Related agents