Agents tagged "benchmark"

5 results
benchmark ×
Dev & Engineering

AppWorld Agent Benchmark Environment

A controllable world of 9 apps and ~100 people for benchmarking function-calling and interactive coding agents with state-based evaluation.

★ 529 FS 56 Major gaps 1mo ago Apache-2.0
Dev & Engineering

OSWorld 2.1 Computer-Use Agent Benchmark

Run long-horizon desktop tasks on real Ubuntu VMs and score computer-use agents reproducibly.

★ 359 FS 53 Major gaps 12d ago Apache-2.0
Dev & Engineering

evo — Autoresearch Orchestrator for Codebases

Turns your codebase into an autoresearch loop that discovers what to measure, sets up the benchmark, and runs tree search with parallel subagents.

★ 1.5k FS 51 Major gaps 4d ago Apache-2.0
Dev & Engineering

Code2Video: Code-Centric Educational Video Generator

Generate high-quality educational videos from knowledge points using executable code as the medium.

★ 2.1k FS 31 Major gaps 1mo ago MIT
Data & Analysis

MemoryOS: A Memory Operating System for Personalized AI Agents

A hierarchically structured memory operating system for personalized AI agents, delivering coherent, context-aware interactions with large gains on the LoCoMo benchmark.

★ 1.6k FS 31 Major gaps 3mo ago Apache-2.0