Agents tagged "benchmarking"

12 results
benchmarking ×
Dev & Engineering

Sprix SAGE Router

Chooses whether an in-flight A2A task should continue, recruit collaborators, or transfer ownership.

★ 3.8k FS 69 Some gaps 26d ago MIT
Dev & Engineering

LongHorizon-Harness

A loop-engineering system that lets Claude Code, Codex, OpenCode, and DeepSeek Harness agents run for hours across desktop apps and the CLI with verified progress.

★ 1.6k FS 66 Some gaps 1mo ago MIT
Dev & Engineering

HUD

A platform for building RL environments and evals for AI agents — define an environment once, then evaluate and train any model on it.

★ 302 FS 63 Some gaps 4d ago MIT
Automation & Ops ✓ Microsoft · Official

OpenRCA Software Failure Analyst

Benchmarks and diagnoses software failures across metrics, traces, and logs.

★ 424 FS 58 Major gaps 5mo ago MIT
Dev & Engineering

Xcode Build Optimization Agent Skill

Speed up Xcode clean and incremental builds through benchmarks and build-setting optimizations.

★ 1.2k FS 57 Major gaps 9d ago MIT
Dev & Engineering

RAI — Embodied AI Agent Framework for Robotics

A vendor-agnostic agent framework that lets developers build multimodal AI capabilities for physical robots on top of ROS 2.

★ 592 FS 55 Major gaps 13d ago Apache-2.0
Data & Analysis

MobileGym

A programmable mobile simulator for deterministic GUI-agent evaluation and online RL training.

★ 794 FS 50 Major gaps 26d ago Apache-2.0
Dev & Engineering

PhyAgentOS — Session-Centered Runtime for Embodied Intelligence

Cognitive-physical decoupling with a session-centered runtime: one codebase, any hardware, with multi-layer safety and full auditability.

★ 2.5k FS 47 Major gaps 2d ago MIT
Dev & Engineering

Multi-SWE-bench

A Docker-based benchmark for validating issue-resolution patches across programming languages.

★ 362 FS 46 Major gaps 9mo ago Apache-2.0
Dev & Engineering ✓ Microsoft · Official

Memora Agent Memory

A structured memory layer that stores rich agent history and retrieves it through abstractions and semantic cues.

★ 250 FS 44 Major gaps 3mo ago MIT
Data & Analysis

ReLE Chinese LLM Benchmark & Defect Library

Continuously updated Chinese LLM evaluation across 7 domains, 300+ dimensions, with leaderboards and a defect library of over 2 million cases.

★ 6.4k FS 14 Major gaps 7d ago
Automation & Ops

Cua: Cross-OS Computer Use and Virtualization Platform

Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks.

★ 26k FS 0 Major gaps today MIT