Sprix SAGE Router
Chooses whether an in-flight A2A task should continue, recruit collaborators, or transfer ownership.
LongHorizon-Harness
A loop-engineering system that lets Claude Code, Codex, OpenCode, and DeepSeek Harness agents run for hours across desktop apps and the CLI with verified progress.
HUD
A platform for building RL environments and evals for AI agents — define an environment once, then evaluate and train any model on it.
OpenRCA Software Failure Analyst
Benchmarks and diagnoses software failures across metrics, traces, and logs.
Xcode Build Optimization Agent Skill
Speed up Xcode clean and incremental builds through benchmarks and build-setting optimizations.
RAI — Embodied AI Agent Framework for Robotics
A vendor-agnostic agent framework that lets developers build multimodal AI capabilities for physical robots on top of ROS 2.
MobileGym
A programmable mobile simulator for deterministic GUI-agent evaluation and online RL training.
PhyAgentOS — Session-Centered Runtime for Embodied Intelligence
Cognitive-physical decoupling with a session-centered runtime: one codebase, any hardware, with multi-layer safety and full auditability.
Multi-SWE-bench
A Docker-based benchmark for validating issue-resolution patches across programming languages.
Memora Agent Memory
A structured memory layer that stores rich agent history and retrieves it through abstractions and semantic cues.
ReLE Chinese LLM Benchmark & Defect Library
Continuously updated Chinese LLM evaluation across 7 domains, 300+ dimensions, with leaderboards and a defect library of over 2 million cases.
Cua: Cross-OS Computer Use and Virtualization Platform
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks.