Agents tagged "benchmark-evaluation"

7 results
benchmark-evaluation ×
Dev & Engineering

Apache Maka

A local-first agent workspace that performs project work while preserving recoverable execution records.

★ 5.6k FS 82 Good today Apache-2.0
Dev & Engineering

FrontierAgent

Open-source terminal agent runtime: a stateful ReAct agent or a coordinator with parallel sub-agents, working in a sandboxed filesystem for long-horizon research and file-based deliverables.

★ 4.5k FS 74 Some gaps today Apache-2.0
Data & Analysis

Tongyi DeepResearch

An open model and inference system for long-horizon web research, evidence gathering, and research question answering.

★ 20k FS 49 Major gaps 6mo ago Apache-2.0
Dev & Engineering

A-Evolve

Evolves an agent’s prompts, skills, and memory from benchmark feedback.

★ 803 FS 38 Major gaps 1mo ago
Data & Analysis

RecursiveMAS

A recursive latent-state framework for training and evaluating collaborating role-specific models.

★ 938 FS 31 Major gaps 2mo ago MIT
Data & Analysis

MiroThinker

A deployable deep-research agent for complex web research, evidence gathering, and prediction tasks.

★ 8.4k FS 27 Major gaps 6mo ago Apache-2.0
Dev & Engineering

rLLM: Reinforcement Learning Framework for LLM Agents

A unified framework for RL training of language agents, supporting any harness, any sandbox, and one-flag backend switching.

★ 5.8k FS 0 Major gaps 11d ago Apache-2.0