Agents tagged "leaderboard"

3 results
leaderboard ×
Dev & Engineering

Tokscale – Track token usage across AI coding agents

Monitor token consumption and costs for multiple AI coding agents from your terminal, with a global leaderboard.

★ 5.5k FS 52 Major gaps 5d ago MIT
Dev & Engineering

PinchBench Agent Benchmark

53 real-world tasks that measure how well an LLM actually performs as an OpenClaw coding agent.

★ 1.4k FS 42 Major gaps 3mo ago MIT
Data & Analysis

ReLE Chinese LLM Benchmark & Defect Library

Continuously updated Chinese LLM evaluation across 7 domains, 300+ dimensions, with leaderboards and a defect library of over 2 million cases.

★ 6.5k FS 14 Major gaps 6d ago