Agents tagged "llm-benchmarking"

2 results
llm-benchmarking ×
Data & Analysis

EDSL (Expected Parrot Domain-Specific Language)

A Python DSL for designing, running, and analyzing AI-powered surveys and experiments with large numbers of AI agents and LLMs, simulating social science and market research.

★ 502 FS 63 Some gaps today MIT
Dev & Engineering

PinchBench Agent Benchmark

53 real-world tasks that measure how well an LLM actually performs as an OpenClaw coding agent.

★ 1.4k FS 42 Major gaps 3mo ago MIT