Agents tagged "agent-evaluation"

18 results
agent-evaluation ×
Dev & Engineering

Strands Evals SDK

Evaluate, diagnose, and improve AI agents and LLM applications.

★ 201 FS 64 Some gaps 5d ago Apache-2.0
Dev & Engineering

HUD

A platform for building RL environments and evals for AI agents — define an environment once, then evaluate and train any model on it.

★ 302 FS 63 Some gaps 4d ago MIT
Dev & Engineering

Designing Multi-Agent Systems (PicoAgents)

Build LLM-enabled multi-agent systems from scratch: a book plus a fully runnable teaching framework covering everything from a single agent to autonomous orchestration.

★ 1.3k FS 59 Major gaps 1mo ago Apache-2.0
Dev & Engineering ✓ Google · Official

Google Agents CLI

CLI commands and coding-agent skills for building, evaluating, and deploying ADK agents on Google Cloud.

★ 6k FS 56 Major gaps 1d ago Apache-2.0
Automation & Ops

iFixAi Agent Auditor

Audit AI-agent governance failures and operational risks in about 120 seconds.

★ 16k FS 56 Major gaps 5d ago Apache-2.0
Dev & Engineering

Fable Method – Distilled Agent Workflow

Converts Claude Fable 5's way of working into executable skills, a loop, and a verifier, kept honest by a trap suite.

★ 2.3k FS 55 Major gaps 2mo ago MIT
Dev & Engineering

Hello-Agents

A hands-on curriculum for learning AI-native agents, from core patterns and framework internals to complete multi-agent projects.

★ 81k FS 54 Major gaps 1d ago NOASSERTION
Dev & Engineering

PenguinHarness

Build, run, and iteratively improve agent applications on a desktop or self-hosted server.

★ 2.3k FS 52 Major gaps 5d ago Apache-2.0
Automation & Ops

Future AGI

One platform to trace, evaluate, simulate, protect, and improve production AI applications.

★ 2.1k FS 51 Major gaps 1d ago Apache-2.0
Dev & Engineering

Email Agents From Scratch

Build an email assistant with triage, approval checkpoints, evaluation, and persistent memory.

★ 2.3k FS 49 Major gaps 1mo ago MIT
Data & Analysis

ClawBench Browser Agent Benchmark

Evaluate browser agents on real-world online tasks with recorded, interceptable browser sessions.

★ 835 FS 48 Major gaps 4d ago Apache-2.0
Dev & Engineering

AnyAgent

A unified Python interface for running and evaluating several agent frameworks.

★ 1.2k FS 41 Major gaps 4mo ago Apache-2.0
Dev & Engineering

PandaProbe

An engineering platform for tracing, evaluating, monitoring, and debugging AI agents.

★ 786 FS 41 Major gaps 1d ago Apache-2.0
Dev & Engineering

Strands Agents Samples

Practical Python and TypeScript examples for learning Strands Agents SDK patterns and deployment options.

★ 852 FS 40 Major gaps 22d ago Apache-2.0
Data & Analysis

Meta ARE

Evaluate AI agents on evolving, real-world tasks that demand multi-step reasoning and adaptation.

★ 557 FS 37 Major gaps 29d ago MIT
Dev & Engineering

Youtu-Agent

A configurable framework for building, evaluating, and improving agents with open-weight models.

★ 4.6k FS 27 Major gaps 6mo ago NOASSERTION
Automation & Ops

Coze Loop

A self-hosted platform for AI agent prompt development, evaluation, and execution tracing.

★ 5.7k FS 26 Major gaps 1d ago Apache-2.0
Productivity & Collaboration

KwaiAgents

A lightweight information-seeking agent system runnable with GPT-3.5 or self-hosted KAgent models.

★ 1.2k FS 0 Major gaps 2y ago NOASSERTION