Agents tagged "evaluation"

13 results
evaluation ×
Dev & Engineering ✓ NVIDIA · Official

NVIDIA NeMo Agent Toolkit

Adds intelligence to AI agents across any framework, enhancing speed, accuracy, and decision-making through enterprise-grade instrumentation, observability, and continuous learning.

★ 2.6k FS 65 Some gaps today Apache-2.0
Dev & Engineering

Yao Meta Skill

Engineering, evaluating, governing, and packaging repeatable workflows as portable agent skills.

★ 2.6k FS 65 Some gaps 1mo ago MIT
Automation & Ops

Arize Phoenix

Open-source AI observability platform for tracing, evaluating, and troubleshooting LLM applications.

★ 12k FS 61 Some gaps today NOASSERTION
Dev & Engineering

ControlKeel

Turns engineering habits into enforceable policy gates, persistent evidence, and reusable memory for coding agents.

★ 11 FS 59 Major gaps today NOASSERTION
Dev & Engineering ✓ Microsoft · Official

RAMPART

A pytest-native safety and security testing framework that brings adversarial and benign-failure coverage to agentic AI apps in your existing test workflow.

★ 413 FS 51 Major gaps 7d ago MIT
Dev & Engineering

Agent Skills for Context Engineering

A comprehensive open collection of skills for context engineering, multi-agent architectures, and production agent systems.

★ 18k FS 49 Major gaps 13d ago MIT
Automation & Ops

RagaAI Catalyst

A Python SDK for evaluating, tracing, debugging, and safety-testing LLM and multi-agent applications.

★ 16k FS 46 Major gaps 7mo ago Apache-2.0
Dev & Engineering

AI Agents — The Definitive Guide (Companion Code)

Official companion repository for the O'Reilly book 'AI Agents - The Definitive Guide', with twelve chapters of runnable Jupyter notebooks covering everything from LLMs to production-grade agents.

★ 2.4k FS 45 Major gaps 1mo ago
Dev & Engineering

Rogue — AI Agent Evaluator & Red Team Platform

Stress-test your AI agents before attackers do, with automated evaluation and red teaming.

★ 1.1k FS 43 Major gaps 4mo ago NOASSERTION
Dev & Engineering

Judgeval

The continuous-improvement stack for agents — detect failures, triage root causes, and ship fixes backed by production data.

★ 1.1k FS 40 Major gaps 1d ago Apache-2.0
Dev & Engineering

IntellAgent: Simulate, Analyze, and Optimize Conversational Agents

Uncover your agent's blind spots through realistic synthetic interactions, diagnose performance gaps, and optimize for reliable deployment.

★ 1.3k FS 33 Major gaps 10d ago Apache-2.0
Dev & Engineering

AI System Design Guide

A living reference for production AI systems, covering RAG, LLM engineering, agentic AI, and interview prep.

★ 3.3k FS 27 Major gaps 1mo ago MIT
Automation & Ops

BISHENG Enterprise LLM DevOps Platform

An open-source platform for building and operating enterprise-grade LLM applications, with workflow orchestration, RAG, agents, model management, and fine-tuning.

★ 12k FS 0 Major gaps 1d ago Apache-2.0