Agents tagged "llm-evaluation"

6 results
llm-evaluation ×
Dev & Engineering

LangWatch

Evaluate, test, trace, and monitor LLM applications and AI agents across development and production.

★ 4.8k FS 78 Good 4d ago Apache-2.0
Dev & Engineering

Generative AI with LangChain

Learn to build production-oriented LLM applications and multi-agent systems with Python, LangChain, and LangGraph.

★ 1.4k FS 52 Major gaps 1mo ago MIT
Dev & Engineering

Giskard

Evaluate, red-team, and generate tests for LLM-powered and multi-turn agent systems.

★ 5.8k FS 47 Major gaps 9d ago Apache-2.0
Productivity & Collaboration

Waku

A local-first personal assistant whose loop, memory, tools, and evals you can inspect and modify.

★ 1.8k FS 47 Major gaps 6d ago MIT
Automation & Ops

Laminar

An observability platform for tracing, evaluating, alerting on, and debugging AI agent runs.

★ 3.3k FS 41 Major gaps 10d ago Apache-2.0
Data & Analysis

ReLE Chinese LLM Benchmark & Defect Library

Continuously updated Chinese LLM evaluation across 7 domains, 300+ dimensions, with leaderboards and a defect library of over 2 million cases.

★ 6.4k FS 14 Major gaps 7d ago