LangWatch
Evaluate, test, trace, and monitor LLM applications and AI agents across development and production.
The sources disclose that the local CLI installs several services, generates local secrets, opens a local endpoint, and runs Langy workers unsandboxed as the user. CI declares read-only contents permission, while the credentialed publishing workflow has target and failure guards. Sensitive-data evidence includes trace secret/PII redaction claims, external Secret handling, residency and self-hosting options, and a detailed vulnerability policy. The license, NOTICE references, and copyright holder provide strong attribution. Deductions: there is no complete least-privilege model or component-by-component data-flow inventory; installation, network access, publishing, and deletion of ~/.langwatch are not shown with interactive confirmation; some security/compliance statements are assertions; versions and action commits are pinned but no dependency-audit result is supplied for this revision; rollback evidence is mainly directory reset, Helm uninstall, and backup facilities rather than granular recovery guidance.
The README, package manifest, Go module, CI workflows, and Helm tests present a coherent product and deployment model. Publishing checks both manifest versions, release tags, and mirrored artifacts; tests cover password preservation, missing Secrets, backup alerts, and deployment modes. Failure paths contain explicit ::error:: messages, FAIL output, and actionable diagnostics, fully supporting failure messaging. Dependency availability is deducted because operation relies on Node, pnpm, uv, databases, containers, Helm/Kubernetes, and a large third-party graph; constraints and frozen-lockfile installation help, but the supplied sources do not establish offline supply, image availability, or comprehensive dependency-outage handling.
The material identifies teams needing regression tests, simulations, evaluations, observability, and governance, and covers cloud, local, Docker Compose, Kubernetes, OnPrem, hybrid, multiple frameworks, and multiple model providers. Feature flags, backup gating, external Secrets, and single/replicated deployment tests show conditional behavior. Capability boundaries are deducted because they are dispersed across setup, security scope, and licensing rather than consolidated into a support matrix. Trigger precision is evidenced for CI path gates and configuration validation, but the complete trigger and tool-action rules of the Agent-facing product are not present.
The README has clear sections for positioning, setup, deployment, integrations, support, contribution, licensing, and security. Local and container commands, endpoint, environment switches, and download sizes are concrete. Publishing enforces consistency among manifests, versions, and tags; licensing and the open-core directory boundary are explicit; security email, issues, Discord, and enterprise support identify maintenance channels. Deductions: examples are chiefly quick-start links and deployment commands rather than a full FAQ; limitations such as unsandboxed workers, resource sizes, and deployment requirements are disclosed but not collected comprehensively; version and release-tag machinery exists, but the actual changelog and a compatibility policy are absent from the supplied files.
The product exposes traces, datasets, evaluations, simulations, alerts, annotations, prompt versions, and governance controls in workflows directly useful for agent testing and operations. Combining observation, evaluation, simulation, and a gateway offers clear marginal value over assembling separate tools. Cost-benefit is deducted because component sizes and budget controls are documented, but the approximately 700 ns overhead claim lacks a supplied benchmark, and total resource requirements and cloud pricing are not shown.
Several claims map to concrete static sources: package metadata supports product identity, workflows support release/version consistency, Helm scripts support Secret, upgrade, backup, and alert behavior, and LICENSE supports licensing claims. The sources provide some cross-corroboration. Deductions arise because major product, security, compliance, no-lock-in, and performance claims appear only in README assertions; linked external documentation is not included and no execution results are supplied. Marketing facts, inference, and independently established conclusions are not systematically separated, particularly for compliance, comprehensive visibility, and performance figures.
- The one-command local setup writes multiple runtimes and data services under ~/.langwatch, opens a local endpoint, and enables Langy workers that run unsandboxed as the current user by default. Review packages and permissions and evaluate it in an isolated environment first.
- Do not treat rm -rf ~/.langwatch as an ordinary recoverable rollback: it removes local configuration, secrets, and data in that directory. Confirm backups and the exact path before use.
- The GDPR, ISO 27001, approximately 700 ns overhead, redaction, and no-lock-in claims are not independently established by the supplied static files. Obtain the relevant reports, benchmarks, and data-flow documentation before procurement or production deployment.
- The dependency graph is large and includes timestamped pseudo-versions. This review did not run vulnerability scans, builds, tests, or provenance/integrity checks.
- The repository uses an open-core split across Apache-2.0 code, MIT SDKs, and commercially licensed enterprise modules. Verify the applicable directory-level license before distribution or production use.
What does this agent do, and when should you use it?
LangWatch is a platform for simulating, evaluating, and monitoring LLM-powered agents before release and in production. It connects OpenTelemetry traces, datasets, offline evaluations, prompt or model optimization, and repeat testing in one workflow, with run review and failure annotation for collaboration. Its scenario runner exercises a full agent stack—including tools, state, a user simulator, and a judge—and identifies failures at the decision level. A separate Go AI Gateway exposes an OpenAI- and Anthropic-compatible proxy with virtual keys, hierarchical budgets, inline guardrails, and provider fallback. Teams can use the hosted service or deploy locally, with Docker Compose, on Kubernetes through Helm, on cloud infrastructure, or in a hybrid data-residency configuration. The platform follows an open-core licensing model: the core is Apache 2.0, designated enterprise modules require a commercial production license, and the SDKs are MIT-licensed.
LangWatch ingests agent and LLM execution traces through its integrations or any OpenTelemetry/OTLP-compatible library. Teams can turn traces into a dataset, run offline evaluation, optimize prompts or models in Optimization Studio, and then re-test. Scenario runs exercise the complete application stack with tools, state, a user simulator, and a judge, producing failure information tied to individual decisions. The platform supports reviewing runs, annotating failures, managing annotation queues, storing prompts in Git through its GitHub integration, and linking prompt versions to traces. The standalone Go service in services/aigateway/ proxies OpenAI- and Anthropic-compatible traffic while applying virtual keys, hierarchical budgets, inline guardrails, automatic provider fallback, and Anthropic cache_control passthrough. LangWatch also provides an MCP integration for Claude Desktop and other MCP clients.
- An agent engineering team runs full-stack scenarios with tools, state, simulated users, and judges before releasing a new agent version.
- A production LLM team collects OpenTelemetry traces, investigates reliability and performance, and converts observed failures into regression datasets.
- A prompt engineering group uses the trace-to-dataset-to-evaluation workflow to compare prompt or model changes and re-test them.
- A platform team operating multiple model providers routes requests through the AI Gateway to enforce virtual keys, budgets, guardrails, and fallback behavior.
- An organization with data-residency requirements deploys LangWatch through Docker Compose, Kubernetes Helm, OnPrem infrastructure, or the hybrid arrangement.
- A quality team invites domain specialists to review runs, annotate failures, and label edge cases through annotation queues.
What are this agent's strengths and limitations?
- It combines traces, datasets, offline evaluation, prompt or model optimization, and repeat testing in a single documented loop.
- Scenario testing covers tools, state, simulated users, and judges rather than evaluating only isolated model responses.
- OpenTelemetry/OTLP foundations and named integrations span multiple frameworks and providers, including OpenAI, Anthropic, Azure OpenAI, Vertex AI, Bedrock, Groq, and Ollama.
- Deployment choices include hosted cloud, a local Node.js launcher, Docker Compose, Kubernetes Helm, OnPrem, and hybrid data residency.
- The AI Gateway adds concrete governance controls such as virtual keys, hierarchical budgets, inline guardrails, and automatic cross-provider fallback.
- The repository is open-core rather than uniformly Apache 2.0: production use of SSO, SCIM, audit logs, gateway webhooks, billing, and back-office modules under platform/app/ee/ requires a commercial license.
- The local launcher installs and operates Postgres, Redis, ClickHouse, the gateway, and additional runtimes under ~/.langwatch/, creating a nontrivial local resource and operations footprint.
- Langy is enabled by default, and its workers run unsandboxed as the current local user, which requires an explicit security assessment.
- Optional evaluators have material storage costs: Presidio adds about 670MB, Lingua about 95MB, and the Langy runtime about 45MB.
- The supplied material lacks a complete SDK setup example, API-key environment variable name, and first trace invocation, so initial instrumentation requires documentation beyond these instructions.
How do you install or deploy this agent?
The shortest local path requires Node.js. Run:
npx @langwatch/serverThe CLI installs uv, Postgres, Redis, ClickHouse, the AI Gateway binary, and the Langy runtime under ~/.langwatch/. It scaffolds ~/.langwatch/.env with locally generated secrets, starts the services in parallel, and opens http://localhost:5560. LANGWATCH_ENABLE_LANGY defaults to true; LANGWATCH_ENABLE_PRESIDIO and LANGWATCH_ENABLE_LINGUA can be enabled when needed. For Docker Compose, use:
git clone https://github.com/langwatch/langwatch.git
cd platform/app
cp platform/app/.env.example platform/app/.env
docker compose up -d --wait --buildThen open http://localhost:5560. For the hosted service, create a free account at https://app.langwatch.ai, create a project, and copy its API key. Kubernetes deployment is available through the documented Helm option.
How do you use this agent?
Start with either the hosted account or a local deployment, create a project, and obtain its API key. Instrument the application with a LangWatch integration or an OpenTelemetry-compatible library; named integrations include LangChain, LangGraph, Vercel AI SDK, Mastra, CrewAI, and Google ADK. After sending traces, create a dataset, run an offline evaluation, adjust prompts or models in Optimization Studio, and repeat the test. To test an agent end to end, configure a Scenario that runs the application with its tools and state against a user simulator and judge, then inspect decision-level failures. For centralized model governance, route compatible OpenAI or Anthropic traffic through the AI Gateway. The supplied material does not include a copyable SDK initialization snippet, the API-key environment variable name, or a first API request, so those integration details cannot be established from this source alone.
How does this agent compare with similar options?
Compared with assembling separate tracing, dataset, evaluation, and prompt-optimization systems, LangWatch explicitly combines them into a trace → dataset → evaluate → optimize → re-test loop. Its OpenTelemetry foundation also allows compatible tracing libraries beyond the directly named framework integrations.