Better Harness
Diagnose and improve coding-agent workflows with verifiable evidence.
The evidence shows strong least-privilege and external-effect controls: CI defaults to contents:read, Pages permissions are isolated by job, the Inspector is described as read-only, preview binds to 127.0.0.1, foreign origins are rejected, artifact compilation blocks package imports and directory escapes, and lifecycle commands produce plans without executing them. Installation paths, local-session inputs, report locations, and coverage gaps are also disclosed. Deductions apply because no comprehensive policy for detecting, redacting, retaining, or deleting sensitive data is shown; confirmation is present for reviewed repairs and lifecycle plans but not as a uniform protocol for every report write or local-session read; and rollback is supported mainly by removal-plan and recovery-path references rather than a concrete, evidenced restoration procedure. Attribution is well supported by the MIT file, package author, repository, and issue metadata; the unverified publisher is treated only as unknown.
The README, package metadata, CI workflow, and tests are internally consistent about supported platforms, Node versions, output forms, error states, and safety boundaries. Tests cover streamed-event reconstruction, preservation of partial results, interrupted and failed states, concurrent build reuse, access restrictions, and actionable diagnostics, justifying full credit for failure messages. Dependency availability is reduced because the Node and npm ranges are narrow, several capabilities rely on changing third-party host contracts, and Cursor, Pi, and WorkBuddy retain unavailable, manual, or partial paths; static evidence also cannot establish actual availability in every target environment.
The material addresses teams, contributors, and many coding-agent hosts with distinct scenarios, adapter matrices, output modes, a no-session path, and host-specific installation instructions. Capability boundaries clearly distinguish configured mechanisms, evidence of use, and evidence of improved outcomes, while partial or missing observations remain visible. Environment coverage includes Windows, macOS, Linux, multiple Node lines, and numerous hosts. Trigger precision is reduced because exact invocation examples are provided, but the supplied files do not fully specify selection, ambiguity handling, or refusal behavior for the core natural-language analysis instruction.
The README organizes quick start, architecture, installation, development, contribution, and licensing well, with routes to models, adapter matrices, references, and contribution guidance. Installation notes include host differences, verification commands, output locations, and unavailable paths; public package, command, and versioned event names appear stable. Limitations are especially explicit, including causal limits, session coverage, the Cursor contract, missing Copilot data, and local-preview boundaries. The MIT text agrees with package metadata. Deductions apply because examples are extensive but there is no explicit FAQ; CHANGELOG is referenced by packaging and generation scripts but its contents are absent, preventing assessment of release-note quality; and author, email, issue, and contribution routes exist without named maintainers, a support commitment, or a clear division of update responsibility.
Self-contained HTML, paired Markdown, findings.json, evidence briefs, prioritized findings, impacts, repair boundaries, and acceptance checks make the output directly usable, with host-specific differences explained. The five-part work-loop analysis plausibly adds value beyond reviewing only the final diff. Marginal value is reduced because the benefit is primarily supported by the project's own explanation rather than an independent comparison or quantified outcome. Cost-benefit is also reduced because no runtime, model-consumption, false-positive, or maintenance-cost measurements are supplied, while strict runtime requirements and multi-host setup create adoption overhead.
Major claims are traceable to named models, workflows, report structures, adapter matrices, scripts, and tests. Key README claims about operation and boundaries are corroborated across package metadata, CI, and test files. The material repeatedly separates the existence of a mechanism, task-linked evidence, missing observation, and demonstrated improvement, and explicitly says historical trends are not causal proof. Full marks here mean the supplied static sources handle these documentary criteria thoroughly; they do not imply execution or independent testing.
- This is a low-confidence static review; no code, tests, installation flow, or host integration was executed.
- The tool may read workspace-matched local agent sessions and write reports by default; confirm first whether those records contain secrets, customer data, personal information, or confidential prompts.
- Studio origin and directory boundaries have test evidence, but the supplied material does not show an end-to-end policy for sensitive-data redaction, retention, or secure deletion.
- Cursor, Pi, WorkBuddy, and some session sources have unavailable, manual, or partial boundaries; the host list should not be interpreted as uniform support.
- Dependencies and GitHub Actions use explicit versions or tags, but there is no supplied evidence of commit-digest pinning, vulnerability scanning, an SBOM, or a dependency-update policy.
- Improvement trends and the project's own value claims are not causal proof; finding quality, false-positive rates, costs, and recovery procedures should be validated separately on a real repository.
What does this agent do, and when should you use it?
Better Harness is an open-source Harness Engineering platform that turns project and supported session evidence into prioritized improvements for coding-agent workflows. It evaluates the Agent Work Loop across Task Understanding, Controlled Execution, Change Validation, Reliable Delivery, and Learning Capture. The repository includes the `/better-harness` workflow, evidence collectors, analyzers, report renderers, Harness Inspector, and thin adapters for multiple coding-agent hosts. Depending on the host, it produces a native Canvas report or a self-contained `report.html` paired with `report.md` and `findings.json`. Its analysis is bounded to relevant Task Episodes and surrounding project mechanisms, with missing or partial evidence shown explicitly instead of converted into unsupported scores. It is a fit for teams that need to assess the delivery system around agent-written code, not merely review the final diff.
When /better-harness runs, it gathers project evidence such as AGENTS.md, specifications, Skills, acceptance criteria, tests, linters, Hooks, and delivery controls; where supported, it also reads workspace-matched session records. Three evidence domains remain independent until a lead agent performs unified analysis. The workflow establishes a task-bounded baseline across the five Agent Work Loop dimensions, detects configured agent assets, and produces prioritized findings with evidence, impact, expected output, repair boundaries, and acceptance checks. Claude Code, Codex, Qwen Code, GitHub Copilot, and Kimi Code produce self-contained HTML with paired Markdown, while Qoder and Cursor use host-native Canvas reports. The read-only Harness Inspector traces product intent through agent activity, sessions, files, and commits, and the history view records movement in the five dimensions without claiming causal improvement. The standalone CLI also exposes better-harness plugin status, doctor, plugin plan, and plugin verify for inspecting and planning host plugin lifecycle changes without executing native installation commands.
- A Claude Code or Codex team seeing repeated misunderstandings can inspect whether specifications,
AGENTS.md, and acceptance criteria gave the agent adequate feedforward guidance. - An engineering lead preparing to expand coding-agent adoption can assess whether tests, linters, Hooks, human review, approvals, and CI/CD form a verifiable delivery path.
- A team operating several coding-agent hosts can compare task reports over time while retaining visible evidence strength, output differences, and coverage gaps.
- A developer facing an unsubstantiated “it works” claim can use Change Validation findings to locate missing tests, diagnostics, or acceptance checks.
- A delivery investigator can use Harness Inspector in a read-only workspace to examine links among intent, prompts, tool calls, sessions, files, and commits.
- A plugin maintainer can inspect installation evidence and prepare host-specific install, update, or removal plans without having Better Harness change host configuration.
What are this agent's strengths and limitations?
- It examines guidance, execution paths, validation sensors, delivery controls, and learning capture instead of limiting review to the final code diff.
- Each finding carries its supporting evidence, impact, expected result, repair boundary, and acceptance route, making remediation reviewable and testable.
- Missing and partial observations remain visible; configured assets are not treated as proof of use, and a current passing check is not treated as proof of longitudinal improvement.
- It documents adapters for Claude Code, Codex, Qoder, Cursor, GitHub Copilot CLI, Qwen Code, Pi, Kimi Code, WorkBuddy, and Grok.
- It offers self-contained HTML, Markdown, structured JSON, native Canvas output, and a read-only inspector for different review needs.
- Installation, invocation, report formats, and session-evidence coverage vary by host; there is no single universal entrypoint.
- The Cursor plugin is not marketplace-published, and its installation plan is currently reported as unavailable while the native command contract is reconciled.
- Some hosts have explicit evidence gaps: GitHub Copilot CLI records no per-response token usage, and VS Code Copilot Chat has no supported durable transcript.
- Historical reports show recorded trends but cannot independently establish that a specific intervention caused improvement.
- Source development has a narrow runtime window: Node.js 22.20 through 24.x and npm 10.9.3 through 11.x.
How do you install or deploy this agent?
No credentials are documented as required. For Codex Desktop, open Settings > Plugins, choose + Add > From Marketplace, enter https://github.com/QoderAI/better-harness.git, set the Git ref to main, leave Sparse paths empty, and install Better Harness from the added marketplace. For Codex CLI, run codex plugin marketplace add 'https://github.com/QoderAI/better-harness.git' --ref main, followed by codex plugin list --marketplace better-harness and codex plugin add better-harness@better-harness. For Claude Code, run /plugin marketplace add QoderAI/better-harness, then /plugin install better-harness@better-harness, and verify that claude plugin details better-harness@better-harness lists Skills (1) better-harness. Qoder Desktop includes Better Harness and needs no marketplace installation. Start a new session or task after installing or updating a plugin. Source development requires Node.js >=22.20.0 <25.0.0 and npm >=10.9.3 <12.0.0; then run npm ci, npm test, and npm run pack:verify.
How do you use this agent?
Start a new host session in the repository to analyze. In Codex Desktop, invoke @better-harness analyze this project's AI coding workflow and generate an evidence-backed report. In Codex CLI, invoke $better-harness:better-harness analyze this project's AI coding workflow and generate an evidence-backed report. In Claude Code, Qoder Desktop, or Qoder CLI, invoke /better-harness analyze this project's AI coding workflow and generate an evidence-backed report. Claude Code defaults to writing report.html, report.md, and findings.json beneath the repository's .claude/better-harness report root; requesting inline or no-files output keeps the result in chat. From a source checkout, run node scripts/better-harness.mjs report --no-sessions to inspect repository evidence without reading local sessions. For lifecycle inspection, use better-harness plugin status --host all and better-harness doctor --platform all.
FAQ
Does Better Harness require a paid account or API key?
Will it modify my project or execute plugin installation plans automatically?
Can it produce a report without session history?
node scripts/better-harness.mjs report --no-sessions to analyze repository evidence only.Does one passing report prove the workflow improved?
Do all supported hosts receive the same output?
findings.json.