LintLang
Deterministic static analysis for agent configs, tool descriptions and system prompts — catch vague tools and missing stop conditions before runtime.
- Source repo
- hermes-labs-ai/lintlang
- Stars
- ★ 125
- Last updated
- today
- License
- Apache-2.0
- Primary language
- Python
- FA score
- 67/100 · Some gaps
At a glance
- How it runs
- Works with
- Universal · cross-platformClaude CodeCodex (Partial support)
- Cost
- Free, no paid service needed
- Setup effort
- Low · running in minutes
- You'll need
- Typical use
- Teams maintaining an MCP server or function-tool set who want to confirm before release that sibling tool descriptions give the model a defensible choice.
- Not a fit if
- Teams wanting an LLM to rewrite or fix prompts, not deterministic checks
- Teams that need runtime proof the agent picks the right tool
- Teams needing tools from separate files merged into one selection namespace
- Source review
- 67/100 · Some gaps
What does this agent do, and when should you use it?
LintLang is a local, deterministic static linter from Hermes Labs aimed at the instructions and tool interfaces an AI agent is given. You point it at a project directory and it finds agent-facing content inside files you already keep: MCP and function-tool definitions nested in JSON/YAML, parameter schemas, system prompts, messages, output contracts, AGENTS.md, CLAUDE.md, GEMINI.md, SKILL.md, and supported Python prompt code. It uses no models at all, so results are reproducible and safe to gate in CI. The checks cover sibling tools that overlap without a clear reason to choose one, retries or loops with no stopping or progress condition, schemas missing required fields, prompts that name more than one output format, long instruction lists with no stated priority, SKILL.md metadata defects, stale context and broken tool-message sequences. Every result states what was inspected, and content with no recognized agent-facing structure is reported as SKIPPED rather than PASS. Reports are available as text, JSON, SARIF and GitLab Code Quality.
Running lintlang scan <path> walks the target directory, recognizes supported agent-facing carriers — MCP and function-tool definitions in JSON/YAML, parameter schemas, system prompts and messages, output contracts, AGENTS.md / CLAUDE.md / GEMINI.md / SKILL.md, and supported Python prompt code — and parses them. It then applies a fixed rule set: the H1.x family compares sibling tools inside one parsed input to expose overlapping choices; H5 flags long instruction lists with no priority order; H6 flags prompts naming two or more recognized output formats even when each is scoped to a case; further checks look for retries, loops or tool use with no explicit stopping or progress condition, schemas missing required fields or unclear parameters, SKILL.md with missing or invalid metadata or a name that does not match its directory, stale project references, unbounded persistence, malformed roles and broken tool-message sequences, plus literal tool definitions and selected pipeline thresholds inside supported Python prompts. The CLI exposes --fail-on fail (gate HIGH/CRITICAL only) and --fail-on review (include MEDIUM); --write-baseline records a reviewed .lintlang-baseline.json and --baseline then gates only new or changed findings; lintlang init --github --path . generates a pinned GitHub Actions workflow that gates HIGH or CRITICAL by default. Output formats are text, JSON, SARIF and GitLab Code Quality.
- Teams maintaining an MCP server or function-tool set who want to confirm before release that sibling tool descriptions give the model a defensible choice.
- Engineering teams that keep agent configs, prompts and skills in a repo and want PR-time blocking of high-risk configuration defects via GitHub Actions.
- Developers already using Claude Code, Cursor or Gemini CLI who want lintlang scan wired into a pre-commit hook for local checking.
- Reviewers of CLAUDE.md, SKILL.md or AGENTS.md who need a deterministic checklist for metadata, usage criteria and naming conventions.
- Projects with a large backlog of existing findings that want to freeze them in .lintlang-baseline.json and gate only newly introduced ones.
- Security and platform engineers who need results surfaced in GitHub Code Scanning or GitLab Code Quality dashboards through SARIF.
How do you install or deploy this agent?
Python 3.10 or newer is required. To try it once without installing:
uvx lintlang scan .Install with pip:
pip install lintlang
lintlang scan .On macOS you can also use Homebrew:
brew install hermes-labs-ai/tap/lintlang
lintlang scan .How do you use this agent?
Point the scanner at a specific configuration source instead of the whole directory:
uvx lintlang scan AGENTS.md
uvx lintlang scan SKILL.md
uvx lintlang scan agent.yamlFindings are advisory by default; to gate on HIGH or CRITICAL:
lintlang scan . --fail-on failInclude MEDIUM findings in the gate:
lintlang scan . --fail-on reviewGenerate a pinned GitHub Actions workflow:
lintlang init --github --path .For an existing repository with known findings, record a baseline and then gate only new or changed findings:
lintlang scan . --write-baseline .lintlang-baseline.jsonlintlang scan . \
--baseline .lintlang-baseline.json \
--fail-on reviewWhat are this agent's strengths and limitations?
- Fully local and model-free: no LLM calls, no API keys, reproducible output, and no need to ship prompts or configs to an external service.
- The detectors target agent-specific failure modes — overlapping sibling tools, missing stopping conditions, schema gaps, mixed output formats, SKILL.md metadata errors — that generic YAML/JSON linters do not cover.
- Honest reporting semantics: unrecognized content is SKIPPED rather than PASS, and every finding states what was inspected, so a clean scan is not overstated.
- Native SARIF, JSON and GitLab Code Quality output make it droppable into GitHub Code Scanning, GitLab CI or MegaLinter, and the baseline mechanism lets legacy repos adopt gradually.
- Low-friction entry: uvx runs it without installation, with pip and Homebrew alternatives, plus documented integrations for Claude Code, Cursor, GitHub Copilot CLI, Gemini CLI and pre-commit.
- No semantic contradiction detection: the project states it will not catch 'always do X' sitting next to 'never do X'; H5 and H6 only cover missing priority and output-format count.
- No runtime verification: it does not run models or observe actual tool choices, so a clean scan only means the selected static checks found no covered defects in recognized content.
- Tool comparison is scoped to one parsed input; a directory scan does not merge tools from separate files into one selection namespace, so cross-file conflicts need separate handling.
- Coverage depends on recognition rules: unsupported structures are skipped rather than flagged, so unusual configuration styles can look clean without being checked.
- Requires a Python 3.10+ runtime, and the docs only give a native macOS install path via Homebrew — Linux and Windows users need pip or uvx.
How does this agent compare with similar options?
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| LintLang This agent | 67 · Some gaps | CLIFree | ★ 125 | today | Python | Claude Code |
| AgentSys | 52 · Major gaps | CLIFree + model costs | ★ 990 | 1d ago | JavaScript | Codex · Claude Code |
| Argot Repository Style Analyzer | 93 · Excellent | CLIFree | ★ 50 | 27d ago | Rust | Claude Code |
| Jev Review | 81 · Good | Agent plugin / skillFree + model costs | ★ 224 | 12d ago | TypeScript | Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
Evidence shows a local, zero-LLM static linter that scans directories read-only and does not run models or observe runtime tool choices, so the permission surface is small (least_privilege=2). However, the README does not state whether scanned content (which may contain secrets in prompts/configs) is read, cached, or uploaded, and there is no data-flow or privacy note (data_flow_transparency=2 only because 'local, zero-LLM' is explicitly claimed; sensitive_data_handling=1 for the missing handling statement). On confirmation, findings are advisory by default, but --fail-on fail/review makes CI fail directly, an external consequence with no interactive confirmation path (user_confirmation=1). Dependencies are limited to pyyaml>=6.0.3 with bounded dev deps, and CI pins GitHub Actions by commit SHA, which is positive (dependency_security=2); no SBOM, audit, or vulnerability-response detail is provided. External effects are limited to writing SARIF/JSON reports and baseline files, with no destructive defaults (external_effects=2). For rollback, --write-baseline and baseline gating provide a form of change control, but there is no documented way to undo generated reports or config (rollback=1). Source attribution is clear (Hermes Labs, Apache-2.0, SECURITY.md contact), though the publisher is unverified (source_attribution=2).
Self-consistency: README's stated boundaries (no semantic-contradiction detection, SKIPPED is not PASS, tool comparison is within one parsed input) align with the pyproject description and the CI self-scan of samples (self_consistency=2). Dependency availability: requires-python>=3.10, CI covers 3.10-3.13, the only runtime dependency is pyyaml, and install paths (uvx/pip/brew) are explicit (dependency_availability=2). Failure messages: the devcontainer test asserts an 'Input error' on malformed input, and README distinguishes SKIPPED from PASS, so failure/skip semantics are identifiable (failure_messages=2); no concrete message format or exit-code table is shown.
Audience and scenarios: aimed at developers and CI, with GitHub Actions, GitLab, pre-commit, and multiple CLI integrations (audience_and_scenarios=2). Capability boundaries: README explicitly lists what it does not do (no model execution, no runtime observation, no production-safety guarantee, no semantic-contradiction detection) and explains SKIPPED semantics, which is a strong boundary statement (capability_boundaries=3). Trigger precision: --fail-on fail/review plus baselines give tiered gating, but rules like H5/H6 are described as flagging any two recognized formats, which may be noisy, and no threshold tuning is documented (trigger_precision=2). Environment fit: multiple Python versions, Homebrew, and devcontainer are supported, but Windows or non-POSIX environments are not addressed (environment_fit=2).
Information architecture: README is well structured (capabilities, quickstart, CI, integrations, docs, contributing, license), but llms-full.txt and docs/ content are not in evidence and cannot be verified (information_architecture=2). Install notes: uvx, pip, brew, and lintlang init all have concrete commands (install_notes=3). Naming stability: rule IDs (H1.1, H3, H5, H6, H1.9) appear in README and CI, but no versioned rule-ID stability commitment is given (naming_stability=2). Examples and FAQ: quickstart examples and devcontainer test fixtures exist, but there is no FAQ (examples_and_faq=2). Known limitations: explicitly lists no semantic-contradiction detection, SKIPPED semantics, and tool-comparison scope (known_limitations=3). License: full Apache-2.0 text matches the pyproject declaration (license=3). Versioning and changelog: pyproject is 0.8.1 while CI references v0.7.1, and CHANGELOG.md is linked but not provided (versioning_changelog=2). Maintenance responsibility: SECURITY.md gives response timelines and a contact email, but the publisher is unverified so the commitment cannot be independently confirmed (maintenance_responsibility=2).
Output usability: JSON, SARIF, and GitLab Code Quality outputs are supported, and CI validates SARIF structure and rule IDs, so output is machine-consumable (output_usability=2). Marginal value: it offers targeted checks for agent configs and tool descriptions that general linters do not, which is differentiated, but no comparison with existing tools or false-positive data is given (marginal_value=2). Cost-benefit: zero-LLM, deterministic, single dependency, runnable via uvx, so cost is low; scan latency or large-repo behavior is not quantified (cost_benefit=2).
Claim traceability: README cites merged PR #5656 and Zenodo DOIs, but those external artifacts are not in evidence and cannot be checked (claim_traceability=2). Cross-source corroboration: README, pyproject, CI, and devcontainer tests agree on versions, commands, and rule IDs (cross_source_corroboration=2). Fact-inference separation: README clearly separates static findings from production-safety claims and states SKIPPED is not PASS (fact_inference_separation=2); however, the screenshot is labeled 'stylized' and the finding it depicts cannot be verified from the evidence.
- The README does not explain how scanned prompts and configs that may contain secrets are handled; confirm local read/cache behavior before adoption.
- Findings are advisory by default, but --fail-on fail/review fails CI directly, an external consequence; teams should evaluate false positives in non-blocking mode first.
- Publisher identity is unverified, so the response timelines and maintenance commitments in SECURITY.md cannot be independently confirmed.
- CHANGELOG.md, llms-full.txt, and docs/ content are not in evidence, so version and rule stability cannot be verified; pyproject is 0.8.1 while CI references v0.7.1.
- Rules such as H5/H6 are described as flagging any two recognized formats, which may be noisy; use a baseline to observe real hit rates first.