Agent Skills
On-demand skills for coding agents — UI audits, typography, docs, PR review, and releases — so you don't ship AI slop.
- Source repo
- mblode/agent-skills
- Stars
- ★ 133
- Last updated
- today
- License
- MIT
- Primary language
- Python
- FA score
- 42/100 · Major gaps
At a glance
- How it runs
- Works with
- Universal · cross-platformCodex · Claude Code
- Cost
- Free software; you pay for model usage
- Setup effort
- Low · running in minutes
- You'll need
- Typical use
- A frontend engineer runs ui-verification before merging a UI change, using nine headless-browser probes to replace subjective judgement with measured defects.
- Not a fit if
- Teams that use neither Claude Code nor Codex
- Users who want a pure GUI with no CLI or config
- Source review
- 42/100 · Major gaps
What does this agent do, and when should you use it?
Agent Skills is a collection of on-demand skills for coding agents, maintained by Matthew Blode (blode.co) and published under the MIT license. Each skill lives as its own SKILL.md file and is installed into Codex and Claude Code with a single npx skills add command. The skills are grouped into six areas — Architecture, Design, Writing, Quality, Shipping, and Authoring — and include codebase-architecture, scaffold-nextjs, ui-design, ui-verification, typography-audit, tidy, pr-creator, pr-babysitter, autoship, agents-md, and more. Rather than running as a service, the repository ships instructions and rule sets that the agent loads when a matching task appears. Distribution goes through skills.sh, and every skill is linked to its own file in the repository.
The repository does not contain a standalone executable; it contains skill definition files such as skills/<name>/SKILL.md that tell an agent what steps and rules to follow for a given task. Skills are installed with npx skills add mblode/agent-skills -g --agent codex claude-code -y and then invoked by the agent during a session. Documented operations include: designing or hardening a code structure with codebase-architecture; scaffolding a Next.js turborepo with Blode UI, Ultracite and Vercel via scaffold-nextjs; booting an app in a headless browser and running nine probes to turn inferred UI defects into measured ones with ui-verification; checking punctuation, fonts, sizing, spacing, hierarchy and pairing against 78 rules with typography-audit; reviewing a diff or PR with file:line findings in confirmed and plausible tiers via tidy; creating, watching and releasing PRs with pr-creator, pr-babysitter and autoship; and wiring Claude Code, Codex and Cursor to a single set of instructions with agents-md.
- A frontend engineer runs ui-verification before merging a UI change, using nine headless-browser probes to replace subjective judgement with measured defects.
- A team maintaining an npm library runs dx-audit to score its library, CLI or SDK against 38 agent-friendliness rules.
- A writer or docs maintainer runs typography-audit to check punctuation, sizing and hierarchy against 78 rules.
- A developer uses agents-md to make Claude Code, Codex and Cursor read one shared AGENTS.md instead of drifting configs.
- A maintainer uses autoship to run the changesets-based npm release flow, covering the fix loop, CI watch and OIDC publish verification.
- A reviewer runs tidy on a PR to get file:line findings, optionally applying fixes and simplifying the diff.
How do you install or deploy this agent?
Prerequisites: Node.js with npx, a network connection, and an installed Codex or Claude Code agent. The README gives exactly one install command:
npx skills add mblode/agent-skills -g --agent codex claude-code -yThis installs globally through the skills.sh CLI and targets both the codex and claude-code agents. The README does not document alternative install paths (manual clone, offline install), required API keys, or permission scopes.
How do you use this agent?
After installation the skills live in the repository as SKILL.md files — the README links every skill to its own skills/<name>/SKILL.md — and the agent loads them on demand when a matching task appears. The runnable invocation shown in the README is the install command itself:
npx skills add mblode/agent-skills -g --agent codex claude-code -ySome skills are documented as explicitly named calls, for example test-audit (for a repo-wide or subsystem-wide pruning campaign) and app-verification ("Call by name"). The README does not show interactive invocation examples or per-skill parameters.
What are this agent's strengths and limitations?
- Skills are finely scoped with checkable rule counts — nine probes in ui-verification, 78 rules in typography-audit, 38 rules in dx-audit.
- One npx skills add command installs across two named agents, codex and claude-code, rather than locking you to a single runtime.
- Setup is a single command plus whatever the underlying agent already requires — no self-hosted service to stand up.
- Coverage spans the whole loop from scaffolding (scaffold-nextjs, scaffold-cli) to npm release (autoship), not just one stage.
- MIT licensed, with skills published as readable Markdown files that can be reviewed before use.
- The repository is skill definitions only; behaviour depends entirely on the host agent (Codex or Claude Code) and nothing runs standalone.
- Installation relies on the skills.sh CLI and network access, with no documented offline or manual path.
- The README does not state required API keys, permission scopes or runtime versions, so first-run cost is not fully predictable from the docs.
- The skill set is broad, but there are no shared invocation parameters or detailed per-skill usage examples.
- Some skills, such as test-audit and app-verification, must be explicitly named and are easy to overlook.
How does this agent compare with similar options?
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Agent Skills This agent | 42 · Major gaps | CLIFree + model costs | ★ 133 | today | Python | Codex · Claude Code |
| Project Cairn | 67 · Some gaps | Agent plugin / skillFree | ★ 234 | 25d ago | Shell | Codex · Claude Code |
| Loop Library: Reusable AI Agent Workflows | 61 · Some gaps | Agent plugin / skillFree | ★ 3.2k | 2mo ago | JavaScript | Codex · Claude Code |
| Caliber — AI Agent Config Sync | 58 · Major gaps | CLIFree + model costs | ★ 1.3k | 5d ago | TypeScript | Codex · Claude Code |
How does FollowAgents rate this agent?
Why each dimension lost points
Evidence is limited to README, LICENSE, and two GitHub Actions workflows; no SKILL.md bodies are provided, so runtime permission scope, confirmation flows, data-flow disclosure, and sensitive-data handling cannot be verified. least_privilege scores 1: workflows explicitly declare permissions: contents: read, a good least-privilege signal, but this covers only CI, not the skills themselves. user_confirmation scores 1: README mentions tidy is report-only by default and autoship has a release flow, hinting at confirmation steps, but no implementation evidence. data_flow_transparency scores 1: rebuild-web.yml comments explain the Vercel deploy hook and CDN cache behavior in detail, which is partial transparency, but skill-level data flows are undisclosed. sensitive_data_handling scores 1: validator tests show a no-absolute-paths check, indicating awareness of path leakage, but no sensitive-data classification or redaction policy. dependency_security scores 1: workflows use mainstream actions (actions/checkout@v4, ruby/setup-ruby@v1, actions/setup-python@v5) but do not pin SHAs, and skill dependencies (npx skills, Ruby, Python packages) are neither listed nor locked. external_effects scores 1: autoship publishes to npm and pr-babysitter modifies PRs, which are external side effects, but no rollback or blast-radius description. rollback scores 1: workflows have concurrency and retries, but no skill-level rollback mechanism is described. source_attribution scores 2: LICENSE clearly attributes Matthew Blode and README names the author and source, so attribution is clear; however, the publisher is not verified by FollowAgents, so no higher score is given.
self_consistency scores 2: tests in maintenance/tests assert consistency between the ghostwriter Surfaces table and references/surfaces.md and exercise the validator CLI contract, showing internal consistency checks; however, the test files have structural issues (test_validator.py has test methods after the if __name__ block that will not execute), weakening reliability. dependency_availability scores 1: workflows depend on ruby 3.3, python 3.12, npx skills, and other external tools, but no version locking or availability guarantees are provided. failure_messages scores 2: rebuild-web.yml emits clear errors and exit codes for missing secrets, HTTP 403/429, and timeouts, and validator tests verify FAIL row formats, so failure messaging is reasonably clear.
audience_and_scenarios scores 2: README categorizes 20+ skills under Architecture, Design, Writing, Quality, Shipping, and Authoring, covering broad scenarios for coding agents and human users. capability_boundaries scores 1: boundaries can only be inferred from one-line descriptions; no SKILL.md body states what each skill does not do or when it is inapplicable. trigger_precision scores 1: README notes app-verification is called by name and test-audit is invoked explicitly, hinting at some trigger conditions, but there is no systematic trigger-word or exclusion documentation. environment_fit scores 1: the install command targets codex and claude-code, but other agent environments, operating systems, and Node version requirements are unspecified.
information_architecture scores 2: README is well structured, organizing skills by category with links to each SKILL.md, and clearly separates install and license sections. install_notes scores 2: provides a concrete npx skills add command with -g, --agent, and -y flags, making installation actionable. naming_stability scores 1: skill names follow a consistent kebab-case style, but there is no rename policy or deprecated-name record (tests reference retired-names.tsv, but the file is not provided). examples_and_faq scores 1: README lists only one-line descriptions with no usage examples or FAQ. known_limitations scores 1: rebuild-web.yml comments disclose operational limits such as CDN caching and 403 rate limiting, but skill-level known limitations are not stated. license scores 3: LICENSE.md contains the full MIT text, and README and badge both mark MIT, with clear copyright year and author. versioning_changelog scores 0: no CHANGELOG, no version numbers, no release history. maintenance_responsibility scores 1: README names the author and links to blode.co, but there is no maintenance commitment, issue-response expectation, or security contact.
output_usability scores 1: skill output formats and usability cannot be judged from the available files; README only describes functionality. marginal_value scores 1: skills cover UI audits, typography, docs, PR review, and releases, which appear useful, but there is no comparison with existing tools or evidence of effect. cost_benefit scores 1: installation is a single npx command, so cost is low; however, runtime token consumption, time cost, and benefit are unsupported by data.
claim_traceability scores 1: README claims 78 typography rules, 27 ax rules, 38 dx rules, etc., but provides no rule list or verifiable source. cross_source_corroboration scores 1: README and test files partially corroborate each other on the ghostwriter Surfaces table and validator contract, but most skill claims have no second source. fact_inference_separation scores 1: README mixes functional descriptions with marketing copy (e.g., 'Nobody ships AI slop on purpose') without separating fact from assertion.
- No SKILL.md bodies are provided, so actual permissions, confirmation flows, and data flows of each skill cannot be verified; this assessment relies only on README, LICENSE, and two workflows.
- Publisher identity is not verified by FollowAgents; confirm source trustworthiness before installation.
- GitHub Actions in workflows are not pinned to commit SHAs, creating supply-chain risk; skill dependencies (npx skills, Ruby, Python packages) are not version-locked.
- autoship publishes to npm and pr-babysitter modifies PRs, which are external side effects, but the repository does not describe rollback or blast radius.
- test_validator.py has test methods after the if __name__ block that will not execute, potentially creating a false impression of coverage.
- Rule counts in README (78, 27, 38, etc.) have no rule list or verifiable source and remain unverified claims.