Agent skill

CLI Audit

by crafter-station in crafter-station/skills

Audit an existing CLI against cli-build: how well an agent can operate it and how well a human can read it.

MITAuto-check passed

Install CLI Audit

skills CLI
$ npx skills add crafter-station/skills --skill cli-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install crafter-station/skills cli-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/crafter-station/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cli-audit .claude/skills/cli-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cli-audit
GitHub stars
112
Token cost
~3k tokens
SKILL.md length
1,810 words
Files
4 (incl. scripts, references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Audit an existing CLI against cli-build: how well an agent can operate it and how well a human can read it.

  • Works in 5 steps: Resolve the conditional gates first → Run the deterministic layer → Run the semi layer → …
  • The user wants to audit a CLI
  • SKILL.md covers Why there is no score, 0. Resolve the conditional…, 1. Run the deterministic layer and 1b. Measure coverage, before…, plus 4 more sections
  • Runs Shell scripts from its folder; calls bun

What it does

CLI Audit is an agent skill from crafter-station/skills. Audit an existing CLI against cli-build: how well an agent can operate it and how well a human can read it. Use when the user wants to audit a CLI, check whether a tool is agent-first, find out what a published CLI got wrong, grade a CLI's human output, verify that named safety features are actually wired, or prepare a patch list before improving an existing command-line tool. Reports findings with evidence; it does not patch.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/checks.md`, `scripts/audit.sh` and `scripts/coverage.sh`).

The repository describes itself as: Agent skills extracted from real work. Each one shipped something first. The licence is MIT.

When your agent uses it

  • The user wants to audit a CLI
  • Check whether a tool is agent-first
  • Find out what a published CLI got wrong
  • Grade a CLIs human output

Example prompts

  • “/cli-audit”

Requirements

  • A Bash shell

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Resolve the conditional gates first
  2. Run the deterministic layer
  3. Run the semi layer
  4. Run the judgment layer
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit f0fe474. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CLI Audit loads about 3k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 1,810 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from crafter-station/skills at commit f0fe474, republished under its MIT licence (© crafter-station). 1,810 words, ~3,005 tokens.

Download SKILL.mdSave it as .claude/skills/cli-audit/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
cli-audit
description
Audit an existing CLI against cli-build: how well an agent can operate it and how well a human can read it. Use when the user wants to audit a CLI, check whether a tool is agent-first, find out what a published CLI got wrong, grade a CLI's human output, verify that named safety features are actually wired, or prepare a patch list before improving an existing command-line tool. Reports findings with evidence; it does not patch.
version
0.2.0

cli-audit

Audit a CLI that already exists against the rules in cli-build. The output is a findings list with evidence, not a score.

cli-build builds. This reads what was built and says where it drifted from what it claims.

Why there is no score

The founding case of cli-build's human-output.md is a CLI that passed every mechanical gate and was unreadable: 275 rows for someone who asked what was playing that evening, a percentage that read backwards, one title repeated fourteen times. No gate failed.

A number over that CLI would have read as healthy. That is the exact failure this skill exists to avoid, so it produces no number. A run yields three streams and they never merge:

LayerDecided byOutput
Deterministicscripts/audit.shPASS / FAIL / NA / SKIP per check, with the evidence inline
Semiscript produces evidence, you read the verdictfinding or cleared, with the evidence quoted
Judgmentyou, reading real outputfinding or cleared, with the output pasted

A clean deterministic run is not a clean CLI. Say that in the report every time, because the reader will otherwise take the passes as the answer.

Weighting these into one number needs thresholds calibrated against a distribution, and no such distribution has been measured. Inventing weights here would repeat the mistake human-output.md names about scale cutoffs: instinct-picked buckets that never discriminate. When enough CLIs have been audited to know the real spread, that is when a score becomes possible, not before.

0. Resolve the conditional gates first

Most checks in cli-build are conditional. Running all of them against every CLI reports absences as defects when they were deliberate decisions, which trains the reader to skim the findings. Resolve these before running anything, and write each answer into the report:

  1. Does any command carry consequence? Money, irreversible writes, third-party side effects, personal data. If no: the entire trust-ladder group (audit log, --dry-run, killswitch, intent tokens, consent gates) is NA, and its absence is compliant. cli-build Phase 4 says so explicitly, and says the decision should be recorded rather than skipped. Look for that record before calling it an omission.
  2. Is there a human-facing surface? If the user scoped the CLI machine-only, every judgment check from human-output.md is void.
  3. What is the distribution target? Source shebang, npm, or native binary. The shebang and build checks only mean something against a declared target.
  4. Is the API asynchronous? Gates --wait and the job ledger.
  5. Is it published? Registry drift, zero-tests, and installer-shebang checks only bite once someone can install it.
  6. Contract origin: discovered, defined, or mixed? Decides what the JSON contract is even measured against.

Everything sourced from compounding-surface.md (--wait, vocabulary in CI, --deliver, feedback) is labeled in cli-build itself as coming from published work rather than the measured corpus. Report a gap there as a gap against a recommendation, and say which kind it is.

Done when: all six are answered in writing, and the check groups they switch off are named.

1. Run the deterministic layer

bash
scripts/audit.sh --repo <package-root> [--bin <installed-command>] [--json]

Without --bin only the static checks run and every runtime check reports SKIP. That is a materially weaker audit: cli-build Phase 6 is explicit that running a source file verifies a file while running the installed name verifies what ships — the bin entry, the shebang, the resolved dependencies, the banner. Link it first:

bash
cd <package-root> && bun link   # or npm link
scripts/audit.sh --repo . --bin <name>

Read references/checks.md for what each id means and where in cli-build it comes from. Every id traces to a line in the source skill; none were invented here.

A FAIL is a fact, not yet a finding. It becomes a finding when you have confirmed it applies given step 0. A FAIL on home-override in a CLI with no tests is noise.

Verify every claim of absence before it reaches the report. The first run of this skill produced four false absences out of five instrument bugs: an audit log with 41 call sites reported as "no audit log in this CLI", twice, under two different name guesses. A FAIL saying a feature is missing is the one output a reader cannot check cheaply, so it is the one that must be checked before it ships. Open the file and look.

The rule that came out of it: a check that resolves a feature by guessing its function names will eventually announce that the feature does not exist. Resolve the module, read the names it exports, then look for those names elsewhere. An export list cannot be wrong about itself.

Run the script against a second, unrelated CLI whenever you change it. Three of those five bugs were invisible against the CLI the check was written for and appeared immediately against another.

Done when: every non-PASS line is either promoted to a finding or dismissed with the conditional gate that excuses it, and every claim of absence has been confirmed by reading the code.

1b. Measure coverage, before and after

bash
scripts/coverage.sh --bin <name> --max-depth 4

This walks the CLI's own --help tree and reports how many leaf commands exist. Run it before the judgment layer so you know the size of the job, and again at the end with the list of what you actually ran:

bash
scripts/coverage.sh --bin <name> --max-depth 4 --exercised exercised.txt
# Coverage of sunat-cli: 27 of 80 leaf commands exercised (33%)

exercised.txt holds one command per line as invoked, without the binary name. Entries matching no command in the surface are counted separately and reported, because a stale note silently inflating coverage is the failure this artifact exists to prevent.

The ratio goes in the report. An audit of four commands and an audit of sixty produce prose that reads identically; this session's first pass covered 4 of 80 and the gap surfaced only because the reader asked. A number the report cannot omit turns partial coverage from a disclaimer into a fact.

Low coverage is often correct — most of what goes unexercised in a consequential CLI is the write path, and not running it is the right call. Say which it is: deliberate, or not reached.

Done when: the surface size is known before judging, and the final ratio plus the not-exercised list are in the report.

Show full SKILL.md (806 more words)Show less

2. Run the semi layer

These need evidence a script can produce and a verdict only a reader can give.

  • Noun-verb consistency. Dump the surface (<cli> --help, or the schema command), list every noun-verb pair, and look for the verb the rest of the surface does not use. info where everything else says get costs an agent a --help round trip.
  • --dry-run exercises the real path. Run it twice with different valid inputs. If the output does not vary with the input, it is a hardcoded shape and proves nothing. The strongest corpus implementation calls the provider's own preview endpoint.
  • Third-party free text is escaped. Anything a remote API returns as prose can carry instructions aimed at the agent reading your output. Find where provider text reaches stdout and check what escapes it.
  • The emitted example runs. When output prints a next command, copy it verbatim into the shell. An emitted example is an executable promise.
  • Tests assert on bad input. A suite that only passes valid input proves the happy path. Check that the input-validating functions have a case passing something unknown.
  • The color path is tested with escapes present. Without a TTY every style function returns plain text, so the escapes that break column arithmetic never appear in the asserted string. Forcing color is not enough either: the test passes trivially if the fixture carries no escapes. The assertion that holds is .length > visibleWidth.

Done when: each applicable item is a finding with its evidence quoted, or cleared with what you ran.

3. Run the judgment layer

Run the CLI and read its real output. Not --help, not the tests: the output a person gets. Read it twice, once for correctness and once as someone who does not know the domain.

Read cli-build's human-output.md with the terminal open, and ask its questions against what is on screen:

  • What does the biggest number mean, and what does a reader assume when it is high? If they assume the opposite, the metric is mispresented even though it is correct.
  • Is there a column where every row says the same thing? That is a heading that has not been promoted.
  • Does the default view answer the question a person asked, or return what the API returned?
  • Are the scale thresholds calibrated against this data, or against round numbers? A bucket that never fills and a bucket holding 84 percent both look like they are working.
  • Is red used for anything that is not an error?
  • Does --help read as the first screen of the product, or as a reference appendix?

One rule here cannot be fixed by writing it down. human-output.md records that the metric-direction inversion came back in new code written after it was documented. The defense that held was a named helper with regression tests. When you find a directional metric, the finding is not "document the direction" — it is "extract a named helper and test it."

Done when: the real output has been read and pasted into the report, and each question above is answered against it.

4. Report

Group findings by what it costs to fix, not by which layer found them. The reader wants to know what to patch, and the layer is provenance.

Every finding carries:

  • what, in one sentence;
  • the evidence, as the command run and its output, or file:line;
  • the rule it violates, cited to cli-build;
  • the class, so the reader knows whether a script re-checks it or a person does.

Two things the report must say out loud:

  1. What was not checked, and why. A conditional gate that switched a group off, a runtime check that reported SKIP because the CLI was not linked. Silence about a skipped group reads as a pass.
  2. That a clean deterministic run is not a clean CLI. With the reason: machine-readable and legible are independent properties and only one of them has rules a script can run.

Write the report to reports/<cli-name>-<date>.md in this skill, so the next audit of the same CLI can diff against it.

Done when: the report exists, names its own blind spots, and every finding cites evidence that can be re-run.

Boundaries

This skill reports; it does not patch. Auditing and fixing in one pass loses the record of what was wrong, and the record is what makes the next audit cheap. Patch in a separate cycle, against the report.

Adoption is not a check. Whether a CLI uses cligentic blocks measures adoption, not quality. cli-build Phase 3 accepts a rejection with a reason, and names the strongest one: a published output contract outranks a shared block. Penalizing a reasoned rejection would contradict the skill being audited.

Findings about someone else's package drop the subject and keep the defect. A case about a CLI Hunter wrote can be specific about what broke. See cli-build's cases/README.md.

© crafter-station, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/cli-audit of crafter-station/skills.

  • SKILL.md
  • references/checks.md
  • scripts/audit.sh
  • scripts/coverage.sh

Open the folder on GitHubat commit f0fe474

Compare with similar skills

CLI Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CLI Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CLI Audit this skillcrafter-station/skills112—~3kAutomated safety check: PassMIT
Health Wellnesssickn33/agentic-awesome-skills47k1 repos~3.5kAutomated safety check: PassMIT
AWS Well Architected Reviewgithub/awesome-copilot40k—~1.9kAutomated safety check: PassMIT
Azure Well Architected Reviewgithub/awesome-copilot40k—~2.5kAutomated safety check: PassMIT
Write Wellbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~4.1kAutomated safety check: PassCustom licence
Wellness Planmohitagw15856/pm-claude-skills1.4k—~902Automated safety check: PassMIT

Similar skills

  • Health Wellness

    sickn33/agentic-awesome-skills

    Wellbeing check-in register: anonymous flag, department, check-in date, wellbeing score, stress and energy levels, support and resource flags, confidentiality and status.

    47k GitHub starsUsed in 1 repo~3.5k tokens
    Productivity & AutomationAuto-check passed
  • AWS Well Architected Review

    github/awesome-copilot

    Official

    Perform an AWS Well-Architected Framework review of the current workload IaC and architecture, generating findings and GitHub issues for improvements.

    40k GitHub stars~1.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Azure Well Architected Review

    github/awesome-copilot

    Official

    Perform an Azure Well-Architected Framework review of the current workload IaC and architecture, generating findings and GitHub issues for improvements.

    40k GitHub stars~2.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Write Well

    brycewang-stanford/Auto-Empirical-Research-Skills

    Prose quality checker for Quarto (.qmd) files, grounded in William Zinsser's On Writing Well (30th Anniversary Edition).

    4.5k GitHub stars~4.1k tokensUpdated 3 days ago
    Writing & ContentAuto-check passed
  • Wellness Plan

    mohitagw15856/pm-claude-skills

    Build a preventive-care (wellness) plan for a pet by species, breed, and life stage.

    1.4k GitHub stars~902 tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • AWS Well Architected Review

    aws/agent-toolkit-for-aws

    Official

    Performs a full AWS Well-Architected Framework review evaluating every framework question across all pillars discovered from the live AWS documentation by analyzing code, IaC, and configurations to…

    2.8k GitHub stars~3k tokensUpdated today
    DevOps & CloudAuto-check passed

More from crafter-station/skills

  • Intent Layer

    crafter-station/skills

    Set up hierarchical Intent Layer (AGENTS.md files) for codebases.

    112 GitHub starsUsed in 1 repo~633 tokens
    Auto-check passed
  • Obsidian Plugin Release

    crafter-station/skills

    Release a new version of an Obsidian community plugin without forgetting steps.

    112 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Skill Gen

    crafter-station/skills

    Deprecated. An agent skill from crafter-station/skills.

    112 GitHub stars~5.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Generate Brand Assets

    crafter-station/skills

    Generate OG images and favicon based on project branding. An agent skill from crafter-station/skills.

    112 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • CLI Build

    crafter-station/skills

    Design and build a CLI that an AI agent can operate safely and a human can supervise.

    112 GitHub stars~5.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Surfacer

    crafter-station/skills

    Compile a mapped surface into working interfaces and keep them alive as the target changes.

    112 GitHub stars~829 tokensUpdated 1 mo ago
    Auto-check passed

Questions about CLI Audit

What does CLI Audit do?

Audit an existing CLI against cli-build: how well an agent can operate it and how well a human can read it. CLI Audit is an agent skill from crafter-station/skills. Audit an existing CLI against cli-build: how well an agent can operate it and how well a human can read it.

When should I use CLI Audit?

CLI Audit fits situations like: the user wants to audit a CLI; check whether a tool is agent-first; find out what a published CLI got wrong; grade a CLIs human output.

How do I install CLI Audit in Claude Code?

Run `npx skills add crafter-station/skills --skill cli-audit -a claude-code`. Or copy the skill folder (skills/cli-audit in crafter-station/skills) into .claude/skills/cli-audit in your project. Claude Code loads it when a task matches its description.

How do I install CLI Audit in Codex?

Run `npx skills add crafter-station/skills --skill cli-audit -a codex`. Or copy the skill folder (skills/cli-audit in crafter-station/skills) into .agents/skills/cli-audit in your project. Codex loads it when a task matches its description.

Can I use CLI Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add crafter-station/skills --skill cli-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cli-audit, .gemini/skills/cli-audit, .github/skills/cli-audit and .opencode/skills/cli-audit in your project.

What does CLI Audit need to run?

Going by SKILL.md and its folder, CLI Audit needs a shell for the scripts in its folder and the command-line tools its instructions call (bun). Our summary lists: A Bash shell.

Does CLI Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is CLI Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does CLI Audit use?

CLI Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CLI Audit use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to CLI Audit?

Skills that share tags, products or a category with CLI Audit: Health Wellness (sickn33/agentic-awesome-skills, 47k stars), AWS Well Architected Review (github/awesome-copilot, 40k stars), Azure Well Architected Review (github/awesome-copilot, 40k stars) and Write Well (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CLI Audit?

crafter-station (a GitHub organization) maintains it in crafter-station/skills, which has 112 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 7, 2026.

Source: crafter-station/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.