Agent skill

CLI Review Runner

by pproenca in pproenca/dot-skills

Black-box CLI grading harness — runs a test suite against a target CLI and reports per-rule pass/fail from the cli-for-agents 45-rule catalog.

MITAuto-check passedTesting & QA

Install CLI Review Runner

skills CLI
$ npx skills add pproenca/dot-skills --skill cli-review-runner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pproenca/dot-skills cli-review-runner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/cli-review-runner .claude/skills/cli-review-runner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cli-review-runner
GitHub stars
214
Token cost
~1.9k tokens
SKILL.md length
641 words
Files
11 (incl. scripts, references)
Skills in repo
182
Repo updated
First seen
Licence
MIT

At a glance

Black-box CLI grading harness — runs a test suite against a target CLI and reports per-rule pass/fail from the cli-for-agents 45-rule catalog.

  • Grading a command-line tool for agent-friendliness
  • SKILL.md covers When to Apply, How to Use, Workflow Overview and Probe Coverage, plus 5 more sections
  • Runs Shell scripts from its folder; calls bash
  • Even if the user doesnt explicitly say agent-friendly — apply whenever they ask is mycli good for agents?

What it does

CLI Review Runner is an agent skill from pproenca/dot-skills. Black-box CLI grading harness — runs a test suite against a target CLI and reports per-rule pass/fail from the cli-for-agents 45-rule catalog. Use when reviewing, auditing, or grading a command-line tool for agent-friendliness. Trigger even if the user doesn't explicitly say "agent-friendly" — apply whenever they ask "is mycli good for agents?", "review this CLI", "grade my cli against the rules", "check if this tool is safe to automate", or "audit command-line design". Companion to the cli-for-agents…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `config.json`, `gotchas.md` and `metadata.json`).

It sits in Testing & QA, covering Test generation. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.

When your agent uses it

  • Grading a command-line tool for agent-friendliness
  • Even if the user doesnt explicitly say agent-friendly — apply whenever they ask is mycli good for agents?
  • Review this CLI
  • Grade my cli against the rules

Example prompts

  • “t explicitly say”
  • “— apply whenever they ask”
  • “review this CLI”
  • “/cli-review-runner”

Requirements

  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CLI Review Runner loads about 1.9k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 136 tokens; SKILL.md has 641 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~136
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 641 words, ~1,939 tokens.

Download SKILL.mdSave it as .claude/skills/cli-review-runner/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
cli-review-runner
description
Black-box CLI grading harness — runs a test suite against a target CLI and reports per-rule pass/fail from the cli-for-agents 45-rule catalog. Use when reviewing, auditing, or grading a command-line tool for agent-friendliness. Trigger even if the user doesn't explicitly say "agent-friendly" — apply whenever they ask "is mycli good for agents?", "review this CLI", "grade my cli against the rules", "check if this tool is safe to automate", or "audit command-line design". Companion to the cli-for-agents distillation skill.

cli-review-runner

Automates the 10-item agent-friendliness audit from cli-for-agents. Runs black-box probes against a target CLI and emits a structured report mapping each finding to a rule ID (e.g., help-examples-in-help, err-non-zero-exit-codes, safe-dry-run-flag). Default mode is read-only - probes never run destructive verbs with real arguments.

When to Apply

  • User asks to review or audit a CLI for agent-friendliness, automation readiness, or CI use
  • User has just finished building a CLI and wants a pre-ship sanity check
  • User is grading their own or a third-party CLI against the cli-for-agents catalog
  • User is asking why a CLI is hanging an agent, blowing up context, or failing to compose in a pipeline
  • PR review for a CLI change - quickly regress-test the --help, errors, and dry-run flags

How to Use

The skill is orchestrated by scripts/review.sh. Point it at the target CLI (absolute path or PATH-resolvable name) and pick an output format.

bash
# Default: text table on stdout, exit 0 if all passed, 1 if any failed
bash scripts/review.sh --target /usr/local/bin/mycli

# Machine-readable output
bash scripts/review.sh --target gh --format json
bash scripts/review.sh --target kubectl --format ndjson

# Supply subcommand list when auto-discovery misses them
bash scripts/review.sh --target gh --subcommands pr,issue,repo

# Preview what would run without touching the target CLI
bash scripts/review.sh --target mycli --dry-run

# Include risky probes on destructive verbs (off by default)
bash scripts/review.sh --target mycli --include-destructive

See bash scripts/review.sh --help for the full flag list.

Workflow Overview

--target <cli>
   │
   ▼
[1] Validate target       fail fast if path missing or not executable
   │
   ▼
[2] Load rule catalog     references/rule-catalog.tsv (45 rules)
   │
   ▼
[3] Discover subcommands  parse top-level --help (gh/kubectl/commander shapes)
   │
   ▼
[4] Run probes P1..P10    each probe emits NDJSON findings to a temp file
   │
   ▼
[5] Render report         scripts/render.sh  ->  text | json | ndjson

Read references/workflow.md when you need the full probe-by-probe breakdown, failure modes, and how to extend the catalog.

Probe Coverage

ProbeRules testedCoverage
P1 Non-interactiveinteract-no-hang-on-stdin, interact-no-input-flag, interact-flags-first, interact-detect-tty, interact-no-timed-prompts, interact-no-arrow-menus, input-no-prompt-fallbackRun under </dev/null with timeout; inspect help for interactive language
P2 Layered helphelp-per-subcommand, help-no-flag-required, help-layered-discoveryTop-level line count; per-subcommand --help; zero-arg invocation
P3 Help exampleshelp-examples-in-help, help-flag-summary, help-suggest-next-stepsGrep each subcommand help for Examples:, short+long flag pairs, "See also"
P4 Actionable errorserr-actionable-fix, err-include-example-invocation, err-exit-fast-on-missing-required, err-no-stack-traces-by-defaultInvoke with bogus flag; grep stderr for fix + example; check for raw stack traces
P5 stderr channelingerr-stderr-not-stdoutError text must land on fd 2
P6 Exit codeserr-non-zero-exit-codesUsage error and runtime error must produce distinct non-zero codes
P7 stdin compositioninput-accept-stdin-dash, input-flags-over-positionalGrep help for - stdin convention; count positionals vs flags
P8 Structured outputoutput-json-flag, output-respect-no-color--json produces JSON; NO_COLOR=1 suppresses ANSI
P9 Destructive safetysafe-dry-run-flag, safe-force-bypass-flag, safe-no-prompts-with-no-inputInspect destructive verbs' --help for --dry-run / --yes / --force / --no-input
P10 Command structurestruct-resource-verb, struct-standard-flag-names, struct-no-hidden-subcommand-catchall, struct-flag-order-independentUniform shape, --help/--version present, unknown subcommand errors, flag position independence

Coverage: 30 of the 45 rules in cli-for-agents are black-box testable. The remaining 15 (idempotency, state reconciliation, NDJSON streaming, bounded output, crash-only recovery, env-var fallback, secret-stdin, confirm-by-typing-name) require either real invocation or source-code inspection - the report lists them as "manual review required".

Show full SKILL.md (267 more words)Show less

Configuration

config.json stores the verb classifier lists and default timeout. The skill works without any setup - defaults are reasonable. Override per-invocation via flags or edit the file for project-wide changes.

json
{
  "timeout_seconds": 5,
  "safe_verbs": ["list", "get", "show", "status", "describe", "help", "version", "config", "ls", "inspect"],
  "destructive_verbs": ["delete", "drop", "destroy", "remove", "reset", "purge", "rm", "del"]
}

Safety Model

All probes are read-only by default:

  • Safe verbs (list, get, show, ...) may be invoked with bogus flags to test error handling
  • Destructive verbs (delete, drop, ...) are ONLY inspected via --help - never executed with arguments
  • Every probe runs with a 5-second wall-clock timeout under </dev/null
  • The target CLI is sandboxed to its own process; no shell metacharacters in arguments

When --include-destructive is passed, probes may invoke destructive verbs with bogus flags too. This exposes the rare case where a CLI does something destructive before validating flags. Only enable this against CLIs you trust, or in a disposable test environment.

Self-test

Before shipping changes to probes, run the self-test - it generates a mock CLI that deliberately violates specific rules and asserts the probes detect them:

bash
bash scripts/selftest.sh

Expected output: Results: 8 passed, 0 failed. Any failure points at a regression in scripts/lib/probes.sh.

Files

FilePurpose
scripts/review.shMain entry point - orchestrates probes, renders the report
scripts/render.shOutput formatter: text / json / ndjson
scripts/selftest.shSanity check against a deliberately-buggy mock CLI
scripts/lib/common.shShared helpers: timeout, verb classifier, JSON escape, catalog loader
scripts/lib/probes.shProbe functions probe_p1..probe_p10
references/rule-catalog.tsv45 rules from cli-for-agents, mapped to probes
references/workflow.mdDetailed probe-by-probe methodology, failure modes, extension guide
gotchas.mdKnown edge cases discovered during use
config.jsonVerb classifier lists and default timeout
  • cli-for-agents - the 45-rule design catalog this skill audits against. Read rule files there when the report flags an issue and you need the full explanation.

© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references) in skills/.experimental/cli-review-runner of pproenca/dot-skills.

  • SKILL.md
  • config.json
  • gotchas.md
  • metadata.json
  • references/rule-catalog.tsv
  • references/workflow.md
  • scripts/lib/common.sh
  • scripts/lib/probes.sh
  • scripts/render.sh
  • scripts/review.sh
  • scripts/selftest.sh

Open the folder on GitHubat commit cf93c57

Compare with similar skills

CLI Review Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CLI Review Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CLI Review Runner this skillpproenca/dot-skills214—~1.9kAutomated safety check: PassMIT
Emcaklofas/kicad-happy1.3k1 repos~2.8kAutomated safety check: PassMIT
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Generate Test Cases342164796/generate-test-cases1191 repos~2.9kAutomated safety check: PassNone
Verify Cc Safety Netkenryu42/cc-safety-net1.6k—~2kAutomated safety check: PassMIT
Wioworkersio/skills180—~5.8kAutomated safety check: PassMIT

Similar skills

  • Emc

    aklofas/kicad-happy

    EMC pre-compliance risk analysis for KiCad PCB designs — 18 check categories, 44 rule IDs covering ground planes, decoupling, I/O filtering, switching harmonics, clock routing, differential pair…

    1.3k GitHub starsUsed in 1 repo~2.8k tokens
    Testing & QAAuto-check passed
  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Generate Test Cases

    342164796/generate-test-cases

    自主学习型测试文档生成器。从需求文档(Markdown)生成测试用例 XMind 文件,支持持久化记忆和持续学习。当用户提到"生成测试用例"、"根据需求生成测试"时触发。

    119 GitHub starsUsed in 1 repo~2.9k tokens
    Testing & QAAuto-check passed
  • Verify Cc Safety Net

    kenryu42/cc-safety-net

    Launch and drive the real cc-safety-net CLI — the hook decision path, explain, status/doctor, logs, and the local policy GUI — against an isolated home, capturing evidence.

    1.6k GitHub stars~2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    180 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • File Server

    microsoft/WindowsProtocolTestSuites

    Official

    ALWAYS LOAD THIS SKILL when working with FileServer, SMB, SMB2, SMB3, CIFS, file sharing, MS-SMB2, MS-FSCC, MS-FSA, MS-DFSC, MS-FSRVP, MS-RSVD, MS-SQOS, or any file server protocol test…

    567 GitHub stars~4.1k tokensUpdated 22 days ago
    Testing & QAAuto-check passed

More from pproenca/dot-skills

All 182 skills in this repo
  • Audio Voice Recovery

    pproenca/dot-skills

    Audio forensics and voice recovery guidelines for CSI-level audio analysis.

    214 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Codemod React Pipeline

    pproenca/dot-skills

    Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.

    214 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dev Rfc

    pproenca/dot-skills

    Create well-structured RFCs and technical proposals for software projects.

    214 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Dx Harness

    pproenca/dot-skills

    Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.

    214 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Language Spec Author

    pproenca/dot-skills

    Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…

    214 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    214 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about CLI Review Runner

What does CLI Review Runner do?

Black-box CLI grading harness — runs a test suite against a target CLI and reports per-rule pass/fail from the cli-for-agents 45-rule catalog. CLI Review Runner is an agent skill from pproenca/dot-skills. Black-box CLI grading harness — runs a test suite against a target CLI and reports per-rule pass/fail from the cli-for-agents 45-rule catalog.

When should I use CLI Review Runner?

CLI Review Runner fits situations like: grading a command-line tool for agent-friendliness; even if the user doesnt explicitly say agent-friendly — apply whenever they ask is mycli good for agents?; review this CLI; grade my cli against the rules.

How do I install CLI Review Runner in Claude Code?

Run `npx skills add pproenca/dot-skills --skill cli-review-runner -a claude-code`. Or copy the skill folder (skills/.experimental/cli-review-runner in pproenca/dot-skills) into .claude/skills/cli-review-runner in your project. Claude Code loads it when a task matches its description.

How do I install CLI Review Runner in Codex?

Run `npx skills add pproenca/dot-skills --skill cli-review-runner -a codex`. Or copy the skill folder (skills/.experimental/cli-review-runner in pproenca/dot-skills) into .agents/skills/cli-review-runner in your project. Codex loads it when a task matches its description.

Can I use CLI Review Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill cli-review-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cli-review-runner, .gemini/skills/cli-review-runner, .github/skills/cli-review-runner and .opencode/skills/cli-review-runner in your project.

What does CLI Review Runner need to run?

Going by SKILL.md and its folder, CLI Review Runner needs a shell for the scripts in its folder and the command-line tools its instructions call (bash). Our summary lists: A Bash shell.

Does CLI Review Runner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is CLI Review Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does CLI Review Runner use?

CLI Review Runner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CLI Review Runner use?

About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to CLI Review Runner?

Skills that share tags, products or a category with CLI Review Runner: Emc (aklofas/kicad-happy, 1.3k stars), Swig Test (swig/swig, 6.3k stars), Generate Test Cases (342164796/generate-test-cases, 119 stars) and Verify Cc Safety Net (kenryu42/cc-safety-net, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CLI Review Runner?

pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 214 GitHub stars. The repository holds 182 skills in this directory. The repository was last updated on August 15, 2026.

Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.