Agent skill

Codewhale Verification Ladder

by codewhale-hq in codewhale-hq/Codewhale

Climbs a focused-to-broad set of verification rungs, from formatting checks to budget scripts CI enforces, so a Codewhale change is called done only with quoted command output behind it.

MITAuto-check passedTesting & QA

Install Codewhale Verification Ladder

skills CLI
$ npx skills add codewhale-hq/Codewhale --skill cw-gates -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install codewhale-hq/Codewhale cw-gates --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/codewhale-hq/Codewhale.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/skills/cw-gates .claude/skills/cw-gates && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cw-gates
GitHub stars
41k
Token cost
~1.4k tokens
SKILL.md length
584 words
Files
1
Skills in repo
63
Repo updated
First seen
Licence
MIT

At a glance

Climbs a focused-to-broad set of verification rungs, from formatting checks to budget scripts CI enforces, so a Codewhale change is called done only with quoted command output behind it.

  • Confirming a Codewhale change is actually safe before calling it done
  • SKILL.md covers When to use, Workflow, Claiming a test passed and Red flags / don't, plus 1 more section
  • Calls python3, cargo and npm
  • Choosing the fastest correct test command for the area a change touches

What it does

This skill exists because the repository has twice been burned by an exit code mistaken for a pass and a test harness whose scoring line silently reported unevaluated rows as green; it insists that assertions without command output are not evidence, and it is stage 3 of a six-stage loop running from orient through slice, gates, dogfood, land and handoff. It applies before saying a change is done, green, passing, fixed or ready to land, before any dogfood build, and whenever asked to run the gates or prove a change is safe; release-specific work instead uses a sibling skill that adds a version-drift gate and manual TUI targets on top of this ladder.

The first rung is always cheap: a formatting check and a diff whitespace check. The second rung runs the fastest correct test invocation for the area a change touches, found through a dev-test script that maps a source path or area name to a filter and applies the project's isolated build-dir topology, preferring cargo nextest when it is on PATH and falling back to libtest otherwise; a compile-only test run can answer a compile question without executing unrelated cases.

The third rung runs the specific budget and drift-check scripts CI enforces for the area the change can move, such as a dead-code budget that may shrink freely but needs a reviewer's sign-off to raise, with a dedicated update flag to lock in an improvement. The skill's instruction throughout is to climb only as far as the actual risk requires and to say plainly where the check stopped and what was skipped.

When your agent uses it

  • Confirming a Codewhale change is actually safe before calling it done
  • Choosing the fastest correct test command for the area a change touches
  • Checking a change against the dead-code or other CI-enforced budget scripts
  • Deciding how far up the verification ladder a given change needs to climb

Example prompts

  • “Run the gates on this change and tell me exactly how far up the ladder you climbed.”
  • “What's the fastest correct test invocation for a change to crates/runtime/src/elapsed.rs?”
  • “Check whether this change moves the dead-code budget and lock in the improvement if so.”
  • “Before I say this is ready to land, show me the actual command output, not just a pass or fail.”

Requirements

  • A Codewhale checkout with Rust, Cargo and the scripts/dev-test.sh tooling
  • cargo-nextest, optionally, for the faster test runner

What it can do on your machine

Read from SKILL.md and the folder at commit 64bb073. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • cargo
    • npm
    • git
    • sh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Codewhale Verification Ladder loads about 1.4k tokens when it runs. Until then it costs about 51 tokens; SKILL.md has 584 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from codewhale-hq/Codewhale at commit 64bb073, republished under its MIT licence (© codewhale-hq). 584 words, ~1,387 tokens.

Download SKILL.mdSave it as .claude/skills/cw-gates/SKILL.md (or your agent's skills folder).
name
cw-gates
description
Use before claiming any Codewhale change is done, green, or ready to land: the focused-to-broad verification ladder, the budget checks CI enforces, and the rules for what counts as a passing test.

cw-gates

Pick the smallest evidence that answers the actual risk, then quote the real output. This repo has been burned twice by the alternative: an exit code mistaken for a pass, and a harness whose scoring line silently reported unevaluated rows as green. Assertions without command output are not evidence.

Stage 3 of the loop: cw-orient → cw-slice → gates → cw-dogfood → cw-land → cw-handoff.

When to use

  • Before saying "done", "green", "passing", "fixed", or "ready to land".
  • Before a dogfood build — never install an ungated binary.
  • When asked to "run the gates" or to prove a change is safe.

For release work specifically, use codew-release-qa-sweep instead — it adds the version-drift gate and the manual TUI QA targets on top of this ladder.

Workflow

Climb only as far as the risk requires. Say where you stopped and what you skipped.

Rung 1 — always, and cheap
bash
cargo fmt --all -- --check
git diff --check
Rung 2 — the area that owns the change

scripts/dev-test.sh maps an area or a source path to the fastest correct invocation, and applies the isolated build-dir topology (docs/BUILD_PERFORMANCE.md):

bash
scripts/dev-test.sh --list
scripts/dev-test.sh crates/command-contract/src/elapsed.rs      # path → area + filter
scripts/dev-test.sh tui tools::                    # area + filter
scripts/dev-test.sh config

It uses cargo nextest run when nextest is on PATH (.config/nextest.toml); CODEWHALE_DEV_NEXTEST=0 forces libtest. For the TUI crate, --lib and --tests are disjoint — choose the target that owns the behavior rather than running both by reflex.

cargo test --no-run answers a compile question without executing unrelated cases. cargo test --doc covers doc examples, and is only worth running when those examples changed.

Rung 3 — the budgets and drift checks CI enforces

Run the ones your change can move. Each fails the build in CI:

bash
python3 scripts/check-dead-code-budget.py          # #[allow(dead_code)] ceiling
python3 scripts/check-runtime-contract-budget.py
python3 scripts/check-persistence-backlog-budget.py
python3 scripts/check-provider-registry.py         # provider registry drift
python3 scripts/check-command-crate-boundaries.py  # command-contract boundary
python3 scripts/check-command-migration-manifest.py
python3 scripts/check-tui-locale-parity.py         # touched crates/localization/locales/
sh scripts/check-tui-product-vocabulary.sh
python3 scripts/check-readme-translations.py       # touched README*.md
./scripts/release/check-versions.sh                # touched a version anywhere

The dead-code budget may go down freely; raising it needs a reviewer to be told why. Lock in a win with python3 scripts/check-dead-code-budget.py --update.

Rung 4 — cross-cutting or release risk only
bash
cargo clippy --workspace --all-targets --all-features --locked -- \
  -D warnings \
  -A clippy::uninlined_format_args \
  -A clippy::too_many_arguments \
  -A clippy::unnecessary_map_or
cargo nextest run --workspace --all-features --locked --profile ci
cargo test --workspace --all-features --locked --doc
git diff --exit-code -- Cargo.lock                 # lockfile drift guard

--all-targets matters: without it, clippy never lints test code. (CI runs the all-targets form since the v0.9.10 gate incident; run the same form locally so there is no weaker subset.)

Rung 5 — website, when web/ changed
bash
cd web && npm ci && npm test && npm run check
Show full SKILL.md (264 more words)Show less

Claiming a test passed

AGENTS.md ("Claiming a test passed") owns the rules: quote the real count line with N > 0 for the tests that cover the change, prefer a regression test that fails without the fix, and audit hand-rolled scorers. Two additions for gate runs:

  • Check the count per filter group. A green run where one required filter matched zero tests is a miss, not a pass.
  • A focused rerun of a failing test distinguishes flake from regression in seconds. Do that before calling anything a flake, and root-cause anything that fails outside a known-flaky name — check for unisolated config-path reads or global timeout knobs first.

Red flags / don't

  • Don't say "tests pass" without the count line. Don't say "CI will catch it".
  • Don't run the full workspace suite as ritual for a leaf change, and don't re-run an unchanged suite to feel more confident.
  • Don't weaken a safety or data-integrity behavior to make a gate go green.
  • Don't call a failure a flake without a focused rerun and a named cause.
  • Don't skip the budget checks because they are "not really tests" — they are required CI contexts, and they encode migrations this repo has already paid for.
  • Don't report a green gate as permission. A passing sweep is readiness evidence; landing, tagging, and publishing need their own approval.

Output

A checklist: each command, pass/fail, and the salient line (test counts, budget numbers, the check-versions.sh verdict). Name explicitly what you did not run and why. If a step could not run in this environment, say so rather than implying coverage you do not have.

© codewhale-hq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/skills/cw-gates of codewhale-hq/Codewhale.

Open the folder on GitHubat commit 64bb073

Compare with similar skills

Codewhale Verification Ladder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Codewhale Verification Ladder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Codewhale Verification Ladder this skillcodewhale-hq/Codewhale41k—~1.4kAutomated safety check: PassMIT
lo2cin4bt Acceptance Reviewlo2cin4/lo2cin4bt289—~1.4kAutomated safety check: PassCustom licence
RuView Result Verificationruvnet/RuView97k—~395Automated safety check: PassMIT
TUI Change Verification for Warpwarpdotdev/warp65k1 repos~5.5kAutomated safety check: NotesAGPL-3.0
Refactor To Rulesopenfootmanager/openfootmanager1.1k—~1.3kAutomated safety check: NotesGPL-3.0
TiDB Verification Profilespingcap/tidb41k—~496Automated safety check: PassApache-2.0

Similar skills

  • Runs a pass, revise or block acceptance review on lo2cin4bt work, checking a deliverable against the request, repo contracts, tests, docs and the public GitHub boundary.

    289 GitHub stars~1.4k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Proves a RuView result is real by running a deterministic SHA-256 proof and a witness bundle, and by checking reports for untagged or unreproducible accuracy claims.

    97k GitHub stars~395 tokensUpdated today
    Testing & QAAuto-check passed
  • Runs Warp's terminal UI for real, drives it with keystrokes and reads the rendered screen back to confirm that a change looks and behaves as intended.

    65k GitHub starsUsed in 1 repo~5.5k tokens
    Testing & QAAuto-check: notes
  • Refactor To Rules

    openfootmanager/openfootmanager

    Decompose OpenFoot Manager code by responsibility when a quality gate is red, a reviewer requests a split or a refactor is authorized.

    1.1k GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check: notes
  • Chooses how much validation a TiDB change needs: scoped checks while iterating, required checks at delivery, and expensive runs only when explicitly needed.

    41k GitHub stars~496 tokensUpdated today
    DevelopmentAuto-check passed
  • Hydra Dev

    streamband/hydra-srt

    Run HydraSRT development workflows: mix q quality gate, Elixir unit/E2E tests, native Rust tests, web Vitest/Playwright, and make dev.

    147 GitHub stars~995 tokensUpdated 24 days ago
    Testing & QAAuto-check passed

More from codewhale-hq/Codewhale

All 63 skills in this repo
  • Codewhale Dogfood Install

    codewhale-hq/Codewhale

    Proves a Codewhale change in the real product: a stamped release build, an atomic local install, fresh-shell verification and manual QA that automated gates cannot cover.

    41k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Codewhale Session Handoff

    codewhale-hq/Codewhale

    Writes a paste-ready handoff for the next agent session, opening with a state-check command block and separating done, suspected and blocked work.

    41k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Codewhale Landing Workflow

    codewhale-hq/Codewhale

    Decides how verified work should reach main, directly, in a worktree or on an integration branch, while keeping contributor credit and respecting merge gates.

    41k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Sub-Agent Delegation

    codewhale-hq/Codewhale

    Guides when and how to split multi-step coding, research or verification work into focused sub-agent runs while the parent keeps integration and final checks.

    41k GitHub stars~790 tokensUpdated today
    Auto-check passed
  • Codewhale Fleet Manager

    codewhale-hq/Codewhale

    Triages and manages Codewhale fleet runs and workers with typed commands, classifying failures and choosing a safe restart, resume or escalation.

    41k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • GitHub Issue Bulk Assigner

    codewhale-hq/Codewhale

    Moves a list of GitHub issues into a milestone or assigns them to owners with the gh CLI, checking each one before and after the change.

    41k GitHub stars~953 tokensUpdated today
    Auto-check passed

Works with

Questions about Codewhale Verification Ladder

What does Codewhale Verification Ladder do?

Climbs a focused-to-broad set of verification rungs, from formatting checks to budget scripts CI enforces, so a Codewhale change is called done only with quoted command output behind it. This skill exists because the repository has twice been burned by an exit code mistaken for a pass and a test harness whose scoring line silently reported unevaluated rows as green; it insists that assertions without command output are not evidence, and it is stage 3 of a six-stage loop running from orient through slice, gates, dogfood, land and handoff. It applies before saying a change is done, green, passing, fixed or ready to land, before any dogfood build, and whenever asked to run the gates or prove a change is safe; release-specific work instead uses a sibling skill that adds a version-drift gate and manual TUI targets on top of this ladder.

When should I use Codewhale Verification Ladder?

Codewhale Verification Ladder fits situations like: confirming a Codewhale change is actually safe before calling it done; choosing the fastest correct test command for the area a change touches; checking a change against the dead-code or other CI-enforced budget scripts; deciding how far up the verification ladder a given change needs to climb.

How do I install Codewhale Verification Ladder in Claude Code?

Run `npx skills add codewhale-hq/Codewhale --skill cw-gates -a claude-code`. Or copy the skill folder (docs/skills/cw-gates in codewhale-hq/Codewhale) into .claude/skills/cw-gates in your project. Claude Code loads it when a task matches its description.

How do I install Codewhale Verification Ladder in Codex?

Run `npx skills add codewhale-hq/Codewhale --skill cw-gates -a codex`. Or copy the skill folder (docs/skills/cw-gates in codewhale-hq/Codewhale) into .agents/skills/cw-gates in your project. Codex loads it when a task matches its description.

Can I use Codewhale Verification Ladder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add codewhale-hq/Codewhale --skill cw-gates -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cw-gates, .gemini/skills/cw-gates, .github/skills/cw-gates and .opencode/skills/cw-gates in your project.

What does Codewhale Verification Ladder need to run?

Going by SKILL.md and its folder, Codewhale Verification Ladder needs the command-line tools its instructions call (python3, cargo, npm, git and sh). Our summary lists: A Codewhale checkout with Rust, Cargo and the scripts/dev-test.sh tooling; cargo-nextest, optionally, for the faster test runner.

Does Codewhale Verification Ladder access the network?

SKILL.md contains no URLs. Its commands use npm and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Codewhale Verification Ladder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Codewhale Verification Ladder use?

Codewhale Verification Ladder is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Codewhale Verification Ladder use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Codewhale Verification Ladder?

Skills that share tags, products or a category with Codewhale Verification Ladder: lo2cin4bt Acceptance Review (lo2cin4/lo2cin4bt, 289 stars), RuView Result Verification (ruvnet/RuView, 97k stars), TUI Change Verification for Warp (warpdotdev/warp, 65k stars) and Refactor To Rules (openfootmanager/openfootmanager, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Codewhale Verification Ladder?

codewhale-hq (a GitHub organization) maintains it in codewhale-hq/Codewhale, which has 41,076 GitHub stars. The repository holds 63 skills in this directory. The repository was last updated on October 10, 2026.

Source: codewhale-hq/Codewhale on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.