Agent skill

Codewhale Release QA Sweep

by codewhale-hq in codewhale-hq/Codewhale

Runs Codewhale's automated verification gates in order on the real release branch, then exercises three manual TUI scenarios in a live terminal before anyone can call release work done.

MITAuto-check passedTesting & QA

Install Codewhale Release QA Sweep

skills CLI
$ npx skills add codewhale-hq/Codewhale --skill codew-release-qa-sweep -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install codewhale-hq/Codewhale codew-release-qa-sweep --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/codewhale-hq/Codewhale.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/skills/codew-release-qa-sweep .claude/skills/codew-release-qa-sweep && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
codew-release-qa-sweep
GitHub stars
41k
Token cost
~1.4k tokens
SKILL.md length
580 words
Files
1
Skills in repo
63
Repo updated
First seen
Licence
MIT

At a glance

Runs Codewhale's automated verification gates in order on the real release branch, then exercises three manual TUI scenarios in a live terminal before anyone can call release work done.

  • Works in 3 steps: Six-worker fanout liveness (regression… → Multi-terminal route isolation… → Running-turn input contract (regression…
  • Confirming release work is actually done before telling a maintainer it is merge-ready
  • SKILL.md covers When to use, Automated gate sweep, Manual QA targets and Reporting format, plus 1 more section
  • Calls git, cargo and npm

What it does

This skill sets the evidence bar for declaring Codewhale release work complete: a green automated gate sweep plus three manual QA targets, with no shortcut for either. It is used before telling a maintainer or a PR thread that release work is merge-ready, after landing pull requests into the release branch, and when verifying a release candidate on the real, often local-only, release branch rather than a main-based assumption.

The automated sweep runs from the repository root in a fixed order and stops at the first failure, starting by confirming the current branch is the real release head with a clean working tree. When validating a PR for landing, it also checks mergeability against the actual release head using git merge-tree, since a PR clean against main can still conflict with the release branch.

Because unit and build gates do not cover the live terminal UI, three manual scenarios are run against a built binary with a sealed local home and loopback fixtures: six-worker fanout liveness, checking that typing, rendering, cancelling and the status bar stay responsive while six sub-agents run and that Escape cancels mid-fanout rather than freezing; multi-terminal route isolation, confirming several terminals on different provider or model routes show no cross-contamination; and a third target the excerpt does not fully describe. Each scenario's dimensions, inputs, visible state and side effects are recorded rather than asserted through a full-screen test harness alone.

When your agent uses it

  • Confirming release work is actually done before telling a maintainer it is merge-ready
  • Verifying a release candidate on the real release branch after landing PRs
  • Testing whether a pull request truly merges cleanly into the release branch, not just main
  • Manually exercising the live TUI for regressions that automated tests cannot see

Example prompts

  • “Run the full release QA sweep before I tell the team this release is ready.”
  • “Check whether this PR actually merges cleanly into the release branch, not just main.”
  • “Walk through the six-worker fanout liveness check in a real terminal and record what you see.”
  • “Confirm multi-terminal route isolation still holds on this build.”

Requirements

  • A local checkout on the real release branch with a clean working tree
  • A built Codewhale binary and a terminal to run the manual QA scenarios in

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Six-worker fanout liveness (regression refs #3216, #2211). Spawn 6 sub-agents. Confirm
  2. Multi-terminal route isolation (regression ref #3227). Open multiple terminals on
  3. Running-turn input contract (regression ref #3203). During a busy turn, confirm Enter

What it can do on your machine

Read from SKILL.md and the folder at commit 8d2fb45. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • cargo
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Codewhale Release QA Sweep loads about 1.4k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 580 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from codewhale-hq/Codewhale at commit 8d2fb45, republished under its MIT licence (© codewhale-hq). 580 words, ~1,357 tokens.

Download SKILL.mdSave it as .claude/skills/codew-release-qa-sweep/SKILL.md (or your agent's skills folder).
name
codew-release-qa-sweep
description
Use before claiming Codewhale release work is done: run the full gate sweep and list the manual QA targets.

Codewhale Release QA Sweep

Run this before claiming any Codewhale release work is "done." A green automated gate sweep plus the three manual QA targets is the evidence bar. No sweep, no "done" — report exactly what was run and the result of each step.

When to use

  • Before telling the maintainer (or a PR thread) that release work is complete or merge-ready.
  • After harvesting/landing PRs into the release branch, before the publish boundary.
  • When verifying a release candidate on the real landing branch (e.g. <release-branch>), which is often local-only.

Automated gate sweep

Run from the repo root, in order. Stop on the first failure and report it.

bash
# 0. Confirm you are on the real release head, not a main-based assumption.
git branch --show-current          # expect e.g. <release-branch>
git status --short                 # working tree should be clean

# 1. Formatting + stray whitespace/conflict markers
cargo fmt --all --check
git diff --check

# Core npm workspace tests plus the shared web gate (install web dependencies first).
# check:web checks committed facts before regeneration, then docs, design tokens,
# lint, TypeScript and the production build. These do not replace the Rust gates.
npm test && npm run check:web

# 2. Library/protocol/cli/flow/state tests, locked
cargo test -p codewhale-config -p codewhale-protocol -p codewhale-cli \
  -p codewhale-workflow -p codewhale-state --locked

# 3. TUI test binaries, locked
cargo test -p codewhale-tui --lib --locked

# 4. TUI debug build, locked
cargo build -p codewhale-cli --bin codewhale --locked

# 5. Release build for the shipped binaries, locked
cargo build --release --locked -p codewhale-cli --bin codewhale

# 6. Version-drift gate (workspace ↔ npm ↔ Cargo.lock ↔ changelog ↔ README)
./scripts/release/check-versions.sh

# 7. Binary smoke
./target/release/codewhale --version

If you are validating a PR for landing, also test mergeability against the actual release head, never the main-based clean flag:

bash
git merge-tree $(git merge-base <release-branch> <pr-head>) <release-branch> <pr-head>

A PR that is clean against main can still conflict with the release branch.

Manual QA targets

Unit/build gates do not cover the live TUI. Exercise all three and record what you saw:

Use the built binary with a sealed local home and loopback fixtures, then exercise the relevant scenarios below in an actual terminal. Record dimensions, inputs, visible state, and side effects. Do not substitute a full-screen assertion harness for looking at and using the product.

  1. Six-worker fanout liveness (regression refs #3216, #2211). Spawn 6 sub-agents. Confirm typing, render, cancel, and the workbar stay live throughout, and that Esc cancels mid-fanout (prompt interrupt, not a wedged ~24s burst or freeze). For the Windows Terminal retest path (ref #3289), start in plan mode, add follow-up input to the plan, press Esc, switch to yolo/accept flow, trigger at least two auto/Fleet worker spawns, and keep typing/cancel/mode-switch checks live for several minutes. Attach logs if the freeze reproduces.
  2. Multi-terminal route isolation (regression ref #3227). Open multiple terminals on distinct provider/model routes. Confirm zero cross-terminal contamination and no provider+model mismatch — each terminal honors its own route.
  3. Running-turn input contract (regression ref #3203). During a busy turn, confirm Enter queues a typed follow-up, the preview advertises Enter send now, and an empty Enter promotes the oldest queued follow-up. Confirm Ctrl+Enter steers typed text directly, Shift+Enter inserts a newline, and Ctrl+G/Ctrl+S only stash drafts.
Show full SKILL.md (220 more words)Show less

Reporting format

Report a checklist: each command, pass/fail, and the salient output line (test counts, the --version string, check-versions.sh verdict). For manual QA, state what you actually observed per target, citing the regression ref where one applies. If a step was skipped or could not be run (e.g. no display for TUI QA), say so explicitly — do not imply coverage you do not have.

Red flags / don't

  • Don't claim "done," "passing," or "merge-ready" without the evidence above. Assertions without command output are not acceptable.
  • Don't trust the main-based mergeability flag for a release branch; use git merge-tree against the real head.
  • Don't skip the manual TUI targets because the build is green — the freeze, route-mismatch, and steering regressions live in the runtime, not the gates.
  • Don't tag, publish, create a GitHub Release, push artifacts, or merge/close any PR or issue without maintainer approval. A green sweep is readiness evidence, not permission.
  • Never harvest/close from a PR title or label alone — review from code, tests, comments, and checks.
  • When the sweep clears a harvested PR, preserve contributor credit: cherry-pick keeps the original author, otherwise add Co-authored-by: Name <email> and a Harvested from PR #N by @handle body line (the spaced form the auto-close-at-main workflow greps for).
  • Keep any contributor-facing comment positive and crediting; gates stay dry-run/advisory unless the maintainer approves enforcement.

© codewhale-hq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/skills/codew-release-qa-sweep of codewhale-hq/Codewhale.

Open the folder on GitHubat commit 8d2fb45

Compare with similar skills

Codewhale Release QA Sweep next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Codewhale Release QA Sweep compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Codewhale Release QA Sweep this skillcodewhale-hq/Codewhale41k—~1.4kAutomated safety check: PassMIT
Code That Fits In Your HeadCodeAlive-AI/ai-driven-development158—~1.7kAutomated safety check: PassMIT
CI Triageandymai/brepjs115—~3.7kAutomated safety check: PassApache-2.0
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
Cherry Studio Regression TestsCherryHQ/cherry-studio52k—~1.2kAutomated safety check: PassAGPL-3.0
Nemoclaw Maintainer Fix E2E FailuresNVIDIA/NemoClaw23k—~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Code That Fits In Your Head

    CodeAlive-AI/ai-driven-development

    Software-engineering heuristics based on Mark Seemann's Code That Fits in Your Head (2021), updated for agent-driven development.

    158 GitHub stars~1.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • CI Triage

    andymai/brepjs

    This skill should be used when a brepjs GitHub Actions job is red or behaving oddly on github.com (a remote CI run, not a local pre-commit/pre-push hook) — "CI failed", "ci-pass is failing", "npm ci…

    115 GitHub stars~3.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • Cherry Studio Regression Tests

    CherryHQ/cherry-studio

    Runs Cherry Studio's critical-path regression suite as deterministic Playwright E2E tests through a GitHub workflow on macOS and Windows runners.

    52k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Continuously maintain automatic NemoClaw main E2E results through coordinated repairs.

    23k GitHub stars~2.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Release

    lycorp-jp/sim-use

    Cut a sim-use release end-to-end. An agent skill from lycorp-jp/sim-use.

    1.4k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed

More from codewhale-hq/Codewhale

All 63 skills in this repo
  • Codewhale Dogfood Install

    codewhale-hq/Codewhale

    Proves a Codewhale change in the real product: a stamped release build, an atomic local install, fresh-shell verification and manual QA that automated gates cannot cover.

    41k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Codewhale Session Handoff

    codewhale-hq/Codewhale

    Writes a paste-ready handoff for the next agent session, opening with a state-check command block and separating done, suspected and blocked work.

    41k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Codewhale Landing Workflow

    codewhale-hq/Codewhale

    Decides how verified work should reach main, directly, in a worktree or on an integration branch, while keeping contributor credit and respecting merge gates.

    41k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Sub-Agent Delegation

    codewhale-hq/Codewhale

    Guides when and how to split multi-step coding, research or verification work into focused sub-agent runs while the parent keeps integration and final checks.

    41k GitHub stars~790 tokensUpdated today
    Auto-check passed
  • Codewhale Fleet Manager

    codewhale-hq/Codewhale

    Triages and manages Codewhale fleet runs and workers with typed commands, classifying failures and choosing a safe restart, resume or escalation.

    41k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • GitHub Issue Bulk Assigner

    codewhale-hq/Codewhale

    Moves a list of GitHub issues into a milestone or assigns them to owners with the gh CLI, checking each one before and after the change.

    41k GitHub stars~953 tokensUpdated today
    Auto-check passed

Works with

Questions about Codewhale Release QA Sweep

What does Codewhale Release QA Sweep do?

Runs Codewhale's automated verification gates in order on the real release branch, then exercises three manual TUI scenarios in a live terminal before anyone can call release work done. This skill sets the evidence bar for declaring Codewhale release work complete: a green automated gate sweep plus three manual QA targets, with no shortcut for either. It is used before telling a maintainer or a PR thread that release work is merge-ready, after landing pull requests into the release branch, and when verifying a release candidate on the real, often local-only, release branch rather than a main-based assumption.

When should I use Codewhale Release QA Sweep?

Codewhale Release QA Sweep fits situations like: confirming release work is actually done before telling a maintainer it is merge-ready; verifying a release candidate on the real release branch after landing PRs; testing whether a pull request truly merges cleanly into the release branch, not just main; manually exercising the live TUI for regressions that automated tests cannot see.

How do I install Codewhale Release QA Sweep in Claude Code?

Run `npx skills add codewhale-hq/Codewhale --skill codew-release-qa-sweep -a claude-code`. Or copy the skill folder (docs/skills/codew-release-qa-sweep in codewhale-hq/Codewhale) into .claude/skills/codew-release-qa-sweep in your project. Claude Code loads it when a task matches its description.

How do I install Codewhale Release QA Sweep in Codex?

Run `npx skills add codewhale-hq/Codewhale --skill codew-release-qa-sweep -a codex`. Or copy the skill folder (docs/skills/codew-release-qa-sweep in codewhale-hq/Codewhale) into .agents/skills/codew-release-qa-sweep in your project. Codex loads it when a task matches its description.

Can I use Codewhale Release QA Sweep in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add codewhale-hq/Codewhale --skill codew-release-qa-sweep -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/codew-release-qa-sweep, .gemini/skills/codew-release-qa-sweep, .github/skills/codew-release-qa-sweep and .opencode/skills/codew-release-qa-sweep in your project.

What does Codewhale Release QA Sweep need to run?

Going by SKILL.md and its folder, Codewhale Release QA Sweep needs the command-line tools its instructions call (git, cargo and npm). Our summary lists: A local checkout on the real release branch with a clean working tree; A built Codewhale binary and a terminal to run the manual QA scenarios in.

Does Codewhale Release QA Sweep access the network?

SKILL.md contains no URLs. Its commands use git and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Codewhale Release QA Sweep safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Codewhale Release QA Sweep use?

Codewhale Release QA Sweep is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Codewhale Release QA Sweep use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Codewhale Release QA Sweep?

Skills that share tags, products or a category with Codewhale Release QA Sweep: Code That Fits In Your Head (CodeAlive-AI/ai-driven-development, 158 stars), CI Triage (andymai/brepjs, 115 stars), OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars) and Cherry Studio Regression Tests (CherryHQ/cherry-studio, 52k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Codewhale Release QA Sweep?

codewhale-hq (a GitHub organization) maintains it in codewhale-hq/Codewhale, which has 41,078 GitHub stars. The repository holds 63 skills in this directory. The repository was last updated on October 9, 2026.

Source: codewhale-hq/Codewhale on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.