Agent skill

Test Confidence

by antiwork in antiwork/gumroad

AI-driven test execution. An agent skill from antiwork/gumroad.

MITAuto-check: notesDevelopment

Install Test Confidence

skills CLI
$ npx skills add antiwork/gumroad --skill test-confidence -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install antiwork/gumroad test-confidence --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/antiwork/gumroad.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/test-confidence .claude/skills/test-confidence && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-confidence
GitHub stars
9.8k
Token cost
~738 tokens
SKILL.md length
342 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

AI-driven test execution. An agent skill from antiwork/gumroad.

  • Works in 7 steps: Finds changed files (branch diff vs main… → Hashes the diff and checks… → Otherwise sends the diff + spec tree +… → …
  • Development work in your project
  • SKILL.md covers Usage, How it works, Pre-existing failure detection and When to use
  • Needs ANTHROPIC_API_KEY

What it does

Test Confidence is an agent skill from antiwork/gumroad. AI-driven test execution. Opus decides what to run and how confident to be, based on your diff.

Its SKILL.md is about 740 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development. The licence is MIT.

When your agent uses it

  • Development work in your project

Example prompts

  • “/test-confidence”

Requirements

  • A credential in ANTHROPIC_API_KEY
  • Pre-approved tools (allowed-tools): Bash(git *), Bash(bundle exec rspec *), Bash(cat *), Bash(find *), Bash(wc *), Bash(head *), Bash(tail *), Bash(grep *), Bash(bin/test-confidence *)

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Finds changed files (branch diff vs main for PR branches, local diff on main)
  2. Hashes the diff and checks tmp/test-confidence/ for a cached plan; reuses if found
  3. Otherwise sends the diff + spec tree + touched directories to Opus 4.7 in one call
  4. Opus returns a plan: risk level, ordered test list, confidence milestones
  5. Script executes the plan, showing yellow progress bar toward 99%
  6. At 99%, bar turns green. Safe to commit.
  7. With --full, continues running remaining tests toward 100%

What it can do on your machine

Read from SKILL.md and the folder at commit 9de1e01. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(git *)
    • Bash(bundle exec rspec *)
    • Bash(cat *)
    • Bash(find *)
    • Bash(wc *)
    • Bash(head *)
    • Bash(tail *)
    • Bash(grep *)
    • Bash(bin/test-confidence *)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Confidence loads about 738 tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 342 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~738

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:20
    ANTHROPIC_API_KEY` is auto-sourced from `.env` if not exported.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from antiwork/gumroad at commit 9de1e01, republished under its MIT licence (© antiwork). 342 words, ~738 tokens.

Download SKILL.mdSave it as .claude/skills/test-confidence/SKILL.md (or your agent's skills folder).
name
test-confidence
description
AI-driven test execution. Opus decides what to run and how confident to be, based on your diff.
allowed-tools
Bash(git *), Bash(bundle exec rspec *), Bash(cat *), Bash(find *), Bash(wc *), Bash(head *), Bash(tail *), Bash(grep *), Bash(bin/test-confidence *)
argument-hint
--full to run to 100% | --strict to halt on pre-existing failures

Test Confidence

Run bin/test-confidence to have Opus 4.7 analyze your diff, decide the risk level, plan which tests to run and in what order, and set confidence milestones. The AI decides the shape of the curve based on this specific diff.

Usage

bash
bin/test-confidence            # Run to 99%, stop. Skips past pre-existing failures.
bin/test-confidence --full     # Run to 100%
bin/test-confidence --strict   # Halt on any failure, including pre-existing

ANTHROPIC_API_KEY is auto-sourced from .env if not exported.

If $ARGUMENTS is provided, pass it through: bin/test-confidence $ARGUMENTS

How it works

  1. Finds changed files (branch diff vs main for PR branches, local diff on main)
  2. Hashes the diff and checks tmp/test-confidence/ for a cached plan; reuses if found
  3. Otherwise sends the diff + spec tree + touched directories to Opus 4.7 in one call
  4. Opus returns a plan: risk level, ordered test list, confidence milestones
  5. Script executes the plan, showing yellow progress bar toward 99%
  6. At 99%, bar turns green. Safe to commit.
  7. With --full, continues running remaining tests toward 100%

Pre-existing failure detection

When a spec fails, the script applies a two-step check:

  1. Path heuristic — does the failing spec file or its source counterpart (e.g., app/models/user.rb for spec/models/user_spec.rb) appear in the diff?
    • In the diff → real regression. Halt immediately.
    • Not in the diff → suspect pre-existing. Verify on merge-base.
  2. Merge-base verify — re-run the failing rspec examples in a temporary worktree at the branch's merge-base with main.
    • Still fails on merge-base → confirmed pre-existing. Continue.
    • Passes on merge-base → cross-file regression caught. Halt.

The verify step is skipped (heuristic verdict trusted) when the diff includes a db/migrate/*.rb (DB schema drift) or Gemfile.lock (bundler drift), or when no merge-base is available. --strict skips both checks and halts on every failure.

Cost profile: zero overhead in the common no-failure case; verify only runs on heuristic-flagged "pre-existing" failures. Catches the dangerous direction (silent regression skipped) while keeping false-alarm investigation costs bounded.

The key insight: Opus decides ad hoc how many tests are needed for each confidence level. A comment-only change might need 2 tests for 99%. A payment model refactor might need 100.

When to use

Run this before every commit. It replaces manually picking which specs to run.

$ARGUMENTS

© antiwork, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/test-confidence of antiwork/gumroad.

Open the folder on GitHubat commit 9de1e01

Compare with similar skills

Test Confidence next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Confidence compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Confidence this skillantiwork/gumroad9.8k—~738Automated safety check: NotesMIT
Vercel Composition Patternssupabase/supabase111k59 repos~726Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers296k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k25 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 59 repos~726 tokens
    DevelopmentAuto-check passed
  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    296k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 25 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed

More from antiwork/gumroad

  • Gumroad Prod Console

    antiwork/gumroad

    Execute read-only Ruby/Rails commands against Gumroad's production database for debugging and investigation.

    9.8k GitHub stars~2.9k tokensUpdated today
    Auto-check: notes
  • Review PR

    antiwork/gumroad

    Review GitHub pull requests for the Gumroad codebase against project guidelines, code quality, and correctness.

    9.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Email Blast

    antiwork/gumroad

    Send one-off email blasts to Gumroad creators directly via production console, no PR or deploy needed.

    9.8k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Commit

    antiwork/gumroad

    Stage and commit changes with a clear, concise commit message.

    9.8k GitHub stars~503 tokensUpdated today
    Auto-check passed
  • Create Issue

    antiwork/gumroad

    Draft GitHub issues for the Gumroad codebase. An agent skill from antiwork/gumroad.

    9.8k GitHub stars~747 tokensUpdated today
    Auto-check passed

Categories

Questions about Test Confidence

What does Test Confidence do?

AI-driven test execution. An agent skill from antiwork/gumroad. Test Confidence is an agent skill from antiwork/gumroad. AI-driven test execution.

When should I use Test Confidence?

Test Confidence fits situations like: development work in your project.

How do I install Test Confidence in Claude Code?

Run `npx skills add antiwork/gumroad --skill test-confidence -a claude-code`. Or copy the skill folder (.agents/skills/test-confidence in antiwork/gumroad) into .claude/skills/test-confidence in your project. Claude Code loads it when a task matches its description.

How do I install Test Confidence in Codex?

Run `npx skills add antiwork/gumroad --skill test-confidence -a codex`. Or copy the skill folder (.agents/skills/test-confidence in antiwork/gumroad) into .agents/skills/test-confidence in your project. Codex loads it when a task matches its description.

Can I use Test Confidence in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add antiwork/gumroad --skill test-confidence -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-confidence, .gemini/skills/test-confidence, .github/skills/test-confidence and .opencode/skills/test-confidence in your project.

What does Test Confidence need to run?

Going by SKILL.md and its folder, Test Confidence needs credentials named ANTHROPIC_API_KEY. Our summary lists: A credential in ANTHROPIC_API_KEY. Its frontmatter pre-approves these tools: Bash(git *), Bash(bundle exec rspec *), Bash(cat *), Bash(find *), Bash(wc *), Bash(head *), Bash(tail *), Bash(grep *), Bash(bin/test-confidence *).

Does Test Confidence access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Confidence safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Test Confidence use?

Test Confidence is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Confidence use?

About 738 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Confidence?

Skills that share tags, products or a category with Test Confidence: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 296k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Confidence?

antiwork (a GitHub organization) maintains it in antiwork/gumroad, which has 9,797 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 8, 2026.

Source: antiwork/gumroad on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.