Agent skill

Test A Skill

by will-ness-ai in will-ness-ai/skills

Field-test a skill by running it cold in parallel agent sessions, then turn what they hit into edits.

MITAuto-check passedAgent Workflows

Install Test A Skill

skills CLI
$ npx skills add will-ness-ai/skills --skill test-a-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install will-ness-ai/skills test-a-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/will-ness-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-a-skill .claude/skills/test-a-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-a-skill
GitHub stars
168
Token cost
~1k tokens
SKILL.md length
687 words
Files
2
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Field-test a skill by running it cold in parallel agent sessions, then turn what they hit into edits.

  • Works in 4 steps: Pick three realistic scenarios → Run it cold → Verify, then rank → …
  • Tasks that involve Subagents
  • SKILL.md covers 1. Pick three realistic…, 2. Run it cold, 3. Verify, then rank and 4. Offer the edits
  • Calls claude

What it does

Test A Skill is an agent skill from will-ness-ai/skills. Field-test a skill by running it cold in parallel agent sessions, then turn what they hit into edits.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Agent Workflows, covering Subagents. The repository describes itself as: Skills for building high quality software. The licence is MIT.

When your agent uses it

  • Tasks that involve Subagents

Example prompts

  • “/test-a-skill”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Pick three realistic scenarios
  2. Run it cold
  3. Verify, then rank
  4. Offer the edits

What it can do on your machine

Read from SKILL.md and the folder at commit 71d8909. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test A Skill loads about 1k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 687 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from will-ness-ai/skills at commit 71d8909, republished under its MIT licence (© will-ness-ai). 687 words, ~1,007 tokens.

Download SKILL.mdSave it as .claude/skills/test-a-skill/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
test-a-skill
description
Field-test a skill by running it cold in parallel agent sessions, then turn what they hit into edits.
disable-model-invocation
true

A skill is a process, so reading it proves nothing. Run it.

/test-a-skill <skill> puts the skill in front of agents that have never seen it, on work you would really bring to it. What they produce shows whether it works. What they had to guess shows what to fix — that is the deliverable.

1. Pick three realistic scenarios

Three inputs you would genuinely hand this skill, spread across the range you would genuinely use it on. Let them differ the way real work differs: in size, in subject, in how much explaining material happens to come with them. Three similar inputs test one thing three times.

Realistic beats adversarial. A pathological input tells you how the skill behaves somewhere it will never go, then tempts you to spend the skill's words fixing that instead of the path every run takes.

Done when each subject is work you would actually bring to this skill, and no two are alike.

2. Run it cold

Pin the skill first — a copy, or a fixed commit. The working tree moves under a long run: another session switches branch and the folder is gone mid-flight.

Then one agent per subject, in parallel, on your strongest model, as a dynamic workflow. Keep the whole fan-out in that one script: its agents are given no Agent or Workflow tool of their own, so a script that expects them to spawn anything quietly comes back empty. Where you are already inside a workflow, spawn claude -p headless instead.

Give each agent a schema, so the gap reports arrive structured rather than parsed out of prose. Give each one the tools it needs rather than only a directory: a permission wall reads back as a defect in the skill, in every report at once.

Each prompt carries the path to the skill file, plus enough to mock the context window this skill would really fire into: the request a human would type, the repo and branch they sit on, what they were doing just before, what they had already decided or read. A skill is never invoked into an empty session, so testing it against one tests a situation that never happens.

The line runs at the skill's own job. Context is what the window would already hold before anyone typed the command. Briefing is what you add because you doubt the agent will do the right thing — and each of those is a gap you just found, so put it in the skill and let the next run prove it landed. A well-briefed agent only tests your briefing.

Show full SKILL.md (256 more words)Show less

Ask each agent to return, alongside its output:

  • what it actually ran: the commands, and the files it wrote
  • where the skill was silent on a decision it had to make
  • what it guessed
  • the hardest part
  • what was wrong, missing, or awkward in the skill, in its reference files, and in the skills it calls

Say the gaps matter more than the deliverable, and that you want them blunt.

Done when every agent has shown what it ran, or is named as having failed. A confident report proves a read, not a run.

3. Verify, then rank

Check every claim against the skill, its reference files, and whatever code or tool it drives. Do this before you believe any of it.

Then rank what survives by convergence: a gap three agents hit independently is real, where one agent's may be one agent's taste. Keep that order and never the reverse — agents sharing one sandbox manufacture identical false positives, so a unanimous gap is as likely to be a property of the harness you gave them as of the skill. The run this skill came from ranked "this file is unreadable" first on a 3/3 vote. The file was fine; the tester's own permission flag was not.

Done when every reported gap is confirmed against the source or dismissed with a reason.

4. Offer the edits

Call /writing-for-agents. Then hand the human one ranked list — the fix, the evidence, and how many agents hit it — and wait for their picks.

Done when the human has chosen.

© will-ness-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/test-a-skill of will-ness-ai/skills.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 71d8909

Compare with similar skills

Test A Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test A Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test A Skill this skillwill-ness-ai/skills168—~1kAutomated safety check: PassMIT
Claude Code Agent Developmentanthropics/claude-plugins-official38k8 repos~2.8kAutomated safety check: PassApache-2.0
Subagent Driven DevelopmentAsvarox/allkaraoke26138 repos~1.2kAutomated safety check: PassNone
Dispatching Parallel Agentsultralisp/ultralisp25841 repos~1.5kAutomated safety check: PassNone
Paseo Advisor Second Opiniongetpaseo/paseo20k1 repos~756Automated safety check: PassCustom licence
Task Observerrebelytics/one-skill-to-rule-them-all3.2k1 repos~12kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 8 repos~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Subagent Driven Development

    Asvarox/allkaraoke

    A skill your agent uses when executing implementation plans with independent tasks in the current session

    261 GitHub starsUsed in 38 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Dispatching Parallel Agents

    ultralisp/ultralisp

    A skill your agent uses when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies

    258 GitHub starsUsed in 41 repos~1.5k tokens
    Agent WorkflowsAuto-check passed
  • Launches one separate agent through Paseo to give a second opinion on the current task, with a self-contained briefing and no permission to edit files.

    20k GitHub starsUsed in 1 repo~756 tokens
    Agent WorkflowsAuto-check passed
  • Task Observer

    rebelytics/one-skill-to-rule-them-all

    Monitors task execution for skill improvement opportunities.

    3.2k GitHub starsUsed in 1 repo~12k tokens
    Agent WorkflowsAuto-check passed
  • O2 Review Loop

    openobserve/openobserve

    Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.

    22k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from will-ness-ai/skills

  • Cmux

    will-ness-ai/skills

    Drive the local cmux app — workspaces, panes, surfaces, terminal input, agent sessions, and browser surfaces.

    168 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Sandcastle

    will-ness-ai/skills

    Run a Sandcastle lane — a sandboxed implement→review loop over a fixed batch of work in its own worktree, ended by an agent-authored PR.

    168 GitHub stars~1.4k tokensUpdated today
    Auto-check: notes
  • Code Story

    will-ness-ai/skills

    Build a wizard-style HTML page that teaches how and why a change works.

    168 GitHub stars~828 tokensUpdated today
    Auto-check passed
  • Flashlight

    will-ness-ai/skills

    Shine a light into a wayfinder map's fog — work one direction now, out of frontier order, or redraw the map itself.

    168 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Grill Design

    will-ness-ai/skills

    Converge on a frontend look through rounds of prototypes and grilling verdicts.

    168 GitHub stars~254 tokensUpdated today
    Auto-check passed
  • Find Standards

    will-ness-ai/skills

    Find how this problem is already solved: standards we could adopt, and the industry's best practices.

    168 GitHub stars~585 tokensUpdated today
    Auto-check passed

Categories

Questions about Test A Skill

What does Test A Skill do?

Field-test a skill by running it cold in parallel agent sessions, then turn what they hit into edits. Test A Skill is an agent skill from will-ness-ai/skills. Field-test a skill by running it cold in parallel agent sessions, then turn what they hit into edits.

When should I use Test A Skill?

Test A Skill fits situations like: tasks that involve Subagents.

How do I install Test A Skill in Claude Code?

Run `npx skills add will-ness-ai/skills --skill test-a-skill -a claude-code`. Or copy the skill folder (skills/test-a-skill in will-ness-ai/skills) into .claude/skills/test-a-skill in your project. Claude Code loads it when a task matches its description.

How do I install Test A Skill in Codex?

Run `npx skills add will-ness-ai/skills --skill test-a-skill -a codex`. Or copy the skill folder (skills/test-a-skill in will-ness-ai/skills) into .agents/skills/test-a-skill in your project. Codex loads it when a task matches its description.

Can I use Test A Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add will-ness-ai/skills --skill test-a-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-a-skill, .gemini/skills/test-a-skill, .github/skills/test-a-skill and .opencode/skills/test-a-skill in your project.

What does Test A Skill need to run?

Going by SKILL.md and its folder, Test A Skill needs the command-line tools its instructions call (claude).

Does Test A Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test A Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test A Skill use?

Test A Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test A Skill use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test A Skill?

Skills that share tags, products or a category with Test A Skill: Claude Code Agent Development (anthropics/claude-plugins-official, 38k stars), Subagent Driven Development (Asvarox/allkaraoke, 261 stars), Dispatching Parallel Agents (ultralisp/ultralisp, 258 stars) and Paseo Advisor Second Opinion (getpaseo/paseo, 20k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test A Skill?

will-ness-ai (a GitHub user) maintains it in will-ness-ai/skills, which has 168 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.

Source: will-ness-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.