Agent skill

Subagent Testing

by athola in athola/claude-night-market

Test skills via TDD in fresh subagents. An agent skill from athola/claude-night-market.

MITAuto-check passedTesting & QA

Install Subagent Testing

skills CLI
$ npx skills add athola/claude-night-market --skill subagent-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install athola/claude-night-market subagent-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/abstract/skills/subagent-testing .claude/skills/subagent-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
subagent-testing
GitHub stars
342
Token cost
~837 tokens
SKILL.md length
307 words
Files
2
Skills in repo
160
Repo updated
First seen
Licence
MIT

At a glance

Test skills via TDD in fresh subagents. An agent skill from athola/claude-night-market.

  • Works in 3 steps: Baseline Testing (RED) → With-Skill Testing (GREEN) → Rationalization Testing (REFACTOR)
  • Validating behavior
  • SKILL.md covers When NOT To Use, Overview, Why Fresh Instances Matter and Testing Methodology, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Subagent Testing is an agent skill from athola/claude-night-market. Test skills via TDD in fresh subagents. Use when validating behavior or preventing bias.

Its SKILL.md is about 840 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `modules/testing-patterns.md`).

It sits in Testing & QA, covering Subagents and Test-driven development. The repository describes itself as: 23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context… The licence is MIT.

When your agent uses it

  • Validating behavior
  • Preventing bias

Example prompts

  • “/subagent-testing”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Baseline Testing (RED)
  2. With-Skill Testing (GREEN)
  3. Rationalization Testing (REFACTOR)

What it can do on your machine

Read from SKILL.md and the folder at commit 9f3eb00. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Subagent Testing loads about 837 tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 307 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~837

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from athola/claude-night-market at commit 9f3eb00, republished under its MIT licence (© athola). 307 words, ~837 tokens.

Download SKILL.mdSave it as .claude/skills/subagent-testing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
subagent-testing
description
Test skills via TDD in fresh subagents. Use when validating behavior or preventing bias.
alwaysApply
false
category
testing
tags
testing, validation, TDD, subagents, fresh-instances
token_budget
30
progressive_loading
true
modules
modules/testing-patterns.md
model_hint
standard

Subagent Testing - TDD for Skills

Test skills with fresh subagent instances to prevent priming bias and validate effectiveness.

When NOT To Use

  • Writing the skill under test (use abstract:skill-authoring)
  • A static quality audit with no execution (use abstract:skills-eval)

Overview

Fresh instances prevent priming: Each test uses a new Claude conversation to verify the skill's impact is measured, not conversation history effects.

Why Fresh Instances Matter

The Priming Problem

Running tests in the same conversation creates bias:

  • Prior context influences responses
  • Skill effects get mixed with conversation history
  • Can't isolate skill's true impact
Fresh Instance Benefits
  • Isolation: Each test starts clean
  • Reproducibility: Consistent baseline state
  • Measurement: Clear before/after comparison
  • Validation: Proves skill effectiveness, not priming

Testing Methodology

Three-phase TDD-style approach:

Phase 1: Baseline Testing (RED)

Test without skill to establish baseline behavior.

Phase 2: With-Skill Testing (GREEN)

Test with skill loaded to measure improvements.

Phase 3: Rationalization Testing (REFACTOR)

Test skill's anti-rationalization guardrails.

Quick Start

bash
# 1. Create baseline tests (without skill)
# Use 5 diverse scenarios
# Document full responses

# 2. Create with-skill tests (fresh instances)
# Load skill explicitly
# Use identical prompts
# Compare to baseline

# 3. Create rationalization tests
# Test anti-rationalization patterns
# Verify guardrails work

Detailed Testing Guide

For complete testing patterns, examples, and templates:

Success Criteria

  • Baseline: Document 5+ diverse baseline scenarios
  • Improvement: ≥50% improvement in skill-related metrics
  • Consistency: Results reproducible across fresh instances
  • Rationalization Defense: Guardrails prevent ≥80% of rationalization attempts

See Also

  • skill-authoring: Creating effective skills
  • test-skill: Automated skill testing command

Exit Criteria

  • Baseline (RED) phase documents at least 5 diverse scenarios run in fresh Claude instances without the skill active, with full response text recorded.
  • With-skill (GREEN) phase uses identical prompts in new fresh instances (not continuations of the baseline conversation) and shows >= 50% improvement on skill-related metrics.
  • Rationalization (REFACTOR) phase shows skill guardrails blocking >= 80% of rationalization attempts tested across at least 3 pressure scenarios.
  • Results are reproducible: the same prompts in a new fresh instance produce consistent outcomes, confirming the effect is not conversation-history priming.

© athola, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/abstract/skills/subagent-testing of athola/claude-night-market.

  • SKILL.md
  • modules/testing-patterns.md

Open the folder on GitHubat commit 9f3eb00

Compare with similar skills

Subagent Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Subagent Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Subagent Testing this skillathola/claude-night-market342—~837Automated safety check: PassMIT
Testing Skills With Subagentsed3dai/ed3d-plugins2502 repos~3.5kAutomated safety check: PassNone
Nv Implementnovuhq/novu40k—~1.7kAutomated safety check: PassCustom licence
Tapd Story ImplementTencentBlueKing/bk-bcs840—~1.2kAutomated safety check: PassCustom licence
Implementopen-octo/octo-agent125—~2.3kAutomated safety check: PassMIT
Ab Start Taskayoubben18/ab-method192—~376Automated safety check: PassMIT

Similar skills

  • A skill your agent uses when creating or editing skills, before deployment, to verify they work under pressure and resist rationalization - applies RED-GREEN-REFACTOR cycle to process documentation…

    250 GitHub starsUsed in 2 repos~3.5k tokens
    Testing & QAAuto-check passed
  • Nv Implement

    novuhq/novu

    Implement planned work by fanning out parallel subagents on isolated worktrees — TDD at pre-agreed seams, per-slice nv-park-and-review, merge back, full suite once at the end.

    40k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Tapd Story Implement

    TencentBlueKing/bk-bcs

    迭代执行流水线代码实现阶段。基于 tasks.md 调用 /speckit.implement 以 TDD 模式完成全部任务。

    840 GitHub stars~1.2k tokensUpdated 13 days ago
    Testing & QAAuto-check passed
  • Implement

    open-octo/octo-agent

    Implement a technical design by decomposing it into dependency-ordered vertical slices, executing each with TDD red-green, reviewing each via an isolated sub-agent, and persisting progress to a…

    125 GitHub stars~2.3k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Ab Start Task

    ayoubben18/ab-method

    Run an existing task autonomously to completion — each remaining mission in a subagent with tdd, tracker updated per mission, a commit after every green mission.

    192 GitHub stars~376 tokensUpdated 6 days ago
    Testing & QAAuto-check passed
  • Writing Skills

    ed3dai/ed3d-plugins

    A skill your agent uses when creating new skills, editing existing skills, or verifying skills work before deployment - applies TDD to process documentation by testing with subagents before writing…

    250 GitHub starsUsed in 1 repo~1.3k tokens
    Agent WorkflowsAuto-check passed

More from athola/claude-night-market

All 160 skills in this repo
  • Night Market Diagnostics Toolkit

    athola/claude-night-market

    Run and interpret repo diagnostic scripts (ratchets, validators, token stats).

    342 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Skills Eval

    athola/claude-night-market

    Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market.

    342 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Agent Teams

    athola/claude-night-market

    Coordinates Claude agent teams via filesystem protocol. An agent skill from athola/claude-night-market.

    342 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Delegation Core

    athola/claude-night-market

    Delegates execution to eight CLIs (Gemini, Qwen, MiniMax, GLM, Muse, Codex, OpenCode, Glimmer).

    342 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Elegant Code

    athola/claude-night-market

    Guide minimal code via a decision ladder with full safety, edge, and negative-case coverage.

    342 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Skill Library Mission

    athola/claude-night-market

    Build a project skill library in .claude/skills/ via discovery, parallel authoring, and review.

    342 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed

Questions about Subagent Testing

What does Subagent Testing do?

Test skills via TDD in fresh subagents. An agent skill from athola/claude-night-market. Subagent Testing is an agent skill from athola/claude-night-market. Test skills via TDD in fresh subagents.

When should I use Subagent Testing?

Subagent Testing fits situations like: validating behavior; preventing bias.

How do I install Subagent Testing in Claude Code?

Run `npx skills add athola/claude-night-market --skill subagent-testing -a claude-code`. Or copy the skill folder (plugins/abstract/skills/subagent-testing in athola/claude-night-market) into .claude/skills/subagent-testing in your project. Claude Code loads it when a task matches its description.

How do I install Subagent Testing in Codex?

Run `npx skills add athola/claude-night-market --skill subagent-testing -a codex`. Or copy the skill folder (plugins/abstract/skills/subagent-testing in athola/claude-night-market) into .agents/skills/subagent-testing in your project. Codex loads it when a task matches its description.

Can I use Subagent Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add athola/claude-night-market --skill subagent-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/subagent-testing, .gemini/skills/subagent-testing, .github/skills/subagent-testing and .opencode/skills/subagent-testing in your project.

What does Subagent Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Subagent Testing is instructions for the agent only.

Does Subagent Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Subagent Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Subagent Testing use?

Subagent Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Subagent Testing use?

About 837 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Subagent Testing?

Skills that share tags, products or a category with Subagent Testing: Testing Skills With Subagents (ed3dai/ed3d-plugins, 250 stars), Nv Implement (novuhq/novu, 40k stars), Tapd Story Implement (TencentBlueKing/bk-bcs, 840 stars) and Implement (open-octo/octo-agent, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Subagent Testing?

athola (a GitHub user) maintains it in athola/claude-night-market, which has 342 GitHub stars. The repository holds 160 skills in this directory. The repository was last updated on October 6, 2026.

Source: athola/claude-night-market on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.