Agent skill

Test Driven Development

by hashgraph-online in hashgraph-online/awesome-codex-plugins

A skill your agent uses when implementing a feature, fixing a bug, or changing behavior — write a failing test first, watch it fail, write minimal code to pass, then refactor.

Apache-2.0Auto-check passedTesting & QA

Install Test Driven Development

skills CLI
$ npx skills add hashgraph-online/awesome-codex-plugins --skill test-driven-development -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hashgraph-online/awesome-codex-plugins test-driven-development --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/yyykf/spellbook-skills/skills/test-driven-development .claude/skills/test-driven-development && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-driven-development
GitHub stars
1.2k
Token cost
~2.6k tokens
SKILL.md length
1,357 words
Files
7 (incl. references)
Skills in repo
736
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when implementing a feature, fixing a bug, or changing behavior — write a failing test first, watch it fail, write minimal code to pass, then refactor.

  • Works in 4 steps: Build a Test List (Canon TDD) → Red-Green-Refactor — one list item at a… → Tidy First — separate structural from… → …
  • Implementing a feature
  • SKILL.md covers Overview, Prerequisites, When to Use and The Iron Law, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Test Driven Development is an agent skill from hashgraph-online/awesome-codex-plugins. Use when implementing a feature, fixing a bug, or changing behavior — write a failing test first, watch it fail, write minimal code to pass, then refactor. Adds Kent Beck's Tidy First (separate structural vs behavioral changes) and a Canon test list. Language-agnostic core; per-language notes under references/languages/. Not for throwaway prototypes, generated code, or pure config changes (ask first).

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `references/languages/go.md`, `references/languages/java.md` and `references/languages/python.md`).

It sits in Testing & QA, covering Test-driven development. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.

When your agent uses it

  • Implementing a feature
  • Changing behavior — write a failing test first
  • Write minimal code to pass

Example prompts

  • “/test-driven-development”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Build a Test List (Canon TDD)
  2. Red-Green-Refactor — one list item at a time
  3. Tidy First — separate structural from behavioral changes
  4. Commit Discipline

What it can do on your machine

Read from SKILL.md and the folder at commit 16b4156. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Driven Development loads about 2.6k tokens when it runs, and up to ~18k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 1,357 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~18k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hashgraph-online/awesome-codex-plugins at commit 16b4156, republished under its Apache-2.0 licence (© hashgraph-online). 1,357 words, ~2,611 tokens.

Download SKILL.mdSave it as .claude/skills/test-driven-development/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
test-driven-development
description
Use when implementing a feature, fixing a bug, or changing behavior — write a failing test first, watch it fail, write minimal code to pass, then refactor. Adds Kent Beck's Tidy First (separate structural vs behavioral changes) and a Canon test list. Language-agnostic core; per-language notes under references/languages/. Not for throwaway prototypes, generated code, or pure config changes (ask first).

Test-Driven Development (TDD)

Overview

Write the test first. Watch it fail. Write minimal code to pass. Then tidy.

Core principle: If you didn't watch the test fail, you don't know if it tests the right thing.

Announce at start: "I'm using the test-driven-development skill to drive this change test-first."

This skill is language-agnostic. The discipline below holds in every language; concrete framework, runner, and idiom examples live in references/languages/ — load the one matching your project.

Prerequisites

  • A working test runner for your stack (see the matching references/languages/<lang>.md)
  • Ability to run a single test and watch its output
  • For mocks/test doubles, read references/testing-anti-patterns.md first

When to Use

Always:

  • New features
  • Bug fixes
  • Refactoring
  • Behavior changes

Exceptions (ask your human partner first):

  • Throwaway prototypes (then throw them away and restart with TDD)
  • Generated code
  • Pure configuration changes

Thinking "skip TDD just this once"? Stop. That's rationalization.

The Iron Law

NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST

Wrote code before the test? Delete it. Start over.

No exceptions: don't keep it as "reference", don't "adapt" it while writing the test, don't even look at it. Delete means delete. Implement fresh from the test.

Workflow

Phase 1: Build a Test List (Canon TDD)

Before writing any test, list the behaviors you need to cover — the basic case plus every variant you can think of ("what if the input is empty? what if the service times out? what if the key is missing?").

  • This is behavioral analysis, not implementation design.
  • Keep the list visible. Add to it whenever you discover a new case mid-cycle.
  • Turn exactly one item into a concrete test at a time — never convert the whole list up front (reworking speculative tests when an early decision changes is wasted effort).
Phase 2: Red-Green-Refactor — one list item at a time
pick one item  →  RED  →  verify red  →  GREEN  →  verify green  →  REFACTOR  →  back to the list

Run tests in the quietest mode that still surfaces failures. Verbose pass logs burn context you don't need — quiet runners stay silent on green and print full detail on red, so you keep the failures without the noise. Switch to verbose (-v / -s / --nocapture) only when isolating one test's output. Per-language quiet flags are in each references/languages/<lang>.md.

RED — write one failing test

One behavior, a name that describes that behavior (e.g. shouldRetryThreeTimesThenSucceed, not test1), real code over mocks. Work backwards from the assertion. See your language file for concrete framework syntax.

Verify RED — watch it fail (MANDATORY, never skip)

Run the single test. Confirm:

  • It fails, not errors (no typos/missing imports)
  • It fails for the expected reason — the behavior is missing, not the setup is broken

Test passes already? You're testing existing behavior. Fix the test.

GREEN — minimal code to pass

Write the simplest code that makes this test (and all previous tests) pass. No speculative APIs, no extra options, no "while I'm here" cleanup. YAGNI.

Verify GREEN — watch it pass (MANDATORY)

Run the test. Confirm it passes, all previous tests still pass, and output is pristine (no warnings/errors). Test fails? Fix the code, not the test.

REFACTOR — only on green

Remove duplication, improve names, extract helpers. Keep every test green. Don't add behavior here — that's the next RED.

Then mark the item done and return to the list. Repeat until the list is empty.

Phase 3: Tidy First — separate structural from behavioral changes

Kent Beck's distinction. Every change is one of two kinds — never mix them in one commit:

KindWhat it isReversible?
StructuralRearranging without changing behavior: rename, extract method, move code, organize importsUsually yes
BehavioralAdding or changing what the system actually doesNo
  • When you need both, make the structural change first, on green.
  • Validate a structural change didn't alter behavior by running the tests before and after — they stay green throughout.
  • If a change is a tangle of both, untangle it (or redo it) into a structural step then a behavioral step.
Phase 4: Commit Discipline

Commit only when:

  • All tests pass and all compiler/linter warnings are resolved
  • The change is a single logical unit
  • The commit message states whether it is structural or behavioral

Prefer small, frequent commits over large, infrequent ones. (For this repo's commit format, the git-commit skill applies — structural vs behavioral maps cleanly onto its refactor: vs feat:/fix: types.)

Why Test-First (Not Test-After)

Tests written after code pass immediately — and passing immediately proves nothing: they may test the wrong thing, test the implementation instead of behavior, or miss the edge case you forgot. You never saw them catch anything.

Test-after answers "what does this do?" Test-first answers "what should this do?" — and forces edge-case discovery before you implement. 30 minutes of tests-after gives you coverage but loses the proof the test works.

Show full SKILL.md (594 more words)Show less

Common Rationalizations

ExcuseReality
"Too simple to test"Simple code breaks. The test takes 30 seconds.
"I'll test after"Tests passing immediately prove nothing.
"Tests after achieve the same goals"Tests-after = "what does this do?" Tests-first = "what should this do?"
"Already manually tested"Ad-hoc ≠ systematic. No record, can't re-run.
"Deleting X hours is wasteful"Sunk cost. Keeping unverified code is the real debt.
"Keep it as reference, write tests first"You'll adapt it. That's testing after. Delete means delete.
"Need to explore first"Fine. Throw away the exploration, restart with TDD.
"Hard to test = need to push through"Hard to test = hard to use. Listen to the test; simplify the design.
"TDD will slow me down"TDD is faster than debugging in production.
"This case is different because…"It isn't. Start over with TDD.

Anti-Cheating Checks

Test-first only works if you don't game it. These failures are especially tempting for AI agents under pressure — treat each as a STOP:

  • Deleting or weakening an assertion to make a test "pass" — make it pass for real.
  • Pasting the computed actual value into the expected slot — that defeats the double-check that gives TDD its value. Derive the expected value independently.
  • Writing a test with no assertion just for coverage.
  • Mixing refactoring into the green step — two hats: make it run, then make it right.
  • Branching production code on test-only artifacts (env flags, test IDs) to fake green.

For bug fixes, add at least one test that reproduces the bug and fails first — never fix a bug without a failing test that proves the fix.

Red Flags — STOP and start over

  • Code written before the test
  • Test passes on the very first run
  • Can't explain why the test failed
  • "I already manually tested it"
  • "It's about the spirit, not the ritual"
  • "Keep as reference" / "adapt the existing code"
  • Tidied and changed behavior in the same commit
  • Edited the test to match the code instead of the code to match the test

Language-Specific Guidance

The discipline above is the same everywhere. For framework choice, runner commands, idioms, and language-specific anti-patterns, load only the file matching your project — and only when you actually need that detail:

No file for your language? The core skill is sufficient — apply the same red-green-refactor and Tidy First discipline with your stack's idiomatic runner.

Testing Anti-Patterns

When adding mocks or test utilities, read references/testing-anti-patterns.md — cross-language pitfalls like testing mock behavior instead of real behavior, test-only methods on production classes, and incomplete mocks.

When Stuck

ProblemSolution
Don't know how to test itWrite the wished-for API in the test first. Ask your human partner.
Test is too complicatedThe design is too complicated. Simplify the interface.
Must mock everythingCode is too coupled. Use dependency injection.
Test setup is hugeExtract helpers; if still complex, the design needs simplifying.

Verification Checklist

Before marking work complete:

  • Every new behavior has a test
  • Watched each test fail first, for the expected reason
  • Wrote minimal code to pass each test
  • All tests pass; output is pristine (no warnings/errors)
  • Tests use real code (mocks only when unavoidable; see anti-patterns)
  • Edge cases from the test list are covered
  • Structural and behavioral changes are in separate commits

Can't check every box? You skipped TDD. Start over.

Attribution

  • Red-green-refactor discipline, the Iron Law, rationalization/anti-pattern tables, and testing-anti-patterns.md are adapted from superpowers (MIT).
  • The Canon test list, Tidy First (structural vs behavioral separation), and commit discipline are from Kent Beck — Canon TDD and Augmented Coding: Beyond the Vibes (tidyfirst.substack.com).

© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in plugins/yyykf/spellbook-skills/skills/test-driven-development of hashgraph-online/awesome-codex-plugins.

  • SKILL.md
  • references/languages/go.md
  • references/languages/java.md
  • references/languages/python.md
  • references/languages/rust.md
  • references/languages/typescript.md
  • references/testing-anti-patterns.md

Open the folder on GitHubat commit 16b4156

Compare with similar skills

Test Driven Development next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Driven Development compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Driven Development this skillhashgraph-online/awesome-codex-plugins1.2k—~2.6kAutomated safety check: PassApache-2.0
TDDfossasia/eventyay-interpretation1.6k28 repos~1.1kAutomated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT
Test Driven Developmentfarm-fe/farm5.6k49 repos~2.5kAutomated safety check: PassMIT
Tapd Story PipelineTencentBlueKing/bk-bcs840—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • TDD

    fossasia/eventyay-interpretation

    Test-driven development. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 28 repos~1.1k tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • A skill your agent uses when implementing any feature or bugfix, before writing implementation code

    5.6k GitHub starsUsed in 49 repos~2.5k tokens
    Testing & QAAuto-check passed
  • Tapd Story Pipeline

    TencentBlueKing/bk-bcs

    单需求实现流水线——把一个 TAPD 需求从零推进到代码提交。自动串联技术澄清、 开发计划、任务拆分、TDD 实现、架构/安全校验、代码提交六个阶段。

    840 GitHub stars~2.6k tokensUpdated 13 days ago
    Testing & QAAuto-check passed
  • Absolute Init

    maddhruv/absolute

    One-time setup for absolute: interview how you want it to behave (output style, autonomy, TDD strictness, spec dir, families) + detect the stack once, then write .absolute.config.json (project…

    218 GitHub starsUsed in 1 repo~3k tokens
    Testing & QAAuto-check passed

More from hashgraph-online/awesome-codex-plugins

All 736 skills in this repo
  • Anime Reaction Gif

    hashgraph-online/awesome-codex-plugins

    Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.

    1.2k GitHub stars~922 tokensUpdated yesterday
    Auto-check passed
  • Calibredb

    hashgraph-online/awesome-codex-plugins

    Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).

    1.2k GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Rust API Test Harness

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…

    1.2k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Art

    hashgraph-online/awesome-codex-plugins

    Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…

    1.2k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Calle

    hashgraph-online/awesome-codex-plugins

    Use CALL-E from Codex through the calle CLI. An agent skill from hashgraph-online/awesome-codex-plugins.

    1.2k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Game Balance Economy

    hashgraph-online/awesome-codex-plugins

    Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.

    1.2k GitHub stars~618 tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Test Driven Development

What does Test Driven Development do?

A skill your agent uses when implementing a feature, fixing a bug, or changing behavior — write a failing test first, watch it fail, write minimal code to pass, then refactor. Test Driven Development is an agent skill from hashgraph-online/awesome-codex-plugins. Use when implementing a feature, fixing a bug, or changing behavior — write a failing test first, watch it fail, write minimal code to pass, then refactor.

When should I use Test Driven Development?

Test Driven Development fits situations like: implementing a feature; changing behavior — write a failing test first; write minimal code to pass.

How do I install Test Driven Development in Claude Code?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill test-driven-development -a claude-code`. Or copy the skill folder (plugins/yyykf/spellbook-skills/skills/test-driven-development in hashgraph-online/awesome-codex-plugins) into .claude/skills/test-driven-development in your project. Claude Code loads it when a task matches its description.

How do I install Test Driven Development in Codex?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill test-driven-development -a codex`. Or copy the skill folder (plugins/yyykf/spellbook-skills/skills/test-driven-development in hashgraph-online/awesome-codex-plugins) into .agents/skills/test-driven-development in your project. Codex loads it when a task matches its description.

Can I use Test Driven Development in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill test-driven-development -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-driven-development, .gemini/skills/test-driven-development, .github/skills/test-driven-development and .opencode/skills/test-driven-development in your project.

What does Test Driven Development need to run?

SKILL.md names no scripts, command-line tools or credentials: Test Driven Development is instructions for the agent only. Our summary lists: Python 3.

Does Test Driven Development access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Test Driven Development safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Driven Development use?

Test Driven Development is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Driven Development use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to Test Driven Development?

Skills that share tags, products or a category with Test Driven Development: TDD (fossasia/eventyay-interpretation, 1.6k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), TDD (sanity-io/sanity, 6.4k stars) and Test Driven Development (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Driven Development?

hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,232 GitHub stars. The repository holds 736 skills in this directory. The repository was last updated on October 6, 2026.

Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.