Implement features using Test-Driven Development (TDD) with Red-Green-Refactor cycle.

MITAuto-check passedTesting & QA

Install TDD

skills CLI
$ npx skills add DeL-TaiseiOzaki/claude-code-orchestra --skill tdd -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DeL-TaiseiOzaki/claude-code-orchestra tdd --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DeL-TaiseiOzaki/claude-code-orchestra.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tdd .claude/skills/tdd && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tdd
GitHub stars
199
Token cost
~2.5k tokens
SKILL.md length
1,006 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Implement features using Test-Driven Development (TDD) with Red-Green-Refactor cycle.

  • Works in 3 steps: Test Design → Red-Green-Refactor → Completion Check
  • Tasks that involve Test-driven development
  • SKILL.md covers TDD Cycle, Who Does What (delegation-first), The Red/Green Invariant and Implementation Steps, plus 2 more sections
  • Calls python3, bash and uv

What it does

TDD is an agent skill from DeL-TaiseiOzaki/claude-code-orchestra. Implement features using Test-Driven Development (TDD) with Red-Green-Refactor cycle.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test-driven development. The licence is MIT.

When your agent uses it

  • Tasks that involve Test-driven development

Example prompts

  • “/tdd”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Test Design
  2. Red-Green-Refactor
  3. Completion Check

What it can do on your machine

Read from SKILL.md and the folder at commit ef0d8f8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • bash
    • uv
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

TDD loads about 2.5k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 1,006 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from DeL-TaiseiOzaki/claude-code-orchestra at commit ef0d8f8, republished under its MIT licence (© DeL-TaiseiOzaki). 1,006 words, ~2,488 tokens.

Download SKILL.mdSave it as .claude/skills/tdd/SKILL.md (or your agent's skills folder).
name
tdd
description
Implement features using Test-Driven Development (TDD) with Red-Green-Refactor cycle.
disable-model-invocation
true

Test-Driven Development

Implement $ARGUMENTS using Test-Driven Development (TDD).

TDD Cycle

Repeat: Red → Green → Refactor

1. Red:    Write a failing test
2. Green:  Write minimal code to pass the test
3. Refactor: Clean up code (tests still pass)

Which test cases matter, and when to refactor, is judgment. The one thing that is not judgment is whether a run came out the way you expected: every test run in this skill goes through .claude/skills/_shared/run_tests.py, which turns "confirm failure" / "confirm success" into an exit code.

Who Does What (delegation-first)

Per .claude/rules/delegation.md, the lead does not write the cycles by hand. The split is fixed:

WorkOwner
Requirement clarification, test-case list, boundary valuesLead (judgment, Self-Handle List item 5)
Writing the tests and the production code, cycle by cyclegeneral-purpose-sonnet
Cycles with ambiguous design, cross-cutting invariants, or security / concurrency / data-integrity riskgeneral-purpose-opus
A Red that is red for the wrong reason, or a Green that will not go green after one retrycodex-debugger
Every run_tests.py / verify.sh gate and the final reportLead — verification never delegates

Only a single trivial cycle on a file already open in context stays with the lead (Self-Handle List item 2). Two or more cycles, or a module the lead has not read, is delegated. Independent modules are delegated in parallel in one message; they share no test file, so nothing serializes them.

The Red/Green Invariant

bash
python3 .claude/skills/_shared/run_tests.py \
  --target tests/test_{module}.py --expect fail --label red-1

--expect fail|pass (required) and --target PATH (repeatable, at least one) are the whole interface; --label names the log file under .claude/logs/ (default run-tests), so a later cycle does not overwrite an earlier log.

Read observed in the JSON — a non-zero exit is not all the same red:

observedpytest exitWhat it actually means
failed1The test ran and failed. This is the only valid Red.
passed0Nothing new is being asserted — the test cannot drive any code.
collection_error2The file did not import (syntax error, bad fixture, missing import): red for the wrong reason. Fix the test, not the production code.
no_tests_collected5Nothing was selected — the test was never written, or the name does not match the discovery pattern.

Exit codes: 0 observed matches --expect; 1 bad arguments — including a --target that does not exist (the mistyped path, caught before pytest runs) and a pytest usage error (exit 4, e.g. an unknown ::node id); 2 observed does not match --expect, or coverage is below --min-coverage; 3 external failure — no pytest runner available, a timeout (--timeout, default 600s), or a pytest internal error. On 1 and 3 the payload's observed is null: the script reports that it has no observation rather than guessing one.

Other payload fields: expected, runner, command, exit_code, summary, failed_tests (node ids from the short summary), coverage_percent, min_coverage, log_file, artifacts, error.

The runner is resolved, not hardcoded: uv run pytest when uv and a pyproject.toml are present, otherwise pytest from PATH, otherwise python -m pytest (probed for importability first). --runner pins one explicitly. If none is available the result is exit 3, never a silent pass.

Implementation Steps

Phase 1: Test Design
  1. Confirm Requirements

    • What is the input
    • What is the output
    • What are the edge cases
  2. List Test Cases — record them with TodoWrite (one todo per case), so the remaining cases survive a context reset instead of living only in this conversation:

    - [ ] Happy path: Basic functionality
    - [ ] Happy path: Boundary values
    - [ ] Error case: Invalid input
    - [ ] Error case: Error handling

Which cases to write, and which boundary values matter, is domain reasoning — never delegate it to a script.

Phase 2: Red-Green-Refactor
Step 0: Hand the Cycles Off

Delegate the Red-Green-Refactor loop with the six-element prompt contract from .claude/rules/delegation.md. One delegation per module, all launched together:

Task tool:
  subagent_type: "general-purpose-sonnet"   # or general-purpose-opus, see the table above
  prompt: |
    Objective: Implement {module} by strict TDD, one test case at a time.

    Scope:
    - Write only tests/test_{module}.py and src/{module}.py. Touch nothing else.
    - Do not modify existing tests, and never skip, delete, or weaken an assertion.

    Inputs:
    - Test cases to implement, in this order: {the Phase 1 list}
    - Read .claude/rules/coding-principles.md and .claude/rules/testing.md first.
    - Existing conventions to follow: {paths of comparable modules}

    Acceptance checks — run these yourself, per cycle, and never skip Red:
      python3 .claude/skills/_shared/run_tests.py \
        --target tests/test_{module}.py --expect fail --label red-{n}
      python3 .claude/skills/_shared/run_tests.py \
        --target tests/test_{module}.py --expect pass --label green-{n}
    A Red that exits 2 with observed=passed/collection_error/no_tests_collected
    is not a Red: fix the test, not the production code, and re-run.

    Output shape:
    ## Cycles completed (case -> red label -> green label)
    ## Files changed
    ## Deviations from the requested test cases, and why
    ## Anything that did not go green

    Context discipline: return this summary only; the run logs stay in
    .claude/logs/ and are referenced by label, not pasted.

Then verify rather than trust: re-run run_tests.py --expect pass yourself and read the diff. A reported cycle with no matching run_tests.py exit 0 did not happen. If a cycle came back unfinished, escalate it (general-purpose-opus or codex-debugger) instead of re-sending the same prompt to the same tier.

The steps below are the contract the delegate follows — and what the lead runs directly in the single-trivial-cycle case.

Show full SKILL.md (373 more words)Show less
Step 1: Write First Test (Red)
python
# tests/test_{module}.py
def test_{function}_basic():
    """Test the most basic case"""
    result = function(input)
    assert result == expected

Confirm the test is red for the right reason:

bash
python3 .claude/skills/_shared/run_tests.py \
  --target tests/test_{module}.py --expect fail --label red-{n}

Exit 0 means it genuinely failed. Do not proceed to Green on any other exit code: on exit 2 read observed (see the table above), on exit 1/3 read error and log_file.

Step 2: Implementation (Green)

Write minimal code to pass the test:

  • Don't aim for perfection
  • Hardcoding is OK
  • Just make the test pass

Confirm success:

bash
python3 .claude/skills/_shared/run_tests.py \
  --target tests/test_{module}.py --expect pass --label green-{n}

Exit 0 means the test now passes. Exit 2 with observed: failed means the implementation is not there yet — read failed_tests.

Step 3: Refactoring (Refactor)

Improve while tests still pass:

  • Remove duplication
  • Improve naming
  • Clean up structure
bash
python3 .claude/skills/_shared/run_tests.py \
  --target tests/test_{module}.py --expect pass --label refactor-{n}

When to refactor, and how far, stays a judgment call.

Step 4: Next Test

Return to Step 1 with the next test case from the Phase 1 list.

Phase 3: Completion Check

Run the full quality gates:

bash
bash .claude/skills/_shared/verify.sh

Read the JSON: overall is pass / fail / no_gates. Exit 0 means overall: pass; exit 2 means a gate failed or no gate could run at all — inspect log_file and tools. no_gates is a failure by default because a code-editing session must not be able to declare done with zero checks executed; if the project genuinely has no configured gates, re-run with --allow-no-gates and verify manually with the project's own commands, and say in the report that you did so. Exit 1 is bad arguments, 3 the log file could not be written.

Then check coverage — with a threshold, so the answer is a gate and not a glance at term-missing output:

bash
python3 .claude/skills/_shared/run_tests.py \
  --target tests/test_{module}.py --expect pass \
  --cov {module} --min-coverage {N} --label coverage

--cov MODULE (repeatable) requires the pytest-cov plugin; without it pytest reports a usage error and run_tests.py exits 1 pointing at log_file. Coverage below --min-coverage is exit 2; --cov with no coverage total in the output is exit 3, never a silent pass. --min-coverage requires --cov. Choosing the threshold is a project decision.

Report Format

markdown
## TDD Complete: {Feature Name}

### Test Cases
- [x] {test1}: {description}
- [x] {test2}: {description}
...

### Coverage
- {coverage_percent}% (threshold {min_coverage}%) — from run_tests.py, label `coverage`

### Quality Gates
- verify.sh: {overall} — `{log_file}`

### Implementation Files
- `src/{module}.py`: {description}
- `tests/test_{module}.py`: {N} tests

Every [x] above must correspond to a run_tests.py --expect pass run that exited 0; mark nothing complete on the strength of having written it.

Notes

  • Write tests first (not after)
  • Keep each cycle small
  • Refactor after tests pass
  • Prioritize working code over perfection
  • Where the production code belongs follows the project's existing layout — that is a design decision, not a derived path

© DeL-TaiseiOzaki, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/tdd of DeL-TaiseiOzaki/claude-code-orchestra.

Open the folder on GitHubat commit ef0d8f8

Compare with similar skills

TDD next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

TDD compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
TDD this skillDeL-TaiseiOzaki/claude-code-orchestra199—~2.5kAutomated safety check: PassMIT
TDDpietheinstrengholt/rssmonster56430 repos~906Automated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT
Test Driven Developmentfarm-fe/farm5.6k52 repos~2.5kAutomated safety check: PassMIT
Tapd Story PipelineTencentBlueKing/bk-bcs840—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • TDD

    pietheinstrengholt/rssmonster

    Test-driven development. An agent skill from pietheinstrengholt/rssmonster.

    564 GitHub starsUsed in 30 repos~906 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • A skill your agent uses when implementing any feature or bugfix, before writing implementation code

    5.6k GitHub starsUsed in 52 repos~2.5k tokens
    Testing & QAAuto-check passed
  • Tapd Story Pipeline

    TencentBlueKing/bk-bcs

    单需求实现流水线——把一个 TAPD 需求从零推进到代码提交。自动串联技术澄清、 开发计划、任务拆分、TDD 实现、架构/安全校验、代码提交六个阶段。

    840 GitHub stars~2.6k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Absolute Init

    maddhruv/absolute

    One-time setup for absolute: interview how you want it to behave (output style, autonomy, TDD strictness, spec dir, families) + detect the stack once, then write .absolute.config.json (project…

    219 GitHub starsUsed in 1 repo~3k tokens
    Testing & QAAuto-check passed

More from DeL-TaiseiOzaki/claude-code-orchestra

All 15 skills in this repo
  • Codex System

    DeL-TaiseiOzaki/claude-code-orchestra

    Codex CLI handles planning, design, and complex code implementation.

    199 GitHub stars~4.2k tokensUpdated 20 days ago
    Auto-check passed
  • Design Tracker

    DeL-TaiseiOzaki/claude-code-orchestra

    Record a project design decision into .claude/docs/DESIGN.md through the shared typed writer.

    199 GitHub stars~2k tokensUpdated 20 days ago
    Auto-check passed
  • Feature

    DeL-TaiseiOzaki/claude-code-orchestra

    Unified feature planning & implementation skill — replaces the old /add-feature and /start-feature skills (both trigger phrases still apply here).

    199 GitHub stars~9.4k tokensUpdated 20 days ago
    Auto-check passed
  • Plan

    DeL-TaiseiOzaki/claude-code-orchestra

    Create a detailed implementation plan for a feature or task.

    199 GitHub stars~2.1k tokensUpdated 20 days ago
    Auto-check passed
  • Catchup

    DeL-TaiseiOzaki/claude-code-orchestra

    Comprehensive onboarding for new or returning contributors. An agent skill from DeL-TaiseiOzaki/claude-code-orchestra.

    199 GitHub stars~1.8k tokensUpdated 20 days ago
    Auto-check passed
  • Checkpointing

    DeL-TaiseiOzaki/claude-code-orchestra

    Save session activity, rebuild rolling PROGRESS.md, and compact stale working blocks in .claude/STATE.md.

    199 GitHub stars~1.8k tokensUpdated 20 days ago
    Auto-check passed

Categories

Questions about TDD

What does TDD do?

Implement features using Test-Driven Development (TDD) with Red-Green-Refactor cycle. TDD is an agent skill from DeL-TaiseiOzaki/claude-code-orchestra. Implement features using Test-Driven Development (TDD) with Red-Green-Refactor cycle.

When should I use TDD?

TDD fits situations like: tasks that involve Test-driven development.

How do I install TDD in Claude Code?

Run `npx skills add DeL-TaiseiOzaki/claude-code-orchestra --skill tdd -a claude-code`. Or copy the skill folder (.claude/skills/tdd in DeL-TaiseiOzaki/claude-code-orchestra) into .claude/skills/tdd in your project. Claude Code loads it when a task matches its description.

How do I install TDD in Codex?

Run `npx skills add DeL-TaiseiOzaki/claude-code-orchestra --skill tdd -a codex`. Or copy the skill folder (.claude/skills/tdd in DeL-TaiseiOzaki/claude-code-orchestra) into .agents/skills/tdd in your project. Codex loads it when a task matches its description.

Can I use TDD in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DeL-TaiseiOzaki/claude-code-orchestra --skill tdd -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tdd, .gemini/skills/tdd, .github/skills/tdd and .opencode/skills/tdd in your project.

What does TDD need to run?

Going by SKILL.md and its folder, TDD needs the command-line tools its instructions call (python3, bash, uv and python). Our summary lists: Python 3.

Does TDD access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is TDD safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does TDD use?

TDD is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does TDD use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to TDD?

Skills that share tags, products or a category with TDD: TDD (pietheinstrengholt/rssmonster, 564 stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), TDD (sanity-io/sanity, 6.4k stars) and Test Driven Development (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains TDD?

DeL-TaiseiOzaki (a GitHub user) maintains it in DeL-TaiseiOzaki/claude-code-orchestra, which has 199 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on September 20, 2026.

Source: DeL-TaiseiOzaki/claude-code-orchestra on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.