MoAI TDD Workflow
modu-ai/moai-adk
Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.
This skill should be used when the user wants to implement features or fix bugs using test-driven development.
$ npx skills add glebis/claude-skills --skill tdd -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install glebis/claude-skills tdd --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tdd .claude/skills/tdd && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tdd" agent skill from https://github.com/glebis/claude-skills/tree/main/tdd into .claude/skills/tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tdd", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/glebis/claude-skills/tree/main/tddType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add glebis/claude-skills --skill tdd -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install glebis/claude-skills tdd --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/tdd .agents/skills/tdd && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tdd" agent skill from https://github.com/glebis/claude-skills/tree/main/tdd into .agents/skills/tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tdd", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add glebis/claude-skills --skill tdd -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install glebis/claude-skills tdd --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/tdd .cursor/skills/tdd && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tdd" agent skill from https://github.com/glebis/claude-skills/tree/main/tdd into .cursor/skills/tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tdd", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/glebis/claude-skills.git --path tdd--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add glebis/claude-skills --skill tdd -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install glebis/claude-skills tdd --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/tdd .gemini/skills/tdd && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tdd" agent skill from https://github.com/glebis/claude-skills/tree/main/tdd into .gemini/skills/tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tdd", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install glebis/claude-skills tddInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add glebis/claude-skills --skill tdd -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/tdd .github/skills/tdd && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tdd" agent skill from https://github.com/glebis/claude-skills/tree/main/tdd into .github/skills/tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tdd", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add glebis/claude-skills --skill tdd -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install glebis/claude-skills tdd --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/tdd .opencode/skills/tdd && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tdd" agent skill from https://github.com/glebis/claude-skills/tree/main/tdd into .opencode/skills/tdd/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tdd", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tddThis skill should be used when the user wants to implement features or fix bugs using test-driven development.
TDD is an agent skill from glebis/claude-skills. This skill should be used when the user wants to implement features or fix bugs using test-driven development. Enforces the RED-GREEN-REFACTOR cycle with vertical slicing, context isolation between test writing and implementation, human checkpoints, and auto-test feedback loops. Uses multi-agent orchestration with the Task tool for architecturally enforced context isolation. Supports Jest, Vitest, pytest, Go test, cargo test, PHPUnit, and RSpec.
Its SKILL.md is about 8.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 37 other files, including scripts and reference files (for example `.claude-plugin/plugin.json`, `README.md` and `references/agent_prompts.md`).
It sits in Testing & QA, covering Test-driven development, Unit testing and Multi-agent orchestration. It works with Jest, pytest and Vitest. The repository describes itself as: Collection of Claude Code skills for enhanced AI workflows. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3b88261. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Shell, Python and Go, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
npxbashcargopythongopytestnpmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx and npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
TDD loads about 8.4k tokens when it runs, and up to ~18k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 3,510 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from glebis/claude-skills at commit 3b88261, republished under its MIT licence (© glebis). 3,510 words, ~8,446 tokens.
.claude/skills/tdd/SKILL.md (or your agent's skills folder). This skill also uses 30 other files; get the full folder from GitHub.Enforce disciplined RED-GREEN-REFACTOR cycles using separate subagents for test writing and implementation. The core innovation: the Test Writer never sees implementation code, and the Implementer never sees the specification. This prevents the LLM from leaking implementation intent into test design.
/tdd with a feature description or bug report| Invocation | Behavior |
|---|---|
/tdd <feature> | Interactive mode — pause for approval at slices and each RED checkpoint |
/tdd --auto <feature> | Autonomous mode — run all slices without pausing; stop ONLY on unrecoverable errors |
/tdd --resume | Resume from .tdd-state.json in project root |
/tdd --dry-run <feature> | Validation mode — runs Phase 0 + Phase 1 fully, renders all prompts, but skips Task() calls. No code is written. |
In --auto mode, skip all [HUMAN CHECKPOINT] steps. Print status lines instead:
[auto] RED slice 1/4: "validates email format" — test failing as expected
[auto] GREEN slice 1/4: passing (attempt 1)
[auto] REFACTOR slice 1/4: 1 suggestion applied, 0 skippedStop and ask the user ONLY when:
In --dry-run mode, validate the entire orchestration pipeline without executing any subagents or writing any code:
Task() calls are made. No test files are written. No implementation code is generated.{UNRESOLVED} placeholders), all scripts execute without error, and the state file is well-formed.DRY RUN COMPLETE: {feature name}
Phase 0:
Framework: {framework}
Language: {language}
Baseline: {pass|greenfield}
API surface: {line count} lines
Doc context: {line count} lines (or "none")
Phase 1:
Slices: {N} ({layer breakdown})
Prompts rendered: {N * 3} (all variables resolved)
Test Writer: {char count} chars
Implementer: {char count} chars
Refactorer: {char count} chars
State file: .tdd-state.json written
No code was modified.This mode is useful for:
ORCHESTRATOR (you, reading this file)
├─ Phase 0: Setup — detect framework, extract API, create state file
├─ Phase 1: Decompose into vertical slices → user approves
│
├─ FOR EACH SLICE:
│ ├─ Phase 2 (RED): Task(Test Writer) ← spec + API only
│ ├─ Phase 3 (GREEN): Task(Implementer) ← failing test + error only
│ └─ Phase 4 (REFACTOR): Task(Refactorer) ← all code + green results
│
└─ Summary| Agent | Sees | Does NOT See |
|---|---|---|
| Test Writer | Slice spec, public API signatures, framework conventions, layer constraints | Implementation code, other slices, implementation plans |
| Implementer | Failing test code, test failure output, file tree, existing source, layer constraints | Original spec, slice descriptions, future plans |
| Refactorer | All implementation + all tests + green results, layers touched | Original spec, decomposition rationale |
Step 1: Detect framework and test runner.
Check for: package.json (jest/vitest), pyproject.toml/pytest.ini (pytest),
go.mod (go test), Cargo.toml (cargo test), Gemfile (rspec), composer.json (phpunit)If ambiguous, ask: "What command runs your tests? (e.g., npm test, pytest)"
Step 2: Detect language from source files (for agent prompts):
TypeScript (.ts/.tsx), JavaScript (.js/.jsx), Python (.py), Go (.go), Rust (.rs), Ruby (.rb), PHP (.php)Step 3: Verify green baseline.
bash ~/.claude/skills/tdd/scripts/run_tests.sh {FRAMEWORK} "{TEST_COMMAND}"Parse the JSON output.
status is "pass": proceed.status is "fail": stop — "Existing tests are failing. TDD starts from a green baseline."status is "error" AND total is 0: greenfield project — no tests exist yet. This is fine. Proceed.Step 4: Extract the public API surface.
bash ~/.claude/skills/tdd/scripts/extract_api.sh {SOURCE_DIR}Save the output — this is what the Test Writer will see. If empty (greenfield), that's expected.
Step 5: Discover project documentation.
bash ~/.claude/skills/tdd/scripts/discover_docs.sh {PROJECT_ROOT} --lang {LANGUAGE}This searches for:
/// commentsSave the output as {DOC_CONTEXT}. This feeds into:
If empty (no docs found), that's fine — proceed without doc context.
Step 6: Create the state file .tdd-state.json in the project root:
{
"feature": "user's feature description",
"framework": "jest|vitest|pytest|go|cargo|rspec|phpunit",
"language": "typescript|javascript|python|go|rust|ruby|php",
"test_command": "the full test command",
"source_dir": "src/",
"doc_context": "output from discover_docs.sh (or empty string)",
"auto_mode": false,
"dry_run": false,
"slices": [],
"current_slice": 0,
"phase": "setup",
"layer_map": {},
"files_modified": [],
"test_files_created": []
}Each slice in the slices array includes a layer field: "domain", "domain-service", "application", or "infrastructure". See Phase 1 for how layers are assigned.
The layer_map maps directory prefixes to layers. Built during Phase 1 from project structure:
{
"layer_map": {
"src/domain/": "domain",
"src/services/": "domain-service",
"src/application/": "application",
"src/infrastructure/": "infrastructure",
"src/adapters/": "infrastructure",
"src/controllers/": "infrastructure"
}
}If the project has no clear directory-layer mapping (flat structure), set layer_map to {} and skip path-based validation.
Step 5a (auto-detect layer_map): If layer_map is empty, scan the source directory for common DDD/layered architecture directory names and auto-populate:
Common directory → layer mappings (check if directories exist):
*/domain/ → "domain"
*/models/ → "domain" (ORM models often serve as domain entities)
*/entities/ → "domain"
*/value_objects/ → "domain"
*/services/ → "application" (unless clearly infrastructure)
*/application/ → "application"
*/use_cases/ → "application"
*/core/ → "application"
*/infrastructure/ → "infrastructure"
*/adapters/ → "infrastructure"
*/controllers/ → "infrastructure"
*/api/ → "infrastructure"
*/bot/ → "infrastructure" (Telegram/Discord bot handlers)
*/handlers/ → "infrastructure"
*/repositories/ → "infrastructure" (concrete repo implementations)Only add entries for directories that actually exist in the source tree. If fewer than 2 directories match, leave layer_map empty (flat project). Present the auto-detected map to the user for confirmation:
Auto-detected layer map from directory structure:
src/models/ → domain
src/services/ → application
src/core/ → application
src/bot/ → infrastructure
src/api/ → infrastructure
Does this mapping look correct? (adjust if needed)Update state: "phase": "setup". Write state file immediately.
Take the user's feature request and decompose into ordered vertical slices. Each slice is one testable behavior.
Use doc context: When decomposing, cross-reference {DOC_CONTEXT} from Phase 0 Step 5. Documentation often describes intended behaviors, edge cases, and API contracts that should inform slice boundaries. If docs mention specific error cases, validation rules, or behavioral requirements, consider them as slice candidates.
After identifying all slices, sort them inside-out by architectural layer. This ensures each slice can build on real (not mocked) implementations from previous slices:
Assign each slice a layer tag: domain, domain-service, application, or infrastructure. Use the heuristics from references/layer_guide.md to classify.
Why inside-out? Domain slices produce real objects that later slices use directly. This minimizes mocking and catches integration issues early. It also ensures business rules are implemented and tested before any infrastructure decisions are made.
For simple projects where all code lives in one layer, all slices get layer: "application" and the ordering doesn't change — the guidance degrades gracefully.
Infrastructure-only features (e.g., "add email provider retry logic", "switch from Postgres to MySQL"):
infrastructure. This is valid — skip the inner layers entirely.Missing port interface (domain-service needs a port that doesn't exist yet):
domain-service slice for RegistrationService creates domain/ports/UserRepository interface as part of GREEN.Cross-cutting slices (a slice touches multiple layers):
application but creates a file in domain/events/.Present to the user:
I've broken this into N vertical slices (ordered inside-out):
Domain:
1. [behavior] — [what the test verifies]
Domain Services:
2. [behavior] — [what the test verifies]
Application:
3. [behavior] — [what the test verifies]
Infrastructure:
4. [behavior] — [what the test verifies]
Each slice follows RED -> GREEN -> REFACTOR before moving to the next.
Does this decomposition look right?If all slices fall in one layer, skip the layer headings and present as a flat list.
Wait for user approval (even in --auto mode — slice decomposition always needs sign-off).
Update state: Write slices array (each with layer field), set "phase": "decomposed".
In --dry-run mode, replace Phases 2–4 entirely with the following for each slice:
extract_api.sh)### Test Writer Prompt (slice N) heading."(dry-run: test code would be generated by Test Writer)" for {FAILING_TEST_CODE} and "(dry-run: no test output)" for {TEST_FAILURE_OUTPUT}."(dry-run: no green output)" for {GREEN_TEST_OUTPUT}, "(dry-run: code from Test Writer)" for {ALL_TEST_CODE}, "(dry-run: code from Implementer)" for {ALL_IMPLEMENTATION_CODE}.{UNRESOLVED_VARIABLE} patterns remain (regex: \{[A-Z][A-Z_]+\}). Report any unresolved variables as errors.Task() calls, no file writes, no test runs).After all slices are processed, print the dry-run summary and exit. Do NOT clean up the state file — it's useful for subsequent --resume.
Step 1: Refresh the API surface (it changes as slices are implemented):
bash ~/.claude/skills/tdd/scripts/extract_api.sh {SOURCE_DIR}Step 2: Read the prompt template from references/agent_prompts.md -> "Test Writer Agent" section. Construct the prompt by filling in:
{SLICE_SPEC}: The current slice's behavior description{LANGUAGE}: Detected language from Phase 0{FRAMEWORK}: Detected framework name{API_SURFACE}: Output from extract_api.sh{DOC_CONTEXT}: Output from discover_docs.sh (Phase 0 Step 5). Include only sections relevant to the current slice — filter by keyword match if the full output is large.{TEST_FILE_PATH}: Where the test should go (follow project conventions){EXISTING_TEST_CONTENT}: Current content of the test file (if it exists), or "No test file exists yet."{FRAMEWORK_SKELETON}: The relevant skeleton from references/framework_configs.md{LAYER}: The slice's layer tag from Phase 1{LAYER_TEST_CONSTRAINTS}: Layer-specific test constraints (see agent_prompts.md -> Layer-Specific Constraint Lookup)Step 3: Launch the Test Writer agent:
Task(subagent_type="general-purpose", prompt=<constructed prompt>)Step 4: Parse the JSON response using the parse_agent_json logic from agent_prompts.md:
{ and last }, try that substringStep 5: Write the test code to the test file. If the file exists, append the test function (and merge imports). If new, create with the agent's imports_needed + test_code.
Step 5a (post-write test smell scan): Scan the test code for common smells before running:
| Smell | Detection | Action |
|---|---|---|
| Assertion Roulette | Multiple bare assert statements without messages in the same test function (3+) | Warn the user (don't block): "Test has N bare assertions — consider adding failure messages for easier debugging." |
| Unknown Test | Test name is generic: matches test_1, test_it, test_works, test_example, test_thing | Re-launch Test Writer with appended: "Use a descriptive test name that reads as a behavior spec (e.g., test_rejects_empty_email)." |
| Tautological assertion | assert True, assert result is not None when function has no None return path, assert isinstance(result, X) as sole assertion | Re-launch Test Writer with appended: "The assertion is tautological — test the actual behavior/value, not just that the function returns something." |
Step 5b (post-write layer lint): Scan the test code for layer-violating patterns:
| Layer | Forbidden patterns in test code |
|---|---|
domain | jest.mock(, vi.mock(, Mock(, mock.patch, unittest.mock, gomock, mockery — domain tests must not use mocking libraries |
domain-service | Same mocking patterns for domain objects (mocking ports/repos is OK) |
application | No forbidden patterns (mocking ports is expected) |
infrastructure | No forbidden patterns |
If forbidden patterns found:
Step 6: Run the test to confirm it FAILS (expect an assertion failure, not a setup error):
bash ~/.claude/skills/tdd/scripts/run_tests.sh {FRAMEWORK} "{TEST_COMMAND_FOR_SPECIFIC_TEST}"Step 7: Evaluate the result with semantic validation:
| Result | Action |
|---|---|
status: "fail", assertion error | Proper RED — test fails because the expected behavior doesn't exist yet. Proceed. |
status: "fail", ImportError / ModuleNotFoundError | Setup problem, not a proper RED. The test can't even import the module under test. Fix: create a minimal stub (empty class/function) so the import resolves, then re-run. The test should now fail on the assertion instead. |
status: "fail", AttributeError on missing method | Similar to import error — the class exists but the method doesn't. This is an acceptable RED if the assertion would also fail. Proceed. |
status: "pass" | Behavior already exists. Log: "Test passes — skipping slice (already implemented)." Increment current_slice, move to next slice. |
status: "error", SyntaxError | Fix: the test has a typo. Read the raw_tail, fix the test file directly. Re-run. If still erroring after 2 fix attempts, ask user. |
status: "error", compile/framework error | Fix: bad import, missing fixture, or framework misconfiguration. Read the raw_tail, fix the test file directly. Re-run. If still erroring after 2 fix attempts, ask user. |
Step 8 (interactive mode only — skip in --auto): Present to the user:
RED: Test written and failing as expected.
Test: {test_name}
File: {test_file_path}
Failure: {failure message from JSON}
This test verifies: {test_description from agent response}
Proceed to GREEN phase? (or adjust the test?)Wait for user approval before proceeding to GREEN.
Update state: "phase": "red", add test file to test_files_created. Write state immediately.
Step 1: Read the failing test file and the test failure output (the full raw_tail from the RED phase run_tests.sh result).
Step 2: Build the file tree of source files (not test files, not node_modules, etc.):
find {SOURCE_DIR} -type f \( -name '*.ts' -o -name '*.js' -o -name '*.py' -o -name '*.go' -o -name '*.rs' -o -name '*.rb' -o -name '*.php' \) | grep -v test | grep -v spec | grep -v node_modules | grep -v __pycache__ | grep -v vendor | grep -v target | grep -v dist | grep -v build | head -50Step 3: Read existing source files that the test imports or references.
Step 4: Read the prompt template from references/agent_prompts.md -> "Implementer Agent" section. Fill in:
{LANGUAGE}: Detected language{FAILING_TEST_CODE}: The complete test file content{TEST_FAILURE_OUTPUT}: The raw_tail from run_tests.sh JSON output{FILE_TREE}: Source file listing from Step 2{EXISTING_SOURCE}: Content of relevant source files (if any — may be empty for greenfield){LAYER}: The slice's layer tag from Phase 1{LAYER_DEPENDENCY_CONSTRAINT}: Layer-specific dependency constraint (see agent_prompts.md -> Layer-Specific Constraint Lookup)On retries (attempt > 1), also fill in the {?PREVIOUS_ATTEMPT} section:
{PREVIOUS_ATTEMPT_DESCRIPTION}: the explanation field from the failed attempt{PREVIOUS_ATTEMPT_ERROR}: the raw_tail from the test run after the failed attemptCRITICAL: Do NOT include the slice specification, feature description, or any future plans. The Implementer works from the test alone.
Step 5: Launch the Implementer agent:
Task(subagent_type="general-purpose", prompt=<constructed prompt>)Step 6: Parse the JSON response. Validate layer boundaries, then apply file changes.
Step 6a (Layer path validation): If layer_map is not empty, check each file path in the response against the current slice's layer:
For each file in response.files:
inferred_layer = lookup file.path against layer_map (longest prefix match)
if inferred_layer exists AND inferred_layer != current_slice.layer:
if inferred_layer is OUTER relative to current_slice.layer:
REJECT: "Implementer created/modified {file.path} which belongs to
the {inferred_layer} layer, but this is a {current_slice.layer} slice.
Inner layers must not depend on outer layers."
→ Re-launch Implementer with appended constraint:
"Do NOT create or modify files in {inferred_layer} directories.
This slice is {current_slice.layer} only."
if inferred_layer is INNER relative to current_slice.layer:
ALLOW: outer layers may touch inner-layer files (e.g., adding a port interface)Layer ordering for "outer" check: domain < domain-service < application < infrastructure.
If layer_map is empty (flat project), skip this validation.
Step 6b: Apply validated file changes:
For each file in the response files array:
action is "create" or "overwrite": Use the Write tool to create or overwrite the file with the complete contentaction is "edit" (used for existing files over 200 lines): Use the Edit tool with old_string → new_string to apply the changes. The Implementer returns only the changed functions with surrounding context — identify the insertion point or the function being replaced, and use Edit tool accordingly. If the edit target is ambiguous, fall back to reading the full file and using Write.Step 7: Run the specific test:
bash ~/.claude/skills/tdd/scripts/run_tests.sh {FRAMEWORK} "{TEST_COMMAND_FOR_SPECIFIC_TEST}"Step 8: RETRY LOOP (if test still fails):
attempt = 1
max_attempts = 5
previous_explanation = null
previous_error = null
while status != "pass" AND attempt <= max_attempts:
previous_explanation = explanation from last Implementer response
previous_error = raw_tail from last test run
Launch FRESH Task(Implementer) with:
- same test code + file tree + existing source (re-read!)
- NEW failure output
- PREVIOUS_ATTEMPT section filled in
Apply changes (Write tool for each file)
Re-run test
attempt += 1
if still failing after max_attempts:
STOP. Present to user:
"Implementation failed after 5 attempts. Last error: {raw_tail}"
Ask: "Adjust the test, try a different approach, or debug manually?"Each retry is a fresh Task call with only the previous attempt's explanation and error. This prevents the Implementer from going down rabbit holes while giving it enough context to try a different strategy.
Step 9: Once the specific test passes, run the FULL test suite:
bash ~/.claude/skills/tdd/scripts/run_tests.sh {FRAMEWORK} "{FULL_TEST_COMMAND}" --allStep 10: Handle regressions:
| Result | Action |
|---|---|
| All pass | Proceed to REFACTOR |
| Regressions found | Auto-fix: launch a fresh Implementer with the regression test failures. Apply. Re-run full suite. Repeat up to 3 times. If still failing after 3 regression-fix attempts, STOP and present to user. |
Step 11 (interactive mode only — skip in --auto): Present to the user:
GREEN: Test passing with minimal implementation.
Implementation: {explanation from agent response}
Files changed: {list}
All tests: {passed} passing, {failed} failing
Proceed to REFACTOR phase? (or adjust?)Update state: "phase": "green", update files_modified. Write state immediately.
Step 12 (domain/domain-service slices only): Layer purity check before REFACTOR:
For each new/modified file in a domain or domain-service layer slice:
layer_map. Flag any import from an outer layer as a violation.Step 13: Full-repo import scan (all layers, runs once per slice):
Scan ALL source files (not just session-modified) for dependency direction violations:
# For each source file, extract imports and check against layer_map
# Language-specific patterns:
# Python: from X import Y, import X
# TypeScript/JS: import ... from 'X', require('X')
# Go: import "X"For each file:
layer_map (skip if no match)layer_mapReport violations to the user before REFACTOR:
Layer scan found N dependency direction violation(s):
- domain/user.py imports infrastructure/db.py (domain → infrastructure)
- domain/services/registration.py imports adapters/email.py (domain-service → infrastructure)In --auto mode: attempt auto-fix (replace concrete import with port interface). In interactive mode: present violations and ask user how to proceed.
This supplements the Refactorer's import checking (which only sees session files) with a repo-wide scan. Static tools miss ~23% of violations (Pruijt et al., 2017) — combining textual + structural checks improves coverage.
Step 1: Gather all context:
Step 2: Read the prompt template from references/agent_prompts.md -> "Refactorer Agent" section. Fill in:
{LANGUAGE}: Detected language{GREEN_TEST_OUTPUT}: Full test output showing all green{ALL_TEST_CODE}: Content of all test files{ALL_IMPLEMENTATION_CODE}: Content of all modified source files{SLICE_LAYERS}: Comma-separated list of unique layers from all slices completed so farStep 3: Launch the Refactorer agent:
Task(subagent_type="general-purpose", prompt=<constructed prompt>)Step 4: Parse the JSON response. If suggestions is empty, skip to Step 6.
Apply suggestions one at a time, in priority order (high first):
For each suggestion:
old_code -> new_code for each file)python -m black --check {files} && python -m flake8 {files} && python -m mypy {files}npx eslint {files} or npx tsc --noEmitgo vet ./...cargo clippyStep 5 (interactive mode only — skip in --auto): Present:
REFACTOR: Code improved, all tests still passing.
Applied: {list of accepted suggestions}
Skipped: {list of reverted suggestions, if any}
All tests: {count} passing
[Moving to slice N of M] or [All slices complete]In --auto mode, print one-liner:
[auto] REFACTOR slice N/M: {applied_count} applied, {skipped_count} skippedUpdate state: "phase": "refactor". Write state immediately.
If more slices remain -> increment current_slice in state, return to Phase 2.
If all slices complete -> present summary:
TDD Complete: {feature name}
Slices implemented: N
Tests written: N
Files created/modified: {list}
All tests passing: yesClean up: remove .tdd-state.json (in --auto mode, remove silently; in interactive, ask user).
When user invokes /tdd --resume:
.tdd-state.json from project rootauto_mode is true in state, continue in auto modeNo source files, no tests, no test configuration. Handle gracefully:
status: "error" with total: 0, check if any test files exist. If none, this is greenfield — proceed."(No existing API — this is a new project)" to the Test Writer.jest.config.js, pytest.ini). If the first test run fails with a framework error (not a test failure), create minimal framework config and retry.If user provides test code:
If a test sometimes passes/fails: stop, report, fix the flaky test before continuing.
| Failure | Phase | Recovery |
|---|---|---|
| Test Writer returns invalid JSON | RED | Parse with fence-stripping + substring extraction. Retry once with "Return ONLY JSON." Fall back to manual extraction. |
| Test passes when it should fail | RED | Log "already implemented", skip slice, move to next. |
| Test has syntax/compile error | RED | Read raw_tail, fix test file directly. Retry up to 2 times. Then ask user. |
| Implementer returns invalid JSON | GREEN | Same JSON recovery as Test Writer. |
| Test still fails after implementation | GREEN | Retry loop: up to 5 fresh Implementer calls with previous-attempt context. Then ask user. |
| Full suite has regressions | GREEN | Auto-fix: fresh Implementer with regression failures. Up to 3 attempts. Then ask user. |
| Refactorer suggestion breaks tests | REFACTOR | Revert immediately, skip suggestion, continue with next. |
| run_tests.sh timeout | Any | Increase timeout. If persistent, ask user about test performance. |
run_tests.sh returns "error" | Any | Read raw_tail for cause. Script error (missing binary, bad path) -> fix and retry. Compilation error -> treat as implementation error. |
| extract_api.sh returns empty | RED | Normal for greenfield. Pass "(No existing API)" message. |
| Agent response is completely empty | Any | Retry the Task call once. If still empty, ask user. |
See references/layer_guide.md for layer definitions, dependency rules, test strategies by layer, and detection heuristics.
See references/anti_patterns.md. Critical ones:
See references/framework_configs.md for setup details.
| Framework | Run single test | Run all | Watch mode |
|---|---|---|---|
| Jest | npx jest --testPathPattern=<file> -t "<name>" | npx jest | npx jest --watch |
| Vitest | npx vitest run <file> -t "<name>" | npx vitest run | npx vitest |
| pytest | pytest <file>::<test_name> -v | pytest -v | pytest-watch |
| Go | go test -run <TestName> ./... | go test ./... | — |
| Cargo | cargo test <test_name> | cargo test | cargo watch -x test |
| RSpec | rspec <file>:<line> | rspec | guard |
| PHPUnit | phpunit --filter <test_name> | phpunit | — |
© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 30 other files (scripts, references) in tdd of glebis/claude-skills.
Open the folder on GitHubat commit 3b88261
TDD next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| TDD this skillglebis/claude-skills | 390 | — | ~8.4k | Automated safety check: Pass | MIT | |
| MoAI TDD Workflowmodu-ai/moai-adk | 1.2k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Test-Driven Development Enforcerzereight/gitlab-mcp | 2k | 1 repos | ~904 | Automated safety check: Pass | MIT | |
| TDD Guidealirezarezvani/claude-skills | 28k | — | ~3.4k | Automated safety check: Pass | MIT | |
| TDD GuideLeoYeAI/openclaw-master-skills | 2.2k | — | ~1.4k | Automated safety check: Pass | MIT | |
| TDD Guideborghei/Claude-Skills | 886 | — | ~1.7k | Automated safety check: Pass | MIT |
modu-ai/moai-adk
Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.
zereight/gitlab-mcp
Enforces strict red-green-refactor, with a failing test first, the minimum code to pass it, then cleanup, and a quick reference for common test runners.
alirezarezvani/claude-skills
Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…
LeoYeAI/openclaw-master-skills
Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…
borghei/Claude-Skills
Guide red-green-refactor TDD with test generation, coverage-gap analysis, and multi- framework support.
aAAaqwq/AGI-Super-Team
Test-driven development workflow with test generation, coverage analysis, and multi-framework support
glebis/claude-skills
Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.
glebis/claude-skills
Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.
glebis/claude-skills
This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.
glebis/claude-skills
This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…
glebis/claude-skills
Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.
glebis/claude-skills
Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.
Categories
This skill should be used when the user wants to implement features or fix bugs using test-driven development. TDD is an agent skill from glebis/claude-skills. This skill should be used when the user wants to implement features or fix bugs using test-driven development.
TDD fits situations like: wants to implement features; fix bugs using test-driven development.
Run `npx skills add glebis/claude-skills --skill tdd -a claude-code`. Or copy the skill folder (tdd in glebis/claude-skills) into .claude/skills/tdd in your project. Claude Code loads it when a task matches its description.
Run `npx skills add glebis/claude-skills --skill tdd -a codex`. Or copy the skill folder (tdd in glebis/claude-skills) into .agents/skills/tdd in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill tdd -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tdd, .gemini/skills/tdd, .github/skills/tdd and .opencode/skills/tdd in your project.
Going by SKILL.md and its folder, TDD needs a shell, Python and Go for the scripts in its folder and the command-line tools its instructions call (npx, bash, cargo, python, go and pytest). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use npx and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
TDD is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.4k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with TDD: MoAI TDD Workflow (modu-ai/moai-adk, 1.2k stars), Test-Driven Development Enforcer (zereight/gitlab-mcp, 2k stars), TDD Guide (alirezarezvani/claude-skills, 28k stars) and TDD Guide (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
glebis (a GitHub user) maintains it in glebis/claude-skills, which has 390 GitHub stars. The repository holds 92 skills in this directory. The repository was last updated on October 8, 2026.
Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.