Agent skill

TDD Repair

by ruvnet in ruvnet/ruflo

Test-Driven Repair — given a failing test, spawn a bounded headless claude -p (Read/Edit/Bash only) that makes the test pass without modifying it.

MITAuto-check: notesTesting & QA

Install TDD Repair

skills CLI
$ npx skills add ruvnet/ruflo --skill tdd-repair -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo tdd-repair --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-testgen/skills/tdd-repair .claude/skills/tdd-repair && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tdd-repair
GitHub stars
74k
Token cost
~1.6k tokens
SKILL.md length
567 words
Files
1
Skills in repo
265
Repo updated
First seen
Licence
MIT

At a glance

Test-Driven Repair — given a failing test, spawn a bounded headless claude -p (Read/Edit/Bash only) that makes the test pass without modifying it.

  • Works in 5 steps: Pre-flight verify — run the test… → Spawn claude -p with a focused prompt → allowedTools Read,Edit,Bash restricts… → …
  • Tasks that involve Test-driven development
  • SKILL.md covers When to use, When NOT to use, Algorithm and Output shape, plus 5 more sections
  • Calls claude, node and git

What it does

TDD Repair is an agent skill from ruvnet/ruflo. Test-Driven Repair — given a failing test, spawn a bounded headless claude -p (Read/Edit/Bash only) that makes the test pass without modifying it. Modeled on agent-harness-generator's ADR-175 Test-Driven Repair mode. Bounded cost via --max-budget-usd, bounded capability via --allowedTools. Closes the loop the TDD plugins didn't — we generate tests, this fixes the code to satisfy them.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test-driven development and Failing and flaky tests. It works with Bash. The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

When your agent uses it

  • Tasks that involve Test-driven development
  • Tasks that involve Failing and flaky tests

Example prompts

  • “/tdd-repair”

Requirements

  • Node.js
  • Pre-approved tools (allowed-tools): Bash

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Pre-flight verify — run the test command. If it already passes, exit 2 (test-already-passes). Repairing a green test is either a no-op or…
  2. Spawn claude -p with a focused prompt
  3. allowedTools Read,Edit,Bash restricts capability. --max-budget-usd caps cost per attempt. --permission-mode acceptEdits auto-accepts file…
  4. Re-run the test to verify. The test's exit code IS the fitness function — no separate sandbox / LLM-as-judge.
  5. If green: emit success: true + per-attempt usage. If red after --max-attempts: emit success: false + receipts. Either way, the workspace…

What it can do on your machine

Read from SKILL.md and the folder at commit 6c04654. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude
    • node
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

TDD Repair loads about 1.6k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 567 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit 6c04654, republished under its MIT licence (© ruvnet). 567 words, ~1,565 tokens.

Download SKILL.mdSave it as .claude/skills/tdd-repair/SKILL.md (or your agent's skills folder).
name
tdd-repair
description
Test-Driven Repair — given a failing test, spawn a bounded headless `claude -p` (Read/Edit/Bash only) that makes the test pass without modifying it. Modeled on agent-harness-generator's ADR-175 Test-Driven Repair mode. Bounded cost via --max-budget-usd, bounded capability via --allowedTools. Closes the loop the TDD plugins didn't — we generate tests, this fixes the code to satisfy them.
allowed-tools
Bash
argument-hint
--repo <path> --test <path> --test-command <cmd> [--max-attempts 1] [--budget 5.00] [--model haiku] [--confirm]

Surfaces the Test-Driven Repair loop as a ruflo skill. Use when you have a failing test and want the source-under-test fixed automatically, with the test's pass/fail as the verification gate (no LLM-as-judge).

When to use

  • Failing CI test from a recent commit — point this at the test file, get a verified fix (or a clear "couldn't repair within budget" receipt).
  • Local TDD workflow — write the failing test first (tdd-workflow skill), then run tdd-repair to drive the green.
  • Regression triage — a previously-green test went red; before opening an issue, spend ~$1 to see if the fix is trivial.

When NOT to use

  • No failing test exists. Conformant mode (--no-test-oracle) is scoped for a follow-up ADR — needs MCTS over repro generation. For now, write a failing test first.
  • Architectural changes. This skill is for tactical "make red green" fixes. Cross-module refactors that incidentally break tests should be done by a human or a swarm.
  • Untrusted code. The headless claude -p runs with --allowedTools Read,Edit,Bash — no MCP, no network, no arbitrary file writes — but Bash can still touch the filesystem. Don't point this at code you wouldn't git checkout . after.

Algorithm

Implementation: scripts/tdd-repair/tdd-repair.mjs.

  1. Pre-flight verify — run the test command. If it already passes, exit 2 (test-already-passes). Repairing a green test is either a no-op or a --test-command typo.
  2. Spawn claude -p with a focused prompt:
    • Failing test file path (read-only intent)
    • Test command (run only)
    • Hard constraint: do NOT modify the test
    • Hard constraint: do NOT add new dependencies
  3. --allowedTools Read,Edit,Bash restricts capability. --max-budget-usd caps cost per attempt. --permission-mode acceptEdits auto-accepts file edits within the allowed set.
  4. Re-run the test to verify. The test's exit code IS the fitness function — no separate sandbox / LLM-as-judge.
  5. If green: emit success: true + per-attempt usage. If red after --max-attempts: emit success: false + receipts. Either way, the workspace is left as claude -p modified it (caller can git diff to review).

Output shape

json
{
  "success": true,
  "data": {
    "repaired": true,
    "attemptsTaken": 1,
    "mode": "test-driven",
    "before": { "passed": false, "exitCode": 1 },
    "after":  { "passed": true,  "exitCode": 0, "durationMs": 4321 },
    "attempts": [
      { "attempt": 1, "claude": { "ok": true, "durationMs": 38421, "usage": { "cost_usd": 0.0234 } }, "verify": { "passed": true } }
    ],
    "totalCostUsd": 0.0234,
    "budgetUsd": 5.0,
    "budgetExhausted": false,
    "shape": { "repo": "...", "test": "...", "testCommand": "...", "maxAttempts": 1, "model": "haiku" }
  }
}
Show full SKILL.md (247 more words)Show less

Exit codes

CodeMeaning
0Test green after repair (success)
1Test still red after --max-attempts
2Config error (test file missing, test already passes, --no-test-oracle unsupported, etc.)
3Claude CLI exited non-zero (infrastructure failure)
99Reserved for safety tripwire (per ADR-153)

Safety posture

LayerMechanism
Cost cap--max-budget-usd default $5, divided across --max-attempts. Hard ceiling — claude exits when reached.
Capability cap--allowedTools Read,Edit,Bash — no MCP, no network, no arbitrary writes.
Scope capPrompt forbids modifying the test or adding dependencies.
Confirmation gate--confirm REQUIRED — without it, returns dry-run plan (mirrors harness-evolve / harness-mint convention).
Hard timeout15 min total wall-clock; per-attempt budget of timeoutMs / maxAttempts.
Pre-flightRefuses to run if the test already passes (catches --test-command typos).

Inspiration

Modeled on the Test-Driven Repair mode from agent-harness-generator/packages/darwin-mode ADR-175. Key design difference: instead of wrapping metaharness-darwin evolve (population-based search), we drive a single claude -p invocation. Rationale:

  • The test command IS the fitness function — no need for variant scoring
  • claude -p is already in our stack — no new optional dep
  • Bounded cost / capability are first-class flags
  • Resumable via --session-id if iteration is needed

Conformant mode (no test, write own repro via MCTS) is deferred to a future ADR.

Example

bash
# Smoke / dry-run (no --confirm yet)
node plugins/ruflo-testgen/scripts/tdd-repair/tdd-repair.mjs \
  --repo /path/to/myrepo \
  --test tests/auth.test.ts \
  --test-command "npx vitest run tests/auth.test.ts"

# Actually repair (Haiku tier, $5 budget, 1 attempt)
node plugins/ruflo-testgen/scripts/tdd-repair/tdd-repair.mjs \
  --repo /path/to/myrepo \
  --test tests/auth.test.ts \
  --test-command "npx vitest run tests/auth.test.ts" \
  --confirm

# Bigger model + more attempts for harder bugs
node plugins/ruflo-testgen/scripts/tdd-repair/tdd-repair.mjs \
  --repo . --test tests/regression-2456.test.ts \
  --test-command "npm test -- tests/regression-2456.test.ts" \
  --model sonnet --max-attempts 3 --budget 15.00 \
  --confirm

Cost ladder

TierModelPer-attempt typicalUse when
1Haiku$0.02 – $0.20First try — most "make red green" bugs are tactical
2Sonnet$0.30 – $2.00Haiku failed, or the bug has multi-file scope
3Opus$1.50 – $8.00Sonnet failed — architectural reasoning required (rarely worth it for a single failing test)

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-testgen/skills/tdd-repair of ruvnet/ruflo.

Open the folder on GitHubat commit 6c04654

Compare with similar skills

TDD Repair next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

TDD Repair compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
TDD Repair this skillruvnet/ruflo74k—~1.6kAutomated safety check: NotesMIT
Opsmill Dev Test Driving Bugsopsmill/infrahub534—~4kAutomated safety check: PassApache-2.0
CI FixFastLED/FastLED7.5k—~897Automated safety check: PassMIT
Bats Shell Testing Patternswshobson/agents40k12 repos~1.3kAutomated safety check: PassMIT
Test Guidelinesgetsentry/sentry-react-native1.8k—~1.3kAutomated safety check: PassMIT
Test Guidelinesgetsentry/sentry-dart873—~3.1kAutomated safety check: PassMIT

Similar skills

  • Writes a single failing test that reproduces a bug after its root-cause analysis is complete, before any fix is written.

    534 GitHub stars~4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • CI Fix

    FastLED/FastLED

    Scan all CI builds and tests, find failures, fetch error logs, and fix the code.

    7.5k GitHub stars~897 tokensUpdated today
    Testing & QAAuto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Testing & QAAuto-check passed
  • Test Guidelines

    getsentry/sentry-react-native

    Official

    Enforce Sentry React Native SDK test conventions for naming, structure, mocking, and fixtures with Jest.

    1.8k GitHub stars~1.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Test Guidelines

    getsentry/sentry-dart

    Official

    Enforce Sentry Dart/Flutter SDK test conventions for naming, structure, and fixtures.

    873 GitHub stars~3.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Agent Harness Testing Methodology

    huiliyi37/Tianshu-harness

    Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.

    1.1k GitHub stars~1k tokensUpdated today
    Testing & QAAuto-check: notes

More from ruvnet/ruflo

All 265 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 1 repo~830 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 1 repo~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 1 repo~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 1 repo~779 tokens
    Auto-check passed
  • Finds models on the Hugging Face router that lack descriptions in chat-ui's prod.yaml and dev.yaml, researches each one and adds short descriptions.

    74k GitHub starsUsed in 1 repo~600 tokens
    Auto-check passed

Works with

Categories

Questions about TDD Repair

What does TDD Repair do?

Test-Driven Repair — given a failing test, spawn a bounded headless claude -p (Read/Edit/Bash only) that makes the test pass without modifying it. TDD Repair is an agent skill from ruvnet/ruflo. Test-Driven Repair — given a failing test, spawn a bounded headless claude -p (Read/Edit/Bash only) that makes the test pass without modifying it.

When should I use TDD Repair?

TDD Repair fits situations like: tasks that involve Test-driven development; tasks that involve Failing and flaky tests.

How do I install TDD Repair in Claude Code?

Run `npx skills add ruvnet/ruflo --skill tdd-repair -a claude-code`. Or copy the skill folder (plugins/ruflo-testgen/skills/tdd-repair in ruvnet/ruflo) into .claude/skills/tdd-repair in your project. Claude Code loads it when a task matches its description.

How do I install TDD Repair in Codex?

Run `npx skills add ruvnet/ruflo --skill tdd-repair -a codex`. Or copy the skill folder (plugins/ruflo-testgen/skills/tdd-repair in ruvnet/ruflo) into .agents/skills/tdd-repair in your project. Codex loads it when a task matches its description.

Can I use TDD Repair in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill tdd-repair -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tdd-repair, .gemini/skills/tdd-repair, .github/skills/tdd-repair and .opencode/skills/tdd-repair in your project.

What does TDD Repair need to run?

Going by SKILL.md and its folder, TDD Repair needs the command-line tools its instructions call (claude, node and git). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Bash.

Does TDD Repair access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is TDD Repair safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does TDD Repair use?

TDD Repair is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does TDD Repair use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to TDD Repair?

Skills that share tags, products or a category with TDD Repair: Opsmill Dev Test Driving Bugs (opsmill/infrahub, 534 stars), CI Fix (FastLED/FastLED, 7.5k stars), Bats Shell Testing Patterns (wshobson/agents, 40k stars) and Test Guidelines (getsentry/sentry-react-native, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains TDD Repair?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,222 GitHub stars. The repository holds 265 skills in this directory. The repository was last updated on October 10, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.