Agent skill

Test Driven Development

by RedWoodOG in RedWoodOG/Hermes-Desktop

A skill your agent uses when implementing any feature or bugfix, before writing implementation code.

MITAuto-check passedTesting & QA

Install Test Driven Development

skills CLI
$ npx skills add RedWoodOG/Hermes-Desktop --skill test-driven-development -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RedWoodOG/Hermes-Desktop test-driven-development --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RedWoodOG/Hermes-Desktop.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/software-development/test-driven-development .claude/skills/test-driven-development && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-driven-development
GitHub stars
177
Used in
6 other repos
Token cost
~2.4k tokens
SKILL.md length
1,033 words
Files
1
Skills in repo
62
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when implementing any feature or bugfix, before writing implementation code.

  • Implementing any feature
  • SKILL.md covers Overview, When to Use, The Iron Law and Red-Green-Refactor Cycle, plus 8 more sections
  • Calls pytest
  • Before writing implementation code

What it does

Test Driven Development is an agent skill from RedWoodOG/Hermes-Desktop. Use when implementing any feature or bugfix, before writing implementation code. Enforces RED-GREEN-REFACTOR cycle with test-first approach.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test-driven development. The licence is MIT.

When your agent uses it

  • Implementing any feature
  • Before writing implementation code

Example prompts

  • “/test-driven-development”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit be46b39. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Driven Development loads about 2.4k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 1,033 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from RedWoodOG/Hermes-Desktop at commit be46b39, republished under its MIT licence (© RedWoodOG). 1,033 words, ~2,398 tokens.

Download SKILL.mdSave it as .claude/skills/test-driven-development/SKILL.md (or your agent's skills folder).
name
test-driven-development
description
Use when implementing any feature or bugfix, before writing implementation code. Enforces RED-GREEN-REFACTOR cycle with test-first approach.
version
1.1.0
author
Hermes Agent (adapted from obra/superpowers)
license
MIT

Test-Driven Development (TDD)

Overview

Write the test first. Watch it fail. Write minimal code to pass.

Core principle: If you didn't watch the test fail, you don't know if it tests the right thing.

Violating the letter of the rules is violating the spirit of the rules.

When to Use

Always:

  • New features
  • Bug fixes
  • Refactoring
  • Behavior changes

Exceptions (ask the user first):

  • Throwaway prototypes
  • Generated code
  • Configuration files

Thinking "skip TDD just this once"? Stop. That's rationalization.

The Iron Law

NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST

Write code before the test? Delete it. Start over.

No exceptions:

  • Don't keep it as "reference"
  • Don't "adapt" it while writing tests
  • Don't look at it
  • Delete means delete

Implement fresh from tests. Period.

Red-Green-Refactor Cycle

RED — Write Failing Test

Write one minimal test showing what should happen.

Good test:

python
def test_retries_failed_operations_3_times():
    attempts = 0
    def operation():
        nonlocal attempts
        attempts += 1
        if attempts < 3:
            raise Exception('fail')
        return 'success'

    result = retry_operation(operation)

    assert result == 'success'
    assert attempts == 3

Clear name, tests real behavior, one thing.

Bad test:

python
def test_retry_works():
    mock = MagicMock()
    mock.side_effect = [Exception(), Exception(), 'success']
    result = retry_operation(mock)
    assert result == 'success'  # What about retry count? Timing?

Vague name, tests mock not real code.

Requirements:

  • One behavior per test
  • Clear descriptive name ("and" in name? Split it)
  • Real code, not mocks (unless truly unavoidable)
  • Name describes behavior, not implementation
Verify RED — Watch It Fail

MANDATORY. Never skip.

bash
# Use terminal tool to run the specific test
pytest tests/test_feature.py::test_specific_behavior -v

Confirm:

  • Test fails (not errors from typos)
  • Failure message is expected
  • Fails because the feature is missing

Test passes immediately? You're testing existing behavior. Fix the test.

Test errors? Fix the error, re-run until it fails correctly.

GREEN — Minimal Code

Write the simplest code to pass the test. Nothing more.

Good:

python
def add(a, b):
    return a + b  # Nothing extra

Bad:

python
def add(a, b):
    result = a + b
    logging.info(f"Adding {a} + {b} = {result}")  # Extra!
    return result

Don't add features, refactor other code, or "improve" beyond the test.

Cheating is OK in GREEN:

  • Hardcode return values
  • Copy-paste
  • Duplicate code
  • Skip edge cases

We'll fix it in REFACTOR.

Verify GREEN — Watch It Pass

MANDATORY.

bash
# Run the specific test
pytest tests/test_feature.py::test_specific_behavior -v

# Then run ALL tests to check for regressions
pytest tests/ -q

Confirm:

  • Test passes
  • Other tests still pass
  • Output pristine (no errors, warnings)

Test fails? Fix the code, not the test.

Other tests fail? Fix regressions now.

REFACTOR — Clean Up

After green only:

  • Remove duplication
  • Improve names
  • Extract helpers
  • Simplify expressions

Keep tests green throughout. Don't add behavior.

If tests fail during refactor: Undo immediately. Take smaller steps.

Repeat

Next failing test for next behavior. One cycle at a time.

Why Order Matters

"I'll write tests after to verify it works"

Tests written after code pass immediately. Passing immediately proves nothing:

  • Might test the wrong thing
  • Might test implementation, not behavior
  • Might miss edge cases you forgot
  • You never saw it catch the bug

Test-first forces you to see the test fail, proving it actually tests something.

"I already manually tested all the edge cases"

Manual testing is ad-hoc. You think you tested everything but:

  • No record of what you tested
  • Can't re-run when code changes
  • Easy to forget cases under pressure
  • "It worked when I tried it" ≠ comprehensive

Automated tests are systematic. They run the same way every time.

"Deleting X hours of work is wasteful"

Sunk cost fallacy. The time is already gone. Your choice now:

  • Delete and rewrite with TDD (high confidence)
  • Keep it and add tests after (low confidence, likely bugs)

The "waste" is keeping code you can't trust.

"TDD is dogmatic, being pragmatic means adapting"

TDD IS pragmatic:

  • Finds bugs before commit (faster than debugging after)
  • Prevents regressions (tests catch breaks immediately)
  • Documents behavior (tests show how to use code)
  • Enables refactoring (change freely, tests catch breaks)

"Pragmatic" shortcuts = debugging in production = slower.

"Tests after achieve the same goals — it's spirit not ritual"

No. Tests-after answer "What does this do?" Tests-first answer "What should this do?"

Tests-after are biased by your implementation. You test what you built, not what's required. Tests-first force edge case discovery before implementing.

Show full SKILL.md (454 more words)Show less

Common Rationalizations

ExcuseReality
"Too simple to test"Simple code breaks. Test takes 30 seconds.
"I'll test after"Tests passing immediately prove nothing.
"Tests after achieve same goals"Tests-after = "what does this do?" Tests-first = "what should this do?"
"Already manually tested"Ad-hoc ≠ systematic. No record, can't re-run.
"Deleting X hours is wasteful"Sunk cost fallacy. Keeping unverified code is technical debt.
"Keep as reference, write tests first"You'll adapt it. That's testing after. Delete means delete.
"Need to explore first"Fine. Throw away exploration, start with TDD.
"Test hard = design unclear"Listen to the test. Hard to test = hard to use.
"TDD will slow me down"TDD faster than debugging. Pragmatic = test-first.
"Manual test faster"Manual doesn't prove edge cases. You'll re-test every change.
"Existing code has no tests"You're improving it. Add tests for the code you touch.

Red Flags — STOP and Start Over

If you catch yourself doing any of these, delete the code and restart with TDD:

  • Code before test
  • Test after implementation
  • Test passes immediately on first run
  • Can't explain why test failed
  • Tests added "later"
  • Rationalizing "just this once"
  • "I already manually tested it"
  • "Tests after achieve the same purpose"
  • "Keep as reference" or "adapt existing code"
  • "Already spent X hours, deleting is wasteful"
  • "TDD is dogmatic, I'm being pragmatic"
  • "This is different because..."

All of these mean: Delete code. Start over with TDD.

Verification Checklist

Before marking work complete:

  • Every new function/method has a test
  • Watched each test fail before implementing
  • Each test failed for expected reason (feature missing, not typo)
  • Wrote minimal code to pass each test
  • All tests pass
  • Output pristine (no errors, warnings)
  • Tests use real code (mocks only if unavoidable)
  • Edge cases and errors covered

Can't check all boxes? You skipped TDD. Start over.

When Stuck

ProblemSolution
Don't know how to testWrite the wished-for API. Write the assertion first. Ask the user.
Test too complicatedDesign too complicated. Simplify the interface.
Must mock everythingCode too coupled. Use dependency injection.
Test setup hugeExtract helpers. Still complex? Simplify the design.

Hermes Agent Integration

Running Tests

Use the terminal tool to run tests at each step:

python
# RED — verify failure
terminal("pytest tests/test_feature.py::test_name -v")

# GREEN — verify pass
terminal("pytest tests/test_feature.py::test_name -v")

# Full suite — verify no regressions
terminal("pytest tests/ -q")
With delegate_task

When dispatching subagents for implementation, enforce TDD in the goal:

python
delegate_task(
    goal="Implement [feature] using strict TDD",
    context="""
    Follow test-driven-development skill:
    1. Write failing test FIRST
    2. Run test to verify it fails
    3. Write minimal code to pass
    4. Run test to verify it passes
    5. Refactor if needed
    6. Commit

    Project test command: pytest tests/ -q
    Project structure: [describe relevant files]
    """,
    toolsets=['terminal', 'file']
)
With systematic-debugging

Bug found? Write failing test reproducing it. Follow TDD cycle. The test proves the fix and prevents regression.

Never fix bugs without a test.

Testing Anti-Patterns

  • Testing mock behavior instead of real behavior — mocks should verify interactions, not replace the system under test
  • Testing implementation details — test behavior/results, not internal method calls
  • Happy path only — always test edge cases, errors, and boundaries
  • Brittle tests — tests should verify behavior, not structure; refactoring shouldn't break them

Final Rule

Production code → test exists and failed first
Otherwise → not TDD

No exceptions without the user's explicit permission.

© RedWoodOG, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/software-development/test-driven-development of RedWoodOG/Hermes-Desktop.

Open the folder on GitHubat commit be46b39

Used in 6 other repositories

We found 7 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 6 other GitHub owners. This page covers the copy in RedWoodOG/Hermes-Desktop, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Test Driven Development next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Driven Development compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Driven Development this skillRedWoodOG/Hermes-Desktop1776 repos~2.4kAutomated safety check: PassMIT
TDDfossasia/eventyay-interpretation1.6k28 repos~1.1kAutomated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT
Test Driven Developmentfarm-fe/farm5.6k49 repos~2.5kAutomated safety check: PassMIT
Tapd Story PipelineTencentBlueKing/bk-bcs840—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • TDD

    fossasia/eventyay-interpretation

    Test-driven development. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 28 repos~1.1k tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • A skill your agent uses when implementing any feature or bugfix, before writing implementation code

    5.6k GitHub starsUsed in 49 repos~2.5k tokens
    Testing & QAAuto-check passed
  • Tapd Story Pipeline

    TencentBlueKing/bk-bcs

    单需求实现流水线——把一个 TAPD 需求从零推进到代码提交。自动串联技术澄清、 开发计划、任务拆分、TDD 实现、架构/安全校验、代码提交六个阶段。

    840 GitHub stars~2.6k tokensUpdated 13 days ago
    Testing & QAAuto-check passed
  • Absolute Init

    maddhruv/absolute

    One-time setup for absolute: interview how you want it to behave (output style, autonomy, TDD strictness, spec dir, families) + detect the stack once, then write .absolute.config.json (project…

    218 GitHub starsUsed in 1 repo~3k tokens
    Testing & QAAuto-check passed

More from RedWoodOG/Hermes-Desktop

All 62 skills in this repo
  • Obliteratus

    RedWoodOG/Hermes-Desktop

    Remove refusal behaviors from open-weight LLMs using OBLITERATUS — mechanistic interpretability techniques (diff-in-means, SVD, whitened SVD, LEACE, SAE decomposition, etc.) to excise guardrails…

    177 GitHub starsUsed in 6 repos~3.8k tokens
    Auto-check passed
  • Excalidraw

    RedWoodOG/Hermes-Desktop

    Create hand-drawn style diagrams using Excalidraw JSON format.

    177 GitHub starsUsed in 5 repos~1.8k tokens
    Auto-check passed
  • Ascii Video

    RedWoodOG/Hermes-Desktop

    Production pipeline for ASCII art video — any format. An agent skill from RedWoodOG/Hermes-Desktop.

    177 GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Systematic Debugging

    RedWoodOG/Hermes-Desktop

    A skill your agent uses when encountering any bug, test failure, or unexpected behavior.

    177 GitHub starsUsed in 6 repos~2.6k tokens
    Auto-check passed
  • Claude Code

    RedWoodOG/Hermes-Desktop

    Delegate coding tasks to Claude Code (Anthropic's CLI agent).

    177 GitHub starsUsed in 4 repos~784 tokens
    Auto-check passed
  • Google Workspace

    RedWoodOG/Hermes-Desktop

    Gmail, Calendar, Drive, Contacts, Sheets, and Docs integration via Python.

    177 GitHub stars~2.1k tokensUpdated 4 mo ago
    Auto-check passed

Categories

Questions about Test Driven Development

What does Test Driven Development do?

A skill your agent uses when implementing any feature or bugfix, before writing implementation code. Test Driven Development is an agent skill from RedWoodOG/Hermes-Desktop. Use when implementing any feature or bugfix, before writing implementation code.

When should I use Test Driven Development?

Test Driven Development fits situations like: implementing any feature; before writing implementation code.

How do I install Test Driven Development in Claude Code?

Run `npx skills add RedWoodOG/Hermes-Desktop --skill test-driven-development -a claude-code`. Or copy the skill folder (skills/software-development/test-driven-development in RedWoodOG/Hermes-Desktop) into .claude/skills/test-driven-development in your project. Claude Code loads it when a task matches its description.

How do I install Test Driven Development in Codex?

Run `npx skills add RedWoodOG/Hermes-Desktop --skill test-driven-development -a codex`. Or copy the skill folder (skills/software-development/test-driven-development in RedWoodOG/Hermes-Desktop) into .agents/skills/test-driven-development in your project. Codex loads it when a task matches its description.

Can I use Test Driven Development in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RedWoodOG/Hermes-Desktop --skill test-driven-development -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-driven-development, .gemini/skills/test-driven-development, .github/skills/test-driven-development and .opencode/skills/test-driven-development in your project.

What does Test Driven Development need to run?

Going by SKILL.md and its folder, Test Driven Development needs the command-line tools its instructions call (pytest). Our summary lists: Python 3.

Does Test Driven Development access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Driven Development safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Driven Development use?

Test Driven Development is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Driven Development use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Driven Development?

Skills that share tags, products or a category with Test Driven Development: TDD (fossasia/eventyay-interpretation, 1.6k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), TDD (sanity-io/sanity, 6.4k stars) and Test Driven Development (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Driven Development?

RedWoodOG (a GitHub user) maintains it in RedWoodOG/Hermes-Desktop, which has 177 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on May 30, 2026.

Source: RedWoodOG/Hermes-Desktop on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.