Agent skill

TDD

by kdlbs in kdlbs/kandev

Implement changes using Test-Driven Development (Red-Green-Refactor).

AGPL-3.0Auto-check passedTesting & QA

Install TDD

skills CLI
$ npx skills add kdlbs/kandev --skill tdd -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kdlbs/kandev tdd --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kdlbs/kandev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/tdd .claude/skills/tdd && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tdd
GitHub stars
909
Token cost
~4.2k tokens
SKILL.md length
2,385 words
Files
3 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Implement changes using Test-Driven Development (Red-Green-Refactor).

  • Works in 5 steps: RED — Write a failing test → GREEN — Minimal code to pass → REFACTOR — Clean up → …
  • Code changes that need coverage
  • SKILL.md covers Execution Context, Related skills, When to use and Test-value gate, plus 4 more sections
  • Calls go and pnpm

What it does

TDD is an agent skill from kdlbs/kandev. Implement changes using Test-Driven Development (Red-Green-Refactor). Use for code changes that need coverage or requested test-value audits.

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/backend-tests.md` and `references/test-audit.md`).

It sits in Testing & QA, covering Test-driven development. The repository describes itself as: AI Kanban & Development Environment. Orchestrate multiple agents, review changes, open PRs. Multi-provider, self-hostable, no telemetry. The licence is AGPL-3.0.

When your agent uses it

  • Code changes that need coverage
  • Requested test-value audits

Example prompts

  • “/tdd”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. RED — Write a failing test
  2. GREEN — Minimal code to pass
  3. REFACTOR — Clean up
  4. Repeat
  5. Final verification

What it can do on your machine

Read from SKILL.md and the folder at commit b734113. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • go
    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

TDD loads about 4.2k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 36 tokens; SKILL.md has 2,385 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~36
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kdlbs/kandev at commit b734113, republished under its AGPL-3.0 licence (© kdlbs). 2,385 words, ~4,179 tokens.

Download SKILL.mdSave it as .claude/skills/tdd/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
tdd
description
Implement changes using Test-Driven Development (Red-Green-Refactor). Use for code changes that need coverage or requested test-value audits.

TDD

Execution Context

Use this procedure directly in the primary conversation for every code change, regardless of size. The user may switch that conversation to the lower-cost implementation model before beginning the Red-Green-Refactor cycle.

Implement code changes using strict Red-Green-Refactor. Iron law: no production code without a failing test first.

Wrote code before a test? Delete it. Start over from a failing test.

  • /e2e — Follow this procedure when the current task needs Playwright E2E coverage.
  • /pr-fixup — After the task-defined checks pass and the PR opens, use it only for CI or actionable reviewer findings.

When to use

  • Bug fixes — write a test that reproduces the bug before fixing
  • New functions, methods, or utilities
  • Refactoring existing logic that lacks tests

Skip for: pure UI components (we don't test React components), config files, generated code.

For UI rendering bugs, prefer extracting or using a pure helper and testing that helper. Add Playwright only when the behavior is truly visual or integration-level. Avoid adding React component tests just to assert DOM output; that does not match this project's testing convention.

Test-value gate

Before adding or changing a test, answer these questions:

  1. What observable behavior, invariant, or independent contract does it protect?
  2. What credible regression makes its assertion fail?
  3. Why does existing coverage miss that regression?
  4. Does it require a production export, flag, or injection hook that only tests use?

Resolve a missing answer before writing the test. Prefer extending an existing case over repeating the same contract. Give each contract a primary test owner at the boundary that can expose its failure. Keep additional layers for distinct transport, lifecycle, platform, or input-handler risks.

For test-only seams, first look for proof through an existing production boundary. Keep necessary performance instrumentation only with a named invariant and no simpler proof.

For a requested test audit or consolidation, load test-audit.md. Do not start a broad cleanup as part of ordinary TDD.

Determine test scope

When implementing a UI work order, read its ASCII UI preview and linked plan before changing the surface. Use the structural requirements to guide the existing rendered checks; do not test ASCII whitespace or treat it as exact pixel geometry. Keep any design revisions synchronized through docs/specs/guide/plans-and-work-orders.md#ascii-ui-previews.

  • Go unit (apps/backend/): test file next to source as *_test.go. Run:
    bash
    cd apps/backend && go test -v -run TestName ./internal/path/to/package/...
  • TypeScript unit (apps/web/lib/): test file next to source as *.test.ts. Run:
    bash
    cd apps && pnpm --filter @kandev/web test -- --run path/to/file.test.ts
  • Web E2E (apps/web/e2e/): follow /e2e when the current task needs Playwright tests.

For Go test fixtures, load backend-tests.md.

Choose the right level:

  • Unit: pure logic or isolated service behavior.
  • Integration: handler/service/repository boundaries, SQLite-backed flows, filesystem behavior, or process boundaries.
  • E2E: critical user-facing browser flows; keep these focused and use /e2e.

Prefer state/output assertions over interaction assertions. Mock only slow, nondeterministic, or external boundaries; use real implementations or fakes when they keep the test deterministic.

When one behavior has separate wheel, keyboard, touch, or pointer handlers, inventory and test every handler with distinct branching. For input at a hard boundary where the underlying position cannot change, dispatch the input event or gesture directly. Cover the eligible direction at the boundary, the wrong direction or off-boundary no-op, one physical gesture firing at most once, and listener cleanup when the change adds subscriptions. An E2E case for one input modality does not cover the other handler paths.

When a work order names an AC-* acceptance criterion, keep the mapping visible in the test name or a nearby @covers AC-... comment. Do not copy the complete requirement into the test.

For failure-path tests, inject the error at the boundary the production code claims to handle and exercise the real downstream call chain; do not short-circuit by mocking the handler under test. When production code has a defensive nil/fallback branch or sanitizes an internal error into a public error, inject that alternate return at the interface boundary and assert the public result plus absence of internal details or sensitive content. For a test asserting that a side effect does not occur, enable every prerequisite that would otherwise cause it and add a positive control with the same setup to prove the side effect is reachable. Observe the actual side effect, not a proxy that can remain unchanged.

For parser, canonicalization, or sanitization tests, build fixtures from the producer's exact serialized shape. Include boundary cases such as optional whitespace, adjacent records, terminators without a newline, and malformed or unclosed input; assert that untrusted content is removed or fails closed while independent following content is preserved. A test-only normalized fixture can pass while the real wire framing still fails.

Cross-layer contracts

When adding or renaming a field that crosses backend, WebSocket, and frontend boundaries, use rg to trace every producer, DTO, store upsert or partial merge, reconnect/readiness handler, and consumer before editing. Add focused coverage at the affected boundaries, including refresh/reconnect and an update that omits the field, so a partial payload cannot silently discard an existing value. For MCP/API tool registration or mode-gated catalogs, add positive tests for intended modes and negative tests asserting tool absence in every excluded mode. For opt-in or feature-flagged implementations, exercise both the explicit enabled path and the disabled/default path; default-path coverage alone does not prove that the opt-in implementation is wired.

When production behavior depends on an optional interface or type assertion, a pass-through fake that omits that capability is invalid coverage. Use the real implementation or a capability-complete fake, and assert the exact accepted content and trusted context at persistence and dispatch, including empty expansions and delayed or created-session paths.

When content passes through multiple canonicalizers or a delayed/created-session path, test each handoff with distinct sentinels and assert exact equality at persistence and dispatch. Carry the exact value produced at acceptance through later stages instead of re-deriving it from mutable source data at launch.

Concurrent and event-driven behavior

Test ordering-sensitive behavior with channels, barriers, or controllable fakes; do not use sleeps to create a race, except for a bounded, named delay that models a known poll-loop schedule when synchronization would alter that relationship. Pause at the ownership boundary, start the competing operation, then release. Exercise the real delivery path where practical, and prove the old interleaving fails before the fix. Assert both the winner state and the untouched replacement state, including relevant buffers, signals, or queue ownership. Cover stale events acting after a replacement operation begins, stale-owner handoffs plus same-owner invalidation/no-replay, cancellation/retry ownership, and at-most-once delivery when they apply. Run affected Go packages with -race.

When a source can deliver either a complete snapshot or a partial event, define the omission semantics and test both forms. Capture a request-start revision or epoch before every asynchronous request, inventory all callers, and use deferred response/event tests to prove ordering. React loading state does not serialize same-tick callbacks; use an immediate ref or shared in-flight promise when request identity must be single-flight.

For fetch effects that depend on live store values such as message count, a dependency rerender is not lifecycle invalidation. Do not let an effect-local cleanup boolean be the only ownership guard: capture a session, connection, or attempt generation for response writes and loading finalization, and invalidate that generation only on the corresponding session/connection change or unmount. Add a red test that defers the response, triggers a same-session live-row rerender, and proves the pending state settles without overwriting newer UI.

For hooks writing shared workspace or global caches, test the ownership matrix: a second consumer preserves valid cached data; an initiator unmount before a deferred response still reaches terminal state; the latest same-key response wins over stale completion; and workspace/key switches remain isolated.

During cancellation or recovery, do not broadly suppress stream frames. Suppress only allowlisted cancellation acknowledgements with immutable operation identity; same-identity message, thinking, and tool frames remain authoritative activity. Add a negative regression for each activity class and a transport-boundary test when ordering depends on subscriber processing.

When delayed state has a lifecycle owner, pair the boundary tests: disposal before the threshold must cancel timers and emit nothing later; disposal or replacement after publication must immediately clear externally observable state; and stale callbacks must not mutate the replacement.

When setup can finish before hydration or restoration, test both readiness states. Cover state that is ready at setup and state that becomes ready after listener registration. Drive the readiness transition explicitly and dispose each added subscription.

When logic rejects ambiguous candidates, filter invalid candidates before the cardinality check. Cover zero valid candidates, one valid candidate mixed with invalid candidates, and multiple valid candidates.

For behavior described as remaining reactive, updating, or responding after initialization, test both the initial state and a deterministic post-initialization transition. An initial snapshot test cannot prove that a subscription, watcher, or reconciliation path remains active; trigger a state, event, or input change and assert the updated outcome. If no public API exposes publication or render counts, a controlled boundary fake may instrument those events, but retain an observable state assertion as the proof of behavior.

Show full SKILL.md (902 more words)Show less

Steps

1. RED — Write a failing test
  1. Identify the single behavior to implement or bug to reproduce
  2. Write the smallest test that asserts the expected behavior — one assertion, clear name
  3. Run the test and confirm it fails with the expected assertion error (not a compile/import error) For a brand-new Go package, create the package directory and minimal test package first, then run the focused package test so RED fails on behavior rather than package-selection or import errors.
  4. If it passes immediately, the test is not testing new behavior — revise it. Exception: requested test-only contract coverage or an audit assertion repair can pass on the current head. Make no permanent production change. For an assertion repair, demonstrate RED with a deliberate mutation through test-audit.md, then restore the source.

For bug fixes, use the Prove-It Pattern: reproduce the bug with a failing test before changing production code. A fix without a regression test is not complete unless the change is explicitly untestable and you say why.

2. GREEN — Minimal code to pass
  1. Write the minimum production code to make the failing test pass
  2. Do not add extra logic, handle other edge cases, or refactor yet
  3. Run the test again and confirm it passes
  4. If it fails, fix the production code (not the test)
3. REFACTOR — Clean up
  1. Improve production code: extract helpers, rename, simplify — without changing behavior
  2. Improve tests: table-driven tests (Go) or describe/it blocks (TS), remove duplication
  3. Run the test after each change to confirm still green

In tests, prefer DAMP over DRY: each test should read like a small specification. Shared helpers are fine when they remove noise, but not when they hide the scenario. For fixture setup that needs a canonicalized identifier, use the production helper instead of copying its algorithm. Keep expected values literal and independent; if a test needs an independent oracle, keep it local and explain why.

4. Repeat

Return to step 1 for the next behavior or edge case. Continue until the feature or fix is complete.

5. Final verification

Run the targeted tests named in the task file and report their results. Commit and open the PR after all affected task checks pass; do not add broad local verification by default.

After the final production-code edit, rerun every new or changed regression test and report the exact command and result. A prior green run does not cover a later patch.

Testing anti-patterns

Don't test implementation details:

  • Assert behavior, state, API response, DB row, emitted event, or UI outcome. Avoid assertions that only prove a helper was called or an internal query string happened to be built a certain way.
  • Custom hooks and async controller helpers are behavior-bearing logic, not pure React markup: add focused tests for success, failure, cancellation/no-op, and busy/loading-state cleanup.
  • When a change adds a bulk or action entry point, test that entry point directly, including its empty, partial, and success paths; coverage of a shared helper or neighboring single-item flow does not prove the new dispatch path works.
  • When an opt-in mode replaces a shared implementation, exercise the real component with that mode enabled and production-shaped callbacks, while keeping separate assertions for the default path. Conversion tests alone do not prove mounted editor reconciliation, selection, undo, or insert-then-submit behavior.

Don't test mock behavior:

  • If your assertion checks a mock element (*-mock test ID, mock return value), you're testing the mock, not the code. Test real behavior or don't mock it.

Don't add test-only methods to production code:

  • destroy(), reset(), _testHelper() that only tests call — put these in test utilities, not production classes.

Mock minimally and understand dependencies:

  • Before mocking, ask: what side effects does the real method have? Does the test depend on any of them?
  • Mock the slow/external part (network, disk), not the method the test depends on.
  • If mock setup is longer than test logic, consider an integration test instead.

Don't use incomplete mocks:

  • Mock the complete data structure as it exists in reality, not just fields your test uses. Partial mocks hide bugs when downstream code accesses omitted fields.
  • When adding a named export to a shared module, search for full-module vi.mock() factories and add the export to each factory. Focused tests can pass while a full suite fails on an out-of-date module shape.
  • For hook, store-selector, action, or data return-shape changes, use rg to find every affected vi.mock() factory and stubbed consumer, update each mock for every consumed field, then run the full owning package test command and the nearest real component or consumer test; typecheck alone does not catch an omitted runtime field.

Never swallow errors in tests:

  • try/catch that silently ignores failures in test helpers or setup — these hide real failures.

Don't repeat unchanged passing commands for reassurance:

  • After a clean targeted run, re-run only after code or test inputs change. Move to the next required verification step instead.

Red flags

  • Writing production code before a failing test exists — delete and start over
  • Test passes on first run — revise it, except for requested test-only contract coverage or an audit assertion repair with mutation proof
  • Fixing a test to make it pass instead of fixing the production code
  • Large jumps — multiple behaviors implemented between test runs
  • Skipping the refactor step
  • Mock setup longer than test logic — consider integration test
  • Asserting on mock elements instead of real behavior
  • "All tests pass" but no relevant test was actually run

© kdlbs, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/tdd of kdlbs/kandev.

  • SKILL.md
  • references/backend-tests.md
  • references/test-audit.md

Open the folder on GitHubat commit b734113

Compare with similar skills

TDD next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

TDD compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
TDD this skillkdlbs/kandev909—~4.2kAutomated safety check: PassAGPL-3.0
TDDpietheinstrengholt/rssmonster56430 repos~906Automated safety check: PassMIT
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT
Test Driven Developmentfarm-fe/farm5.6k51 repos~2.5kAutomated safety check: PassMIT
Tapd Story PipelineTencentBlueKing/bk-bcs840—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • TDD

    pietheinstrengholt/rssmonster

    Test-driven development. An agent skill from pietheinstrengholt/rssmonster.

    564 GitHub starsUsed in 30 repos~906 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • A skill your agent uses when implementing any feature or bugfix, before writing implementation code

    5.6k GitHub starsUsed in 51 repos~2.5k tokens
    Testing & QAAuto-check passed
  • Tapd Story Pipeline

    TencentBlueKing/bk-bcs

    单需求实现流水线——把一个 TAPD 需求从零推进到代码提交。自动串联技术澄清、 开发计划、任务拆分、TDD 实现、架构/安全校验、代码提交六个阶段。

    840 GitHub stars~2.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Absolute Init

    maddhruv/absolute

    One-time setup for absolute: interview how you want it to behave (output style, autonomy, TDD strictness, spec dir, families) + detect the stack once, then write .absolute.config.json (project…

    218 GitHub starsUsed in 1 repo~3k tokens
    Testing & QAAuto-check passed

More from kdlbs/kandev

All 45 skills in this repo
  • PR Walkthrough

    kdlbs/kandev

    Generate a single-file HTML walkthrough that explains a PR's purpose, user impact, interface changes, compatibility risks, and implementation.

    909 GitHub stars~6.2k tokensUpdated today
    Auto-check passed
  • Debug

    kdlbs/kandev

    Diagnose Kandev bugs, running-instance issues, UI/browser failures, and runtime behavior.

    909 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Improve Kandev's AI harness from session learnings or explicit requests.

    909 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Diagram Design

    kdlbs/kandev

    Create branded architecture, IT current-state, flowchart, sequence, state machine, ER/data model, timeline, swimlane, quadrant, radar/spider, polar chart (polar/radial lollipop), loop/flywheel…

    909 GitHub starsUsed in 1 repo~8k tokens
    Auto-check passed
  • Verify

    kdlbs/kandev

    Run a broad local verification audit only when the user explicitly requests it or PR/CI remediation requires it.

    909 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Acp Debug

    kdlbs/kandev

    Debug an ACP agent CLI by spawning it, speaking raw JSON-RPC, and capturing every frame to a JSONL file.

    909 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Categories

Questions about TDD

What does TDD do?

Implement changes using Test-Driven Development (Red-Green-Refactor). TDD is an agent skill from kdlbs/kandev. Implement changes using Test-Driven Development (Red-Green-Refactor).

When should I use TDD?

TDD fits situations like: code changes that need coverage; requested test-value audits.

How do I install TDD in Claude Code?

Run `npx skills add kdlbs/kandev --skill tdd -a claude-code`. Or copy the skill folder (.agents/skills/tdd in kdlbs/kandev) into .claude/skills/tdd in your project. Claude Code loads it when a task matches its description.

How do I install TDD in Codex?

Run `npx skills add kdlbs/kandev --skill tdd -a codex`. Or copy the skill folder (.agents/skills/tdd in kdlbs/kandev) into .agents/skills/tdd in your project. Codex loads it when a task matches its description.

Can I use TDD in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kdlbs/kandev --skill tdd -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tdd, .gemini/skills/tdd, .github/skills/tdd and .opencode/skills/tdd in your project.

What does TDD need to run?

Going by SKILL.md and its folder, TDD needs the command-line tools its instructions call (go and pnpm).

Does TDD access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is TDD safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does TDD use?

TDD is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does TDD use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.

What are the alternatives to TDD?

Skills that share tags, products or a category with TDD: TDD (pietheinstrengholt/rssmonster, 564 stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), TDD (sanity-io/sanity, 6.4k stars) and Test Driven Development (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains TDD?

kdlbs (a GitHub organization) maintains it in kdlbs/kandev, which has 909 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on October 8, 2026.

Source: kdlbs/kandev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.