Agent skill

Write Tests

by dzhng in dzhng/skills

Write tests that pin real behavior instead of implementation details, config values, or lucky samples.

MITAuto-check passedTesting & QA

Install Write Tests

skills CLI
$ npx skills add dzhng/skills --skill write-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dzhng/skills write-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dzhng/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/engineering/write-tests .claude/skills/write-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
write-tests
GitHub stars
1k
Token cost
~2.4k tokens
SKILL.md length
1,430 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Write tests that pin real behavior instead of implementation details, config values, or lucky samples.

  • Works in 3 steps: Write ONE test at a time. Assert first,… → Iterate on the fastest focused runner… → Before calling it done, prove the test…
  • Adding tests for new behavior
  • SKILL.md covers Workflow: tracer bullets, not…, What to assert, Seams and mocks and Control variables and probes, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Write Tests is an agent skill from dzhng/skills. Write tests that pin real behavior instead of implementation details, config values, or lucky samples. Use when adding tests for new behavior, writing a regression test, fixing a brittle or flaky test, reviewing a test diff, or when a test breaks after a refactor or config change that didn't change behavior.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test generation, Failing and flaky tests and Refactoring. The repository describes itself as: Reusable AI agent skills for software factories: explore ideas, write specs, implement, review, and run autonomous research. Works with Claude Code, Codex, and other… The licence is MIT.

When your agent uses it

  • Adding tests for new behavior
  • Writing a regression test
  • Fixing a brittle
  • Reviewing a test diff

Example prompts

  • “/write-tests”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Write ONE test at a time. Assert first, watch it go red on the un-fixed
  2. Iterate on the fastest focused runner (one file, one test name), and run
  3. Before calling it done, prove the test can fail (below) and walk the

What it can do on your machine

Read from SKILL.md and the folder at commit d513228. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Write Tests loads about 2.4k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 1,430 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dzhng/skills at commit d513228, republished under its MIT licence (© dzhng). 1,430 words, ~2,357 tokens.

Download SKILL.mdSave it as .claude/skills/write-tests/SKILL.md (or your agent's skills folder).
name
write-tests
description
Write tests that pin real behavior instead of implementation details, config values, or lucky samples. Use when adding tests for new behavior, writing a regression test, fixing a brittle or flaky test, reviewing a test diff, or when a test breaks after a refactor or config change that didn't change behavior.

Write Tests

A good test fails only when real behavior breaks, and passes through every refactor or config change that preserves it. Most bad tests fail the opposite way: red on harmless changes, green while the real path is broken. Every rule below serves that one goal.

For deciding whether coverage adds independent proof, where it belongs, or which existing tests can go, use audit-tests.

Workflow: tracer bullets, not a batch

  1. Write ONE test at a time. Assert first, watch it go red on the un-fixed code, make the code earn green, learn, then write the next. Never a batch up front: a batch written against imagined behavior pins what you guessed — those tests pass when the mechanism breaks and fail when it's fine. Each green cycle tells you what the next test should actually assert.
  2. Iterate on the fastest focused runner (one file, one test name), and run the full suite only as a final gate before handing off. Check the exit code, not just the output — a runner that prints nothing and a green run look the same. On failures, read EVERY red test before fixing one; they often share a root cause.
  3. Before calling it done, prove the test can fail (below) and walk the review checklist.

What to assert

  • Observable behavior through the outermost practical entry point — return values, exit codes, persisted rows, HTTP responses, rendered output — never which internal functions ran or how a value is computed. A test on the public surface survives a rewrite of everything underneath; a test that reaches into internals breaks on every refactor and pins implementation, not behavior. Reserve isolated unit tests for genuinely tricky pure logic (parsers, schedulers, state machines).
  • Nothing the compiler already guarantees. A test that re-asserts a type signature — field shapes, rejected argument types — can only fail if the compiler failed first. Spend the budget on business rules, arithmetic, branching, ordering, edge cases, side effects.
  • Actual values, not collection sizes. For dedup/normalize/idempotency paths, length == 1 passes even when normalization is broken; also assert the stored value equals the expected canonical form.
  • The smallest scale that can show the behavior. Two entities and one mechanism before crowds and integration; small tests fail fast with readable state and don't entangle five behaviors in one assert.
  • The design contract, not current behavior. When a test goes red, the reflex is to re-measure and pin the new number — resist it: a bar calibrated to whatever the code currently does silently encodes bugs as baseline. Write the assert from the stated contract and make the code earn it; if no contract exists, that's a question for the owner, not a number to measure-and-pin. Recalibrating is legitimate only when the contract itself changed.
  • Deterministic claims tight, stochastic claims as distributions. An invariant that holds every run gets an exact threshold. Anything noisy or tuning-dependent ("A usually beats B", "load balances evenly") must be asserted over a set of runs/seeds as a band — a single-sample pin on a stochastic outcome is not a weak test, it is a blind one: it certifies whatever the lucky sample did and can mask a systematic bias for months. If you must assert a noisy differential, widen the margin and name it chaos-marginal in a comment.

Seams and mocks

  • Don't couple tests to config — mock the seam. A test keyed to a live config value (which flag is on, which tier is default) breaks when someone legitimately edits config. The fix is never skipIf/conditionals — a skipped assertion hides the coupling and stops covering the path. Feed the function fixed inputs, or stub the lookup so fixture ids resolve to fixed values. Litmus: "would this break if a config value changed with no logic change?"
  • Don't over-mock — every mock is a frozen assumption. Mock at the system's edges (network, clock, filesystem, third-party SDK, config lookup), never internal collaborators. Stubbing an internal hard-codes its current contract; refactor it and the test lies. Never make "was called with" the primary assertion when an observable outcome exists. Mocking three internal modules to test one function means: test one layer up, where they're real.
  • Harnesses must wire the system the way production does. If a bug only showed up against a real environment, the harness skipped an input production always sets — fix the harness, don't just fix the bug.

Control variables and probes

  • One variable per comparison. Pin everything else — same seed, same counts, same configuration on both arms; disable mechanisms that aren't the subject. When a comparison test breaks, first ask whether an unrelated mechanism leaked into the experiment before touching the code under test.
  • Validate, don't assume. Never reason your way to a conclusion about cause — which path fired, where the failures come from — and act on it. Write a throwaway probe that reads public state and prints; the plausible story is wrong often enough to burn a session, and one probe redirects the whole effort. Probes are scratch: put them where they cannot be committed, delete them the moment the question is answered, or promote them into a real test if the behavior deserves a permanent guard.
Show full SKILL.md (578 more words)Show less

Tests run on someone's live machine

A suite that grabs the desktop is a suite people stop running. Whatever the test needs, it takes the least intrusive form of it:

  • Never take focus. No activating the app, no makeKeyAndOrderFront, no raising, no moving the pointer, no simulated clicks into the session, no audio playback. A UI a person must not be interrupted by is still fully testable: build the window, drive its model, and assert on the rendered view. Where a screenshot is the evidence, capture the window's own image offscreen rather than photographing the screen. Keep activation on the code path a person triggers, and exercise it by calling the handler, not by launching the app repeatedly.
  • Write nowhere the person keeps their own state. Point every test at a scratch home, a scratch preferences domain or suite, and scratch install locations. A test must never modify the real config, library, login items or installed applications. Production may read an override the tests set; it must behave normally when that override is absent.
  • Batch anything unavoidably visible. If a check genuinely needs a real launch, do all of its observations in one run instead of one relaunch per assertion, and say in the evidence that it was visible.
  • Leave nothing running. Every process, window, server and temporary file the test started is stopped and removed, including on failure.

Prove the test can fail

A test you never saw fail is decoration. For any regression test — especially one written after the fix — falsify it once: revert or break the production code the way the bug would, confirm red for the expected reason, restore, confirm green. Verify the revert actually took: a stash or checkout with a wrong pathspec reverts nothing, silently, and the "red" run quietly tests the fixed code. The tell: the "red" numbers equal the green numbers.

When a test goes red after a change

Triage in order of likelihood before editing anything:

  1. The test encodes deleted semantics — it pinned an accident of the old implementation. Re-spec it to the contract's real claim.
  2. The experiment lost control of a variable — a new mechanism leaks in. Pin the variable.
  3. The margin was chaos-tight. Widen it, with a comment saying so.
  4. The code is actually wrong. Probe the state, find the mechanism.

Never tune constants to make one test pass without rerunning the neighbors: coupled systems reshuffle. If two consecutive tweaks each break different tests, stop poking — the control surface is wrong; find the mechanism.

Review checklist

Walk this on any test diff, apply fixes in the same pass, re-run the suite:

  1. Does an assertion restate a type guarantee? → delete it.
  2. Would it break on a config edit with no logic change, or is it skipIf-conditional on config? → mock the seam.
  3. Are assertions mostly mocked-internal "called with"? → test a layer up against observable output.
  4. Does it pin a single sample of a stochastic outcome? → assert a distribution.
  5. Does it assert a count without the value on a dedup/normalize path? → add the value assertion.
  6. Is the function under test still called in production? → if the only callers are tests, delete the function and the test together.
  7. Can it fail for the right reason? → falsify once to confirm.
  8. Does it take focus, move the pointer, play sound, or write the person's real config, library or applications? → drive the model and a scratch location instead, and capture windows offscreen.

© dzhng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/engineering/write-tests of dzhng/skills.

Open the folder on GitHubat commit d513228

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in dzhng/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Write Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Write Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Write Tests this skilldzhng/skills1k—~2.4kAutomated safety check: PassMIT
Selector Drift Recoverypetrkindlmann/qa-skills163—~5.1kAutomated safety check: PassMIT
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Wioworkersio/skills180—~5.8kAutomated safety check: PassMIT
Offloadimbue-ai/offload125—~3.1kAutomated safety check: PassMIT
Test Fixingdavila7/claude-code-templates32k7 repos~739Automated safety check: PassMIT

Similar skills

  • Selector Drift Recovery

    petrkindlmann/qa-skills

    Bulk-regenerate broken test selectors after a UI refactor or redesign.

    163 GitHub stars~5.1k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    180 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Offload

    imbue-ai/offload

    Activate when you see offload.toml in a repo, offload referenced in build targets (justfile, Makefile, scripts), or when you need to run a large test suite in parallel.

    125 GitHub stars~3.1k tokensUpdated 12 days ago
    Testing & QAAuto-check passed
  • Test Fixing

    davila7/claude-code-templates

    Run tests and systematically fix all failing tests using smart error grouping.

    32k GitHub starsUsed in 7 repos~739 tokens
    Testing & QAAuto-check passed
  • Memstack Development Test Writer

    cwinvestments/memstack

    A skill your agent uses when the user says 'write tests', 'add tests', 'test coverage', 'unit tests', 'integration tests', 'component tests', 'mocking', 'edge cases', or needs to generate tests with…

    423 GitHub stars~3.7k tokensUpdated 11 days ago
    Testing & QAAuto-check passed

More from dzhng/skills

All 27 skills in this repo
  • Compare screenshots against the intended design, distinguishing approved references from historical baselines.

    1k GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Claude

    dzhng/skills

    Use Claude Code as an independent claude -p subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude.

    1k GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Refactor Clean

    dzhng/skills

    Refactor cleanly instead of layering sediment. An agent skill from dzhng/skills.

    1k GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Write Skills

    dzhng/skills

    Create or revise agent skills. An agent skill from dzhng/skills.

    1k GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed
  • Codex

    dzhng/skills

    Use the local Codex CLI as an independent second agent. An agent skill from dzhng/skills.

    1k GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check: warnings
  • Run [implement-spec](../implement-spec/SKILL.md) with Codex doing the implementation passes while you orchestrate, integrate, and review.

    1k GitHub starsUsed in 1 repo~543 tokens
    Auto-check passed

Categories

Questions about Write Tests

What does Write Tests do?

Write tests that pin real behavior instead of implementation details, config values, or lucky samples. Write Tests is an agent skill from dzhng/skills. Write tests that pin real behavior instead of implementation details, config values, or lucky samples.

When should I use Write Tests?

Write Tests fits situations like: adding tests for new behavior; writing a regression test; fixing a brittle; reviewing a test diff.

How do I install Write Tests in Claude Code?

Run `npx skills add dzhng/skills --skill write-tests -a claude-code`. Or copy the skill folder (skills/engineering/write-tests in dzhng/skills) into .claude/skills/write-tests in your project. Claude Code loads it when a task matches its description.

How do I install Write Tests in Codex?

Run `npx skills add dzhng/skills --skill write-tests -a codex`. Or copy the skill folder (skills/engineering/write-tests in dzhng/skills) into .agents/skills/write-tests in your project. Codex loads it when a task matches its description.

Can I use Write Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dzhng/skills --skill write-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/write-tests, .gemini/skills/write-tests, .github/skills/write-tests and .opencode/skills/write-tests in your project.

What does Write Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Write Tests is instructions for the agent only.

Does Write Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Write Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Write Tests use?

Write Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Write Tests use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Write Tests?

Skills that share tags, products or a category with Write Tests: Selector Drift Recovery (petrkindlmann/qa-skills, 163 stars), Swig Test (swig/swig, 6.3k stars), Wio (workersio/skills, 180 stars) and Offload (imbue-ai/offload, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Write Tests?

dzhng (a GitHub user) maintains it in dzhng/skills, which has 1,013 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 5, 2026.

Source: dzhng/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.