Agent skill

Writing Tests

by caarlos0 in caarlos0/dotfiles

Write deterministic tests that fail for real product defects.

MITAuto-check passedTesting & QA

Install Writing Tests

skills CLI
$ npx skills add caarlos0/dotfiles --skill writing-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install caarlos0/dotfiles writing-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/caarlos0/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/writing-tests .claude/skills/writing-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
writing-tests
GitHub stars
220
Token cost
~1.9k tokens
SKILL.md length
1,051 words
Files
1
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

Write deterministic tests that fail for real product defects.

  • Works in 7 steps: Read the original failure text, stack,… → Separate infrastructure failures such as… → Compare the failure window with the… → …
  • Reviewing test coverage and reliability
  • SKILL.md covers Prove the behavior, Control every input, Synchronize; do not sleep and Own resources and cleanup, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Writing Tests is an agent skill from caarlos0/dotfiles. Write deterministic tests that fail for real product defects. Use when adding tests, fixing flakes, or reviewing test coverage and reliability.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test coverage. The licence is MIT.

When your agent uses it

  • Reviewing test coverage and reliability
  • Tasks that involve Test coverage

Example prompts

  • “/writing-tests”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Read the original failure text, stack, stderr, seed, logs, and artifacts.
  2. Separate infrastructure failures such as OOM, disk, runner, or container
  3. Compare the failure window with the lifetime of the failing code. Use history
  4. Inspect every CI attempt because rerun-to-green summaries hide failures.
  5. Reproduce on parents when needed to distinguish a landed regression or
  6. Identify the mechanism: ordering, async completion, data race, resource leak,
  7. Fix the owning layer and write a deterministic regression test controlling

What it can do on your machine

Read from SKILL.md and the folder at commit 892360f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Writing Tests loads about 1.9k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 1,051 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from caarlos0/dotfiles at commit 892360f, republished under its MIT licence (© caarlos0). 1,051 words, ~1,935 tokens.

Download SKILL.mdSave it as .claude/skills/writing-tests/SKILL.md (or your agent's skills folder).
name
writing-tests
description
Write deterministic tests that fail for real product defects. Use when adding tests, fixing flakes, or reviewing test coverage and reliability.
user_invocable
true

Writing Tests

A useful test fails when behavior is wrong, passes when it is right, and explains the failure without a debugger. Follow repository instructions first.

Prove the behavior

  • Choose the lowest layer that proves the contract: unit for local logic, integration for real boundaries, CLI/UI/E2E only for behavior unavailable below them.
  • Exercise production code. Do not copy its logic into the test or mock the unit under test.
  • A bug fix needs a focused regression test. Where practical, revert or mutate the fix and confirm that exact test fails.
  • Assert observable results, not private calls or incidental structure.
  • Print expected and actual values, variants, exit status, stderr, and case labels needed to diagnose a rare CI failure.
  • Cover material success, failure, boundary, default, and state transitions. Do not add combinations without a concrete regression risk.

Control every input

  • Inject clocks; fix timezone and locale when formatting matters.
  • Seed randomness and print the replay seed on failure.
  • Use semantic unordered equality when order is not a contract. Do not hide a required order by sorting expected output.
  • Give each test an isolated environment, home, working directory, config root, database state, cache, and temporary directory.
  • Bind servers to port 0 and retain the listener. Never find, close, and rebind a "free" port.
  • Avoid process-global mutation from parallel tests. Use helpers that restore state automatically.

Synchronize; do not sleep

Replace sleeps and grace periods with a happens-before edge: channel receive, barrier, latch, event, condition, queue drain, file close, EOF, process exit, or application readiness.

Use fake time only when time is the contract. Know which clock it controls. Use a timeout only around an observable wait as an outer deadlock detector; expiry must fail with diagnostics. Never use a sleep, retry, timeout increase, or looser assertion as a race fix.

For processes, establish the required order: start output readers, close stdin when EOF is required, drain stdout and stderr, observe exit, then clean up. Preserve output even when finalization fails.

Own resources and cleanup

Register cleanup immediately after acquisition. Teardown in dependency order: stop producers, cancel work, drain or join consumers, flush, close resources, then delete storage. Surface cleanup failures.

Parallel tests need separate resources. A shared temp path, port, database row, clock, fake timer, module cache, or mutable global is a race even if it usually passes.

Use a local implementation of the real protocol at boundaries: an in-process HTTP server, transaction-isolated database, or controlled fake implementing the real interface. Do not put live third-party availability, credentials, mutable data, or rate limits in required tests.

Snapshots, properties, and fuzzing

Normalize only fields proven volatile, such as generated IDs or temp roots. Broad normalization can erase the bug. Golden updates must be explicit and reviewed.

Use property, fuzz, differential, or metamorphic tests when examples cannot cover the input space or no simple expected value exists. Retain the minimized input or seed so every discovery becomes a deterministic regression.

E2E, browser, CLI, and TUI

Wait for the exact readiness needed by the next action: health response, UI state, log event, or prompt. Do not use arbitrary sleep or generic network idle. Prefer user-visible and accessibility-based locators plus retrying assertions.

Control browser, viewport, fonts, animation, and dynamic content for visual tests. Assert CLI/TUI exit status and relevant stdout/stderr after complete drain. For AI-backed behavior, assert an objective observable contract, not one exact natural-language answer unless the text itself is the contract.

Language mechanics

Go
  • Use t.TempDir, t.Setenv, t.Cleanup, httptest.Server, the race detector, and shuffled runs where relevant. Record shuffle seeds.
  • Process-global environment and working-directory changes are incompatible with parallel tests.
  • FailNow and testify require must run on the test goroutine. From handlers or spawned goroutines, report errors to the test goroutine and synchronize.
  • Use channels and WaitGroup; use testing/synctest only when supported by the repository's Go version.
  • Compare maps semantically or sort a copied key set for golden output.
Show full SKILL.md (406 more words)Show less
Rust
  • Keep test-only helpers behind #[cfg(test)].
  • Use channels and barriers for ordering; use #[tokio::test(start_paused = true)] for timer behavior. Prefer time::sleep(d).await over tokio::time::advance(d), which does not wait for tasks to be polled and can produce false passes.
  • When already present, Loom or Shuttle can explore schedules and Proptest can shrink and persist failures.
  • Ensure all process pipe owners close so EOF and exit are observable.
TypeScript and JavaScript
  • Always await promise assertions and async timer advancement.
  • Restore spies, modules, and fake timers after each test; clearing calls does not restore implementations.
  • Do not combine process-global fake timers or mutable module state with concurrent tests.
  • Use Playwright locators and web assertions for readiness.
  • For subprocess output assertions, await the 'close' event, not 'exit'; 'exit' fires before stdio streams finish flushing.
Python
  • Use tmp_path and monkeypatch for automatically restored state.
  • Cancel and await asyncio tasks before ending a test: task.cancel(); with contextlib.suppress(asyncio.CancelledError): await task.
  • Use subprocess.communicate() for bidirectional process I/O.
  • When order or property tools are already present, retain their seed or minimized example.

Flake triage

A passing rerun on the same commit proves nondeterminism, not correctness.

  1. Read the original failure text, stack, stderr, seed, logs, and artifacts.
  2. Separate infrastructure failures such as OOM, disk, runner, or container loss from product assertions.
  3. Compare the failure window with the lifetime of the failing code. Use history to name the introducing or fixing commit; absence of recent failures is not evidence of a fix.
  4. Inspect every CI attempt because rerun-to-green summaries hide failures.
  5. Reproduce on parents when needed to distinguish a landed regression or semantic merge conflict from a race.
  6. Identify the mechanism: ordering, async completion, data race, resource leak, state pollution, clock, randomness, platform, external dependency, or infrastructure.
  7. Fix the owning layer and write a deterministic regression test controlling that mechanism.

Prioritize flaky required checks because they block merges. Retries or quarantine may temporarily unblock work, but they must retain failed attempts, an owner, and an issue. Do not weaken assertions or delete safety coverage to make CI green.

  • code-review: coverage and determinism review. When invoked from code-review, do not invoke it again.
  • runtime-process-debugging: process, pipe, lifecycle, shutdown, and ordering mechanisms. When invoked from it, do not invoke it again.
  • go-conventions and rust-specialist: language-specific production changes.
  • change-impact-auditor: tests for configuration, protocol, default, or shared model changes.

Do not recurse. Keep the fix to one concern and run the smallest decisive validation.

© caarlos0, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/writing-tests of caarlos0/dotfiles.

Open the folder on GitHubat commit 892360f

Compare with similar skills

Writing Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Writing Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Writing Tests this skillcaarlos0/dotfiles220—~1.9kAutomated safety check: PassMIT
Requirementsrizsotto/Bear6.5k—~2kAutomated safety check: PassGPL-3.0
Crap Analysisardalis/RiverBooks1352 repos~3.4kAutomated safety check: PassNone
Code Coverages3s-project/s3s311—~789Automated safety check: PassApache-2.0
Project Statusbactopia/bactopia522—~787Automated safety check: PassMIT
Check Coverageldayton/Dippy243—~403Automated safety check: PassMIT

Similar skills

  • Requirements

    rizsotto/Bear

    Write, modify, or review a requirement file under docs/requirements -- pick the single owning file, keep the text contract-only, name IDs so they need no explanation, and verify cross-references and…

    6.5k GitHub stars~2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Crap Analysis

    ardalis/RiverBooks

    Analyze code coverage and CRAP (Change Risk Anti-Patterns) scores to identify high-risk code.

    135 GitHub starsUsed in 2 repos~3.4k tokens
    Testing & QAAuto-check passed
  • Code Coverage

    s3s-project/s3s

    Measure and grow the line coverage of the s3s crate. An agent skill from s3s-project/s3s.

    311 GitHub stars~789 tokensUpdated today
    Testing & QAAuto-check passed
  • Project Status

    bactopia/bactopia

    Show a live snapshot of the Bactopia project state — component counts, GroovyDoc coverage, nf-test coverage, and structural issues.

    522 GitHub stars~787 tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Check Coverage

    ldayton/Dippy

    Ensure comprehensive test coverage for a CLI handler. An agent skill from ldayton/Dippy.

    243 GitHub stars~403 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Guarding Destructive Operations

    kajisho5/ffmpeg-skill

    Add and review preconditions on operations that delete, overwrite, rewrite history, or resolve a caller-supplied name to a filesystem path — refusing instead of warning, placing the guard ahead of…

    1.9k GitHub stars~2.6k tokensUpdated 4 days ago
    Testing & QAAuto-check passed

More from caarlos0/dotfiles

All 20 skills in this repo
  • CLI Design

    caarlos0/dotfiles

    Design and review command-line interfaces for usability, automation, safety, accessibility, and long-term compatibility.

    220 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Dependabot Merge

    caarlos0/dotfiles

    Review and merge open dependency pull requests from Dependabot, Renovate and similar bots across the goreleaser organization and the caarlos0 user.

    220 GitHub stars~5k tokensUpdated today
    Auto-check passed
  • Gh CLI

    caarlos0/dotfiles

    Use GitHub CLI efficiently for pull requests, CI checks, workflow runs, logs, and merge status.

    220 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Tui Design

    caarlos0/dotfiles

    Design terminal user interfaces and interactive CLIs that stay usable, accessible, and scriptable.

    220 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Dashboard

    caarlos0/dotfiles

    Design and review dashboards that are informative, honest, accessible, and visually polished, independent of any tool.

    220 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Gh Doc Author

    caarlos0/dotfiles

    Author and revise clear GitHub internal documentation, including design docs, proposals, decision records, runbooks, status updates, and handoffs.

    220 GitHub stars~2.2k tokensUpdated today
    Auto-check passed

Categories

Questions about Writing Tests

What does Writing Tests do?

Write deterministic tests that fail for real product defects. Writing Tests is an agent skill from caarlos0/dotfiles. Write deterministic tests that fail for real product defects.

When should I use Writing Tests?

Writing Tests fits situations like: reviewing test coverage and reliability; tasks that involve Test coverage.

How do I install Writing Tests in Claude Code?

Run `npx skills add caarlos0/dotfiles --skill writing-tests -a claude-code`. Or copy the skill folder (skills/writing-tests in caarlos0/dotfiles) into .claude/skills/writing-tests in your project. Claude Code loads it when a task matches its description.

How do I install Writing Tests in Codex?

Run `npx skills add caarlos0/dotfiles --skill writing-tests -a codex`. Or copy the skill folder (skills/writing-tests in caarlos0/dotfiles) into .agents/skills/writing-tests in your project. Codex loads it when a task matches its description.

Can I use Writing Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add caarlos0/dotfiles --skill writing-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/writing-tests, .gemini/skills/writing-tests, .github/skills/writing-tests and .opencode/skills/writing-tests in your project.

What does Writing Tests need to run?

SKILL.md names no scripts, command-line tools or credentials: Writing Tests is instructions for the agent only. Our summary lists: Python 3.

Does Writing Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Writing Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Writing Tests use?

Writing Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Writing Tests use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Writing Tests?

Skills that share tags, products or a category with Writing Tests: Requirements (rizsotto/Bear, 6.5k stars), Crap Analysis (ardalis/RiverBooks, 135 stars), Code Coverage (s3s-project/s3s, 311 stars) and Project Status (bactopia/bactopia, 522 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Writing Tests?

caarlos0 (a GitHub user) maintains it in caarlos0/dotfiles, which has 220 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 9, 2026.

Source: caarlos0/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.