Agent skill

Testing

by WrongStack in WrongStack/WrongStack

A skill your agent uses when writing, fixing, reviewing, or planning tests in any project, in whatever runner the project already uses.

MITAuto-check passedTesting & QA

Install Testing

skills CLI
$ npx skills add WrongStack/WrongStack --skill testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install WrongStack/WrongStack testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/WrongStack/WrongStack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/core/skills/testing .claude/skills/testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing
GitHub stars
368
Token cost
~2k tokens
SKILL.md length
983 words
Files
2
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when writing, fixing, reviewing, or planning tests in any project, in whatever runner the project already uses.

  • Works in 7 steps: Match the project before writing… → See it fail first. A regression test… → Test behaviour through the public… → …
  • Planning tests in any project
  • SKILL.md covers Overview, Rules, Workflow and From proof to regression test, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Testing is an agent skill from WrongStack/WrongStack. Use this skill when writing, fixing, reviewing, or planning tests in any project, in whatever runner the project already uses. Also use it to write the failing proof for a suspected bug and to promote that proof into a durable regression test. Triggers: user says "test", "unit test", "integration test", "e2e", "mock", "coverage", "flaky", "failing test", "regression test", "write tests", "vitest", "jest", "pytest", "go test", "proof", "red/green".

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `SKILL.save.md`).

It sits in Testing & QA, covering Unit testing and Failing and flaky tests. It works with Jest, pytest and Vitest. The repository describes itself as: An AI coding agent that reads your code, edits files, runs commands, and reasons through bugs — across a terminal REPL, a full-screen TUI, and a browser UI, while you keep your… The licence is MIT.

When your agent uses it

  • Planning tests in any project
  • In whatever runner the project already uses
  • Write the failing proof for a suspected bug and to promote that proof into a durable regression test

Example prompts

  • “unit test”
  • “integration test”
  • “coverage”
  • “/testing”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Match the project before writing anything. Find the runner (package.json
  2. See it fail first. A regression test must fail without the fix — a test that
  3. Test behaviour through the public surface. Don't assert on private helpers or
  4. Mock the boundaries you don't own (network, clock, randomness, third-party
  5. Keep every test isolated: no order dependence, no shared mutable state;
  6. Bound every wait that can hang (network, sockets, child processes, polling),
  7. Report exactly what ran: the command, the files, pass/fail/skip counts. A run

What it can do on your machine

Read from SKILL.md and the folder at commit ec76a20. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing loads about 2k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 983 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from WrongStack/WrongStack at commit ec76a20, republished under its MIT licence (© WrongStack). 983 words, ~1,972 tokens.

Download SKILL.mdSave it as .claude/skills/testing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
testing
description
Use this skill when writing, fixing, reviewing, or planning tests in any project, in whatever runner the project already uses. Also use it to write the failing proof for a suspected bug and to promote that proof into a durable regression test. Triggers: user says "test", "unit test", "integration test", "e2e", "mock", "coverage", "flaky", "failing test", "regression test", "write tests", "vitest", "jest", "pytest", "go test", "proof", "red/green".
version
2.1.0
required-capabilities
filesystem.read, verification.run
optional-capabilities
execution.shell, code.inspect

Testing

Overview

Write tests that fail for the right reason and pass for the right reason, in the project's own runner, layout, and style. A test earns its place by catching a regression someone could plausibly introduce; everything else is maintenance cost.

Rules

  1. Match the project before writing anything. Find the runner (package.json scripts, vitest/jest config, pytest.ini or pyproject, go.mod, Cargo.toml), the test layout, the naming, and the helpers neighbouring tests already use.
  2. See it fail first. A regression test must fail without the fix — a test that was never red proves nothing.
  3. Test behaviour through the public surface. Don't assert on private helpers or internal structure a legitimate refactor would change.
  4. Mock the boundaries you don't own (network, clock, randomness, third-party SDKs, slow I/O), not the collaborator next door.
  5. Keep every test isolated: no order dependence, no shared mutable state; restore mocks, timers, and environment in teardown.
  6. Bound every wait that can hang (network, sockets, child processes, polling), and drive time-based logic with fake timers instead of real sleeps.
  7. Report exactly what ran: the command, the files, pass/fail/skip counts. A run that matched no tests is not a pass.

Workflow

  1. Locate the code under test and its existing tests. With a codebase index, the codebase-context and codebase-search tools find them faster than grep.
  2. Pick the level. Unit for pure logic; integration where the bug lives in the wiring; end-to-end only for a user-visible flow nothing cheaper covers.
  3. Write the smallest failing test that pins the behaviour. Run it and confirm it fails on the assertion, not on an import or setup error.
  4. Make it pass, or confirm the fix makes it pass.
  5. Widen. Run the covering suites (the codebase-targeted-test tool finds them for a symbol or file), then the full suite when the change touches shared code.

From proof to regression test

A bug fix usually starts with a throwaway proof — a script or scratch test that went red against the unfixed code. It is not done until that case lives in the project's normal suite.

  1. Place it where the suite runs it. Same runner, same layout, next to the existing tests for that module. Check the runner's include and exclude patterns: a test file the config never picks up passes forever by not running.
  2. Keep the exact trigger from the proof, then add what the proof skipped: the important boundary (empty, exact limit, last item), the secondary branch the fix touched, and the control case that must keep passing.
  3. Name the behaviour, not the ticket. "keeps the abort listener count flat across retries" survives; "fixes bug 42" tells the next reader nothing.
  4. See it red against the unfixed code. If the proof already went red with the same assertions, that counts. Otherwise run a mutation check: back up the fixed file, restore the old code in place, run the test and watch it fail, then restore the backup and watch it pass. Never use stash, checkout, or reset for this in a shared working tree — they carry other people's edits away with yours.
  5. Make it a good citizen. No real sleeps, no leaked timers, handles, listeners, or temp files; every wait bounded. A regression test that is itself flaky will get skipped, and the bug comes back.
Show full SKILL.md (427 more words)Show less

Choosing what to assert

SituationAssertAvoid
Pure functionOutputs across normal, boundary, and invalid inputs (table-driven)Intermediate variables
Error pathThe error type, code, or message the caller relies onA bare "throws" with no matcher
Async flowFinal state and outputs after completion is awaitedArbitrary sleeps
Bug fixThe exact input from the bug reportA paraphrase that already passed before the fix
UI componentWhat the user sees and can do (roles, text, events)Whole-tree snapshots, class names

Flaky tests

A flaky test is a bug in the test or in the code, never background noise.

SymptomUsual causeFix
Fails under load or in CI onlyReal time, timers, racesFake timers; inject the clock; await the real completion signal
Fails depending on orderLeaked state between testsReset in teardown; run the file alone and shuffled
Fails intermittently with no errorUnawaited promiseAwait it; enable floating-promise lint
Fails when suites run in parallelShared ports, files, envEphemeral ports, per-test temp dirs, scoped env

Don't "fix" flakiness with retries or larger timeouts until the cause is found, and say what the cause was.

Patterns

Examples use vitest/jest syntax; translate to the project's runner.

ts
describe('parseDuration', () => {
  it.each([
    ['90s', 90_000],
    ['2m', 120_000],
    ['0s', 0],
  ])('parses %s', (input, expected) => {
    expect(parseDuration(input)).toBe(expected);
  });

  it('rejects an unknown unit', () => {
    expect(() => parseDuration('5y')).toThrow(/unknown unit/);
  });
});

describe('withRetry', () => {
  afterEach(() => {
    vi.useRealTimers();
    vi.restoreAllMocks();
  });

  it('retries once after the backoff delay', async () => {
    vi.useFakeTimers();
    const call = vi.fn().mockRejectedValueOnce(new Error('503')).mockResolvedValue('ok');
    const pending = withRetry(call, { delayMs: 1_000 });
    await vi.advanceTimersByTimeAsync(1_000);
    await expect(pending).resolves.toBe('ok');
    expect(call).toHaveBeenCalledTimes(2); // the retry count is the contract here
  });
});

Anti-patterns

  • Tests that mirror the implementation line by line — they break on every refactor and catch nothing.
  • Mocking the unit under test, or mocking so much that the test checks the mock.
  • Loosening or deleting an assertion to make a failing test pass. Find out why it fails first.
  • Skipping tests or lowering coverage thresholds to get a green run.
  • Snapshots as the only assertion on logic.
  • Claiming "tests pass" from a filtered or partial run without saying so.
  • Leaving the proof only in a scratch directory — deleted with the cleanup, so nothing guards the fix.
  • A red run that failed for the wrong reason — import error, missing fixture, timeout — counted as the bug reproducing.

Before returning

  • Runner, layout, naming, and helpers match the project's existing tests
  • Every new regression test was seen failing before the fix, on its assertion rather than on setup
  • The test file is inside the runner's include patterns and actually ran
  • Assertions target behaviour with specific matchers
  • Mocks, timers, and environment restored; no order dependence
  • Commands and results reported exactly, including skips and filters

Skills in scope

  • debugging — when a failing test's cause is unknown
  • bug-hunter — when a failing test points at a real defect to locate
  • typescript-strict — for type-safe fixtures and assertions
  • verify-before-done — for the evidence to report once tests pass
  • git-flow — for committing tests together with the change they cover

© WrongStack, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in packages/core/skills/testing of WrongStack/WrongStack.

  • SKILL.md
  • SKILL.save.md

Open the folder on GitHubat commit ec76a20

Compare with similar skills

Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing this skillWrongStack/WrongStack368—~2kAutomated safety check: PassMIT
Designing TestsCloudAI-X/claude-workflow-v21.4k1 repos~1.5kAutomated safety check: PassMIT
Test GuardamElnagdy/guard-skills1.3k2 repos~2.1kAutomated safety check: PassMIT
Agent Harness Testing Methodologyhuiliyi37/Tianshu-harness1.1k—~1kAutomated safety check: NotesApache-2.0
Playwright Testingchongdashu/vibejam-starter-pack149—~2.1kAutomated safety check: PassNone
Playwright Testingchongdashu/vibejam-starter-pack149—~2.2kAutomated safety check: PassNone

Similar skills

  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed
  • Test Guard

    amElnagdy/guard-skills

    Reviews newly written or edited tests against nine rules that cut test bloat, such as mock-heavy checks and near-duplicate cases, before they are committed.

    1.3k GitHub starsUsed in 2 repos~2.1k tokens
    Testing & QAAuto-check passed
  • Agent Harness Testing Methodology

    huiliyi37/Tianshu-harness

    Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.

    1.1k GitHub stars~1k tokensUpdated 2 days ago
    Testing & QAAuto-check: notes
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.1k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Playwright Testing

    chongdashu/vibejam-starter-pack

    Plan, implement, and debug frontend tests: unit/integration/E2E/visual/a11y.

    149 GitHub stars~2.2k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • MoAI TDD Workflow

    modu-ai/moai-adk

    Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.

    1.2k GitHub stars~3.1k tokensUpdated today
    Testing & QAAuto-check passed

More from WrongStack/WrongStack

All 38 skills in this repo
  • Design Craft

    WrongStack/WrongStack

    Design or substantially improve user-facing interfaces with a product-specific visual direction, content hierarchy, and rendered critique.

    368 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Design Critique

    WrongStack/WrongStack

    A skill your agent uses to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states…

    368 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Mailbox Bridge

    WrongStack/WrongStack

    A skill your agent uses when external coding agents (Claude Code, Aider, custom scripts) need to participate in the project's shared WrongStack mailbox, or when a user asks to "expose the mailbox"…

    368 GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Multi Agent

    WrongStack/WrongStack

    A skill your agent uses whenever work can be split across multiple AI agents running in parallel, or when orchestrating leader/worker patterns in WrongStack.

    368 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Web Platform Baseline

    WrongStack/WrongStack

    Use this skill before asserting that a CSS, HTML or accessibility capability is available, unavailable, or the right tool — it carries dated, refreshable platform facts and refuses to let stale…

    368 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Wrongstack Mailbox

    WrongStack/WrongStack

    A skill your agent uses when the user wants to communicate with WrongStack's shared project mailbox from outside WrongStack — read messages sent by WrongStack agents, send replies, broadcast to all…

    368 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Testing

What does Testing do?

A skill your agent uses when writing, fixing, reviewing, or planning tests in any project, in whatever runner the project already uses. Testing is an agent skill from WrongStack/WrongStack. Use this skill when writing, fixing, reviewing, or planning tests in any project, in whatever runner the project already uses.

When should I use Testing?

Testing fits situations like: planning tests in any project; in whatever runner the project already uses; write the failing proof for a suspected bug and to promote that proof into a durable regression test.

How do I install Testing in Claude Code?

Run `npx skills add WrongStack/WrongStack --skill testing -a claude-code`. Or copy the skill folder (packages/core/skills/testing in WrongStack/WrongStack) into .claude/skills/testing in your project. Claude Code loads it when a task matches its description.

How do I install Testing in Codex?

Run `npx skills add WrongStack/WrongStack --skill testing -a codex`. Or copy the skill folder (packages/core/skills/testing in WrongStack/WrongStack) into .agents/skills/testing in your project. Codex loads it when a task matches its description.

Can I use Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add WrongStack/WrongStack --skill testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing, .gemini/skills/testing, .github/skills/testing and .opencode/skills/testing in your project.

What does Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Testing is instructions for the agent only.

Does Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing use?

Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing?

Skills that share tags, products or a category with Testing: Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars), Test Guard (amElnagdy/guard-skills, 1.3k stars), Agent Harness Testing Methodology (huiliyi37/Tianshu-harness, 1.1k stars) and Playwright Testing (chongdashu/vibejam-starter-pack, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing?

WrongStack (a GitHub organization) maintains it in WrongStack/WrongStack, which has 368 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on October 6, 2026.

Source: WrongStack/WrongStack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.