Agent skill

Test The Tests

by werf in werf/werf

Verify a test actually falsifies the behavior it claims to cover, via real mutation.

Apache-2.0Auto-check passedDevOps & Cloud

Install Test The Tests

skills CLI
$ npx skills add werf/werf --skill test-the-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install werf/werf test-the-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/werf/werf.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/test-the-tests .claude/skills/test-the-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-the-tests
GitHub stars
4.7k
Token cost
~1k tokens
SKILL.md length
599 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Verify a test actually falsifies the behavior it claims to cover, via real mutation.

  • Works in 3 steps: Pick the smallest plausible fault:… → Apply it to the implementation, run the… → Revert immediately, confirm the tree is…
  • Tasks that involve CI/CD
  • SKILL.md covers The mutation loop, Common ways a test looks… and Output
  • Calls git

What it does

Test The Tests is an agent skill from werf/werf. Verify a test actually falsifies the behavior it claims to cover, via real mutation. Use right after writing or changing a test, when auditing an existing suite's real coverage, or as a step inside a code review.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering CI/CD. It works with Docker and Git. The repository describes itself as: A solution for implementing efficient and consistent software delivery to Kubernetes facilitating best practices. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve CI/CD

Example prompts

  • “/test-the-tests”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Pick the smallest plausible fault: invert a condition, change < to <=, remove a
  2. Apply it to the implementation, run the test, confirm it fails for the right reason —
  3. Revert immediately, confirm the tree is clean (git status/git diff), confirm the

What it can do on your machine

Read from SKILL.md and the folder at commit fa73c7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test The Tests loads about 1k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 599 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from werf/werf at commit fa73c7a, republished under its Apache-2.0 licence (© werf). 599 words, ~1,034 tokens.

Download SKILL.mdSave it as .claude/skills/test-the-tests/SKILL.md (or your agent's skills folder).
name
test-the-tests
description
Verify a test actually falsifies the behavior it claims to cover, via real mutation. Use right after writing or changing a test, when auditing an existing suite's real coverage, or as a step inside a code review.

Test the Tests

Passing and high coverage are not evidence, whoever wrote the test — including you, just now. A test proves it runs; it does not prove it discriminates correct from incorrect behavior. Ask:

What fault would this test fail to catch?

The mutation loop

For each test that carries real weight (guards a fix, a regression, an invariant — not a trivial getter):

  1. Pick the smallest plausible fault: invert a condition, change < to <=, remove a validation, suppress an error, skip a side effect, hardcode a return value, revert to the prior (buggy) behavior.
  2. Apply it to the implementation, run the test, confirm it fails for the right reason — not an unrelated crash.
  3. Revert immediately, confirm the tree is clean (git status/git diff), confirm the suite passes again. NEVER leave mutated code in the repository between steps.

Copy the file before mutating it (cp f f.bak) and restore from the copy. NEVER restore with git checkout/git restore while the work under test is still uncommitted — those discard the work along with the mutation. Commit first, or use the copy.

If running the mutation isn't practical, name the smallest mutation that should be tried instead of skipping the exercise — that name is itself the finding.

"Not practical" is a conclusion, not an assumption — establish it as AGENTS.md requires (is the runtime actually missing, is a Linux host available) before falling back to naming the mutation. A "can't run it here" that turns out to be wrong ships tests nobody has ever seen fail.

Common ways a test looks strong but isn't

  • The assertion holds under the bug too. A chain assertion like index(a) < index(b) < index(c) can pass under both the fix and the regression it's meant to catch if the scenario doesn't force them to disagree. Reshape the scenario (add a case whose correct answer differs from what the bug would produce) instead of trusting a single run.
  • It re-asserts a second recording of the same event instead of a property true by construction — comparing two independently-recorded orderings is racier and weaker than asserting a monotonic counter or a guaranteed dependency order.
  • It exercises the unit in isolation, never the real wiring a regression would actually break. Prefer driving the real entry point end-to-end when the risk is in the wiring.
  • It asserts implementation trivia (mock called N times, internal helper ran) instead of externally observable behavior.
  • Coverage number, not falsifiability. A line being executed says nothing about whether a wrong value on that line would be caught.
  • It proves the lab, not production wiring. A probe on a bare context "proved" that cancellation panics past a branch — but production wraps commands in graceful.WithTermination, where Terminate returns normally and the branch executes. Before concluding "X cannot reach this code", rebuild the exact wiring production uses (context construction, env, entry point); a conclusion drawn from simplified wiring describes the probe, not the system.
  • An extended test quietly drops what the old one proved. Adding cases to a fixture can remove the conflict that made an earlier property observable. After editing an existing test, re-run the mutations the previous version caught, not only the new ones.
  • Ordered behavior tested without conflicts. For precedence lists and fallback chains, a fixture with one candidate per level survives swapping adjacent candidates. Give every level two candidates whose effects differ.
Show full SKILL.md (44 more words)Show less

Output

For each test verified this way: the fault tried, whether it failed as expected, and if not, the smallest change that would make it discriminate. For tests not mutated, name the smallest fault that should be tried next — don't assert confidence without it.

© werf, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/test-the-tests of werf/werf.

Open the folder on GitHubat commit fa73c7a

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in werf/werf, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Test The Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test The Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test The Tests this skillwerf/werf4.7k—~1kAutomated safety check: PassApache-2.0
Infrastructure Setuppavel-molyanov/molyanov-ai-dev296—~1.9kAutomated safety check: NotesMIT
Deploy To Tempsgotempsh/temps822—~1.3kAutomated safety check: NotesApache-2.0
Scanning Containers With Trivy In Cicdmukul975/Anthropic-Cybersecurity-Skills34k—~2.7kAutomated safety check: PassApache-2.0
Modern Web GuidanceJetBrains/skills3633 repos~1.4kAutomated safety check: PassApache-2.0
Performing Container Security Scanning With Trivymukul975/Anthropic-Cybersecurity-Skills34k—~818Automated safety check: PassApache-2.0

Similar skills

  • Infrastructure Setup

    pavel-molyanov/molyanov-ai-dev

    Provides project infrastructure conventions and review criteria for local setup, Docker, Git hooks, CI/CD, service delivery, release artifacts, monitoring, backups, and operations.

    296 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check: notes
  • Deploy To Temps

    gotempsh/temps

    Deploy applications to the Temps platform with automatic framework detection, Dockerfile generation, and container orchestration.

    822 GitHub stars~1.3k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Scanning Containers With Trivy In Cicd

    mukul975/Anthropic-Cybersecurity-Skills

    Integrates Aqua Security's Trivy scanner into CI/CD pipelines to detect OS package and application dependency CVEs, Dockerfile misconfigurations, and issues in filesystems or git repositories, and…

    34k GitHub stars~2.7k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Modern Web Guidance

    JetBrains/skills

    Official

    Search tool for modern web development best practices. An agent skill from JetBrains/skills.

    363 GitHub starsUsed in 3 repos~1.4k tokens
    DevOps & CloudAuto-check passed
  • Performing Container Security Scanning With Trivy

    mukul975/Anthropic-Cybersecurity-Skills

    Runs Trivy across every target type it supports - container images, filesystems, Git repositories, and Kubernetes clusters - for OS and dependency vulnerabilities, IaC misconfiguration, exposed…

    34k GitHub stars~818 tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Devops Automation

    rohitg00/awesome-claude-code-toolkit

    CI/CD pipeline design with GitHub Actions, Docker, Kubernetes, Helm, and GitOps patterns

    2.7k GitHub stars~1.6k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed

More from werf/werf

  • werf conventions for branch names and commit messages. An agent skill from werf/werf.

    4.7k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Pull Request

    werf/werf

    Generates Pull Request titles and descriptions according to werf conventions.

    4.7k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Review

    werf/werf

    Code review of a pull request, branch, or diff. An agent skill from werf/werf.

    4.7k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Independent challenge pass for a non-trivial or high-risk change.

    4.7k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • How to treat conclusions inherited from an earlier session — handover notes, prepared comments, verdict files, plans.

    4.7k GitHub stars~669 tokensUpdated yesterday
    Auto-check passed
  • Session Retro

    werf/werf

    Analyze the current session for harness-worthy lessons — repeated corrections, discovered conventions, skill bugs — and turn them into concrete repo changes: docs, skills, task targets, linter…

    4.7k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Test The Tests

What does Test The Tests do?

Verify a test actually falsifies the behavior it claims to cover, via real mutation. Test The Tests is an agent skill from werf/werf. Verify a test actually falsifies the behavior it claims to cover, via real mutation.

When should I use Test The Tests?

Test The Tests fits situations like: tasks that involve CI/CD.

How do I install Test The Tests in Claude Code?

Run `npx skills add werf/werf --skill test-the-tests -a claude-code`. Or copy the skill folder (.agents/skills/test-the-tests in werf/werf) into .claude/skills/test-the-tests in your project. Claude Code loads it when a task matches its description.

How do I install Test The Tests in Codex?

Run `npx skills add werf/werf --skill test-the-tests -a codex`. Or copy the skill folder (.agents/skills/test-the-tests in werf/werf) into .agents/skills/test-the-tests in your project. Codex loads it when a task matches its description.

Can I use Test The Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add werf/werf --skill test-the-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-the-tests, .gemini/skills/test-the-tests, .github/skills/test-the-tests and .opencode/skills/test-the-tests in your project.

What does Test The Tests need to run?

Going by SKILL.md and its folder, Test The Tests needs the command-line tools its instructions call (git).

Does Test The Tests access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test The Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test The Tests use?

Test The Tests is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test The Tests use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test The Tests?

Skills that share tags, products or a category with Test The Tests: Infrastructure Setup (pavel-molyanov/molyanov-ai-dev, 296 stars), Deploy To Temps (gotempsh/temps, 822 stars), Scanning Containers With Trivy In Cicd (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Modern Web Guidance (JetBrains/skills, 363 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test The Tests?

werf (a GitHub organization) maintains it in werf/werf, which has 4,728 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.

Source: werf/werf on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.