Agent skill

Mutation Test Strength Audit

by buildfastwithai in buildfastwithai/gen-ai-experiments

Measures how well a Python pytest suite catches behavior changes through diff-scoped mutation testing, then proposes and verifies tests for surviving mutants.

MITAuto-check passedTesting & QA

Install Mutation Test Strength Audit

skills CLI
$ npx skills add buildfastwithai/gen-ai-experiments --skill test-strength -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install buildfastwithai/gen-ai-experiments test-strength --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/buildfastwithai/gen-ai-experiments.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-strength-codex-test-strength .claude/skills/test-strength && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-strength
GitHub stars
785
Token cost
~641 tokens
SKILL.md length
245 words
Files
16 (incl. scripts)
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Measures how well a Python pytest suite catches behavior changes through diff-scoped mutation testing, then proposes and verifies tests for surviving mutants.

  • Works in 3 steps: Confirm that Python and pytest are… → Run the bundled script, using its… → Read strength_report.json in the target…
  • Auditing whether a pytest suite really covers recently changed code
  • SKILL.md covers Run the audit, Analyze survivors and Report results
  • Runs Python scripts from its folder; calls python and pytest

What it does

The skill runs a bundled script, scripts/strength.py, from the repository under review. It mutates changed Python production files, comparing the current worktree with the merge-base of the default branch or a chosen --base, and runs the full test suite for every mutant so indirect, fixture-based and integration coverage all count. The runner copies the repo to a temporary directory for each mutant run, so your working tree stays unchanged. Results land in strength_report.json, where a killed mutant means the suite caught the change and a survivor means it did not.

For each survivor the agent inspects the diff, discards only genuinely equivalent mutants with a stated reason, and writes one focused test that passes on the original code and fails on the mutant. The test is checked with pytest and with the script's verify command, which must print VERIFICATION SUCCESS. The final report gives test, mutant and changed-function counts and lists only real survivors with their verified tests. Fixtures with strong and weak sample suites are included.

When your agent uses it

  • Auditing whether a pytest suite really covers recently changed code
  • Investigating surviving mutants after a refactor
  • Adding targeted tests for behavior the suite missed

Example prompts

  • “Run the test strength audit against main and show me the survivors.”
  • “Check whether my tests catch changes in src/billing.py.”
  • “For each surviving mutant, write a focused test and verify that it fails on the mutant.”

Requirements

  • Python with pytest
  • A git repository with a default branch to diff against

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Confirm that Python and pytest are available. Run from the target repository root.
  2. Run the bundled script, using its absolute skill path
  3. Read strength_report.json in the target repository root. Treat a killed mutant as behavior caught by the suite and a survived mutant as…

What it can do on your machine

Read from SKILL.md and the folder at commit 7b62043. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mutation Test Strength Audit loads about 641 tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 245 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~641

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from buildfastwithai/gen-ai-experiments at commit 7b62043, republished under its MIT licence (© buildfastwithai). 245 words, ~641 tokens.

Download SKILL.mdSave it as .claude/skills/test-strength/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
test-strength
description
Measure whether Python tests detect behavior changes through diff-scoped mutation testing. Use when Codex needs to audit the strength of a pytest suite, evaluate whether tests cover changed code, investigate surviving mutants, or propose and verify targeted tests for missed behavior.

Test Strength

Run the bundled runner from the repository being assessed. Keep the repository working tree unchanged: the runner copies it to a temporary directory before each mutant run.

Run the audit

  1. Confirm that Python and pytest are available. Run from the target repository root.

  2. Run the bundled script, using its absolute skill path:

    powershell
    python <skill-path>\scripts\strength.py [--base <ref>] [path]

    Omit both arguments to compare the current worktree with the merge-base of the default branch. Use --base main (or a commit) to choose the diff base. Pass an optional file or directory path to limit which changed Python production files are mutated.

  3. Read strength_report.json in the target repository root. Treat a killed mutant as behavior caught by the suite and a survived mutant as behavior that the suite did not catch. The runner executes the full test suite for every mutant so indirect, fixture-based, and integration coverage are included.

Analyze survivors

For each item in survived_mutants:

  1. Inspect its diff and decide whether it changes observable behavior for valid inputs.

  2. Discard only genuinely equivalent mutants; state why they are equivalent.

  3. For every real survivor, add one focused test that passes on the original code and fails for the mutant.

  4. Verify that test before suggesting it:

    powershell
    pytest <temporary-test-file>
    python <skill-path>\scripts\strength.py verify "<mutant-id>" <temporary-test-file>

    The first command must pass. The verification command must print VERIFICATION SUCCESS and exit with code 0. Remove the temporary verification file afterward unless the user asks to keep it.

Report results

Use this structure and include only real survivors:

markdown
<Total Tests> tests | <Total Mutants> mutants across <Changed Functions Count> changed functions
<Killed Count> killed | <Survived Count> survived | <Equivalent Count> equivalent (discarded)

### SURVIVED

#### <File Path>:<Line Number> <Function Name>
`<Original Expression>` -> `<Mutated Expression>`

Reason for survival: <why the current tests miss this behavior>

Suggested test:
```diff
+ <verified test change>

© buildfastwithai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (scripts) in skills/test-strength-codex-test-strength of buildfastwithai/gen-ai-experiments.

  • SKILL.md
  • README.md
  • agents/openai.yaml
  • fixtures/strong/auth.py
  • fixtures/strong/billing.py
  • fixtures/strong/calc.py
  • fixtures/strong/logger.py
  • fixtures/strong/test_strong.py
  • fixtures/strong/utils.py
  • fixtures/weak/auth.py
  • fixtures/weak/billing.py
  • fixtures/weak/calc.py
  • fixtures/weak/logger.py
  • fixtures/weak/test_weak.py
  • fixtures/weak/utils.py
  • scripts/strength.py

Open the folder on GitHubat commit 7b62043

Compare with similar skills

Mutation Test Strength Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mutation Test Strength Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mutation Test Strength Audit this skillbuildfastwithai/gen-ai-experiments785—~641Automated safety check: PassMIT
Test Coverage Reviewareed1192/finance-news-aggregator149—~2.6kAutomated safety check: PassMIT
TDD Guidealirezarezvani/claude-skills28k—~3.4kAutomated safety check: PassMIT
TDD GuideLeoYeAI/openclaw-master-skills2.2k—~1.4kAutomated safety check: PassMIT
Pytest Patternscohen-liel/hivemind110—~806Automated safety check: PassApache-2.0
Test Oracle GeneratorArabelaTso/Skills-4-SE253—~3.5kAutomated safety check: PassApache-2.0

Similar skills

  • Test Coverage Review

    areed1192/finance-news-aggregator

    Audit, plan, write, and verify unit tests for Python projects using pytest.

    149 GitHub stars~2.6k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • TDD Guide

    alirezarezvani/claude-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • TDD Guide

    LeoYeAI/openclaw-master-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    2.2k GitHub stars~1.4k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Pytest Patterns

    cohen-liel/hivemind

    pytest best practices for writing comprehensive test suites.

    110 GitHub stars~806 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Test Oracle Generator

    ArabelaTso/Skills-4-SE

    Generates automated test oracles to verify correct software behavior.

    253 GitHub stars~3.5k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed

More from buildfastwithai/gen-ai-experiments

All 13 skills in this repo
  • Landing Page Generator

    buildfastwithai/gen-ai-experiments

    Builds a single self-contained HTML landing page from a template with copy frameworks, four themes and SEO meta, then runs bundled audit scripts for conversion and speed.

    785 GitHub stars~1.5k tokensUpdated 16 days ago
    Auto-check passed
  • Talking Avatar Voice Chat App

    buildfastwithai/gen-ai-experiments

    Builds a realtime voice-chat app around a talking character portrait made from your photo or a text description, with mouth sprites driven by the audio.

    785 GitHub stars~1.7k tokensUpdated 16 days ago
    Auto-check passed
  • LaunchAudit

    buildfastwithai/gen-ai-experiments

    Audits a startup, app or landing page from a URL, localhost, repository, screenshots or copy, then gives a launch-readiness verdict, prioritized fixes and an HTML report.

    785 GitHub stars~1.8k tokensUpdated 16 days ago
    Auto-check passed
  • Startupblueprint

    buildfastwithai/gen-ai-experiments

    Turn a startup, SaaS, app, developer tool, website, repository, or product idea into an evidence-backed business plan, monetization strategy, pricing architecture, editable 12-month financial model…

    785 GitHub stars~2.2k tokensUpdated 16 days ago
    Auto-check passed
  • Eazyr Repo Onboarding Kit

    buildfastwithai/gen-ai-experiments

    Creates a QUICKSTART, doctor scripts and a documented .env.example for an unfamiliar repo from facts scanned in the repo, and audits how easy it is to onboard.

    785 GitHub stars~1.4k tokensUpdated 16 days ago
    Auto-check: notes
  • Customer Finder

    buildfastwithai/gen-ai-experiments

    Find and qualify evidence-backed potential first customers, early adopters, design partners, or beta users for a startup using recent public pain and buying signals.

    785 GitHub stars~1.4k tokensUpdated 16 days ago
    Auto-check passed

Works with

Categories

Questions about Mutation Test Strength Audit

What does Mutation Test Strength Audit do?

Measures how well a Python pytest suite catches behavior changes through diff-scoped mutation testing, then proposes and verifies tests for surviving mutants. py, from the repository under review. It mutates changed Python production files, comparing the current worktree with the merge-base of the default branch or a chosen --base, and runs the full test suite for every mutant so indirect, fixture-based and integration coverage all count.

When should I use Mutation Test Strength Audit?

Mutation Test Strength Audit fits situations like: auditing whether a pytest suite really covers recently changed code; investigating surviving mutants after a refactor; adding targeted tests for behavior the suite missed.

How do I install Mutation Test Strength Audit in Claude Code?

Run `npx skills add buildfastwithai/gen-ai-experiments --skill test-strength -a claude-code`. Or copy the skill folder (skills/test-strength-codex-test-strength in buildfastwithai/gen-ai-experiments) into .claude/skills/test-strength in your project. Claude Code loads it when a task matches its description.

How do I install Mutation Test Strength Audit in Codex?

Run `npx skills add buildfastwithai/gen-ai-experiments --skill test-strength -a codex`. Or copy the skill folder (skills/test-strength-codex-test-strength in buildfastwithai/gen-ai-experiments) into .agents/skills/test-strength in your project. Codex loads it when a task matches its description.

Can I use Mutation Test Strength Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add buildfastwithai/gen-ai-experiments --skill test-strength -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-strength, .gemini/skills/test-strength, .github/skills/test-strength and .opencode/skills/test-strength in your project.

What does Mutation Test Strength Audit need to run?

Going by SKILL.md and its folder, Mutation Test Strength Audit needs Python for the scripts in its folder and the command-line tools its instructions call (python and pytest). Our summary lists: Python with pytest; A git repository with a default branch to diff against.

Does Mutation Test Strength Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Mutation Test Strength Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Mutation Test Strength Audit use?

Mutation Test Strength Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mutation Test Strength Audit use?

About 641 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Mutation Test Strength Audit?

Skills that share tags, products or a category with Mutation Test Strength Audit: Test Coverage Review (areed1192/finance-news-aggregator, 149 stars), TDD Guide (alirezarezvani/claude-skills, 28k stars), TDD Guide (LeoYeAI/openclaw-master-skills, 2.2k stars) and Pytest Patterns (cohen-liel/hivemind, 110 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mutation Test Strength Audit?

buildfastwithai (a GitHub organization) maintains it in buildfastwithai/gen-ai-experiments, which has 785 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on September 22, 2026.

Source: buildfastwithai/gen-ai-experiments on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.