Agent skill

Mutation Test

by jmagly in jmagly/aiwg

Run mutation testing to validate test quality beyond code coverage.

MITAuto-check passedTesting & QA

Install Mutation Test

skills CLI
$ npx skills add jmagly/aiwg --skill mutation-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jmagly/aiwg mutation-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jmagly/aiwg.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agentic/code/addons/testing-quality/skills/mutation-test .claude/skills/mutation-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mutation-test
GitHub stars
220
Token cost
~3.2k tokens
SKILL.md length
933 words
Files
3 (incl. scripts)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Run mutation testing to validate test quality beyond code coverage.

  • Works in 5 steps: Detect Project and Resolve Tool → Configure Mutation Testing → Preflight Python Native Extensions → …
  • Assessing test effectiveness
  • SKILL.md covers Purpose, Research Foundation, When This Skill Applies and Trigger Phrases, plus 7 more sections
  • Runs Python scripts from its folder; calls python, npx and mvn

What it does

Mutation Test is an agent skill from jmagly/aiwg. Run mutation testing to validate test quality beyond code coverage. Use when assessing test effectiveness, finding weak tests, or validating test suite quality.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/aiwg_mutation_native_probe.py` and `scripts/native_extension_preflight.py`).

It sits in Testing & QA, covering Test coverage and Test generation. The repository describes itself as: Cognitive architecture for AI-augmented software development. Specialized agents, structured workflows, and multi-platform deployment. Claude Code · Codex · Copilot · Cursor ·… The licence is MIT.

When your agent uses it

  • Assessing test effectiveness
  • Finding weak tests
  • Validating test suite quality

Example prompts

  • “/mutation-test”

Requirements

  • Python 3
  • Node.js

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Detect Project and Resolve Tool
  2. Configure Mutation Testing
  3. Preflight Python Native Extensions
  4. Run Mutation Analysis
  5. Classify and Report Results

What it can do on your machine

Read from SKILL.md and the folder at commit dda238f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • npx
    • mvn
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • conf.researchr.org
    • stryker-mutator.io
    • pitest.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mutation Test loads about 3.2k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 933 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jmagly/aiwg at commit dda238f, republished under its MIT licence (© jmagly). 933 words, ~3,188 tokens.

Download SKILL.mdSave it as .claude/skills/mutation-test/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
mutation-test
description
Run mutation testing to validate test quality beyond code coverage. Use when assessing test effectiveness, finding weak tests, or validating test suite quality.
namespace
aiwg
version
1.1.0
platforms
all

Mutation Test Skill

Purpose

Run mutation testing to measure test suite effectiveness. Mutation testing introduces small changes (mutants) to code and checks if tests catch them. High coverage with low mutation score indicates weak tests.

Research Foundation

ConceptSourceReference
Mutation Testing TheoryIEEE TSE (2019)Papadakis et al. "Mutation Testing Advances"
ICST Mutation WorkshopIEEE AnnualMutation 2024
Stryker MutatorIndustry Toolstryker-mutator.io
PITestJava Toolpitest.org
mutmutPython Toolgithub.com/boxed/mutmut

When This Skill Applies

  • User asks to "validate test quality" or "check test effectiveness"
  • User mentions "mutation testing" or "mutation score"
  • User wants to know if tests are "actually testing anything"
  • High coverage but bugs still escaping
  • Assessing test suite health
  • Pre-release quality validation

Trigger Phrases

Natural LanguageAction
"Run mutation testing"Execute mutation analysis
"Check if my tests are effective"Run mutation + analyze
"Validate test quality"Mutation score report
"Are my tests catching real bugs?"Mutation analysis
"Find weak tests"Identify low-score tests
"Why did this bug escape tests?"Mutation analysis on module

Mutation Testing Concepts

What is a Mutant?

A mutant is a small code change that should cause tests to fail:

javascript
// Original
if (age >= 18) { return "adult"; }

// Mutant 1: Changed >= to >
if (age > 18) { return "adult"; }

// Mutant 2: Changed >= to ==
if (age == 18) { return "adult"; }

// Mutant 3: Changed "adult" to ""
if (age >= 18) { return ""; }
Mutation Operators
OperatorExampleTests
Arithmetic+ → -Math operations
Relational>= → >Boundary conditions
Logical&& → ||Boolean logic
Literaltrue → falseConstant handling
Returnreturn x → return nullReturn value handling
Mutation Score
Mutation Score = (Killed Mutants / Total Mutants) × 100
ScoreQualityInterpretation
90%+ExcellentTests are highly effective
80-89%GoodTarget for production
60-79%AdequateRoom for improvement
<60%PoorTests need significant work

Implementation Process

1. Detect Project and Resolve Tool
python
def mutation_tool(project_type):
    if project_type == "javascript":
        return "stryker"
    elif project_type == "python":
        return "mutmut"
    elif project_type == "java":
        return "pitest"

Use the project's existing lockfile or approved dependency workflow. Record the resolved tool and version in the evidence; do not install an unpinned tool as an implicit side effect of this skill.

2. Configure Mutation Testing

Stryker (JavaScript):

json
// stryker.config.json
{
  "mutate": ["src/**/*.ts", "!src/**/*.test.ts"],
  "testRunner": "vitest",
  "reporters": ["html", "progress"],
  "coverageAnalysis": "perTest",
  "thresholds": {
    "high": 80,
    "low": 60,
    "break": 50
  }
}

mutmut 3 (Python):

toml
# pyproject.toml
[tool.mutmut]
source_paths = ["src/package/"]
only_mutate = ["src/package/critical_contract.py"]
pytest_add_cli_args = ["--no-cov", "-q"]
pytest_add_cli_args_test_selection = ["tests/test_critical_contract.py"]
mutate_only_covered_lines = true

Keep both only_mutate and the selected tests explicit. A directory-wide fallback is not an acceptable substitute for a target that was selected to fit the runtime budget.

PITest (Java):

xml
<!-- pom.xml -->
<plugin>
    <groupId>org.pitest</groupId>
    <artifactId>pitest-maven</artifactId>
    <version>1.15.0</version>
    <configuration>
        <targetClasses>
            <param>com.example.*</param>
        </targetClasses>
        <mutationThreshold>80</mutationThreshold>
    </configuration>
</plugin>
3. Preflight Python Native Extensions

mutmut 3's covered-line mode runs coverage and stats in the same parent Python process. Its coverage phase snapshots sys.modules, runs the selected tests, then unloads every newly imported module. A later stats pass can re-import a native extension that cannot safely be initialized twice. This is an upstream mutmut boundary, not a killed/survived mutant and not evidence that the project tests failed. See mutmut #528.

Before any Python run with mutate_only_covered_lines = true, execute the selected tests through the bundled subprocess probe:

bash
python scripts/native_extension_preflight.py \
  --test-selection tests/test_critical_contract.py \
  --pytest-arg=--no-cov \
  --pytest-arg=-q \
  --mutation-target src/package/critical_contract.py \
  --estimated-mutants 120 \
  --max-children 4 \
  --runtime-budget-seconds 900 \
  --format json

The probe runs the test selection in a disposable subprocess and records native modules imported after its pytest plugin loads. It never unloads or re-imports a module. Use --import-file or --import-module for import-only checks when a pytest selection is not yet available.

ExitClassificationRequired action
0preflight_safeCovered-line mode may proceed; preserve the report.
2harness_native_extension_reload_riskDo not run mutmut covered-line mode. Use only an allowed bounded fallback or a project-approved subprocess-isolated tool.
3project_test_or_import_failureFix or clarify the direct project test/import failure before mutation testing.
4harness_preflight_timeoutReduce the selection or increase the explicitly approved preflight budget.

Detection is conservative. A safe result proves only that this selected run did not load a native extension after the probe boundary; change the selection or environment and the preflight must be repeated.

Show full SKILL.md (397 more words)Show less
Bounded mutmut fallback

When a native extension is found, the supported mutmut fallback is mutate_only_covered_lines = false, but only when the JSON report says fallback.allowed: true. The fallback gate requires all of:

  • one or more explicit --mutation-target values matching only_mutate;
  • an observed or tool-generated --estimated-mutants count (do not guess);
  • the intended --max-children value;
  • an explicit --runtime-budget-seconds value; and
  • the conservative estimate baseline_test_seconds × estimated_mutants ÷ max_children × 1.5 to fit the budget.

If any bound is missing or the estimate exceeds budget, shrink the mutation target or use a project-approved mutation harness whose coverage and stats phases start in fresh subprocesses. Never silently widen only_mutate to compensate for disabling covered-line selection.

4. Run Mutation Analysis
bash
# JavaScript
npx stryker run

# Python, only after the preflight decision is recorded
python -m mutmut run --max-children 4

# Java
mvn org.pitest:pitest-maven:mutationCoverage
5. Classify and Report Results

Keep harness/tool failures, direct project-test failures, and mutant outcomes in separate fields. For an existing mutmut log, the preflight script can classify known native reload signatures:

bash
python scripts/native_extension_preflight.py \
  --classify-mutmut-log mutation-run.log \
  --format json

Running stats combined with cannot load module more than once per process, module functions cannot set METH_CLASS or METH_STATIC, or a fatal segmentation fault whose report lists extension modules is harness_tool_failure_native_extension_reload. It must set counts_as_mutant_outcome: false; do not compute a mutation score from that run. failed to collect stats is supporting context, not a sufficient signature by itself. Only a direct test run that fails independently of the mutation harness is a project_test_failure. Killed, survived, timeout, and no-tests results are mutant outcomes only after preflight, baseline tests, stats collection, and the mutation harness complete successfully.

python
def parse_mutation_results(report_path):
    """Parse mutation testing report"""
    return {
        "execution_mode": "mutmut-covered-lines-preflight-safe",
        "harness_status": "passed",
        "project_test_status": "passed",
        "total_mutants": 150,
        "killed": 120,
        "survived": 25,
        "timeout": 5,
        "mutation_score": 80.0,
        "survivors": [
            {
                "file": "src/auth/validate.ts",
                "line": 45,
                "mutator": "RelationalOperator",
                "original": "age >= 18",
                "mutant": "age > 18",
                "status": "survived"
            }
            # ... more survivors
        ]
    }

Output Format

markdown
## Mutation Testing Report

**Module**: src/auth/
**Test Suite**: test/auth/
**Execution mode**: mutmut covered-line mode (native-extension preflight safe)
**Harness status**: passed
**Project baseline tests**: passed

### Summary

| Metric | Value |
|--------|-------|
| Total Mutants | 150 |
| Killed | 120 (80%) |
| Survived | 25 (17%) |
| Timeout | 5 (3%) |
| **Mutation Score** | **80%** |

### Status: PASSED (threshold: 80%)

### Survived Mutants (Highest Priority)

#### 1. `src/auth/validate.ts:45`
```diff
- if (age >= 18) { return "adult"; }
+ if (age > 18) { return "adult"; }

Problem: Boundary condition not tested Fix: Add test case for age = 18

2. src/auth/login.ts:23
diff
- if (attempts < maxAttempts) { allow(); }
+ if (attempts <= maxAttempts) { allow(); }

Problem: Off-by-one boundary not tested Fix: Add test for attempts = maxAttempts

  1. Add boundary tests for validate.ts (3 survivors)
  2. Add error path tests for login.ts (2 survivors)
  3. Test null/undefined cases in session.ts (1 survivor)
Coverage vs Mutation Score
FileLine CoverageMutation ScoreGap
validate.ts95%72%23%
login.ts88%85%3%
session.ts100%91%9%

High coverage with low mutation score indicates weak assertions


## Integration with CI

### GitHub Actions Integration

```yaml
- name: Run mutation testing
  run: npx stryker run --reporters json

- name: Check mutation threshold
  run: |
    SCORE=$(jq '.metrics.mutationScore' reports/mutation/stryker-incremental.json)
    if (( $(echo "$SCORE < 80" | bc -l) )); then
      echo "::error::Mutation score $SCORE% below 80% threshold"
      exit 1
    fi

Optimization Tips

Incremental Mutation Testing

Only test changed code:

bash
# Stryker incremental
npx stryker run --incremental

# PITest history
mvn pitest:mutationCoverage -DwithHistory
Target Critical Modules First
json
{
  "mutate": [
    "src/auth/**/*.ts",
    "src/payment/**/*.ts",
    "src/validation/**/*.ts"
  ]
}
  • tdd-enforce - Enforce test-first development
  • flaky-detect - Identify unreliable tests
  • test-sync - Maintain test-code alignment

Script Reference

mutation_runner.py

Run mutation testing for project:

bash
python scripts/mutation_runner.py --module src/auth
mutation_analyzer.py

Analyze and prioritize survivors:

bash
python scripts/mutation_analyzer.py --report stryker-report.json
native_extension_preflight.py

Run the isolated Python import/test preflight and emit machine-readable evidence:

bash
python scripts/native_extension_preflight.py \
  --test-selection tests/test_critical_contract.py \
  --mutation-target src/package/critical_contract.py \
  --estimated-mutants 120 --max-children 4 \
  --runtime-budget-seconds 900 --format json

References

  • @$AIWG_ROOT/agentic/code/addons/testing-quality/README.md — Testing quality addon overview
  • @$AIWG_ROOT/agentic/code/frameworks/sdlc-complete/README.md — SDLC framework context for quality gates
  • @$AIWG_ROOT/agentic/code/addons/aiwg-utils/rules/vague-discretion.md — Measurable quality thresholds and gate criteria
  • @$AIWG_ROOT/docs/cli-reference.md — CLI reference

© jmagly, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in agentic/code/addons/testing-quality/skills/mutation-test of jmagly/aiwg.

  • SKILL.md
  • scripts/aiwg_mutation_native_probe.py
  • scripts/native_extension_preflight.py

Open the folder on GitHubat commit dda238f

Compare with similar skills

Mutation Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mutation Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mutation Test this skilljmagly/aiwg220—~3.2kAutomated safety check: PassMIT
E2E Test ThinkerUniClipboard/UniClipboard1.8k—~1.7kAutomated safety check: PassAGPL-3.0
Ralph Coveragejvm-skills/jvm-skills140—~683Automated safety check: PassApache-2.0
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Mutation Testingproffesor-for-testing/agentic-qe494—~1.7kAutomated safety check: PassMIT
Write Testsgnomeria/usbtree688—~622Automated safety check: PassMIT

Similar skills

  • E2E Test Thinker

    UniClipboard/UniClipboard

    Analyze the current branch's diff against main and determine which changes are testable via CLI-based end-to-end tests.

    1.8k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Ralph Coverage

    jvm-skills/jvm-skills

    Run Ralph in coverage mode — iteratively write tests for untested classes until coverage targets are met.

    140 GitHub stars~683 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Mutation Testing

    proffesor-for-testing/agentic-qe

    Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

    494 GitHub stars~1.7k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Write Tests

    gnomeria/usbtree

    Author tests that match the repo's stack and existing test style, at the cheapest level that catches the regression.

    688 GitHub stars~622 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Write Frontend Tests

    Elite588/AUTOGPT

    Analyze the current branch diff against dev, plan integration tests for changed frontend pages/components, and write them.

    103 GitHub stars~1.9k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed

More from jmagly/aiwg

All 12 skills in this repo
  • Review editorial phrase patterns and suggest contextual alternatives; legacy name does not imply authorship detection.

    220 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Voice Apply

    jmagly/aiwg

    Apply a voice profile to transform content. An agent skill from jmagly/aiwg.

    220 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • TDD Enforce

    jmagly/aiwg

    Configure TDD enforcement via pre-commit hooks and CI coverage gates.

    220 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • LLMs Txt Support

    jmagly/aiwg

    Detect and use llms.txt files for LLM-optimized documentation.

    220 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Detect project type, AIWG framework state, team configuration, and active work to summarize status and recommend next actions

    220 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Manage artifact metadata, versioning, ownership, and review history across the SDLC lifecycle

    220 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Mutation Test

What does Mutation Test do?

Run mutation testing to validate test quality beyond code coverage. Mutation Test is an agent skill from jmagly/aiwg. Run mutation testing to validate test quality beyond code coverage.

When should I use Mutation Test?

Mutation Test fits situations like: assessing test effectiveness; finding weak tests; validating test suite quality.

How do I install Mutation Test in Claude Code?

Run `npx skills add jmagly/aiwg --skill mutation-test -a claude-code`. Or copy the skill folder (agentic/code/addons/testing-quality/skills/mutation-test in jmagly/aiwg) into .claude/skills/mutation-test in your project. Claude Code loads it when a task matches its description.

How do I install Mutation Test in Codex?

Run `npx skills add jmagly/aiwg --skill mutation-test -a codex`. Or copy the skill folder (agentic/code/addons/testing-quality/skills/mutation-test in jmagly/aiwg) into .agents/skills/mutation-test in your project. Codex loads it when a task matches its description.

Can I use Mutation Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jmagly/aiwg --skill mutation-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mutation-test, .gemini/skills/mutation-test, .github/skills/mutation-test and .opencode/skills/mutation-test in your project.

What does Mutation Test need to run?

Going by SKILL.md and its folder, Mutation Test needs Python for the scripts in its folder and the command-line tools its instructions call (python, npx, mvn and jq). Our summary lists: Python 3; Node.js.

Does Mutation Test access the network?

SKILL.md names 4 domains. As links in the text: github.com, conf.researchr.org, stryker-mutator.io and pitest.org. This is read from the text; nothing was executed.

Is Mutation Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Mutation Test use?

Mutation Test is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mutation Test use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Mutation Test?

Skills that share tags, products or a category with Mutation Test: E2E Test Thinker (UniClipboard/UniClipboard, 1.8k stars), Ralph Coverage (jvm-skills/jvm-skills, 140 stars), Designing Tests (CloudAI-X/opencode-workflow, 275 stars) and Mutation Testing (proffesor-for-testing/agentic-qe, 494 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mutation Test?

jmagly (a GitHub user) maintains it in jmagly/aiwg, which has 220 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 5, 2026.

Source: jmagly/aiwg on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.