Agent skill

Mutation Testing

by jvm-skills in jvm-skills/jvm-skills

Bootstrap pitest via the info.solidsoft.pitest Gradle plugin with Kotlin-sane defaults, run mutation tests scoped to changed classes, interpret surviving mutants from mutations.xml, triage…

Apache-2.0Auto-check passedTesting & QA

Install Mutation Testing

skills CLI
$ npx skills add jvm-skills/jvm-skills --skill mutation-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jvm-skills/jvm-skills mutation-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jvm-skills/jvm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.junie/skills/mutation-testing .claude/skills/mutation-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mutation-testing
GitHub stars
140
Token cost
~6.7k tokens
SKILL.md length
2,537 words
Files
2 (incl. scripts)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Bootstrap pitest via the info.solidsoft.pitest Gradle plugin with Kotlin-sane defaults, run mutation tests scoped to changed classes, interpret surviving mutants from mutations.xml, triage…

  • Works in 5 steps: Bootstrap → Scoped run (on a change) → Interpret survivors → …
  • The user asks to add mutation testing
  • SKILL.md covers When to use this skill, Prerequisites, Preflight — run before any… and Capabilities, plus 4 more sections
  • Runs Python scripts from its folder; calls git and python3

What it does

Mutation Testing is an agent skill from jvm-skills/jvm-skills. Bootstrap pitest via the info.solidsoft.pitest Gradle plugin with Kotlin-sane defaults, run mutation tests scoped to changed classes, interpret surviving mutants from mutations.xml, triage likely-equivalent mutants out of the kill queue, and drive a kill-survivor workflow. Use when the user asks to add mutation testing, check mutation score, investigate surviving mutants, or strengthen test quality beyond line coverage.

Its SKILL.md is about 6.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/triage.py`).

It sits in Testing & QA, covering Test coverage and Android development. It works with Gradle, Kotlin and JUnit. The licence is Apache-2.0.

When your agent uses it

  • The user asks to add mutation testing
  • Check mutation score
  • Investigate surviving mutants
  • Strengthen test quality beyond line coverage

Example prompts

  • “/mutation-testing”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Bootstrap
  2. Scoped run (on a change)
  3. Interpret survivors
  4. Triage — classify survivors before trying to kill them
  5. Kill a survivor

What it can do on your machine

Read from SKILL.md and the folder at commit c6d477f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mutation Testing loads about 6.7k tokens when it runs. Until then it costs about 110 tokens; SKILL.md has 2,537 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jvm-skills/jvm-skills at commit c6d477f, republished under its Apache-2.0 licence (© jvm-skills). 2,537 words, ~6,684 tokens.

Download SKILL.mdSave it as .claude/skills/mutation-testing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
mutation-testing
description
Bootstrap pitest via the info.solidsoft.pitest Gradle plugin with Kotlin-sane defaults, run mutation tests scoped to changed classes, interpret surviving mutants from mutations.xml, triage likely-equivalent mutants out of the kill queue, and drive a kill-survivor workflow. Use when the user asks to add mutation testing, check mutation score, investigate surviving mutants, or strengthen test quality beyond line coverage.

Mutation Testing (pitest)

Mutation testing for Kotlin/Java Gradle projects using pitest via the info.solidsoft.pitest plugin. Non-intrusive: mutates bytecode, requires no production-code changes.

When to use this skill

  • User asks: "add mutation testing", "check mutation score", "why does this mutant survive", "strengthen tests", "is this class well-tested".
  • A survivor appeared in a pitest report and needs killing.
  • A project wants a CI gate on mutation score.

Prerequisites

  • Gradle project (Kotlin DSL build.gradle.kts primary, Groovy build.gradle supported via analogous blocks).
  • Tests already exist. If not, run /ralph-coverage or /kotest-create first — mutation testing on an empty suite only produces NO_COVERAGE noise.
  • Clean or near-clean working tree.

Preflight — run before any pitest invocation

Pitest fails cryptically (silent UNKNOWN_ERROR from the coverage minion) when the project itself doesn't compile or when the test runtime has misaligned JUnit jars. Always run these checks before ./gradlew pitest so failures surface with their real cause, not pitest's.

bash
# 1. Test sources must compile — pitest can't help if the project itself is broken.
./gradlew compileTestKotlin compileTestJava --quiet

# 2. Find the project's actual junit-platform-launcher version.
./gradlew dependencyInsight --configuration testRuntimeClasspath \
  --dependency org.junit.platform:junit-platform-launcher 2>/dev/null \
  | head -1

# 3. Check that Kotlin bytecode keeps line-number debug info (pitest needs it).
grep -RnE 'Xno-source-debug-extension|-g:none' build.gradle.kts settings.gradle.kts 2>/dev/null

If step 1 fails, abort and surface the compile error — do not attempt pitest. If step 2 reports a version, use it to pick junit5PluginVersion (see next section). If step 3 finds either flag set, pitest will produce bogus line numbers or silently fail on mutators that need source info — remove the flag or exclude those modules.

Capabilities

1. Bootstrap

Detect whether pitest is already configured:

bash
grep -n "info.solidsoft.pitest" build.gradle.kts settings.gradle.kts **/build.gradle.kts 2>/dev/null

If not configured, propose this diff to the root build.gradle.kts:

kotlin
plugins {
    id("info.solidsoft.pitest") version "1.19.0"
}

pitest {
    pitestVersion.set("1.19.0")
    targetClasses.set(setOf("<group.package>.*"))          // derive from rootProject.group
    targetTests.set(setOf("<group.package>.*"))
    mutators.set(setOf("STRONGER"))                        // do NOT use ALL — NPE- and equivalent-mutation-prone
    features.set(listOf(
        "+FLOGCALL",                                       // built-in: silence many logger-call mutations
        // `+fkotlin` is auto-enabled by the junit5 plugin when detected — filters Kotlin bytecode-junk mutations
    ))
    avoidCallsTo.set(setOf(
        // FLOGCALL and avoidCallsTo overlap partially but neither fully covers the other on real Kotlin+SLF4J code.
        // Verified empirically: removing these re-introduces 7+ logger-call survivors per scoped run.
        "kotlin.jvm.internal", "kotlin.Metadata",
        "org.slf4j.Logger",
        "org.apache.logging.log4j.Logger",
        "java.util.logging.Logger",
    ))
    excludedMethods.set(setOf(
        // Kotlin data-class synthesized members — generated, behavior guaranteed
        "component*", "copy", "hashCode", "equals", "toString",
    ))
    junit5PluginVersion.set("1.2.1")                       // JUnit 5 projects only — pin to project's Platform version
    outputFormats.set(setOf("HTML", "XML"))
    threads.set(Runtime.getRuntime().availableProcessors())
    timestampedReports.set(false)                          // stable path for parsing
    fullMutationMatrix.set(true)                           // names covering tests for SURVIVED mutants too — see §3
    exportLineCoverage.set(true)                           // emits coverage.xml alongside mutations.xml
    historyInputLocation.set(file("build/pitest/history.bin"))     // incremental analysis — read + write same file to skip re-running killed/survived mutants on unchanged classes
    historyOutputLocation.set(file("build/pitest/history.bin"))    // the gradle-pitest-plugin exposes explicit paths, not the `withHistory` CLI flag
    jvmArgs.set(listOf("-Xmx2g"))                          // avoid MEMORY_ERROR noise on forked minion; raise to 4g for large classpaths
    mutationThreshold.set(60)
    coverageThreshold.set(60)
    testStrengthThreshold.set(60)
}

Key points when proposing:

  • Derive targetClasses from rootProject.group (e.g. com.example.foo → "com.example.foo.*"), show the detected value in the diff and ask for confirmation.
  • Omit junit5PluginVersion for JUnit 4–only projects. Detect by searching for junit-jupiter in dependencies.
  • Pin junit5PluginVersion to the project's JUnit Platform line. Pitest-junit5-plugin bundles its own junit-platform-launcher. If it targets a different Platform than the one in testRuntimeClasspath, the minion dies with OutputDirectoryCreator not available (a silent UNKNOWN_ERROR at the plugin level). Rough matrix (verify against the plugin's release notes before pinning):
    Project junit-platform-launcherSafe junit5PluginVersion
    1.8.x1.0.0
    1.9.x1.1.0
    1.10.x1.2.0
    1.11.x1.2.1
    1.12.x1.2.2
    Newer Platform / Jupiter (6.x / Platform 2.x) may require the pitest configuration to also pin a matching launcher explicitly:
    kotlin
    dependencies { pitest("org.junit.platform:junit-platform-launcher:<version>") }
  • Respect existing config — if a pitest { } block already exists, only patch missing fields (e.g. add XML to outputFormats if only HTML is set). Never overwrite thresholds or targetClasses.
  • Groovy DSL — emit the equivalent pitest { ... } using = assignment instead of .set().
  • Forward tasks.test settings. Pitest's minion does not inherit jvmArgs, systemProperty(...), systemProperties(...), minHeapSize, or maxHeapSize from tasks.test. Inspect the test task and mirror whatever it sets into pitest { jvmArgs = [...] } — system properties become -D entries, heap becomes -Xmx/-Xms. Missing this often manifests as the silent UNKNOWN_ERROR (e.g. JUnit can't instantiate a custom @TestClassOrder because the test task sets it via system property and the minion doesn't see it).
    kotlin
    // If tasks.test has:
    //   maxHeapSize = "4096m"
    //   systemProperty("junit.jupiter.testclass.order.default", "com.example.MyOrderer")
    //
    // Then pitest needs:
    pitest {
        jvmArgs.set(listOf(
            "-Xmx4g",
            "-Dkotlin.jupiter.testclass.order.default=com.example.MyOrderer",
        ))
    }
  • Turn on verbose.set(true) during bootstrap. Silent minion crashes produce actionable output only when verbose is on. Can be switched off after the first clean run.

After applying the diff, run ./gradlew pitest once to generate the baseline. First-run failure modes and what they usually mean:

SymptomLikely cause
OutputDirectoryCreator not available ... unaligned versions of the junit-platform-engine and junit-platform-launcherjunit5PluginVersion doesn't match the project's junit-platform-launcher — see matrix above
Silent UNKNOWN_ERROR with no stackEnable verbose.set(true) and rerun — real cause will then surface in PIT >> SEVERE : MINION lines
NoClassDefFoundError / custom orderer / listener failstasks.test system property or jvmArg not forwarded to pitest
Unresolved reference in test sourcesPre-existing compile error — run the preflight compileTestKotlin check first
2. Scoped run (on a change)

Derive target classes from git diff and pass them to the wired pitestScope property (see §5 for the wiring). The info.solidsoft.pitest plugin does NOT honor -Ppitest.targetClasses natively — a raw CLI flag with that name will silently be ignored and pitest will run against whatever the pitest { } block has hard-coded.

bash
BASE=$(git merge-base HEAD origin/main 2>/dev/null || git merge-base HEAD main)
CHANGED=$(git diff --name-only "$BASE" HEAD -- '*.kt' '*.java' \
  | grep '^src/main/' \
  | sed -E 's#^src/main/(kotlin|java)/##; s#\.(kt|java)$##; s#/#.#g' \
  | paste -sd, -)

if [ -z "$CHANGED" ]; then
  echo "No production classes changed — aborting scoped run."
  exit 1
fi

./gradlew pitest -PpitestScope="$CHANGED"

Aborting on empty match is deliberate — never silently fall back to whole-codebase on a scoped run.

For Groovy-DSL projects, the wiring and flag are identical (same findProperty API).

For multi-module, prefix with the module path: ./gradlew :module-name:pitest -PpitestScope=.... Wire the property in each module's pitest { } block.

3. Interpret survivors

Pitest writes build/reports/pitest/mutations.xml (stable path because timestampedReports.set(false)). Parse it:

bash
# Quick survivor extraction
xmllint --xpath '//mutation[@status="SURVIVED"]' build/reports/pitest/mutations.xml 2>/dev/null

Each <mutation> element carries:

  • status — SURVIVED / KILLED / NO_COVERAGE / TIMED_OUT / MEMORY_ERROR / RUN_ERROR / NON_VIABLE
  • <sourceFile>, <mutatedClass>, <mutatedMethod>, <lineNumber>, <mutator>, <description>
  • <killingTest> — present and populated only for KILLED mutants
  • numberOfTestsRun (attribute) — how many tests reached this line

Important: with the default pitest configuration, SURVIVED mutants do not list the names of the tests that reached them — only a count via numberOfTestsRun. The killing-test field in the XML is empty for anything that survived. To identify which test to strengthen, the bootstrap sets fullMutationMatrix.set(true), which adds a <killingTests> / <succeedingTests> matrix to each <mutation> element. Without this flag, the skill can only say "2 tests ran against this line and missed" — it cannot name them.

Side-effect: fullMutationMatrix = true increases report size and slightly increases runtime. Acceptable on scoped runs (one package / one class); consider turning off for whole-codebase CI runs where survivor diagnosis isn't the goal.

exportLineCoverage.set(true) emits a separate coverage.xml listing which tests reach each line. Combined with mutations.xml, you can reconstruct the covering-test set even without fullMutationMatrix, but it requires joining two files. Enabled by default in the bootstrap because it is cheap and useful for the aggregator.

Print a compact table:

STATUS       FILE:LINE              MUTATOR                   DESCRIPTION
SURVIVED     Calculator.kt:42       MATH                      replaced + with -
             covered by: CalculatorTest.`positive number is positive`
SURVIVED     Calculator.kt:43       CONDITIONALS_BOUNDARY     changed > to >=
             covered by: CalculatorTest.`zero is not positive`
NO_COVERAGE  PriceCalc.kt:88        VOID_METHOD_CALLS         removed call to log.debug

Distinguish SURVIVED (test exists but missed the behavior) from NO_COVERAGE (no test reaches the line) and TIMED_OUT (loop mutant caught by pitest's timeout heuristic — not a real survivor).

4. Triage — classify survivors before trying to kill them

This step runs between interpretation and kill-a-survivor. Its job: separate real test-quality gaps from equivalent mutants — mutations that change the bytecode but not the observable behavior in any test a human would reasonably write. Without triage, a kill-a-survivor loop wastes iterations (and, under unattended automation, produces tautological tests) trying to kill things that shouldn't be killed.

Classify each SURVIVED mutant into one of three buckets:

BucketMeaningFed to kill-survivor / autoresearch?
LIKELY_KILLABLEBehavioral change a sensible test could catchYes
LIKELY_EQUIVALENTNo realistic test can distinguish mutant from originalNo. Written to suspected-equivalent.md for the record
AMBIGUOUSCan't tell without reading more contextSurface for human decision; do not auto-feed to autoresearch

Read the source line at file:line for each survivor and match against these archetypes (expand via the project overlay):

Archetype A — logger-gate conditionals. The line is an if (<numeric> <comparator> <constant>) { ... } whose body contains only calls to loggers (logger.info, .warn, .error, .debug, .trace) or trace/metric sinks. Any ConditionalsBoundary / RemoveConditional_ORDER_* / RemoveConditional_EQUAL_* survivor on such a line is LIKELY_EQUIVALENT.

kotlin
if (deleted > 0) {          // <— mutations on `> 0` all survive
    logger.info("...")      //   because tests don't capture log output
}

Archetype B — unreachable loop bounds. The line is inside a while / do..while / for condition whose opposing operand is a constant larger than any realistic test input (threshold: default 1000; configurable via overlay). ConditionalsBoundary / RemoveConditional_ORDER_* survivors on such lines are LIKELY_EQUIVALENT — the mutation is theoretically killable but only with infeasible data volumes.

kotlin
} while (batch.isNotEmpty() && batchCount < maxBatches)  // maxBatches = 10, batch = 10000 rows

Archetype C — null-elvis on non-nullable-in-practice values. The line is a <value> ?: throw <X> pattern where <value> is the non-nullable-by-contract result of a DB primary-key fetch, an Optional.get(), or similar. RemoveConditional_EQUAL_IF / RemoveConditional_EQUAL_ELSE survivors are LIKELY_EQUIVALENT — null is unreachable in any test a DB schema permits.

Kotlin-specific note: pitest's NULL_RETURNS mutator already skips @NotNull-annotated methods, and Kotlin compiles non-nullable return types to @NotNull in bytecode. So null-return mutants only surface on nullable Kotlin returns (Foo?) — archetype C is tuned to that reality and does not over-trigger on plain non-nullable returns.

kotlin
eventCleanup.hardDeleteEvent(eventId ?: throw IllegalStateException("eventId is null"))
//                           ^-- eventId is jOOQ-fetched primary key, never null in practice

Archetype D — branch-identity (partial-kill residual). An if (cond) { X } else { Y } where X and Y produce the same observable output when cond is false. Classic example: if (list.isNotEmpty()) { repo.fetch(list).associateBy { it.id } } else { emptyMap() } — when list is empty, fetch([]) returns empty and associateBy on empty yields emptyMap, so the IF-branch is indistinguishable from the ELSE-branch for the "empty" case. RemoveConditional_EQUAL_IF mutations on this line survive any test because flipping the branch doesn't change the observable output.

Detecting this archetype statically is hard — it requires knowing that the IF-branch operation is idempotent over the identity input of the ELSE-branch. The script can't match it reliably from source alone. This archetype is surfaced dynamically: when a single strengthened test kills the EQUAL_ELSE mutation on a line but leaves the EQUAL_IF mutation (or vice versa) surviving, the surviving mutation is a strong candidate for LIKELY_EQUIVALENT. Promote it after 2 failed targeted kill attempts (not 3, since the signal is stronger than blind failure) and skip it in future iterations.

Anything that matches no archetype is AMBIGUOUS by default — bias toward ambiguity, not toward equivalence, so real gaps don't get silently excluded.

Implementation. A ready-to-use triage script lives at scripts/triage.py in this skill directory. Run:

bash
python3 .claude/skills/mutation-testing/scripts/triage.py \
  build/reports/pitest/mutations.xml \
  .

It reads the mutations XML, opens each survivor's source line, applies archetypes A/B/C, and writes triage.md + triage.json next to the mutations file.

Output: triage.md (human) and triage.json (machine) next to mutations.xml. Each survivor carries the source pointer, archetype reason, and the names of the tests that covered but didn't kill it (extracted from <succeedingTests> / <coveringTests>). Covering-test names require fullMutationMatrix=true in the bootstrap block; otherwise the script prints a hint to enable it.

Shape:

markdown
# Triage of <N> survivors

## LIKELY_KILLABLE (<count>)

- `UserCleanupService.kt:58` **VoidMethodCall** — removed call to deleteS3Files
  - in `UserCleanupService.hardDeleteUser()`
  - reason: no archetype matched
  - covering tests:
    - `UserCleanupServiceTest.hardDeleteUser removes user and all owned data`

## LIKELY_EQUIVALENT (<count>)

- `AnonymousUserCleanupJob.kt:70` **ConditionalsBoundary** — changed > to >=
  - reason: archetype A — logger-gate conditional
- `AnonymousUserCleanupJob.kt:68` **ConditionalsBoundary** — changed < to <=
  - reason: archetype B — loop-bound comparison

The suspected-equivalent.md set is durable — it accumulates across sessions, not per-run. A mutant that's triaged as equivalent once stays in that list unless the source line changes (detected by re-running triage when the line's content differs).

Project overlay extensions. references/project.md can add project-specific archetypes:

  • Custom logger types beyond the stdlib set
  • Known-non-null value types beyond DB primary keys (e.g. "values returned from MyContext.requireUser()")
  • Loop-bound constants specific to the project
  • Explicit allowlist of mutants always classified as equivalent (by file:line:mutator key)

Do not skip triage before autoresearch. The mutation-autoresearch skill requires a current triage output as its input — it only iterates LIKELY_KILLABLE. Running autoresearch on the raw survivor list would burn iterations on equivalents and produce tautological tests under pressure.

Show full SKILL.md (940 more words)Show less
5. Kill a survivor

Inverted TDD cycle. Unlike a normal /tdd-task RED → GREEN cycle, a mutation-kill test is written against production code that is already correct. The test should pass on the first run. The RED → GREEN proof comes from the scoped pitest rerun (see verify step below): the mutant flips from SURVIVED to KILLED. A /tdd-task agent that tries to force a red-first state here will waste iterations — pass this expectation in the prompt.

For each survivor, delegate to /tdd-task with a prompt like:

Kill mutation: <mutator> at <file>:<line> in <mutatedClass>.<mutatedMethod>. Description: <description>. Currently covered by: <coveringTests> (from triage.md). Spell out both behaviors: what the original code returns and what the mutated code would return for the chosen input. The assertion must distinguish the two. Do not assert the raw mutated value or operator — assert the observable behavior on a specific input. Prefer strengthening assertions in the covering test; if that isn't natural, add a new sibling test case. The test is expected to pass against current production code — don't force a red-first state. RED proof comes from the pitest rerun below.

After /tdd-task returns, verify with a scoped pitest rerun. Override scope via Gradle -P flags — do not edit build.gradle.kts for per-iteration scope changes. The info.solidsoft.pitest plugin does NOT read -Ppitest.targetClasses natively; you must wire a custom property into the pitest { } block. Pick a name like pitestScope:

kotlin
pitest {
    pitestVersion.set("1.19.0")
    val scopeProp: String? = project.findProperty("pitestScope") as String?
    val testsProp: String? = project.findProperty("pitestTests") as String?
    val defaultScope = setOf("com.example.*")
    val scopeFqns = scopeProp?.split(",")?.map { it.trim() }?.filter { it.isNotEmpty() }
    targetClasses.set(
        // Auto-expand each FQN to `{FQN, FQN$*}` so Kotlin-emitted nested
        // and synthetic classes (`Foo$Page`, `Foo$methodName$1`) are included.
        scopeFqns?.flatMap { listOf(it, "$it\$*") }?.toSet() ?: defaultScope
    )
    targetTests.set(
        testsProp?.split(",")?.map { it.trim() }?.filter { it.isNotEmpty() }?.toSet()
            ?: scopeFqns?.toSet()
            ?: defaultScope
    )
    // ... rest of the block stays stable
}

Then the per-iteration rerun is:

bash
./gradlew pitest -PpitestScope='com.example.pkg.MediaSubmission'

Scoping rules:

  • Include nested/synthetic classes. Kotlin emits synthetic bytecode for nested data classes (Foo$Page) and lambdas (Foo$methodName$1, .let { }, .map { }). The bare FQN alone doesn't match them — without $* you will miss mutations inside inner classes and lambda bodies.
  • Don't use the prefix glob FQN* (without $). It also matches sibling top-level classes (MediaSubmissionService) and inflates scope.
  • Keep the pitest { } block pinned to a stable, broad default (the module-level package you CI against). Treat per-iteration scopes as transient CLI overrides that never land in git.
bash
./gradlew pitest

Parse build/reports/pitest/mutations.xml directly — do not rely on the Gradle exit code. Pitest exits non-zero when line coverage falls below coverageThreshold (typical on narrow single-class scopes), but the XML is still valid. The skill only cares about three facts:

  1. The target mutation's status is now KILLED.
  2. No previously-KILLED mutation in the same class flipped to SURVIVED.
  3. The class's killed_count strictly increased.

If all three hold, delegate to /commit. If the new test asserts a tautology (literally references the mutated operator/constant, or only checks a value trivially equal to the mutation's replacement), reject and retry with a stronger prompt. If the rerun shows the mutation still alive, read the committed assertions and ask: did the test actually exercise an input where mutated vs original diverge? Tightening the input data usually kills it.

A quick awk / xmllint / Python one-liner on mutations.xml is enough — pitest 1.19 writes attributes with single quotes (status='SURVIVED'), so grep patterns using double quotes will silently match nothing.

Fallbacks when dependencies aren't installed
  • No /tdd-task: follow the inline inverted-TDD checklist — (1) identify the behavioral input where original and mutated code diverge, (2) add a test asserting the correct behavior on that input (it passes immediately against current code), (3) run ./gradlew test --tests <pattern> and confirm green, (4) run scoped pitest and confirm the mutant moved SURVIVED → KILLED. The initial RED step from classic TDD does not apply — production code is correct; RED proof comes from the pitest rerun.
  • No /commit: git add src/test && git commit -m "test(mutation): kill <mutator> at <file>:<line>".
  • No /test-gradle: ./gradlew test --tests <fully-qualified-test-pattern>.

Kotlin-specific gotchas

GotchaMitigation
kotlin.jvm.internal null-check noiseavoidCallsTo.set(setOf("kotlin.jvm.internal", "kotlin.Metadata")) (already in bootstrap)
Logger call survivors (org.slf4j.Logger::info/debug/warn)Already in bootstrap via avoidCallsTo. These are noise unless logging itself is the feature under test
Data class component* / copy / equals / hashCode / toString mutantsAlready in bootstrap via excludedMethods. Generated methods, semantics guaranteed by Kotlin
Inline classes / value classesAdd to excludedClasses; test against the underlying type
Extension functions (compile to FileNameKt)Ensure targetClasses glob covers the file-level class (e.g. com.example.UtilsKt)
Coroutines state-machine bytecodeNoisy generated-bridge mutants. If survivors are all synthetic, narrow mutators to DEFAULTS for coroutine-heavy modules
Generated code (ksp, jOOQ, Dagger)Exclude via excludedClasses.set(setOf("com.example.generated.*"))
Mixed JUnit 4 + JUnit 5Set junit5PluginVersion and verify junit-vintage-engine is on the test classpath
IntelliJ Gradle DSL warningsPrefer .set(...) over direct = assignment to silence lazy-property warnings
History cache poisoning (incremental analysis)withHistory = true reads java.io.tmpdir; bust it after dependency bumps, kapt/ksp config changes, Spring AOT changes, and across branches. In CI, key history per-branch (e.g. archive build/pitest/history.bin with the branch name) — dependency changes are invisible to pitest's dep graph, so a stale cache silently reports wrong results.
Pitest mutator groupsDefault is STRONGER. Do not use ALL — includes CONSTRUCTOR_CALLS, NON_VOID_METHOD_CALLS, UOI, ROR, AOR/AOD, CRCR; docs flag them as "fairly unstable" and equivalent-prone, producing false survivor noise
+FLOGCALL featureBuilt into pitest via the feature language. Already in the bootstrap block. If the project uses a custom logger type not auto-detected by FLOGCALL, supplement via avoidCallsTo

Status-symbol legend for summaries

SymbolStatus
✓KILLED
✗SURVIVED
○NO_COVERAGE
⏱TIMED_OUT
⚠MEMORY_ERROR / RUN_ERROR / NON_VIABLE

What this skill does NOT do (v0.1.0)

  • Does not modify production code. Survivors that indicate a real bug are surfaced for manual decision; route through /fix explicitly.
  • Does not run unattended overnight — that's the mutation-autoresearch skill.
  • Does not auto-ratchet thresholds. Ratcheting is a separate capability; this skill only advises on values.
  • Does not patch CI. CI wiring is documented but not automated in v0.1.0.
  • Does not handle Maven. Pitest has a Maven plugin; this skill is Gradle-only.

Project overlay

Read references/project.md in this skill's directory if it exists. The overlay provides:

  • Project-specific targetClasses / excludedClasses patterns
  • Base test classes and when to use each
  • Known-equivalent mutants to skip
  • Preferred commit-message style

© jvm-skills, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .junie/skills/mutation-testing of jvm-skills/jvm-skills.

  • SKILL.md
  • scripts/triage.py

Open the folder on GitHubat commit c6d477f

Compare with similar skills

Mutation Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mutation Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mutation Testing this skilljvm-skills/jvm-skills140—~6.7kAutomated safety check: PassApache-2.0
Java TestingHoangNguyen0403/agent-skills-standard571—~1kAutomated safety check: PassMIT
Jvm Helpershepherdjerred/monorepo112—~1.9kAutomated safety check: PassGPL-3.0
Native Testing Strategybladeofgod/flutter-ai-harness116—~345Automated safety check: PassMIT
Quality Test Pilotnekomangaorg/Neko2.8k—~1.1kAutomated safety check: PassApache-2.0
Android Audio E2Ehyochan/react-native-nitro-sound961—~814Automated safety check: PassMIT

Similar skills

  • Java Testing

    HoangNguyen0403/agent-skills-standard

    Testing standards using JUnit 5, AssertJ, Mockito, Cucumber, and Spring Boot integration tests for Java.

    571 GitHub stars~1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Jvm Helper

    shepherdjerred/monorepo

    Current Java, Kotlin, Gradle, Maven, JUnit, JVM diagnostics, packaging, and performance guidance.

    112 GitHub stars~1.9k tokensUpdated today
    MobileAuto-check passed
  • Native Testing Strategy

    bladeofgod/flutter-ai-harness

    适用:设计或审查 Kotlin/Swift 原生模块、Bridge Adapter、Host 编译、模拟器/设备和系统能力验证。不适用:纯 Dart/Flutter 测试或用 Fake 代替相机和权限真机验证。触发词:JUnit、XCTest、Robolectric、instrumented test、Swift Testing、Framework Fake、Gradle…

    116 GitHub stars~345 tokensUpdated 1 mo ago
    MobileAuto-check passed
  • Quality Test Pilot

    nekomangaorg/Neko

    Writes high-quality unit tests for Kotlin Android codebases to increase meaningful test coverage.

    2.8k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Android Audio E2E

    hyochan/react-native-nitro-sound

    Build and run device-backed Android end-to-end checks for react-native-nitro-sound permissions, recording, playback, listeners, pause/resume, seek, speed, audio focus, lifecycle behavior, and rapid…

    961 GitHub stars~814 tokensUpdated 8 days ago
    MobileAuto-check passed
  • Improve Code Coverage

    alexvanyo/composelife

    Helps increment code coverage in this Kotlin Multiplatform project.

    267 GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed

More from jvm-skills/jvm-skills

All 16 skills in this repo
  • Ralph Coverage

    jvm-skills/jvm-skills

    Run Ralph in coverage mode — iteratively write tests for untested classes until coverage targets are met.

    140 GitHub stars~683 tokensUpdated 1 mo ago
    Auto-check passed
  • Skill Scout

    jvm-skills/jvm-skills

    Run the skill-scout loop — scan the next batch of unscanned JVM-conference rosters for speaker-created AI skills and apply results to the CSV store via the overnight Workflow.

    140 GitHub stars~727 tokensUpdated 1 mo ago
    Auto-check passed
  • Spec

    jvm-skills/jvm-skills

    Generate a feature spec with user stories directly from conversation context and codebase exploration — no interview needed.

    140 GitHub stars~832 tokensUpdated 1 mo ago
    Auto-check passed
  • UI Review

    jvm-skills/jvm-skills

    Verify UI/UX by running Playwright tests and reviewing screenshots.

    140 GitHub stars~545 tokensUpdated 1 mo ago
    Auto-check passed
  • Kotest

    jvm-skills/jvm-skills

    Write a new Kotlin test, or modernize an existing one, using Kotest matchers and idiomatic Kotlin test style — backticked names, apply/assertSoftly blocks, and existing object mothers and helpers.

    140 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Voice

    jvm-skills/jvm-skills

    Make prose sound like Thomas, not like a generic technical writer.

    140 GitHub stars~592 tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Mutation Testing

What does Mutation Testing do?

Bootstrap pitest via the info.solidsoft.pitest Gradle plugin with Kotlin-sane defaults, run mutation tests scoped to changed classes, interpret surviving mutants from mutations.xml, triage…. Mutation Testing is an agent skill from jvm-skills/jvm-skills.xml, triage likely-equivalent mutants out of the kill queue, and drive a kill-survivor workflow.

When should I use Mutation Testing?

Mutation Testing fits situations like: the user asks to add mutation testing; check mutation score; investigate surviving mutants; strengthen test quality beyond line coverage.

How do I install Mutation Testing in Claude Code?

Run `npx skills add jvm-skills/jvm-skills --skill mutation-testing -a claude-code`. Or copy the skill folder (.junie/skills/mutation-testing in jvm-skills/jvm-skills) into .claude/skills/mutation-testing in your project. Claude Code loads it when a task matches its description.

How do I install Mutation Testing in Codex?

Run `npx skills add jvm-skills/jvm-skills --skill mutation-testing -a codex`. Or copy the skill folder (.junie/skills/mutation-testing in jvm-skills/jvm-skills) into .agents/skills/mutation-testing in your project. Codex loads it when a task matches its description.

Can I use Mutation Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jvm-skills/jvm-skills --skill mutation-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mutation-testing, .gemini/skills/mutation-testing, .github/skills/mutation-testing and .opencode/skills/mutation-testing in your project.

What does Mutation Testing need to run?

Going by SKILL.md and its folder, Mutation Testing needs Python for the scripts in its folder and the command-line tools its instructions call (git and python3). Our summary lists: Python 3.

Does Mutation Testing access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Mutation Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Mutation Testing use?

Mutation Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mutation Testing use?

About 6.7k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Mutation Testing?

Skills that share tags, products or a category with Mutation Testing: Java Testing (HoangNguyen0403/agent-skills-standard, 571 stars), Jvm Helper (shepherdjerred/monorepo, 112 stars), Native Testing Strategy (bladeofgod/flutter-ai-harness, 116 stars) and Quality Test Pilot (nekomangaorg/Neko, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mutation Testing?

jvm-skills (a GitHub organization) maintains it in jvm-skills/jvm-skills, which has 140 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on August 31, 2026.

Source: jvm-skills/jvm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.