Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

MITAuto-check passedTesting & QA

Install Mutation Testing

skills CLI
$ npx skills add proffesor-for-testing/agentic-qe --skill mutation-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install proffesor-for-testing/agentic-qe mutation-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/proffesor-for-testing/agentic-qe.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assets/skills/mutation-testing .claude/skills/mutation-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mutation-testing
GitHub stars
494
Token cost
~1.7k tokens
SKILL.md length
355 words
Files
8 (incl. scripts, references)
Skills in repo
111
Repo updated
First seen
Licence
MIT

At a glance

Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

  • Works in 5 steps: MUTATE code (change + to -, >= to >,… → RUN tests against each mutant → VERIFY tests catch mutations (kill… → …
  • Evaluating test quality
  • SKILL.md covers Quick Reference Card, How Mutation Testing Works, Using Stryker and Fixing Surviving Mutants, plus 7 more sections
  • Calls npx, npm and node

What it does

Mutation Testing is an agent skill from proffesor-for-testing/agentic-qe. Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate. Use when evaluating test quality, identifying weak tests, or proving tests actually catch bugs.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `config.json`, `evals/mutation-testing.yaml` and `references/mutation-operators.md`).

It sits in Testing & QA, covering Test coverage and Test generation. The repository describes itself as: Agentic QE Fleet is an open-source AI-powered QA/QE platform designed for use with Coding Agents (works best with Claude Code) featuring specialized agents and skills to support… The licence is MIT.

When your agent uses it

  • Evaluating test quality
  • Identifying weak tests
  • Proving tests actually catch bugs

Example prompts

  • “/mutation-testing”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. MUTATE code (change + to -, >= to >, remove statements)
  2. RUN tests against each mutant
  3. VERIFY tests catch mutations (kill mutants)
  4. IDENTIFY surviving mutants (tests need improvement)
  5. STRENGTHEN tests to kill surviving mutants

What it can do on your machine

Read from SKILL.md and the folder at commit 829d030. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • npm
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mutation Testing loads about 1.7k tokens when it runs, and up to ~2.1k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 355 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from proffesor-for-testing/agentic-qe at commit 829d030, republished under its MIT licence (© proffesor-for-testing). 355 words, ~1,742 tokens.

Download SKILL.mdSave it as .claude/skills/mutation-testing/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
mutation-testing
description
Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate. Use when evaluating test quality, identifying weak tests, or proving tests actually catch bugs.
category
specialized-testing
priority
high
tokenEstimate
900
agents
qe-test-generator, qe-coverage-analyzer, qe-quality-analyzer, qe-mutation-tester
implementation_status
optimized
optimization_version
1
last_optimized
2025-12-02
quick_reference_card
true
tags
mutation, stryker, test-quality, kill-rate, assertions, effectiveness
trust_tier
3

Mutation Testing

<default_to_action> When validating test quality or improving test effectiveness:

  1. MUTATE code (change + to -, >= to >, remove statements)
  2. RUN tests against each mutant
  3. VERIFY tests catch mutations (kill mutants)
  4. IDENTIFY surviving mutants (tests need improvement)
  5. STRENGTHEN tests to kill surviving mutants

Quick Mutation Metrics:

  • Mutation Score = Killed / (Killed + Survived)
  • Target: > 80% mutation score
  • Surviving mutants = weak tests

Critical Success Factors:

  • High coverage ≠ good tests (100% coverage, 0% assertions)
  • Mutation testing proves tests actually catch bugs
  • Focus on critical code paths first </default_to_action>

Quick Reference Card

When to Use
  • Evaluating test suite quality
  • Finding gaps in test assertions
  • Proving tests catch bugs
  • Before critical releases
Mutation Score Interpretation
ScoreInterpretation
90%+Excellent test quality
80-90%Good, minor improvements
60-80%Needs attention
< 60%Significant gaps
Common Mutation Operators
CategoryOriginalMutant
Arithmetica + ba - b
Relationalx >= 18x > 18
Logicala && ba || b
Conditionalif (x)if (true)
Statementreturn x(removed)

How Mutation Testing Works

javascript
// Original code
function isAdult(age) {
  return age >= 18; // ← Mutant: change >= to >
}

// Strong test (catches mutation)
test('18 is adult', () => {
  expect(isAdult(18)).toBe(true); // Kills mutant!
});

// Weak test (mutation survives)
test('19 is adult', () => {
  expect(isAdult(19)).toBe(true); // Doesn't catch >= vs >
});
// Surviving mutant → Test needs boundary value

Using Stryker

bash
# Install
npm install --save-dev @stryker-mutator/core @stryker-mutator/jest-runner

# Initialize
npx stryker init

Configuration:

json
{
  "packageManager": "npm",
  "reporters": ["html", "clear-text", "progress"],
  "testRunner": "jest",
  "coverageAnalysis": "perTest",
  "mutate": [
    "src/**/*.ts",
    "!src/**/*.spec.ts"
  ],
  "thresholds": {
    "high": 90,
    "low": 70,
    "break": 60
  }
}

Run:

bash
npx stryker run

Output:

Mutation Score: 87.3%
Killed: 124
Survived: 18
No Coverage: 3
Timeout: 1

Fixing Surviving Mutants

javascript
// Surviving mutant: >= changed to >
function calculateDiscount(quantity) {
  if (quantity >= 10) { // Mutant survives!
    return 0.1;
  }
  return 0;
}

// Original weak test
test('large order gets discount', () => {
  expect(calculateDiscount(15)).toBe(0.1); // Doesn't test boundary
});

// Fixed: Add boundary test
test('exactly 10 gets discount', () => {
  expect(calculateDiscount(10)).toBe(0.1); // Kills mutant!
});

test('9 does not get discount', () => {
  expect(calculateDiscount(9)).toBe(0); // Tests below boundary
});

Agent-Driven Mutation Testing

typescript
// Analyze mutation score and generate fixes
await Task("Mutation Analysis", {
  targetFile: 'src/payment.ts',
  generateMissingTests: true,
  minScore: 80
}, "qe-test-generator");

// Returns:
// {
//   mutationScore: 0.65,
//   survivedMutations: [
//     { line: 45, operator: '>=', mutant: '>', killedBy: null }
//   ],
//   generatedTests: [
//     'test for boundary at line 45'
//   ]
// }

// Coverage + mutation correlation
await Task("Coverage Quality Analysis", {
  coverageData: coverageReport,
  mutationData: mutationReport,
  identifyWeakCoverage: true
}, "qe-coverage-analyzer");

Agent Coordination Hints

Memory Namespace
aqe/mutation-testing/
├── mutation-results/*   - Stryker reports
├── surviving/*          - Surviving mutants
├── generated-tests/*    - Tests to kill mutants
└── trends/*             - Mutation score over time
Fleet Coordination
typescript
const mutationFleet = await FleetManager.coordinate({
  strategy: 'mutation-testing',
  agents: [
    'qe-test-generator',     // Generate tests for survivors
    'qe-coverage-analyzer',  // Coverage correlation
    'qe-quality-analyzer'    // Quality assessment
  ],
  topology: 'sequential'
});


Show full SKILL.md (160 more words)Show less

Remember

High code coverage ≠ good tests. 100% coverage but weak assertions = useless. Mutation testing proves tests actually catch bugs.

Focus on critical paths first. Don't mutation test everything - prioritize payment, authentication, data integrity code.

With Agents: Agents run mutation analysis, identify surviving mutants, and generate missing test cases to kill them. Automated improvement of test quality.

Run History

After each mutation test run, append results to run-history.json in this skill directory:

bash
node -e "
const fs = require('fs');
const h = JSON.parse(fs.readFileSync('.claude/skills/mutation-testing/run-history.json'));
h.runs.push({date: new Date().toISOString().split('T')[0], mutation_score_pct: SCORE, killed: KILLED, survived: SURVIVED});
fs.writeFileSync('.claude/skills/mutation-testing/run-history.json', JSON.stringify(h, null, 2));
"

Read run-history.json before each run to track score improvements over time.

Skill Composition

  • Before mutation testing → Run /qe-test-generation to ensure tests exist
  • After mutation results → Use /qe-coverage-analysis to prioritize improvement areas
  • Quality gate → Feed results into /qe-quality-assessment for ship/no-ship decision

Gotchas

  • Stryker requires --testRunner jest explicitly if both jest and vitest are installed
  • Mutating >= to > in date comparisons rarely gets killed — add boundary tests
  • Running on files >500 LOC will timeout; use --mutate to target specific functions
  • --concurrency defaults to CPU count which OOMs in containers — set to 2

© proffesor-for-testing, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in assets/skills/mutation-testing of proffesor-for-testing/agentic-qe.

  • SKILL.md
  • config.json
  • evals/mutation-testing.yaml
  • references/mutation-operators.md
  • run-history.json
  • schemas/output.json
  • scripts/validate-config.json
  • test-data/sample-output.json

Open the folder on GitHubat commit 829d030

Compare with similar skills

Mutation Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mutation Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mutation Testing this skillproffesor-for-testing/agentic-qe494—~1.7kAutomated safety check: PassMIT
E2E Test ThinkerUniClipboard/UniClipboard1.8k—~1.7kAutomated safety check: PassAGPL-3.0
Ralph Coveragejvm-skills/jvm-skills140—~683Automated safety check: PassApache-2.0
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Write Testsgnomeria/usbtree688—~622Automated safety check: PassMIT
Mutation Testjmagly/aiwg220—~3.2kAutomated safety check: PassMIT

Similar skills

  • E2E Test Thinker

    UniClipboard/UniClipboard

    Analyze the current branch's diff against main and determine which changes are testable via CLI-based end-to-end tests.

    1.8k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Ralph Coverage

    jvm-skills/jvm-skills

    Run Ralph in coverage mode — iteratively write tests for untested classes until coverage targets are met.

    140 GitHub stars~683 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Write Tests

    gnomeria/usbtree

    Author tests that match the repo's stack and existing test style, at the cheapest level that catches the regression.

    688 GitHub stars~622 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Mutation Test

    jmagly/aiwg

    Run mutation testing to validate test quality beyond code coverage.

    220 GitHub stars~3.2k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Write Frontend Tests

    Elite588/AUTOGPT

    Analyze the current branch diff against dev, plan integration tests for changed frontend pages/components, and write them.

    103 GitHub stars~1.9k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed

More from proffesor-for-testing/agentic-qe

All 111 skills in this repo
  • Contract Testing

    proffesor-for-testing/agentic-qe

    Consumer-driven contract testing for microservices using Pact, schema validation, API versioning, and backward compatibility testing.

    494 GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check passed
  • Performance Testing

    proffesor-for-testing/agentic-qe

    Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates.

    494 GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check passed
  • Code Review Quality

    proffesor-for-testing/agentic-qe

    Conduct context-driven code reviews focusing on quality, testability, and maintainability.

    494 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Security Testing

    proffesor-for-testing/agentic-qe

    Scans for security vulnerabilities including XSS, SQL injection, CSRF, and auth flaws using OWASP Top 10 methodology.

    494 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check: notes
  • Database Testing

    proffesor-for-testing/agentic-qe

    Database schema validation, data integrity testing, migration testing, transaction isolation, and query performance.

    494 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Accessibility Testing

    proffesor-for-testing/agentic-qe

    WCAG 2.2 compliance testing, screen reader validation, and inclusive design verification.

    494 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about Mutation Testing

What does Mutation Testing do?

Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate. Mutation Testing is an agent skill from proffesor-for-testing/agentic-qe. Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

When should I use Mutation Testing?

Mutation Testing fits situations like: evaluating test quality; identifying weak tests; proving tests actually catch bugs.

How do I install Mutation Testing in Claude Code?

Run `npx skills add proffesor-for-testing/agentic-qe --skill mutation-testing -a claude-code`. Or copy the skill folder (assets/skills/mutation-testing in proffesor-for-testing/agentic-qe) into .claude/skills/mutation-testing in your project. Claude Code loads it when a task matches its description.

How do I install Mutation Testing in Codex?

Run `npx skills add proffesor-for-testing/agentic-qe --skill mutation-testing -a codex`. Or copy the skill folder (assets/skills/mutation-testing in proffesor-for-testing/agentic-qe) into .agents/skills/mutation-testing in your project. Codex loads it when a task matches its description.

Can I use Mutation Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add proffesor-for-testing/agentic-qe --skill mutation-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mutation-testing, .gemini/skills/mutation-testing, .github/skills/mutation-testing and .opencode/skills/mutation-testing in your project.

What does Mutation Testing need to run?

Going by SKILL.md and its folder, Mutation Testing needs the command-line tools its instructions call (npx, npm and node). Our summary lists: Node.js.

Does Mutation Testing access the network?

SKILL.md contains no URLs. Its commands use npx and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Mutation Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Mutation Testing use?

Mutation Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mutation Testing use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 377 tokens, read only when the agent opens those files.

What are the alternatives to Mutation Testing?

Skills that share tags, products or a category with Mutation Testing: E2E Test Thinker (UniClipboard/UniClipboard, 1.8k stars), Ralph Coverage (jvm-skills/jvm-skills, 140 stars), Designing Tests (CloudAI-X/opencode-workflow, 275 stars) and Write Tests (gnomeria/usbtree, 688 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mutation Testing?

proffesor-for-testing (a GitHub user) maintains it in proffesor-for-testing/agentic-qe, which has 494 GitHub stars. The repository holds 111 skills in this directory. The repository was last updated on October 4, 2026.

Source: proffesor-for-testing/agentic-qe on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.