Agent skill

Behavioral Mutation Analyzer

by ArabelaTso in ArabelaTso/Skills-4-SE

Analyzes surviving mutants from mutation testing to identify why tests failed to detect them.

Apache-2.0Auto-check passedTesting & QA

Install Behavioral Mutation Analyzer

skills CLI
$ npx skills add ArabelaTso/Skills-4-SE --skill behavioral-mutation-analyzer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ArabelaTso/Skills-4-SE behavioral-mutation-analyzer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ArabelaTso/Skills-4-SE.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/behavioral-mutation-analyzer .claude/skills/behavioral-mutation-analyzer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
behavioral-mutation-analyzer
GitHub stars
253
Token cost
~2.5k tokens
SKILL.md length
853 words
Files
5 (incl. references, assets)
Skills in repo
150
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyzes surviving mutants from mutation testing to identify why tests failed to detect them.

  • Works in 6 steps: Input Collection and Validation → Surviving Mutant Extraction → Root Cause Classification → …
  • Analyzing mutation testing results
  • SKILL.md covers Overview, Analysis Workflow, Mutation Operators Reference and Tool Integration, plus 3 more sections
  • Calls mvn and npx

What it does

Behavioral Mutation Analyzer is an agent skill from ArabelaTso/Skills-4-SE. Analyzes surviving mutants from mutation testing to identify why tests failed to detect them. Takes repository code, test suite, and mutation testing results as input. Identifies root causes including insufficient coverage, equivalent mutants, weak assertions, and missed edge cases. Automatically generates actionable test improvements and new test cases. Use when analyzing mutation testing results, improving test suite effectiveness, investigating low mutation scores, generating tests to kill surviving mutants…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files and assets (for example `assets/mutation_analysis_report.md`, `references/mutation_operators.md` and `references/test_patterns.md`).

It sits in Testing & QA, covering Test generation, Test coverage and Root cause analysis. The repository describes itself as: A curated list of 180+ useful Claude Skills for Software Engineering and resources for customizing AI for SE workflows. The licence is Apache-2.0.

When your agent uses it

  • Analyzing mutation testing results
  • Improving test suite effectiveness
  • Investigating low mutation scores
  • Generating tests to kill surviving mutants

Example prompts

  • “Use the behavioral-mutation-analyzer skill to analyz surviving mutants from mutation testing to identify why tests failed to detect them”
  • “/behavioral-mutation-analyzer”

Requirements

  • Python 3
  • Node.js

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Input Collection and Validation
  2. Surviving Mutant Extraction
  3. Root Cause Classification
  4. Test Generation Strategy
  5. Automated Test Generation
  6. Report Generation

What it can do on your machine

Read from SKILL.md and the folder at commit 4f38503. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • mvn
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Behavioral Mutation Analyzer loads about 2.5k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 150 tokens; SKILL.md has 853 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~150
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ArabelaTso/Skills-4-SE at commit 4f38503, republished under its Apache-2.0 licence (© ArabelaTso). 853 words, ~2,500 tokens.

Download SKILL.mdSave it as .claude/skills/behavioral-mutation-analyzer/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
behavioral-mutation-analyzer
description
Analyzes surviving mutants from mutation testing to identify why tests failed to detect them. Takes repository code, test suite, and mutation testing results as input. Identifies root causes including insufficient coverage, equivalent mutants, weak assertions, and missed edge cases. Automatically generates actionable test improvements and new test cases. Use when analyzing mutation testing results, improving test suite effectiveness, investigating low mutation scores, generating tests to kill surviving mutants, or enhancing test quality based on mutation analysis.

Behavioral Mutation Analyzer

Overview

This skill systematically analyzes surviving mutants from mutation testing to understand test suite weaknesses and automatically generate improvements. It identifies why mutants survived, categorizes root causes, and produces actionable test enhancements to increase mutation detection rates.

Analysis Workflow

Step 1: Input Collection and Validation

Gather required inputs and verify completeness:

Required Inputs:

  • Repository source code (path or files)
  • Test suite (test files and framework)
  • Mutation testing results (report file or data)

Mutation Result Formats:

  • PIT (Java): XML or HTML reports
  • Stryker (JavaScript/TypeScript): JSON reports
  • mutmut (Python): result files
  • Pitest, Infection (PHP), Cosmic Ray, etc.

Validation checklist:

  • Source code accessible
  • Test suite runnable
  • Mutation results parseable
  • Mutation tool and version identified
Step 2: Surviving Mutant Extraction

Parse mutation results to identify all surviving mutants:

Extract for each mutant:

  • Mutant ID
  • Source file and line number
  • Mutation operator (e.g., boundary change, negation)
  • Original code
  • Mutated code
  • Status (survived/killed/timeout/error)

Focus on survived mutants: Filter out killed mutants and focus analysis on survivors that indicate test weaknesses.

Step 3: Root Cause Classification

Analyze each surviving mutant to determine why it survived:

Category 1: Insufficient Coverage

Indicators:

  • Mutated line not executed by any test
  • Mutated method/function never called
  • Conditional branch not taken

Analysis:

  • Check code coverage data
  • Identify uncovered code paths
  • Trace execution from test entry points

Example:

java
// Original
public int calculate(int x) {
    if (x > 0) {
        return x * 2;  // Line 3: Covered
    }
    return 0;  // Line 5: NOT covered
}

// Mutant: Line 5 changed to "return 1;"
// Survives because no test calls calculate() with x <= 0
Category 2: Equivalent Mutants

Indicators:

  • Mutation produces semantically identical behavior
  • Mathematical or logical equivalence
  • Dead code or unreachable state

Analysis:

  • Compare control flow graphs
  • Check for mathematical identities
  • Identify redundant operations

Example:

python
# Original
result = x * 1

# Mutant: changed to "result = x"
# Equivalent: multiplying by 1 has no effect
Category 3: Weak Assertions

Indicators:

  • Test executes mutated code but doesn't verify output
  • Assertions too broad or generic
  • Only checking for exceptions, not correctness

Analysis:

  • Review test assertions
  • Check what properties are verified
  • Identify missing postconditions

Example:

javascript
// Test
test('calculate returns a number', () => {
    const result = calculate(5);
    expect(typeof result).toBe('number');  // Weak: doesn't check value
});

// Mutant: "return x * 2" → "return x * 3"
// Survives because test only checks type, not value
Category 4: Missed Edge Cases

Indicators:

  • Mutation affects boundary conditions
  • Special values not tested (null, zero, empty, max/min)
  • Error handling paths not verified

Analysis:

  • Identify boundary values in mutated code
  • Check test inputs for edge case coverage
  • Review exception handling tests

Example:

java
// Original
public int divide(int a, int b) {
    return a / b;
}

// Mutant: added "if (b == 0) return 0;"
// Survives because no test checks division by zero
Category 5: Timing and Concurrency Issues

Indicators:

  • Mutant affects timing, delays, or synchronization
  • Race conditions or thread safety
  • Asynchronous behavior changes

Analysis:

  • Check for concurrent code
  • Identify timing-dependent logic
  • Review async/await patterns
Category 6: State-Dependent Behavior

Indicators:

  • Mutant affects state transitions
  • Order-dependent operations
  • Side effects not verified

Analysis:

  • Trace state changes
  • Check for stateful objects
  • Verify side effect assertions
Step 4: Test Generation Strategy

For each surviving mutant, determine the appropriate test enhancement:

Strategy 1: Add Missing Test Cases

  • When: Insufficient coverage
  • Action: Generate new test that executes mutated code
  • Focus: Cover the uncovered path

Strategy 2: Strengthen Assertions

  • When: Weak assertions
  • Action: Add specific value checks
  • Focus: Verify exact expected behavior

Strategy 3: Add Edge Case Tests

  • When: Missed edge cases
  • Action: Generate boundary value tests
  • Focus: Test special inputs (null, zero, empty, max, min)

Strategy 4: Mark as Equivalent

  • When: Equivalent mutant
  • Action: Document equivalence reasoning
  • Focus: No test needed, update mutation config to ignore

Strategy 5: Add Integration Tests

  • When: State or timing issues
  • Action: Create tests verifying end-to-end behavior
  • Focus: Observable effects and state transitions
Show full SKILL.md (346 more words)Show less
Step 5: Automated Test Generation

Generate concrete test code to kill surviving mutants:

Test Generation Process:

  1. Identify test framework (JUnit, pytest, Jest, etc.)
  2. Analyze existing test patterns and style
  3. Generate test following project conventions
  4. Include descriptive test names
  5. Add comments explaining what mutant is targeted

Example Generated Test:

python
def test_calculate_with_negative_input():
    """
    Test to kill mutant #42: calculate() with x <= 0
    Mutant changed 'return 0' to 'return 1' on line 5
    """
    result = calculate(-5)
    assert result == 0, "calculate() should return 0 for negative input"

    result = calculate(0)
    assert result == 0, "calculate() should return 0 for zero input"
Step 6: Report Generation

Create comprehensive analysis report using template in assets/mutation_analysis_report.md:

Report Sections:

  1. Executive summary (mutation score, survival rate)
  2. Surviving mutants by category
  3. Root cause analysis for each mutant
  4. Generated test enhancements
  5. Equivalent mutant documentation
  6. Recommendations for test suite improvement

Mutation Operators Reference

Common mutation operators and their implications:

Arithmetic Operators:

  • + ↔ -, * ↔ /, % ↔ *
  • Tests should verify exact numeric results

Relational Operators:

  • > ↔ >=, < ↔ <=, == ↔ !=
  • Tests should cover boundary conditions

Logical Operators:

  • && ↔ ||, ! insertion/removal
  • Tests should verify boolean logic

Conditional Boundaries:

  • < ↔ <=, > ↔ >=
  • Tests should include boundary values

Return Values:

  • Return value changes, void method calls removed
  • Tests should assert return values

Statement Deletion:

  • Remove method calls, assignments
  • Tests should verify side effects

For detailed mutation operator catalog, see references/mutation_operators.md.

Tool Integration

PIT (Java)

Parse PIT XML reports:

bash
# Run PIT
mvn org.pitest:pitest-maven:mutationCoverage

# Report location
target/pit-reports/YYYYMMDDHHMI/mutations.xml
Stryker (JavaScript/TypeScript)

Parse Stryker JSON reports:

bash
# Run Stryker
npx stryker run

# Report location
reports/mutation/mutation.json
mutmut (Python)

Parse mutmut results:

bash
# Run mutmut
mutmut run

# Show results
mutmut results
mutmut show [mutant-id]

For tool-specific parsing guidance, see references/tool_integration.md.

Practical Examples

Example 1: Insufficient Coverage

Surviving mutant:

java
// Line 15: return defaultValue; → return null;

Analysis: No test calls this method with conditions triggering line 15.

Generated test:

java
@Test
public void testGetValueWithMissingKey() {
    // Kills mutant on line 15
    String result = config.getValue("nonexistent");
    assertEquals("default", result);
}

Example 2: Weak Assertion

Surviving mutant:

python
# Line 8: return items[:5] → return items[:4]

Analysis: Test only checks len(result) > 0, not exact length.

Enhanced test:

python
def test_get_top_items_returns_five():
    # Kills mutant on line 8
    items = create_test_items(10)
    result = get_top_items(items)
    assert len(result) == 5, "Should return exactly 5 items"

Example 3: Equivalent Mutant

Surviving mutant:

javascript
// Original: if (x > 0 && x < 100)
// Mutant: if (0 < x && 100 > x)

Analysis: Logically equivalent, no behavioral difference.

Action: Mark as equivalent in mutation config, no test needed.

Best Practices

Prioritize mutants:

  1. High-impact code (critical business logic)
  2. Frequently executed paths
  3. Security-sensitive operations
  4. Public API methods

Test quality over quantity:

  • Focus on meaningful assertions
  • Avoid brittle tests
  • Test behavior, not implementation

Iterative improvement:

  • Start with easiest mutants to kill
  • Gradually tackle complex cases
  • Re-run mutation testing after improvements

Document equivalent mutants:

  • Maintain list of known equivalent mutants
  • Configure mutation tool to skip them
  • Explain equivalence reasoning

References

For detailed information on specific topics:

  • Mutation operators: See references/mutation_operators.md
  • Tool integration: See references/tool_integration.md
  • Test patterns: See references/test_patterns.md

© ArabelaTso, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references, assets) in skills/behavioral-mutation-analyzer of ArabelaTso/Skills-4-SE.

  • SKILL.md
  • assets/mutation_analysis_report.md
  • references/mutation_operators.md
  • references/test_patterns.md
  • references/tool_integration.md

Open the folder on GitHubat commit 4f38503

Compare with similar skills

Behavioral Mutation Analyzer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Behavioral Mutation Analyzer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Behavioral Mutation Analyzer this skillArabelaTso/Skills-4-SE253—~2.5kAutomated safety check: PassApache-2.0
E2E Test ThinkerUniClipboard/UniClipboard1.9k—~1.7kAutomated safety check: PassAGPL-3.0
Ralph Coveragejvm-skills/jvm-skills140—~683Automated safety check: PassApache-2.0
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Mutation Testingproffesor-for-testing/agentic-qe494—~1.7kAutomated safety check: PassMIT
Write Testsgnomeria/usbtree690—~622Automated safety check: PassMIT

Similar skills

  • E2E Test Thinker

    UniClipboard/UniClipboard

    Analyze the current branch's diff against main and determine which changes are testable via CLI-based end-to-end tests.

    1.9k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Ralph Coverage

    jvm-skills/jvm-skills

    Run Ralph in coverage mode — iteratively write tests for untested classes until coverage targets are met.

    140 GitHub stars~683 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Mutation Testing

    proffesor-for-testing/agentic-qe

    Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

    494 GitHub stars~1.7k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Write Tests

    gnomeria/usbtree

    Author tests that match the repo's stack and existing test style, at the cheapest level that catches the regression.

    690 GitHub stars~622 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Mutation Test

    jmagly/aiwg

    Run mutation testing to validate test quality beyond code coverage.

    220 GitHub stars~3.2k tokensUpdated 2 days ago
    Testing & QAAuto-check passed

More from ArabelaTso/Skills-4-SE

All 150 skills in this repo
  • Framework Migration Assistant

    ArabelaTso/Skills-4-SE

    Automatically migrate Python web applications between frameworks (Flask → FastAPI, Django → FastAPI).

    253 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Metamorphic Test Generator

    ArabelaTso/Skills-4-SE

    Generate test cases using metamorphic testing by applying transformations based on metamorphic properties.

    253 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Reproduction Trace Instrumenter

    ArabelaTso/Skills-4-SE

    Instruments programs to capture execution traces specifically for reproducing reported bugs, enabling consistent replay and diagnosis of failures.

    253 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spring Mvc To Boot Migrator

    ArabelaTso/Skills-4-SE

    Automatically migrate Spring MVC applications to Spring Boot.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • State Snapshot Instrumenter

    ArabelaTso/Skills-4-SE

    Instrument programs (Python, C/C++, Java) to capture snapshots of key program states at runtime, including variables, memory, and call stacks.

    253 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Behavioral Mutation Analyzer

What does Behavioral Mutation Analyzer do?

Analyzes surviving mutants from mutation testing to identify why tests failed to detect them. Behavioral Mutation Analyzer is an agent skill from ArabelaTso/Skills-4-SE. Analyzes surviving mutants from mutation testing to identify why tests failed to detect them.

When should I use Behavioral Mutation Analyzer?

Behavioral Mutation Analyzer fits situations like: analyzing mutation testing results; improving test suite effectiveness; investigating low mutation scores; generating tests to kill surviving mutants.

How do I install Behavioral Mutation Analyzer in Claude Code?

Run `npx skills add ArabelaTso/Skills-4-SE --skill behavioral-mutation-analyzer -a claude-code`. Or copy the skill folder (skills/behavioral-mutation-analyzer in ArabelaTso/Skills-4-SE) into .claude/skills/behavioral-mutation-analyzer in your project. Claude Code loads it when a task matches its description.

How do I install Behavioral Mutation Analyzer in Codex?

Run `npx skills add ArabelaTso/Skills-4-SE --skill behavioral-mutation-analyzer -a codex`. Or copy the skill folder (skills/behavioral-mutation-analyzer in ArabelaTso/Skills-4-SE) into .agents/skills/behavioral-mutation-analyzer in your project. Codex loads it when a task matches its description.

Can I use Behavioral Mutation Analyzer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ArabelaTso/Skills-4-SE --skill behavioral-mutation-analyzer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/behavioral-mutation-analyzer, .gemini/skills/behavioral-mutation-analyzer, .github/skills/behavioral-mutation-analyzer and .opencode/skills/behavioral-mutation-analyzer in your project.

What does Behavioral Mutation Analyzer need to run?

Going by SKILL.md and its folder, Behavioral Mutation Analyzer needs the command-line tools its instructions call (mvn and npx). Our summary lists: Python 3; Node.js.

Does Behavioral Mutation Analyzer access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Behavioral Mutation Analyzer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Behavioral Mutation Analyzer use?

Behavioral Mutation Analyzer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Behavioral Mutation Analyzer use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.3k tokens, read only when the agent opens those files.

What are the alternatives to Behavioral Mutation Analyzer?

Skills that share tags, products or a category with Behavioral Mutation Analyzer: E2E Test Thinker (UniClipboard/UniClipboard, 1.9k stars), Ralph Coverage (jvm-skills/jvm-skills, 140 stars), Designing Tests (CloudAI-X/opencode-workflow, 275 stars) and Mutation Testing (proffesor-for-testing/agentic-qe, 494 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Behavioral Mutation Analyzer?

ArabelaTso (a GitHub user) maintains it in ArabelaTso/Skills-4-SE, which has 253 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on August 21, 2026.

Source: ArabelaTso/Skills-4-SE on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.