Official agent skill

Mutation Testing Triage

by trailofbits in trailofbits/skills

Triages survived mutants and unnecessary test statements using Trailmark call-graph data, sorting them into false positives, missing unit tests and fuzzing targets.

OfficialCC-BY-SA-4.0Auto-check passedTesting & QA

Install Mutation Testing Triage

skills CLI
$ npx skills add trailofbits/skills --skill genotoxic -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install trailofbits/skills genotoxic --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/trailofbits/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/trailmark/skills/genotoxic .claude/skills/genotoxic && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
genotoxic
GitHub stars
7.4k
Token cost
~3.2k tokens
SKILL.md length
1,113 words
Files
6 (incl. references, assets)
Skills in repo
79
Repo updated
First seen
Licence
CC-BY-SA-4.0

At a glance

Triages survived mutants and unnecessary test statements using Trailmark call-graph data, sorting them into false positives, missing unit tests and fuzzing targets.

  • Works in 3 steps: Build Code Graph and Run Pre-Analysis → Run Mutation Testing → Triage Findings
  • Triaging survived mutants after a mutation testing run
  • SKILL.md covers When to Use, When NOT to Use, Prerequisites and Rationalizations to Reject, plus 11 more sections
  • Calls uv and cargo

What it does

After a mutation testing run, the agent combines the survived mutants with necessist results, which flag test statements that can be removed without failing the test, and with a Trailmark code graph. Data flow context then decides where a mutant is harmless, where a unit test would add the most, and which functions would be better served by a fuzz harness.

Prerequisites are stated firmly: Trailmark installed with uv, a mutation testing framework for the language (install steps live in a reference file), a test suite that already passes, and optionally necessist for languages such as Go, Rust, Solidity and TypeScript. The agent must install a missing tool or report the error rather than switch to manual analysis. On macOS, ulimit -n 1024 is set before running mull-runner.

When your agent uses it

  • Triaging survived mutants after a mutation testing run
  • Deciding where new unit tests would have the highest impact
  • Finding functions that need fuzz harnesses instead of unit tests
  • Spotting tests with unnecessary statements that point to weak assertions

Example prompts

  • “Triage the survived mutants in the mutation report and separate the harmless ones.”
  • “Run necessist on the Rust crate and tell me which tests have weak assertions.”
  • “Which parsing functions should get fuzz harnesses, based on the mutation results?”

Requirements

  • Trailmark, installed with uv
  • A mutation testing framework for the target language
  • A test suite that already passes
  • necessist, optional, installed with cargo

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Build Code Graph and Run Pre-Analysis
  2. Run Mutation Testing
  3. Triage Findings

What it can do on your machine

Read from SKILL.md and the folder at commit 82fe822. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mutation Testing Triage loads about 3.2k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 1,113 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from trailofbits/skills at commit 82fe822, republished under its CC-BY-SA-4.0 licence (© trailofbits). 1,113 words, ~3,226 tokens.

Download SKILL.mdSave it as .claude/skills/genotoxic/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
genotoxic
description
Graph-informed mutation testing triage. Parses codebases with Trailmark, runs mutation testing and necessist, then uses survived mutants, unnecessary test statements, and call graph data to identify false positives, missing test coverage, and fuzzing targets. Use when triaging survived mutants, analyzing mutation testing results, identifying test gaps, finding fuzzing targets from weak tests, running mutation frameworks (including circomvent and cairo-mutants), or using necessist.

Genotoxic

Combines mutation testing and necessist (test statement removal) with code graph analysis to triage findings into actionable categories: false positives, missing unit tests, and fuzzing targets.

When to Use

  • After mutation testing reveals survived mutants that need triage
  • Identifying where unit tests would have the highest impact
  • Finding functions that need fuzz harnesses instead of unit tests
  • Prioritizing test improvements using data flow context
  • Filtering out harmless mutants from actionable ones
  • Finding unnecessary test statements that indicate weak assertions (necessist)

When NOT to Use

  • Codebase has no existing test suite (write tests first)
  • Pure documentation or configuration changes
  • Single-file scripts with trivial logic

Prerequisites

  • trailmark installed — if uv run trailmark fails, run:
    bash
    uv tool install trailmark

Python snippets: uv run --with trailmark python - (a tool env is not importable)

**DO NOT** fall back to "manual verification" or "manual analysis"
as a substitute for running trailmark. Install it first. If installation
fails, report the error instead of switching to manual analysis.
- A **mutation testing framework** for the target language — if the framework
command fails (not found, not installed), install it using the instructions
in [references/mutation-frameworks.md](references/mutation-frameworks.md).
**DO NOT** fall back to "manual mutation analysis" or skip mutation testing.
Install the framework first. If installation fails, report the error
instead of switching to manual mutation analysis.
- **necessist** (optional, recommended) — if the target language is
supported (Go, Rust, Solidity/Foundry, TypeScript/Hardhat,
TypeScript/Vitest, Rust/Anchor), install with `cargo install necessist`.
See [references/mutation-frameworks.md](references/mutation-frameworks.md)
for details.
- An existing test suite that passes
- **macOS environment**: Run `ulimit -n 1024` before any `mull-runner`
invocation. macOS Tahoe (26+) sets unlimited file descriptors by
default, which crashes Mull's subprocess spawning. See
[references/mutation-frameworks.md](references/mutation-frameworks.md)
for details.

---

## Rationalizations to Reject

| Rationalization | Why It's Wrong | Required Action |
|-----------------|----------------|-----------------|
| "All survived mutants need tests" | Many are harmless or equivalent | Triage before writing tests |
| "Mutation testing is too noisy" | Noise means you're not triaging | Use graph data to filter |
| "Unit tests cover everything" | Complex data flows need fuzzing | Check entrypoint reachability |
| "Dead code mutants don't matter" | Dead code should be removed | Flag for cleanup |
| "Low complexity = low risk" | Boundary bugs hide in simple code | Check mutant location |
| "Tool isn't installed, I'll do it manually" | Manual analysis misses what tooling catches | Install the tool first |
| "Necessist isn't mutation testing, skip it" | Necessist finds what mutation testing misses: weak tests | Run both when the language supports it |

---

## Quick Start

```bash
# 1. Build the code graph
uv run trailmark analyze --language auto --summary {targetDir}

# 2. Run mutation testing (language-dependent)
# Python:
uv run mutmut run --paths-to-mutate {targetDir}/src
uv run mutmut results

# 2b. Run necessist (if language supported)
necessist

# 3. Analyze results with this skill's workflow (Phase 3)

Workflow Overview

Phase 1: Graph Build      → Parse codebase with trailmark
      ↓
Phase 2: Mutation Run     → Execute mutation testing framework
Phase 2b: Necessist Run   → Remove test statements (optional, parallel)
      ↓
Phase 3: Triage           → Classify findings using graph data
      ↓
Output: Categorized Report
  ├── Corroborated         (both tools flag same function — highest value)
  ├── False Positives      (harmless, skip)
  ├── Missing Tests        (write unit tests)
  └── Fuzzing Targets      (set up fuzz harnesses)

Decision Tree

├─ Need to set up mutation testing for a language?
│  └─ Read: references/mutation-frameworks.md
│
├─ Need to set up necessist or find weak test statements?
│  └─ Read: references/mutation-frameworks.md (Necessist section)
│
├─ Need to understand the triage criteria in depth?
│  └─ Read: references/triage-methodology.md
│
├─ Need to understand how graph data informs triage?
│  └─ Read: references/graph-analysis.md
│
└─ Already have results + graph? Use Phase 3 below.

Phase 1: Build Code Graph and Run Pre-Analysis

Parse the target codebase with trailmark and run pre-analysis before mutation testing. Pre-analysis computes blast radius, entry points, privilege boundaries, and taint propagation, which Phase 3 uses for triage.

bash
uv run trailmark analyze --language auto --summary {targetDir}

Use the QueryEngine API to build the graph and run pre-analysis:

  1. QueryEngine.from_directory("{targetDir}", language="auto")
  2. Call engine.preanalysis() — mandatory before triage
  3. Export with engine.to_json() for cross-referencing with mutation results

If auto-detection is wrong for the target, rerun with an explicit language or comma-separated list such as python,rust.

See references/graph-analysis.md for the full API: node mapping, reachability queries, blast radius, and pre-analysis subgraph lookups.


Phase 2: Run Mutation Testing

Select and run the appropriate framework. See references/mutation-frameworks.md for language-specific setup.

Capture survived mutants. Each framework reports differently, but extract these fields per mutant:

FieldDescription
File pathSource file containing the mutant
Line numberLine where mutation was applied
Mutation typeWhat was changed (operator, value, etc.)
Statussurvived, killed, timeout, error

Filter to survived mutants only for Phase 3.


Phase 2b: Run Necessist (Optional)

If the target language is supported (Go, Rust, Solidity/Foundry, TypeScript/Hardhat, TypeScript/Vitest, Rust/Anchor), run necessist to find unnecessary test statements. This runs independently of Phase 2 and can execute in parallel.

bash
# Auto-detect framework
necessist

# Or target specific test files
necessist tests/test_parser.rs

# Export results
necessist --dump

Filter to findings where the test passed after removal. See references/mutation-frameworks.md for framework-specific configuration and the normalized record format.

Map each removal to a production function using the algorithm in references/graph-analysis.md.


Phase 3: Triage Findings

For each survived mutant and each necessist removal, determine its triage bucket using graph data. Necessist removals must first be mapped to a production function (see references/graph-analysis.md).

Quick Classification (Mutation Testing)
SignalBucketReasoning
No callers in graphFalse PositiveDead code, mutant is unreachable
Only test callersFalse PositiveTest infrastructure, not production
Logging/display stringFalse PositiveCosmetic, no behavioral impact
Equivalent mutantFalse PositiveBehavior unchanged despite mutation
Simple function, low CC, no entrypoint pathMissing TestsUnit test is straightforward
Error handling pathMissing TestsShould have negative test cases
Boundary condition (off-by-one)Missing TestsProperty-based test candidate
Pure function, deterministicMissing TestsEasy to test, high value
High CC (>10), entrypoint reachableFuzzing TargetComplex + exposed = fuzz it
Parser/validator/deserializerFuzzing TargetStructured input handling
Many callers (>10) + moderate CCFuzzing TargetHigh blast radius
Binary/wire protocol handlingFuzzing TargetFuzzers excel at format testing
Show full SKILL.md (354 more words)Show less
Quick Classification (Necessist)
SignalBucketReasoning
Redundant setup or debug callFalse PositiveStatement genuinely unnecessary
Cannot map to production functionFalse PositiveNo graph context for triage
Call removed, no assertion checks its effectMissing TestsTest has weak assertions
Assertion removed, test still passesMissing TestsRedundant or insufficient coverage
Maps to high-CC entrypoint-reachable functionFuzzing TargetComplex + exposed + weak test

When both mutation testing and necessist flag the same production function, mark as corroborated — highest confidence finding.

For detailed criteria, see references/triage-methodology.md.

Graph Queries for Triage

For each mutant, map it to its containing graph node and use pre-analysis subgraphs (tainted, high_blast_radius, privilege_boundary) from Phase 1 to classify it. The classification logic checks: no callers → false positive, privilege boundary → fuzzing, high CC + tainted → fuzzing, high blast radius → fuzzing, otherwise → missing tests.

See references/graph-analysis.md for the batch_triage implementation and node mapping functions.


Output Format

Generate a markdown report:

markdown
# Genotoxic Triage Report

## Summary
- Total survived mutants: N
- Total necessist removals: N
- Corroborated findings: N
- False positives: N (N%)
- Missing test coverage: N (N%)
- Fuzzing targets: N (N%)

## Corroborated Findings
| File | Line | Function | Mutation Signal | Necessist Signal | Action |
|------|------|----------|----------------|------------------|--------|

## False Positives
| File | Line | Mutation | Reason | Source |
|------|------|----------|--------|--------|

## Missing Test Coverage
| File | Line | Function | CC | Callers | Suggested Test | Source |
|------|------|----------|----|---------|----------------|--------|

## Fuzzing Targets
| File | Line | Function | CC | Entrypoint Path | Blast Radius | Source |
|------|------|----------|----|-----------------|--------------|--------|

The Source column is mutation, necessist, or corroborated.

Write the report to GENOTOXIC_REPORT.md in the working directory.


Quality Checklist

Before delivering:

  • Trailmark graph built for target language
  • Mutation framework ran to completion
  • Necessist ran (if language supported) or noted as not applicable
  • All survived mutants triaged (none unclassified)
  • All necessist removals triaged (if applicable)
  • Corroborated findings identified (if both tools ran)
  • False positives have clear justifications
  • Missing test items include suggested test type
  • Fuzzing targets include entrypoint paths and blast radius
  • Report file written to GENOTOXIC_REPORT.md
  • User notified with summary statistics

Integration

trailmark skill:

  • Phase 1: Build code graph, query complexity and entrypoints
  • Phase 3: Caller analysis, reachability, blast radius

property-based-testing skill:

  • Missing test coverage items involving boundary conditions
  • Roundtrip/idempotence properties for serialization mutants

testing-handbook-skills (fuzzing):

  • Fuzzing target items: use harness-writing, cargo-fuzz, atheris

Supporting Documentation


First-time users: Start with Phase 1 (graph build), then run mutations, then use the Quick Classification table in Phase 3.

Experienced users: Jump to Phase 3 and use the Decision Tree to load specific reference material.

© trailofbits, CC-BY-SA-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references, assets) in plugins/trailmark/skills/genotoxic of trailofbits/skills.

  • SKILL.md
  • agents/openai.yaml
  • assets/trail-of-bits-mark.svg
  • references/graph-analysis.md
  • references/mutation-frameworks.md
  • references/triage-methodology.md

Open the folder on GitHubat commit 82fe822

Compare with similar skills

Mutation Testing Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mutation Testing Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mutation Testing Triage this skilltrailofbits/skills7.4k—~3.2kAutomated safety check: PassCC-BY-SA-4.0
Test Writing WorkflowiOfficeAI/AionUi33k1 repos~1.2kAutomated safety check: PassApache-2.0
Golang Testingantoniopaya22/go-rest-template1729 repos~4.2kAutomated safety check: PassNone
Supercovsupercorp-ai/supercov1481 repos~415Automated safety check: PassMIT
Coverage Toolszkldi/Tachi237—~580Automated safety check: PassNone
Supercov Securitysupercorp-ai/supercov1481 repos~236Automated safety check: PassMIT

Similar skills

  • Test Writing Workflow

    iOfficeAI/AionUi

    Sets the test-writing workflow for the repository: risk-first scenario lists, behavior-focused Vitest tests, a full run before each commit and a coverage target.

    33k GitHub starsUsed in 1 repo~1.2k tokens
    Testing & QAAuto-check passed
  • Golang Testing

    antoniopaya22/go-rest-template

    Go testing patterns including table-driven tests, subtests, benchmarks, fuzzing, and test coverage.

    172 GitHub starsUsed in 9 repos~4.2k tokens
    Testing & QAAuto-check passed
  • Supercov

    supercorp-ai/supercov

    Measures test coverage and code quality in a repository with the supercov CLI, and turns what it finds into small, focused tests or fixes.

    148 GitHub starsUsed in 1 repo~415 tokens
    Testing & QAAuto-check passed
  • Coverage Tools

    zkldi/Tachi

    Aggregates Vitest v8/Istanbul coverage across Tachi workspaces via tachi-coverage-tools (manifest, CLI, optional programmatic API).

    237 GitHub stars~580 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Supercov Security

    supercorp-ai/supercov

    Scans a repository's source for security vulnerabilities with the supercov CLI, pointing to the line of each finding and mapping it to CWE classes.

    148 GitHub starsUsed in 1 repo~236 tokens
    Testing & QAAuto-check passed
  • Testing Strategies

    ancoleman/ai-design-components

    Strategic guidance for choosing and implementing testing approaches across the test pyramid.

    526 GitHub stars~3.8k tokensUpdated 10 mo ago
    Testing & QAAuto-check passed

More from trailofbits/skills

All 79 skills in this repo
  • CodeQL Security Scan

    trailofbits/skills

    Official

    Scans a codebase for vulnerabilities with CodeQL's data flow and taint tracking in run-all or important-only modes, including data extensions for project-specific sources and sinks.

    7.4k GitHub stars~4.6k tokensUpdated 5 days ago
    Auto-check: notes
  • Code Graph Mermaid Diagrams

    trailofbits/skills

    Official

    Generates Mermaid diagrams from Trailmark code graphs, including call graphs, class hierarchies, module dependency maps, complexity heatmaps and attack surface data flows.

    7.4k GitHub stars~1.7k tokensUpdated 5 days ago
    Auto-check passed
  • Trailmark Graph Evolution

    trailofbits/skills

    Official

    Compares Trailmark code graphs at two snapshots, such as commits, tags or directories, to surface attack paths, blast radius and taint changes that text diffs miss.

    7.4k GitHub stars~3.4k tokensUpdated 5 days ago
    Auto-check passed
  • Let Fate Decide

    trailofbits/skills

    Official

    Draws a 12 Houses tarot spread to break ties when a request is vague or casually delegated, then reads the cards to pick the next step.

    7.4k GitHub stars~2.5k tokensUpdated 5 days ago
    Auto-check: notes
  • Semgrep Security Scan

    trailofbits/skills

    Official

    Detects languages, proposes rulesets for approval, then runs the approved Semgrep scan across a codebase and merges the output into one SARIF file.

    7.4k GitHub stars~3.7k tokensUpdated 5 days ago
    Auto-check: notes
  • Burp Suite Project Parser

    trailofbits/skills

    Official

    Searches and extracts data from Burp Suite project files on the command line: regex searches over responses, audit findings, proxy history and site map data.

    7.4k GitHub starsUsed in 3 repos~4.2k tokens
    Auto-check: notes

Questions about Mutation Testing Triage

What does Mutation Testing Triage do?

Triages survived mutants and unnecessary test statements using Trailmark call-graph data, sorting them into false positives, missing unit tests and fuzzing targets. After a mutation testing run, the agent combines the survived mutants with necessist results, which flag test statements that can be removed without failing the test, and with a Trailmark code graph. Data flow context then decides where a mutant is harmless, where a unit test would add the most, and which functions would be better served by a fuzz harness.

When should I use Mutation Testing Triage?

Mutation Testing Triage fits situations like: triaging survived mutants after a mutation testing run; deciding where new unit tests would have the highest impact; finding functions that need fuzz harnesses instead of unit tests; spotting tests with unnecessary statements that point to weak assertions.

How do I install Mutation Testing Triage in Claude Code?

Run `npx skills add trailofbits/skills --skill genotoxic -a claude-code`. Or copy the skill folder (plugins/trailmark/skills/genotoxic in trailofbits/skills) into .claude/skills/genotoxic in your project. Claude Code loads it when a task matches its description.

How do I install Mutation Testing Triage in Codex?

Run `npx skills add trailofbits/skills --skill genotoxic -a codex`. Or copy the skill folder (plugins/trailmark/skills/genotoxic in trailofbits/skills) into .agents/skills/genotoxic in your project. Codex loads it when a task matches its description.

Can I use Mutation Testing Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add trailofbits/skills --skill genotoxic -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/genotoxic, .gemini/skills/genotoxic, .github/skills/genotoxic and .opencode/skills/genotoxic in your project.

What does Mutation Testing Triage need to run?

Going by SKILL.md and its folder, Mutation Testing Triage needs the command-line tools its instructions call (uv and cargo). Our summary lists: Trailmark, installed with uv; A mutation testing framework for the target language; A test suite that already passes; necessist, optional, installed with cargo.

Does Mutation Testing Triage access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Mutation Testing Triage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Mutation Testing Triage use?

Mutation Testing Triage is published under the CC-BY-SA-4.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mutation Testing Triage use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Mutation Testing Triage?

Skills that share tags, products or a category with Mutation Testing Triage: Test Writing Workflow (iOfficeAI/AionUi, 33k stars), Golang Testing (antoniopaya22/go-rest-template, 172 stars), Supercov (supercorp-ai/supercov, 148 stars) and Coverage Tools (zkldi/Tachi, 237 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mutation Testing Triage?

trailofbits (a GitHub organization, an official publisher) maintains it in trailofbits/skills, which has 7,400 GitHub stars. The repository holds 79 skills in this directory. The repository was last updated on October 2, 2026.

Source: trailofbits/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.