Test Writing Workflow
iOfficeAI/AionUi
Sets the test-writing workflow for the repository: risk-first scenario lists, behavior-focused Vitest tests, a full run before each commit and a coverage target.
Triages survived mutants and unnecessary test statements using Trailmark call-graph data, sorting them into false positives, missing unit tests and fuzzing targets.
$ npx skills add trailofbits/skills --skill genotoxic -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install trailofbits/skills genotoxic --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/trailofbits/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/trailmark/skills/genotoxic .claude/skills/genotoxic && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "genotoxic" agent skill from https://github.com/trailofbits/skills/tree/main/plugins/trailmark/skills/genotoxic into .claude/skills/genotoxic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genotoxic", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/trailofbits/skills/tree/main/plugins/trailmark/skills/genotoxicType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add trailofbits/skills --skill genotoxic -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install trailofbits/skills genotoxic --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/trailofbits/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/trailmark/skills/genotoxic .agents/skills/genotoxic && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "genotoxic" agent skill from https://github.com/trailofbits/skills/tree/main/plugins/trailmark/skills/genotoxic into .agents/skills/genotoxic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genotoxic", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add trailofbits/skills --skill genotoxic -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install trailofbits/skills genotoxic --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/trailofbits/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/trailmark/skills/genotoxic .cursor/skills/genotoxic && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "genotoxic" agent skill from https://github.com/trailofbits/skills/tree/main/plugins/trailmark/skills/genotoxic into .cursor/skills/genotoxic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genotoxic", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/trailofbits/skills.git --path plugins/trailmark/skills/genotoxic--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add trailofbits/skills --skill genotoxic -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install trailofbits/skills genotoxic --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/trailofbits/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/trailmark/skills/genotoxic .gemini/skills/genotoxic && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "genotoxic" agent skill from https://github.com/trailofbits/skills/tree/main/plugins/trailmark/skills/genotoxic into .gemini/skills/genotoxic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genotoxic", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install trailofbits/skills genotoxicInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add trailofbits/skills --skill genotoxic -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/trailofbits/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/trailmark/skills/genotoxic .github/skills/genotoxic && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "genotoxic" agent skill from https://github.com/trailofbits/skills/tree/main/plugins/trailmark/skills/genotoxic into .github/skills/genotoxic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genotoxic", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add trailofbits/skills --skill genotoxic -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install trailofbits/skills genotoxic --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/trailofbits/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/trailmark/skills/genotoxic .opencode/skills/genotoxic && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "genotoxic" agent skill from https://github.com/trailofbits/skills/tree/main/plugins/trailmark/skills/genotoxic into .opencode/skills/genotoxic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "genotoxic", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
genotoxicTriages survived mutants and unnecessary test statements using Trailmark call-graph data, sorting them into false positives, missing unit tests and fuzzing targets.
After a mutation testing run, the agent combines the survived mutants with necessist results, which flag test statements that can be removed without failing the test, and with a Trailmark code graph. Data flow context then decides where a mutant is harmless, where a unit test would add the most, and which functions would be better served by a fuzz harness.
Prerequisites are stated firmly: Trailmark installed with uv, a mutation testing framework for the language (install steps live in a reference file), a test suite that already passes, and optionally necessist for languages such as Go, Rust, Solidity and TypeScript. The agent must install a missing tool or report the error rather than switch to manual analysis. On macOS, ulimit -n 1024 is set before running mull-runner.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82fe822. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvcargoFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Mutation Testing Triage loads about 3.2k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 1,113 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from trailofbits/skills at commit 82fe822, republished under its CC-BY-SA-4.0 licence (© trailofbits). 1,113 words, ~3,226 tokens.
.claude/skills/genotoxic/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Combines mutation testing and necessist (test statement removal) with code graph analysis to triage findings into actionable categories: false positives, missing unit tests, and fuzzing targets.
uv run trailmark fails, run:uv tool install trailmark**DO NOT** fall back to "manual verification" or "manual analysis"
as a substitute for running trailmark. Install it first. If installation
fails, report the error instead of switching to manual analysis.
- A **mutation testing framework** for the target language — if the framework
command fails (not found, not installed), install it using the instructions
in [references/mutation-frameworks.md](references/mutation-frameworks.md).
**DO NOT** fall back to "manual mutation analysis" or skip mutation testing.
Install the framework first. If installation fails, report the error
instead of switching to manual mutation analysis.
- **necessist** (optional, recommended) — if the target language is
supported (Go, Rust, Solidity/Foundry, TypeScript/Hardhat,
TypeScript/Vitest, Rust/Anchor), install with `cargo install necessist`.
See [references/mutation-frameworks.md](references/mutation-frameworks.md)
for details.
- An existing test suite that passes
- **macOS environment**: Run `ulimit -n 1024` before any `mull-runner`
invocation. macOS Tahoe (26+) sets unlimited file descriptors by
default, which crashes Mull's subprocess spawning. See
[references/mutation-frameworks.md](references/mutation-frameworks.md)
for details.
---
## Rationalizations to Reject
| Rationalization | Why It's Wrong | Required Action |
|-----------------|----------------|-----------------|
| "All survived mutants need tests" | Many are harmless or equivalent | Triage before writing tests |
| "Mutation testing is too noisy" | Noise means you're not triaging | Use graph data to filter |
| "Unit tests cover everything" | Complex data flows need fuzzing | Check entrypoint reachability |
| "Dead code mutants don't matter" | Dead code should be removed | Flag for cleanup |
| "Low complexity = low risk" | Boundary bugs hide in simple code | Check mutant location |
| "Tool isn't installed, I'll do it manually" | Manual analysis misses what tooling catches | Install the tool first |
| "Necessist isn't mutation testing, skip it" | Necessist finds what mutation testing misses: weak tests | Run both when the language supports it |
---
## Quick Start
```bash
# 1. Build the code graph
uv run trailmark analyze --language auto --summary {targetDir}
# 2. Run mutation testing (language-dependent)
# Python:
uv run mutmut run --paths-to-mutate {targetDir}/src
uv run mutmut results
# 2b. Run necessist (if language supported)
necessist
# 3. Analyze results with this skill's workflow (Phase 3)Phase 1: Graph Build → Parse codebase with trailmark
↓
Phase 2: Mutation Run → Execute mutation testing framework
Phase 2b: Necessist Run → Remove test statements (optional, parallel)
↓
Phase 3: Triage → Classify findings using graph data
↓
Output: Categorized Report
├── Corroborated (both tools flag same function — highest value)
├── False Positives (harmless, skip)
├── Missing Tests (write unit tests)
└── Fuzzing Targets (set up fuzz harnesses)├─ Need to set up mutation testing for a language?
│ └─ Read: references/mutation-frameworks.md
│
├─ Need to set up necessist or find weak test statements?
│ └─ Read: references/mutation-frameworks.md (Necessist section)
│
├─ Need to understand the triage criteria in depth?
│ └─ Read: references/triage-methodology.md
│
├─ Need to understand how graph data informs triage?
│ └─ Read: references/graph-analysis.md
│
└─ Already have results + graph? Use Phase 3 below.Parse the target codebase with trailmark and run pre-analysis before mutation testing. Pre-analysis computes blast radius, entry points, privilege boundaries, and taint propagation, which Phase 3 uses for triage.
uv run trailmark analyze --language auto --summary {targetDir}Use the QueryEngine API to build the graph and run pre-analysis:
QueryEngine.from_directory("{targetDir}", language="auto")engine.preanalysis() — mandatory before triageengine.to_json() for cross-referencing with mutation resultsIf auto-detection is wrong for the target, rerun with an explicit language or
comma-separated list such as python,rust.
See references/graph-analysis.md for the full API: node mapping, reachability queries, blast radius, and pre-analysis subgraph lookups.
Select and run the appropriate framework. See references/mutation-frameworks.md for language-specific setup.
Capture survived mutants. Each framework reports differently, but extract these fields per mutant:
| Field | Description |
|---|---|
| File path | Source file containing the mutant |
| Line number | Line where mutation was applied |
| Mutation type | What was changed (operator, value, etc.) |
| Status | survived, killed, timeout, error |
Filter to survived mutants only for Phase 3.
If the target language is supported (Go, Rust, Solidity/Foundry, TypeScript/Hardhat, TypeScript/Vitest, Rust/Anchor), run necessist to find unnecessary test statements. This runs independently of Phase 2 and can execute in parallel.
# Auto-detect framework
necessist
# Or target specific test files
necessist tests/test_parser.rs
# Export results
necessist --dumpFilter to findings where the test passed after removal. See references/mutation-frameworks.md for framework-specific configuration and the normalized record format.
Map each removal to a production function using the algorithm in references/graph-analysis.md.
For each survived mutant and each necessist removal, determine its triage bucket using graph data. Necessist removals must first be mapped to a production function (see references/graph-analysis.md).
| Signal | Bucket | Reasoning |
|---|---|---|
| No callers in graph | False Positive | Dead code, mutant is unreachable |
| Only test callers | False Positive | Test infrastructure, not production |
| Logging/display string | False Positive | Cosmetic, no behavioral impact |
| Equivalent mutant | False Positive | Behavior unchanged despite mutation |
| Simple function, low CC, no entrypoint path | Missing Tests | Unit test is straightforward |
| Error handling path | Missing Tests | Should have negative test cases |
| Boundary condition (off-by-one) | Missing Tests | Property-based test candidate |
| Pure function, deterministic | Missing Tests | Easy to test, high value |
| High CC (>10), entrypoint reachable | Fuzzing Target | Complex + exposed = fuzz it |
| Parser/validator/deserializer | Fuzzing Target | Structured input handling |
| Many callers (>10) + moderate CC | Fuzzing Target | High blast radius |
| Binary/wire protocol handling | Fuzzing Target | Fuzzers excel at format testing |
| Signal | Bucket | Reasoning |
|---|---|---|
| Redundant setup or debug call | False Positive | Statement genuinely unnecessary |
| Cannot map to production function | False Positive | No graph context for triage |
| Call removed, no assertion checks its effect | Missing Tests | Test has weak assertions |
| Assertion removed, test still passes | Missing Tests | Redundant or insufficient coverage |
| Maps to high-CC entrypoint-reachable function | Fuzzing Target | Complex + exposed + weak test |
When both mutation testing and necessist flag the same production function, mark as corroborated — highest confidence finding.
For detailed criteria, see references/triage-methodology.md.
For each mutant, map it to its containing graph node and use pre-analysis subgraphs (tainted, high_blast_radius, privilege_boundary) from Phase 1 to classify it. The classification logic checks: no callers → false positive, privilege boundary → fuzzing, high CC + tainted → fuzzing, high blast radius → fuzzing, otherwise → missing tests.
See references/graph-analysis.md for
the batch_triage implementation and node mapping functions.
Generate a markdown report:
# Genotoxic Triage Report
## Summary
- Total survived mutants: N
- Total necessist removals: N
- Corroborated findings: N
- False positives: N (N%)
- Missing test coverage: N (N%)
- Fuzzing targets: N (N%)
## Corroborated Findings
| File | Line | Function | Mutation Signal | Necessist Signal | Action |
|------|------|----------|----------------|------------------|--------|
## False Positives
| File | Line | Mutation | Reason | Source |
|------|------|----------|--------|--------|
## Missing Test Coverage
| File | Line | Function | CC | Callers | Suggested Test | Source |
|------|------|----------|----|---------|----------------|--------|
## Fuzzing Targets
| File | Line | Function | CC | Entrypoint Path | Blast Radius | Source |
|------|------|----------|----|-----------------|--------------|--------|The Source column is mutation, necessist, or corroborated.
Write the report to GENOTOXIC_REPORT.md in the working directory.
Before delivering:
GENOTOXIC_REPORT.mdtrailmark skill:
property-based-testing skill:
testing-handbook-skills (fuzzing):
harness-writing, cargo-fuzz, atherisFirst-time users: Start with Phase 1 (graph build), then run mutations, then use the Quick Classification table in Phase 3.
Experienced users: Jump to Phase 3 and use the Decision Tree to load specific reference material.
© trailofbits, CC-BY-SA-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references, assets) in plugins/trailmark/skills/genotoxic of trailofbits/skills.
Open the folder on GitHubat commit 82fe822
Mutation Testing Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Mutation Testing Triage this skilltrailofbits/skills | 7.4k | — | ~3.2k | Automated safety check: Pass | CC-BY-SA-4.0 | |
| Test Writing WorkflowiOfficeAI/AionUi | 33k | 1 repos | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Golang Testingantoniopaya22/go-rest-template | 172 | 9 repos | ~4.2k | Automated safety check: Pass | None | |
| Supercovsupercorp-ai/supercov | 148 | 1 repos | ~415 | Automated safety check: Pass | MIT | |
| Coverage Toolszkldi/Tachi | 237 | — | ~580 | Automated safety check: Pass | None | |
| Supercov Securitysupercorp-ai/supercov | 148 | 1 repos | ~236 | Automated safety check: Pass | MIT |
iOfficeAI/AionUi
Sets the test-writing workflow for the repository: risk-first scenario lists, behavior-focused Vitest tests, a full run before each commit and a coverage target.
antoniopaya22/go-rest-template
Go testing patterns including table-driven tests, subtests, benchmarks, fuzzing, and test coverage.
supercorp-ai/supercov
Measures test coverage and code quality in a repository with the supercov CLI, and turns what it finds into small, focused tests or fixes.
zkldi/Tachi
Aggregates Vitest v8/Istanbul coverage across Tachi workspaces via tachi-coverage-tools (manifest, CLI, optional programmatic API).
supercorp-ai/supercov
Scans a repository's source for security vulnerabilities with the supercov CLI, pointing to the line of each finding and mapping it to CWE classes.
ancoleman/ai-design-components
Strategic guidance for choosing and implementing testing approaches across the test pyramid.
trailofbits/skills
Scans a codebase for vulnerabilities with CodeQL's data flow and taint tracking in run-all or important-only modes, including data extensions for project-specific sources and sinks.
trailofbits/skills
Generates Mermaid diagrams from Trailmark code graphs, including call graphs, class hierarchies, module dependency maps, complexity heatmaps and attack surface data flows.
trailofbits/skills
Compares Trailmark code graphs at two snapshots, such as commits, tags or directories, to surface attack paths, blast radius and taint changes that text diffs miss.
trailofbits/skills
Draws a 12 Houses tarot spread to break ties when a request is vague or casually delegated, then reads the cards to pick the next step.
trailofbits/skills
Detects languages, proposes rulesets for approval, then runs the approved Semgrep scan across a codebase and merges the output into one SARIF file.
trailofbits/skills
Searches and extracts data from Burp Suite project files on the command line: regex searches over responses, audit findings, proxy history and site map data.
Works with
Categories
Triages survived mutants and unnecessary test statements using Trailmark call-graph data, sorting them into false positives, missing unit tests and fuzzing targets. After a mutation testing run, the agent combines the survived mutants with necessist results, which flag test statements that can be removed without failing the test, and with a Trailmark code graph. Data flow context then decides where a mutant is harmless, where a unit test would add the most, and which functions would be better served by a fuzz harness.
Mutation Testing Triage fits situations like: triaging survived mutants after a mutation testing run; deciding where new unit tests would have the highest impact; finding functions that need fuzz harnesses instead of unit tests; spotting tests with unnecessary statements that point to weak assertions.
Run `npx skills add trailofbits/skills --skill genotoxic -a claude-code`. Or copy the skill folder (plugins/trailmark/skills/genotoxic in trailofbits/skills) into .claude/skills/genotoxic in your project. Claude Code loads it when a task matches its description.
Run `npx skills add trailofbits/skills --skill genotoxic -a codex`. Or copy the skill folder (plugins/trailmark/skills/genotoxic in trailofbits/skills) into .agents/skills/genotoxic in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add trailofbits/skills --skill genotoxic -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/genotoxic, .gemini/skills/genotoxic, .github/skills/genotoxic and .opencode/skills/genotoxic in your project.
Going by SKILL.md and its folder, Mutation Testing Triage needs the command-line tools its instructions call (uv and cargo). Our summary lists: Trailmark, installed with uv; A mutation testing framework for the target language; A test suite that already passes; necessist, optional, installed with cargo.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Mutation Testing Triage is published under the CC-BY-SA-4.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Mutation Testing Triage: Test Writing Workflow (iOfficeAI/AionUi, 33k stars), Golang Testing (antoniopaya22/go-rest-template, 172 stars), Supercov (supercorp-ai/supercov, 148 stars) and Coverage Tools (zkldi/Tachi, 237 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
trailofbits (a GitHub organization, an official publisher) maintains it in trailofbits/skills, which has 7,400 GitHub stars. The repository holds 79 skills in this directory. The repository was last updated on October 2, 2026.
Source: trailofbits/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.