Agent skill

Review Tests

by bactopia in bactopia/bactopia

Review nf-test run results and present a diagnostic summary with grouped error analysis.

MITAuto-check passedDevOps & Cloud

Install Review Tests

skills CLI
$ npx skills add bactopia/bactopia --skill review-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bactopia/bactopia review-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bactopia/bactopia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/review-tests .claude/skills/review-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review-tests
GitHub stars
522
Token cost
~3.3k tokens
SKILL.md length
1,445 words
Files
2 (incl. scripts)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Review nf-test run results and present a diagnostic summary with grouped error analysis.

  • Works in 4 steps: Run bactopia-review-tests via the… → Present the text output directly to the… → Add interpretation and context after… → …
  • Asked to review tests
  • SKILL.md covers Steps, Status reference (per-cell…, Multi-profile layout &… and Diagnostic files (per profile), plus 6 more sections
  • Runs Shell scripts from its folder; calls bash

What it does

Review Tests is an agent skill from bactopia/bactopia. Review nf-test run results and present a diagnostic summary with grouped error analysis. Use when asked to review tests, check test results, show test failures, analyze test output, investigate why tests failed, see what's broken, or check test status. Runs are multi-profile (docker/conda/singularity) with docker as the reference baseline. Accepts an optional timestamp argument to review a specific run.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/run-bactopia-review-tests.sh`).

It sits in DevOps & Cloud, covering Containers and Failing and flaky tests. It works with Docker. The repository describes itself as: A flexible pipeline for complete analysis of bacterial genomes. The licence is MIT.

When your agent uses it

  • Asked to review tests
  • Check test results
  • Show test failures
  • Analyze test output

Example prompts

  • “/review-tests”

Requirements

  • A Bash shell
  • Docker

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run bactopia-review-tests via the wrapper script using the default text output
  2. Present the text output directly to the user. The CLI produces a clean summary with a
  3. Add interpretation and context after showing the output. Interpret by status
  4. If the text output is too large for a single response, summarize the key sections

What it can do on your machine

Read from SKILL.md and the folder at commit 29fb741. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Review Tests loads about 3.3k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 1,445 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bactopia/bactopia at commit 29fb741, republished under its MIT licence (© bactopia). 1,445 words, ~3,266 tokens.

Download SKILL.mdSave it as .claude/skills/review-tests/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
review-tests
description
Review nf-test run results and present a diagnostic summary with grouped error analysis. Use when asked to review tests, check test results, show test failures, analyze test output, investigate why tests failed, see what's broken, or check test status. Runs are multi-profile (docker/conda/singularity) with docker as the reference baseline. Accepts an optional timestamp argument to review a specific run.

Review Tests

Run the review-tests CLI and present the results to the user.

Runs are multi-profile: each component/tier is tested across up to four profiles -- docker (the reference baseline), conda, singularity_galaxy, singularity_pull. A cell is one (component, tier, profile). Most interpretation is a comparison of each profile against docker.

Steps

  1. Run bactopia-review-tests via the wrapper script using the default text output (do NOT use --json):

    bash .agents/skills/review-tests/scripts/run-bactopia-review-tests.sh --bactopia-path /home/rpetit3/repos/bactopia/bactopia --silent

    If the user provided a timestamp argument (e.g., /review-tests 20260324_081306), add --run 20260324_081306.

  2. Present the text output directly to the user. The CLI produces a clean summary with a "Status Breakdown by Profile" matrix and one section per failing status. Do NOT parse JSON or write extra code to reformat -- just relay the output with your interpretation.

  3. Add interpretation and context after showing the output. Interpret by status (see the status reference below), and always frame failures as "which profiles differ from docker, and why." Summarize actionable items and next steps.

  4. If the text output is too large for a single response, summarize the key sections (overview, status-by-profile matrix, failures) and note that per-file detail is in summary.json. Use --json/--pretty or read summary.json directly for structured detail.

Status reference (per-cell status)

  • passed -- cell matched the committed snapshot / assertions.
  • version_drift / output_drift / version+output_drift -- this profile's outputs differ from docker: version_drift = versions.yml (runtime resolved a different tool version than the docker container pin), output_drift = tool output file(s), version+output_drift = both. reason names the divergent fields. When docker passed and only conda/singularity drift, this is genuine dependency-solve divergence, not a bug. Fix = the sccmec-pattern test migration (md5 the profile-stable files, existence-check the divergent ones, versions -> contains('<tool>')) -- NOT snapshot regeneration. See files[] / suggested_edit in summary.json for the exact bucketing. (A pure version_drift may instead warrant updating the container version pin.)
  • snapshot_mismatch -- the snapshot didn't match but the files matrix could not attribute it to specific fields (drift not subclassified). Inspect files[] and {profile}/stderr.txt.
  • snapshot_stale -- the committed .snap no longer matches the reference runtime (docker); it shows on all profiles including docker. Fix: re-run with --generate under docker to re-baseline. NOT a content or tool problem (on generate=false, docker's own mismatch is promoted to this). Typical cause: a test's snapshot() shape was edited but .snap was never regenerated.
  • assertion_failed -- a non-snapshot assertion failed (no output divergence detected). A test logic/assertion issue, not drift -- read {profile}/stderr.txt.
  • non_reproducible -- two docker runs produced different snapshots; docker's own output is non-deterministic. Investigate the tool/test; regeneration will not fix it.
  • build_failed -- the Conda env or Singularity image failed to build before testing. Infra: build the env/image, then re-run (it blocks triage of that profile).
  • no_ground_truth -- docker established no snapshot for the non-docker profiles to validate against (usually docker itself failed to produce one).
  • syntax_error -- the Nextflow script failed to compile. Fix the .nf.
  • Housekeeping statuses you may also see: skipped, timeout (exceeded the per-run timeout), no_snapshot, n/a.
  • undeclared_outputs -- the tool produced files not declared in the module's results, logs, versions, or nf_logs. For each file help the user route it:
    • results: a real tool output users want (report, summary, data file)
    • logs: stdout/stderr from the tool
    • .outputs-ignore: staging artifact, intermediate, version-info side effect, or DB file .outputs-ignore lives at modules/{name}/tests/.outputs-ignore (one glob per line; # comments and blanks allowed; staging/** is ignored by default). NOTE: this check only runs when the tool succeeds, so it is masked on a profile that tool_error'd -- use undeclared_outputs_union to see the full set.
  • tool_error -- the tool crashed at runtime. Read error_class:
    • env_dependency -- conda/singularity re-solved a too-new interpreter/dependency (e.g. py>=3.12 pkg_resources, numpy2 newshape, biopython SeqFeature.strand, R readr/lifecycle deprecate_stop). Fix the env/recipe, NOT the test.
    • tool_crash / staging_bug / fs_permission / unknown -- fix the module/upstream or the workspace; when unknown, read the Command error: block. reason carries the real tool error (from Command error:), not the downstream nf-test NullPointerException.

generate gates interpretation (shown in Run Parameters and the # generate=<bool> header of summary.tsv):

  • generate=true: the .snap was regenerated under docker first, so docker passing is the re-baseline; any drift shown is genuine (docker vs profile). snapshot_stale cannot occur.
  • generate=false: docker also failing => snapshot_stale (run --generate). docker passing while a profile drifts => genuine content drift.

Multi-profile layout & summary.json

Structured results live at logs/run-tests/{ts}/summary.json (plus summary.tsv, whose first line is # generate=<bool>). Prefer summary.json for machine-readable detail; the CLI text is the human summary. Key fields:

  • profiles[], reference_profile ("docker").
  • results[].cells.{profile}: status, duration, reason, error_class (tool_error only), undeclared_outputs[].
  • results[].undeclared_outputs_union: undeclared files unioned across profiles (unmasks profiles that tool_error'd).
  • results[].files[]: per output file, the cross-profile md5 matrix -- process, scope (sample/run; subworkflow multi-record), field, name, md5:{profile -> hash|null}, verdict, divergent_profiles[], plus:
    • verdict: stable (equal across all profiles that ran) | divergent | indeterminate (a profile didn't produce it) | skip.
    • comparable: false = intrinsically non-hashable (gz / normalized -> byte md5 is meaningless) => bucket existence-only; verdict:"skip".
    • incomplete[]: profiles that produced no file (e.g. a tool_error'd conda) => verdict is indeterminate, NOT a false stable; re-check after fixing that profile.
    • kind:"versions" + tool_key: a versions.yml -> bucket to contains('<tool_key>'). This matrix is computed from actual runtime outputs, so it is populated even on passing or stale cells -- an always-on divergence diagnostic (also useful for add-* at creation time).
  • results[].suggested_edit (module/subworkflow only): the exact test change implied by the verdicts -- snapshot:[fields] (stable), existence:[fields] (divergent content), contains:[{field,value}] (divergent versions). Subworkflow fields are scope-prefixed (sample./run., e.g. sample.blast, run.versions). Directly consumable and self-verifying (diff against the committed test). Workflow tier is intentionally .nftignore-only, so it has no suggested_edit; add the divergent globs to workflows/**/tests/.nftignore instead.
Show full SKILL.md (555 more words)Show less

Diagnostic files (per profile)

Layout: logs/run-tests/{ts}/{tier}/{component}/{profile}/:

  • stdout.txt -- nf-test console, including the tool's own Command error: block. Read this for tool_error root cause.
  • stderr.txt -- nf-test assertions, including the Different Snapshot per-file md5 diff. Read this for drift / assertion detail.
  • outputs.txt -- # Undeclared outputs: list, or # OK.
  • .nf-test/** -- preserved work tree (present for all cells, passing included), including meta/output_0.json (record field -> output file paths) and meta/nextflow.log.

Both stdout.txt and stderr.txt matter now, split by class (this replaces the old "read stdout, not stderr" rule).

Progressive Disclosure

Keep the initial summary compact and scannable. Do NOT open stdout/stderr/nextflow.log during the initial summary -- the status matrix, reason, error_class, and files[] usually suffice. When the user asks for deeper detail:

  • Specific component: read its summary.json results[] entry first (cells, reason, error_class, files[], suggested_edit). Then, if needed, open {tier}/{component}/{profile}/stdout.txt (tool_error) or that same dir's stderr.txt (drift diff).
  • Undeclared outputs: use undeclared_outputs_union (or a cell's undeclared_outputs[]), then read the module's main.nf output block to advise results / logs / .outputs-ignore.
  • Tool / abort errors: read the Command error: block in {profile}/stdout.txt; the full Nextflow log is at {tier}/{component}/{profile}/.nf-test/tests/*/meta/nextflow.log (focus on ERROR/WARN and the last ~50 lines).
  • Drift bucketing: use files[] + suggested_edit; cross-check with the Different Snapshot block in {profile}/stderr.txt.

Important Reminders

  • CRITICAL: NEVER suggest --update-snapshots / snapshot regeneration for the drift statuses (output_drift/version_drift/version+output_drift) or env-drift tool_errors. Regen does NOT fix profile divergence -- migrate the test (sccmec pattern) or fix the env. --generate is the fix only for snapshot_stale.
  • error_class: env_dependency => fix the conda/singularity env or bioconda recipe, NOT the test.
  • Read {profile}/stdout.txt for tool_error root cause and {profile}/stderr.txt for drift/assertion diffs -- both matter.
  • undeclared_outputs can be masked on a tool_error'd profile -- always check undeclared_outputs_union.
  • files[] verdicts: comparable:false (gz/normalized) -> existence-only; verdict:indeterminate
    • incomplete:[...] -> a profile didn't run, re-check after fixing (never a clean bill).
  • Not all tiers/profiles appear in every run; a component with no Galaxy image has galaxy:false and no singularity_galaxy cell.
  • .nf-test/ work dirs are preserved per profile for all cells (passing included), so you can inspect any profile's meta/output_0.json or work tree -- not just failures.

Updating Baselines

Baselines file: conf/test-times.json. Durations are docker-based (the CLI reports "Docker duration").

To update baselines after a clean all-pass run, add --update-baselines:

bash .agents/skills/review-tests/scripts/run-bactopia-review-tests.sh --bactopia-path /home/rpetit3/repos/bactopia/bactopia --silent --update-baselines

This writes actual runtimes from the current run into the baselines file and updates the _meta.updated timestamp. Only entries for tested components are updated; other tiers are left unchanged. After updating, re-run without --update-baselines to confirm anomalies are resolved.

Interpreting Timing Anomalies

Timing is measured against the docker profile.

  • generate=true vs generate=false: a generate=true run executes tests twice (generate snapshots, then test against them). If baselines were recorded from a generate=true run but the current run uses generate=false, tests run at ~0.5x baseline -- expected, not suspicious.
  • Slow tests: may reflect newly added test cases rather than regressions. Check recent commits to the component's test file before flagging.
  • Only flag anomalies as concerning when the generate parameter matches between the baseline run and the current run.

Self-Improvement

If you find yourself writing ad-hoc Python or bash to parse, explore, or extract data from the CLI output or summary.json, that logic should be added to this skill or the underlying bactopia-review-tests CLI instead. Update the skill so future sessions don't reinvent it.

JSON Output

logs/run-tests/{ts}/summary.json is the primary structured source (schema above). The CLI can also emit it with --json (add --pretty for readable output). See bactopia-review-tests --help for details.

© bactopia, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .agents/skills/review-tests of bactopia/bactopia.

  • SKILL.md
  • scripts/run-bactopia-review-tests.sh

Open the folder on GitHubat commit 29fb741

Compare with similar skills

Review Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Review Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Review Tests this skillbactopia/bactopia522—~3.3kAutomated safety check: PassMIT
Swig CI Reproswig/swig6.3k—~1.2kAutomated safety check: PassCustom licence
Troubleshootserithemage/serverless-openclaw196—~1.6kAutomated safety check: NotesNone
Debug CIweb-infra-dev/rslint460—~2.8kAutomated safety check: PassMIT
Iron Proxy Gateway for NanoClawnanocoai/nanoclaw31k—~4.6kAutomated safety check: NotesMIT
GreptimeDB Dev Docker ImageGreptimeTeam/greptimedb6.7k—~4kAutomated safety check: NotesApache-2.0

Similar skills

  • Swig CI Repro

    swig/swig

    Reproduce a GitHub Actions Linux CI failure locally when it does not happen on your machine: a podman/docker image that mirrors the ubuntu-22.04 runner by reusing the real Tools/CI-linux-.sh install…

    6.3k GitHub stars~1.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Troubleshoot

    serithemage/serverless-openclaw

    Troubleshoots common issues. An agent skill from serithemage/serverless-openclaw.

    196 GitHub stars~1.6k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check: notes
  • Debug CI

    web-infra-dev/rslint

    Reproduce Linux CI failures locally using Docker when the same tests pass on the host, especially Go platform differences and VS Code extension tests requiring xvfb.

    460 GitHub stars~2.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Installs or refreshes Iron Proxy and its Iron Control web console for NanoClaw, with a local Docker setup, database, credentials and a human approval bridge.

    31k GitHub stars~4.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • GreptimeDB Dev Docker Image

    GreptimeTeam/greptimedb

    Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.

    6.7k GitHub stars~4k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    260 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes

More from bactopia/bactopia

All 15 skills in this repo
  • Add Bactopia Tool

    bactopia/bactopia

    Scaffold a complete Bactopia Tool across all three tiers -- module, subworkflow, and workflow entry point under workflows/bactopia-tools/.

    522 GitHub stars~4.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Bump Versions

    bactopia/bactopia

    Propagate the Bactopia and nf-bactopia versions declared in versions.yml into the hand-maintained files that carry a literal version (conf/testbase.config, CITATION.cff, bin/bactopia…

    522 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Merge Schemas

    bactopia/bactopia

    Regenerate nextflow.config and nextflowschema.json for Bactopia workflows by running bactopia-merge-schemas.

    522 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Project Status

    bactopia/bactopia

    Show a live snapshot of the Bactopia project state — component counts, GroovyDoc coverage, nf-test coverage, and structural issues.

    522 GitHub stars~787 tokensUpdated 2 mo ago
    Auto-check passed
  • Release Checklist

    bactopia/bactopia

    Audit whether Bactopia is ready for a version release and produce a GO / NO-GO recommendation report.

    522 GitHub stars~4.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Review Citations

    bactopia/bactopia

    Review citation integrity across data/citations.yml and @citation tags using bactopia-citations --validate.

    522 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Categories

Questions about Review Tests

What does Review Tests do?

Review nf-test run results and present a diagnostic summary with grouped error analysis. Review Tests is an agent skill from bactopia/bactopia. Review nf-test run results and present a diagnostic summary with grouped error analysis.

When should I use Review Tests?

Review Tests fits situations like: asked to review tests; check test results; show test failures; analyze test output.

How do I install Review Tests in Claude Code?

Run `npx skills add bactopia/bactopia --skill review-tests -a claude-code`. Or copy the skill folder (.agents/skills/review-tests in bactopia/bactopia) into .claude/skills/review-tests in your project. Claude Code loads it when a task matches its description.

How do I install Review Tests in Codex?

Run `npx skills add bactopia/bactopia --skill review-tests -a codex`. Or copy the skill folder (.agents/skills/review-tests in bactopia/bactopia) into .agents/skills/review-tests in your project. Codex loads it when a task matches its description.

Can I use Review Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bactopia/bactopia --skill review-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-tests, .gemini/skills/review-tests, .github/skills/review-tests and .opencode/skills/review-tests in your project.

What does Review Tests need to run?

Going by SKILL.md and its folder, Review Tests needs a shell for the scripts in its folder and the command-line tools its instructions call (bash). Our summary lists: A Bash shell; Docker.

Does Review Tests access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Review Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Review Tests use?

Review Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Review Tests use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Review Tests?

Skills that share tags, products or a category with Review Tests: Swig CI Repro (swig/swig, 6.3k stars), Troubleshoot (serithemage/serverless-openclaw, 196 stars), Debug CI (web-infra-dev/rslint, 460 stars) and Iron Proxy Gateway for NanoClaw (nanocoai/nanoclaw, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Review Tests?

bactopia (a GitHub organization) maintains it in bactopia/bactopia, which has 522 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on August 5, 2026.

Source: bactopia/bactopia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.