Agent skill

Bio Workflow Management Snakemake Workflows

by GPTomics in GPTomics/bioSkills

Authors reproducible bioinformatics pipelines with Snakemake - rules wired by output-file pattern, wildcards and expand() for sample fan-out, checkpoints for runtime-unknown outputs, resource/retry…

MITAuto-check passedResearch & Science

Install Bio Workflow Management Snakemake Workflows

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-workflow-management-snakemake-workflows -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-workflow-management-snakemake-workflows --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workflow-management/snakemake-workflows .claude/skills/bio-workflow-management-snakemake-workflows && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-workflow-management-snakemake-workflows
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.9k tokens
SKILL.md length
1,959 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Authors reproducible bioinformatics pipelines with Snakemake - rules wired by output-file pattern, wildcards and expand() for sample fan-out, checkpoints for runtime-unknown outputs, resource/retry…

  • Deciding rule-based (Snakemake) vs channel/dataflow (Nextflow) authoring
  • SKILL.md covers Version Compatibility, The governing principle:…, Decision: Snakemake vs… and Decision: rerun triggers - why…, plus 11 more sections
  • Runs R scripts from its folder; calls pip
  • Wiring rules by OUTPUT-file pattern rather than imperative order

What it does

Bio Workflow Management Snakemake Workflows is an agent skill from GPTomics/bioSkills. Authors reproducible bioinformatics pipelines with Snakemake - rules wired by output-file pattern, wildcards and expand() for sample fan-out, checkpoints for runtime-unknown outputs, resource/retry escalation, and conda/container software deployment on HPC and cloud. Use when deciding rule-based (Snakemake) vs channel/dataflow (Nextflow) authoring; wiring rules by OUTPUT-file pattern rather than imperative order; using wildcards + expand() for sample fan-out and constraining them to stop silent mis-routing…

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Reproducible research. It works with Nextflow and Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Deciding rule-based (Snakemake) vs channel/dataflow (Nextflow) authoring
  • Wiring rules by OUTPUT-file pattern rather than imperative order
  • Using wildcards + expand() for sample fan-out and constraining them to stop silent mis-routing
  • Adding checkpoints when the set of outputs is unknown until a step runs (dynamic DAG)

Example prompts

  • “Use the bio-workflow-management-snakemake-workflows skill to author reproducible bioinformatics pipelines with Snakemake - rules wired by…”
  • “/bio-workflow-management-snakemake-workflows”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Workflow Management Snakemake Workflows loads about 4.9k tokens when it runs. Until then it costs about 243 tokens; SKILL.md has 1,959 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~243
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,959 words, ~4,930 tokens.

Download SKILL.mdSave it as .claude/skills/bio-workflow-management-snakemake-workflows/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-workflow-management-snakemake-workflows
description
Authors reproducible bioinformatics pipelines with Snakemake - rules wired by output-file pattern, wildcards and expand() for sample fan-out, checkpoints for runtime-unknown outputs, resource/retry escalation, and conda/container software deployment on HPC and cloud. Use when deciding rule-based (Snakemake) vs channel/dataflow (Nextflow) authoring; wiring rules by OUTPUT-file pattern rather than imperative order; using wildcards + expand() for sample fan-out and constraining them to stop silent mis-routing; adding checkpoints when the set of outputs is unknown until a step runs (dynamic DAG); diagnosing why a job reran (or did not) under the mtime-plus-provenance trigger set; escalating memory on retry for OOM-killed jobs; and porting a Snakemake 7 `--cluster`/remote-provider command to the Snakemake 8+ executor-plugin and storage-plugin model (snakemake-executor-plugin-slurm) with `--software-deployment-method`.
tool_type
python
primary_tool
Snakemake
goal_approach_exempt
true

Version Compatibility

Reference examples tested with: Snakemake 8.0+, Python 3.11+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Note: Snakemake 8 (Jan 2024) removed --cluster, --drmaa, and the *RemoteProvider classes from core and moved them to pip-installable EXECUTOR plugins (--executor slurm, package snakemake-executor-plugin-slurm) and STORAGE plugins (storage.s3(...), snakemake-storage-plugin-s3). --use-conda/--use-singularity became --software-deployment-method / --sdm conda apptainer. A Snakemake 7 command line does not run unchanged on 8/9. Run snakemake --version FIRST and branch all execution guidance on 7 vs 8/9.

Snakemake Workflows

"Build a reproducible bioinformatics pipeline with Snakemake" -> Declare each step as a rule that says "a file matching THIS output pattern is produced FROM those inputs", let the engine resolve the DAG backward from requested targets, fan out over samples with wildcards, and pin the software environment so the result reproduces next year.

  • Python: Snakefile rule/checkpoint blocks with expand(), wildcards, config, resources, and conda:/container: (Snakemake)

The governing principle: Snakemake is pull/goal-oriented - it builds a STATIC DAG backward from requested target files

Snakemake is a pull, make-like engine. An author does NOT describe a forward flow of data. Each rule is a pattern-matched recipe ("a file that looks like THIS can be produced FROM that"), and the engine takes the requested target files and works BACKWARD, unifying wildcards by string-matching output filename patterns, until it reaches files already on disk. The whole plan - a static DAG - is computed at parse time, before a single job runs (Köster & Rahmann 2012 Bioinformatics 28:2520-2522). Almost every Snakemake bug a biologist hits is a downstream consequence of this one model:

  • Rules are wired by OUTPUT-FILE PATTERN, not call order. A missing or typo'd output path silently drops a rule from the DAG - there is no error, the job just never runs. Debugging "why didn't it run" means tracing the backward dependency from the target, not reading top-to-bottom.
  • Because the plan is fully known up front, snakemake -n (dry run), --dag, and --report are first-class. This is the payoff of the static model. Nextflow's reactive-dataflow model (processes connected by asynchronous channels, DAG emerges at runtime) has no true dry-run - hold both models in mind and most "why did/didn't it run" questions answer themselves.
  • Data-dependent branching is impossible in the base model. If the NUMBER or identity of outputs is unknown until a step runs (split into one file per detected cluster, scatter over however many contigs an assembler emits), the static DAG cannot represent it -> that is exactly what CHECKPOINTS exist for. A biologist who thinks "the pipeline decides at runtime how many chunks" is fighting the paradigm and needs a checkpoint, not a clever run: block.
  • A workflow manager buys reproducible LOGIC and nothing else automatically. The DAG being deterministic says nothing about tool versions. A rule with no conda:/container: runs against whatever is on $PATH; "reproducible" is unearned until the software environment is pinned (Grüning et al. 2018 Cell Syst 6:631-635). Pin containers by DIGEST and conda by LOCKFILE - see Software Deployment below.

Decision: Snakemake vs Nextflow (pick by team and infrastructure, not benchmarks)

DimensionSnakemakeNextflowBest when
Modelpull/make, static DAG at parse timepush/reactive dataflow, dynamic DAGSnakemake: the plan must be visible before committing an allocation
LanguagePython DSL (real Python + pandas in the Snakefile)Groovy DSLSnakemake: Python-native lab, file-pattern logic
Dry run / DAG vizfirst-class (-n, --dag, --report)no true dry-run (-stub/-preview check wiring only)Snakemake: HPC where a bad plan is expensive
Data-dependent branchingneeds checkpoints (escape hatch)native (channels)Nextflow: shape depends on runtime data
Community pipelinesWorkflow Catalog / wrappers (smaller)nf-core (large, curated)Nextflow: run a maintained pipeline as-is
Sweet spotsingle-lab reproducible research, HPC, tight Python integrationcloud/production, multi-institution, nf-core stackschoose by the ecosystem to integrate with

Reuse before authoring: for a mainstream analysis (RNA-seq, variant calling, ATAC-seq), a curated community pipeline already encodes years of QC and edge cases. Adopting one means RUNNING it (e.g. nf-core/rnaseq via workflow-management/nf-core-pipelines), not authoring Groovy - so a Python-shop preference for Snakemake only decides the authoring case, not whether to build at all. Author from scratch only for a novel method or an unsupported combination of steps.

Decision: rerun triggers - why a job reran, or did not

Since Snakemake 7.8 the default is NOT pure mtime. A rerun fires on a SET of triggers: {mtime, params, input, code, software-env} (Mölder et al. 2021 F1000Research 10:33). This surprises everyone upgrading from old Snakemake.

WantUse
classic Make behavior, minimize surprise reruns--rerun-triggers mtime
max reproducibility (default)all five triggers
ignore a stable reference's timestampancient("ref.fa") on that input
mark results current without recompute--touch
force specific rules--forcerun rule / -R
see what WOULD rerun and whysnakemake -n -R / --list-changes code

The code trigger catches shell/script/run body changes - reformatting whitespace or editing a comment counts as a code change and reruns the job. On very large DAGs or multi-TB inputs the provenance triggers add a hashing/stat storm; --rerun-triggers mtime skips it.

Decision: run vs script vs shell vs notebook vs wrapper

SituationPickWhy
call a CLI tool (samtools, bwa)shell:subprocess, conda/container-isolated
reusable Python/R analysis needing isolationscript:separate process, snakemake object injected
standard tool, do not want to write shellwrapper: (PINNED tag)maintained, ships its own env
exploratory, want a re-runnable notebooknotebook:params injected, --edit-notebook
trivial in-Snakefile glue onlyrun:NEVER heavy work

run: executes IN the main Snakemake process - it shares the interpreter and GIL, cannot be conda/container-isolated (the conda: directive is disallowed with run:), blocks the scheduler, and an OOM in it takes down the whole workflow. Move anything beyond trivial glue to script:.

Decision: execution backend (Snakemake 8/9)

TargetCommand
laptop/workstationsnakemake --cores N --sdm conda
SLURM (native)pip install snakemake-executor-plugin-slurm then --executor slurm --jobs N --default-resources
SLURM (legacy sbatch string)pip install snakemake-executor-plugin-cluster-generic then --executor cluster-generic --cluster-generic-submit-cmd "sbatch ..."
S3/GCS I/Opip install snakemake-storage-plugin-s3 then --default-storage-provider s3 --default-storage-prefix s3://.../
thousands of tiny jobsadd group: / --group-components to collapse scheduler overhead

Porting a v7 --cluster "sbatch --account=X --partition=Y --mem=Z --time=T" command: for the native slurm executor, map those sbatch flags to resource keys (slurm_account, slurm_partition, mem_mb, runtime in minutes) set per-rule in resources: or globally via --default-resources; for a drop-in port keep the old string under cluster-generic (its own plugin, above). --cores = local cores; --jobs/-j = number of concurrent cluster/cloud jobs (in v8 these are separate). Profiles are versioned: the file is config/config.v8+.yaml, every long option becomes a YAML key.

Rules, wildcards, and expand

expand() returns a LIST of strings by combinatorial substitution - it does NOT touch the filesystem. Use it to enumerate targets in rule all.

python
configfile: 'config/config.yaml'
SAMPLES = config['samples']

rule all:                                              # the requested targets; the DAG is built backward from here
    input:
        expand('results/{sample}.bam', sample=SAMPLES)

rule align:
    input:
        r1 = 'data/{sample}_R1.fq.gz',
        r2 = 'data/{sample}_R2.fq.gz',
        index = 'ref/genome.fa'
    output:
        bam = 'aligned/{sample}.bam'                   # this OUTPUT PATTERN, matched against the target, wires the rule in
    threads: 8
    log:
        'logs/align/{sample}.log'
    shell:
        'bwa mem -t {threads} {input.index} {input.r1} {input.r2} | '
        'samtools sort -@ {threads} -o {output.bam} 2> {log}'

Wildcard constraints - stop silent mis-routing

Wildcards are greedy regex string-unification ({sample} compiles to .+), not typed parameters. An unconstrained wildcard swallows path separators and adjacent tokens: data/{sample}.txt matches data/a/b.txt as sample=a/b, and {a}.{b}.txt on 101.B.normal.txt has no unique parse. The failure is silent mis-routing, not an error. Constrain whenever a value can contain /, ., or _, or a filename has multiple variable tokens.

python
wildcard_constraints:
    sample = '[^/]+',                                  # no path separators
    chrom = r'\d+|X|Y|MT'                              # only real chromosome tokens

# two rules whose output patterns can both produce a requested file raise AmbiguousRuleException;
# prefer non-overlapping constraints to disambiguate, and fall back to `ruleorder: a > b` only if needed.
Show full SKILL.md (771 more words)Show less

Checkpoints - the ONLY data-dependent-DAG mechanism

Goal: produce downstream jobs for a set of files whose number and identity are unknown until a step runs (split a FASTA into one file per detected cluster; scatter over an assembler's contigs).

Approach: declare the producing step a checkpoint with a directory() output; in an input function on the AGGREGATING rule, call checkpoints.<name>.get(**wildcards) FIRST - its exception is what forces the engine to run the checkpoint and RE-EVALUATE the DAG - then glob_wildcards the checkpoint's declared output dir and expand() the real targets.

python
checkpoint split_fasta:
    input:
        'data/all.fasta'
    output:
        directory('split/{sample}')                    # directory() because the file set is unknowable at parse time
    shell:
        'split_by_cluster.py {input} split/{wildcards.sample}'

def gather_clusters(wildcards):
    # .get() RAISES until the checkpoint has run; that exception drives DAG re-evaluation.
    # Omitting it globs at parse time (empty), so the aggregation silently gets zero inputs - the classic bug.
    ckpt_dir = checkpoints.split_fasta.get(**wildcards).output[0]
    ids = glob_wildcards(f'{ckpt_dir}/{{id}}.fasta').id
    return expand('processed/{sample}/{id}.done', sample=wildcards.sample, id=ids)

rule aggregate:                                         # the input function MUST be attached to the rule that consumes the set
    input:
        gather_clusters
    output:
        'results/{sample}_summary.txt'
    shell:
        'cat {input} > {output}'

Point glob_wildcards at the checkpoint's declared directory() output (a fresh dir) so stale files do not leak into the glob. Prefer one scatter->gather to chains of nested checkpoints.

Resources, escalating retries, and grouping

Resource callables differ by directive: resources is callable(wildcards [, input] [, threads] [, attempt]); threads is callable(wildcards [, input]) only; getting the signature wrong is a top error source. runtime is in MINUTES. attempt starts at 1 and increments per retry - the canonical fix for OOM-killed jobs.

python
rule call_variants:
    input:
        bam = 'aligned/{sample}.bam'
    output:
        'results/{sample}.vcf'
    threads: 4
    retries: 3                                          # or global --retries 3
    resources:
        mem_mb = lambda wildcards, attempt: 8000 * attempt,   # 8 GB, doubling to 16/24 on OOM restart
        runtime = 240                                   # MINUTES, not seconds; SLURM wall-time
    log:
        'logs/call/{sample}.log'
    shell:
        'variant_caller --threads {threads} {input.bam} > {output} 2> {log}'

For thousands of tiny jobs on HPC, per-job scheduler latency dominates: assign rules a group: (or --group-components rule=N) so they submit as one job. temp('x.bam') deletes an intermediate once all consumers are done (huge for disk); ancient('ref.fa') excludes an input from mtime-based rerun decisions.

Software deployment - the layer that makes it reproducible

A clean DAG over unpinned tools is not reproducible. The engine gives layer 1 (logic); the author must pin the software environment. Declare conda: (a pinnable file) or container: per rule, and activate deployment at run time.

python
rule fastqc:
    input:
        'data/{sample}.fq.gz'
    output:
        'qc/{sample}_fastqc.html'
    conda:
        'envs/qc.yaml'                                  # a FILE (pinnable), not a bare named env
    container:
        # PIN BY DIGEST, never a mutable tag - :latest or a re-pushed :0.7.17 silently changes the tool and busts the cache
        'docker://quay.io/biocontainers/fastqc@sha256:<digest>'
    shell:
        'fastqc {input} -o qc/'
bash
snakemake --sdm conda --cores 8                         # build per-rule conda envs (was --use-conda in v7)
snakemake --sdm apptainer --cores 8                     # run each rule in its container (was --use-singularity)
snakemake --sdm conda apptainer --cores 8               # containerized conda: build the env INSIDE the pinned image

A bare environment.yml with samtools (no version) resolves differently over time; pin exact builds with a lockfile (conda-lock) for bit-reproducibility. --containerize auto-generates a Dockerfile baking all conda envs into one image. Between-workflow caching (cache: True + SNAKEMAKE_OUTPUT_CACHE) reuses results across workflows but ONLY for deterministic rules - a nondeterministic tool poisons the shared cache with wrong results silently.

Modularization and reuse

include: 'rules/x.smk' is textual inclusion sharing one namespace. For genuine composition of published workflows, use the module system (module other: snakefile: '...'; use rule * from other as other_*), which can import, prefix, and override rules. wrapper: 'v5.0.2/bio/bwa/mem' pulls a maintained, conda-shipping wrapper - PIN the leading version tag; an unpinned wrapper drifts silently.

python
include: 'rules/qc.smk'
include: 'rules/align.smk'

rule all:
    input:
        rules.qc_all.input,
        rules.call_all.input

Common Errors

SymptomCauseFix
A rule silently never runsits output pattern does not match any requested target (typo/path); the DAG dropped ittrace backward from rule all; run snakemake -n and inspect the DAG; fix the output path
Wildcard captures too much / wrong sampleunconstrained greedy .+ swallowed a delimiter or path separatoradd wildcard_constraints (e.g. sample='[^/]+')
Aggregation after a split has zero inputsforgot checkpoints.X.get(), or globbed at parse time, or input function on the wrong rulecall .get(**wildcards) first, glob the checkpoint's directory() output, attach the function to the consumer
Everything reruns after a cosmetic editthe code trigger - editing the shell/script body (even whitespace) countsexpected under provenance triggers; use --rerun-triggers mtime to opt out
Expected a rerun, got noneold mtime mental model, or output is newer than inputcheck triggers; --forcerun rule / -R
--cluster/S3RemoteProvider errors on v8removed from core in Snakemake 8install the executor/storage plugin; use --executor slurm and storage.s3(...)
OOM-killed job (exit 137)static mem_mb too low for the largest sampleretries + mem_mb=lambda wildcards, attempt: base*attempt
Per-job scheduler meltdown on HPCthousands of tiny jobs, submission overhead dominatesgroup: / --group-components to batch into one submission
MissingOutputException on NFS/Lustrenetworked filesystem lags after a job finishesraise --latency-wait
Workflow refuses to continue after a killed job ("Incomplete files")a job died mid-write (SLURM kill, node crash), so its outputs are flagged incompletere-run with --rerun-incomplete (--ri); this is Snakemake's crash-resume, distinct from the rerun triggers
Heavy run: block hangs the workflowruns in the main process, shares the GIL, no isolationmove to script:
"Reproducible" but results differ on a colleague's clusterno --sdm, or a mutable :latest container tagdeclare conda:/container:, pin by digest + conda lockfile
  • workflow-management/nextflow-pipelines - The reactive-dataflow alternative; author here when the DAG must emerge from runtime data
  • workflow-management/nf-core-pipelines - Run a curated community Nextflow pipeline instead of authoring from scratch
  • workflows/rnaseq-to-de - End-to-end RNA-seq-to-differential-expression pipeline this engine can orchestrate
  • read-qc/quality-reports - The QC step a pipeline wraps as an early rule
  • read-alignment/bwa-alignment - The alignment tool invoked inside a Snakemake shell: rule

References

  • Köster J, Rahmann S. 2012. Snakemake - a scalable bioinformatics workflow engine. Bioinformatics 28(19):2520-2522.
  • Mölder F, Jablonski KP, Letcher B, et al. 2021. Sustainable data analysis with Snakemake. F1000Research 10:33.
  • Grüning B, Chilton J, Köster J, et al. 2018. Practical computational reproducibility in the life sciences. Cell Systems 6(6):631-635.
  • Di Tommaso P, Chatzou M, Floden EW, et al. 2017. Nextflow enables reproducible computational workflows. Nature Biotechnology 35(4):316-319.

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in workflow-management/snakemake-workflows of GPTomics/bioSkills.

  • SKILL.md
  • examples/rnaseq_snakefile
  • examples/scripts/run_deseq2.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Workflow Management Snakemake Workflows next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Workflow Management Snakemake Workflows compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Workflow Management Snakemake Workflows this skillGPTomics/bioSkills1.2k1 repos~4.9kAutomated safety check: PassMIT
LaminDB Biological Data Managementdavila7/claude-code-templates32k12 repos~3.6kAutomated safety check: PassMIT
Latchbio Integrationdavila7/claude-code-templates32k11 repos~2.4kAutomated safety check: PassMIT
Latchbio IntegrationK-Dense-AI/scientific-agent-skills48k1 repos~2.5kAutomated safety check: NotesMIT
Add Bactopia Toolbactopia/bactopia522—~4.1kAutomated safety check: PassMIT
Modeling Code and Result Contractsyushui2022/MathModel-Skill453—~1.4kAutomated safety check: PassMIT

Similar skills

  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    32k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Latchbio Integration

    davila7/claude-code-templates

    Latch platform for bioinformatics workflows. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~2.4k tokens
    Research & ScienceAuto-check passed
  • Latchbio Integration

    K-Dense-AI/scientific-agent-skills

    Builds, registers, debugs, and operates bioinformatics workflows on Latch using the Python SDK, CLI, Latch Data and Registry, Nextflow, Snakemake, programmatic execution, and Latch MCP.

    48k GitHub starsUsed in 1 repo~2.5k tokens
    Research & ScienceAuto-check: notes
  • Add Bactopia Tool

    bactopia/bactopia

    Scaffold a complete Bactopia Tool across all three tiers -- module, subworkflow, and workflow entry point under workflows/bactopia-tools/.

    522 GitHub stars~4.1k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Modeling Code and Result Contracts

    yushui2022/MathModel-Skill

    Generates result-evidence contracts, tables and runnable q1 to q3 modeling code scaffolds for a math modeling paper from a model route, a data plan and cleaned data.

    453 GitHub stars~1.4k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • HypoGeniC Hypothesis Generation

    K-Dense-AI/scientific-agent-skills

    Plans and audits runs of the HypoGeniC and HypoRefine packages, which propose hypotheses from labeled text datasets, with local checks before any model call.

    48k GitHub starsUsed in 1 repo~3.6k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Workflow Management Snakemake Workflows

What does Bio Workflow Management Snakemake Workflows do?

Authors reproducible bioinformatics pipelines with Snakemake - rules wired by output-file pattern, wildcards and expand() for sample fan-out, checkpoints for runtime-unknown outputs, resource/retry…. Bio Workflow Management Snakemake Workflows is an agent skill from GPTomics/bioSkills. Authors reproducible bioinformatics pipelines with Snakemake - rules wired by output-file pattern, wildcards and expand() for sample fan-out, checkpoints for runtime-unknown outputs, resource/retry escalation, and conda/container software deployment on HPC and cloud.

When should I use Bio Workflow Management Snakemake Workflows?

Bio Workflow Management Snakemake Workflows fits situations like: deciding rule-based (Snakemake) vs channel/dataflow (Nextflow) authoring; wiring rules by OUTPUT-file pattern rather than imperative order; using wildcards + expand() for sample fan-out and constraining them to stop silent mis-routing; adding checkpoints when the set of outputs is unknown until a step runs (dynamic DAG).

How do I install Bio Workflow Management Snakemake Workflows in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-workflow-management-snakemake-workflows -a claude-code`. Or copy the skill folder (workflow-management/snakemake-workflows in GPTomics/bioSkills) into .claude/skills/bio-workflow-management-snakemake-workflows in your project. Claude Code loads it when a task matches its description.

How do I install Bio Workflow Management Snakemake Workflows in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-workflow-management-snakemake-workflows -a codex`. Or copy the skill folder (workflow-management/snakemake-workflows in GPTomics/bioSkills) into .agents/skills/bio-workflow-management-snakemake-workflows in your project. Codex loads it when a task matches its description.

Can I use Bio Workflow Management Snakemake Workflows in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-workflow-management-snakemake-workflows -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-workflow-management-snakemake-workflows, .gemini/skills/bio-workflow-management-snakemake-workflows, .github/skills/bio-workflow-management-snakemake-workflows and .opencode/skills/bio-workflow-management-snakemake-workflows in your project.

What does Bio Workflow Management Snakemake Workflows need to run?

Going by SKILL.md and its folder, Bio Workflow Management Snakemake Workflows needs R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Workflow Management Snakemake Workflows access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Workflow Management Snakemake Workflows safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Workflow Management Snakemake Workflows use?

Bio Workflow Management Snakemake Workflows is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Workflow Management Snakemake Workflows use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Workflow Management Snakemake Workflows?

Skills that share tags, products or a category with Bio Workflow Management Snakemake Workflows: LaminDB Biological Data Management (davila7/claude-code-templates, 32k stars), Latchbio Integration (davila7/claude-code-templates, 32k stars), Latchbio Integration (K-Dense-AI/scientific-agent-skills, 48k stars) and Add Bactopia Tool (bactopia/bactopia, 522 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Workflow Management Snakemake Workflows?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.