Agent skill

Bio Workflows Genome Assembly Pipeline

by GPTomics in GPTomics/bioSkills

Orchestrates an end-to-end de novo genome assembly project, routing each step to the right genome-assembly skill rather than restating it.

MITAuto-check passedResearch & Science

Install Bio Workflows Genome Assembly Pipeline

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-workflows-genome-assembly-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-workflows-genome-assembly-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workflows/genome-assembly-pipeline .claude/skills/bio-workflows-genome-assembly-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-workflows-genome-assembly-pipeline
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.1k tokens
SKILL.md length
1,125 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Orchestrates an end-to-end de novo genome assembly project, routing each step to the right genome-assembly skill rather than restating it.

  • Works in 7 steps: Profile the Genome (do this first) → QC Reads → Assemble (route by data type) → …
  • Assembling a genome from raw reads and deciding which assembler
  • SKILL.md covers Version Compatibility, The Single Most Important…, Decision Flow (Step 0 -> 6) and Routing Table by Scenario, plus 11 more sections
  • Runs Shell scripts from its folder; calls python3 and sh

What it does

Bio Workflows Genome Assembly Pipeline is an agent skill from GPTomics/bioSkills. Orchestrates an end-to-end de novo genome assembly project, routing each step to the right genome-assembly skill rather than restating it. Profiles the genome first (k-mer spectrum - size, heterozygosity, ploidy), QCs reads, chooses an assembly path by data type (SPAdes for Illumina, Flye for noisy long reads, hifiasm for HiFi, metaFlye for communities), polishes only when needed, decontaminates, scaffolds with Hi-C, and finishes with three-axis QC (contiguity + completeness + correctness). Use when assembling a…

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/bacterial_assembly.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Assembling a genome from raw reads and deciding which assembler
  • Whether to polish
  • How to prove the result is good

Example prompts

  • “Use the bio-workflows-genome-assembly-pipeline skill to orchestrate an end-to-end de novo genome assembly project, routing each step to the right…”
  • “/bio-workflows-genome-assembly-pipeline”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Profile the Genome (do this first)
  2. QC Reads
  3. Assemble (route by data type)
  4. Polish IF Needed
  5. Decontaminate (route by sample type)
  6. Scaffold IF Hi-C Is Available
  7. Three-Axis QC (contiguity + completeness + correctness)

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • sh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Workflows Genome Assembly Pipeline loads about 4.1k tokens when it runs. Until then it costs about 166 tokens; SKILL.md has 1,125 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~166
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,125 words, ~4,147 tokens.

Download SKILL.mdSave it as .claude/skills/bio-workflows-genome-assembly-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-workflows-genome-assembly-pipeline
description
Orchestrates an end-to-end de novo genome assembly project, routing each step to the right genome-assembly skill rather than restating it. Profiles the genome first (k-mer spectrum -> size, heterozygosity, ploidy), QCs reads, chooses an assembly path by data type (SPAdes for Illumina, Flye for noisy long reads, hifiasm for HiFi, metaFlye for communities), polishes only when needed, decontaminates, scaffolds with Hi-C, and finishes with three-axis QC (contiguity + completeness + correctness). Use when assembling a genome from raw reads and deciding which assembler, whether to polish, and how to prove the result is good.
tool_type
cli
primary_tool
Flye
workflow
true
depends_on
genome-assembly/genome-profiling, read-qc/fastp-workflow, long-read-sequencing/long-read-qc, genome-assembly/short-read-assembly…

Version Compatibility

Reference examples tested with: GenomeScope2 2.0+, meryl 1.4+, Merqury 1.3+, fastp 0.23+, SPAdes 4.0+, Flye 2.9+, hifiasm 0.25+, metaFlye 2.9+, Racon 1.5+, medaka 2.0+, minimap2 2.26+, FCS-GX 0.5+, CheckM2 1.0+, GUNC 1.0+, YaHS 1.2+, QUAST 5.2+, BUSCO 5.5+, samtools 1.19+. Each owning genome-assembly skill is the source of truth for its tool's pinned version.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

Tool outputs are driven by more than the binary version: medaka consensus quality depends on the basecaller MODEL string (must match the basecaller, e.g. -m r1041_e82_400bps_sup_v5.0.0); BUSCO/compleasm results depend on the lineage dataset and OrthoDB generation (record them); CheckM2/GTDB-Tk results track the reference DATABASE release; hifiasm output filenames and default purge behaviour change across versions (verify against the installed build). If a command errors, introspect the installed tool and adapt rather than retrying.

Genome Assembly Pipeline

"Assemble a genome from my sequencing reads and prove it is good" -> Profile the genome, QC reads, route to the right assembler by data type, polish only if needed, decontaminate, scaffold if Hi-C exists, and finish with three-axis QC. This skill ORCHESTRATES the genome-assembly category; it routes each step to the owning skill and encodes the cross-cutting decisions, not each tool's full option set.

The Single Most Important Modern Insight -- Assembly Is Three Orthogonal Questions, and Each Step Answers One

A genome project fails when one number stands in for the whole. Profiling sets expectations (how big, how heterozygous, how many haplotypes) BEFORE assembling, the assembler answers contiguity, polishing answers per-base accuracy, decontamination answers provenance, scaffolding answers arrangement, and QC must independently address all three of contiguity, completeness, and correctness. The orchestration job is to keep these separate and route each to its skill: a high N50 says nothing about whether the bases are right (Merqury QV) or whether the sequence is the organism's (contamination), and skipping profiling means the assembler guesses the parameters that profiling would have set.

Decision Flow (Step 0 -> 6)

Raw reads (+ optional Hi-C, trio, short reads)
    |
    v
[0. Profile the genome] --> genome-assembly/genome-profiling
    |   k-mer spectrum (GenomeScope2) -> genome size, heterozygosity, ploidy.
    |   Sets NG50 denominator, expected haplotype count, hifiasm purge level,
    |   and which assembly path is even sensible. Do this BEFORE assembling.
    v
[1. QC reads] -----------> short: read-qc/fastp-workflow
    |                       long:  long-read-sequencing/long-read-qc
    |   Garbage-in caps assembly quality; record platform + basecaller era
    |   (it is an assembly PARAMETER, see step 2), trim internal adapters.
    v
[2. Choose path BY DATA TYPE]
    |  Illumina-only small/isolate -> genome-assembly/short-read-assembly (SPAdes)
    |  noisy ONT/CLR              -> genome-assembly/long-read-assembly (Flye --nano-hq for R10)
    |  PacBio HiFi                -> genome-assembly/hifi-assembly (hifiasm, phased)
    |  community sample           -> genome-assembly/metagenome-assembly (metaFlye/metaSPAdes + binning)
    |  large/heterozygous euk     -> long-read or HiFi, NOT short reads
    v
[3. Polish IF needed] ---> genome-assembly/assembly-polishing
    |   noisy long-read assemblies: Racon -> medaka (model MUST match basecaller).
    |   Do NOT polish HiFi reflexively (often net-harmful). Measure with Merqury QV,
    |   not the reads polished with. Skip entirely for SPAdes/HiFi when QV is already high.
    v
[4. Decontaminate] ------> genome-assembly/contamination-detection
    |   single organism: FCS-GX (GenBank-mandatory) + BlobToolKit blob plot.
    |   MAG:              CheckM2 + GUNC (chimerism). Two disjoint problems (see below).
    v
[5. Scaffold IF Hi-C] ---> genome-assembly/scaffolding
    |   automated YaHS produces a DRAFT; manual contact-map curation is the standard.
    |   Scaffold N50 != contig N50 (gaps are Ns). Skip if no Hi-C.
    v
[6. Three-axis QC] ------> genome-assembly/assembly-qc
        contiguity (auN/NG50 vs profiled size) + completeness (BUSCO/compleasm)
        + correctness (Merqury QV). Report the triad; NEVER N50 alone.

Routing Table by Scenario

ScenarioPathRoutes to
Bacterial isolate, ONT R10 onlyprofile -> QC -> Flye --nano-hq -> medaka -> FCS-GX -> QClong-read-assembly, assembly-polishing, contamination-detection
Bacterial isolate, Illumina onlyprofile -> fastp -> SPAdes --isolate -> FCS-GX -> QCshort-read-assembly
Small genome, ONT, max qualityprofile -> QC -> multi-assembler consensus (Trycycler/Autocycler) -> medaka -> QClong-read-assembly
Diploid eukaryote, HiFi (+Hi-C/trio)profile -> QC -> hifiasm (hap1/hap2) -> purge check -> decontam -> scaffold -> QChifi-assembly, scaffolding, contamination-detection
Large heterozygous eukaryote, ONTprofile -> QC -> Flye -> purge_dups -> medaka -> decontam -> scaffold -> QClong-read-assembly, scaffolding
Community / microbiome sampleQC -> metaFlye/metaSPAdes -> binning -> CheckM2 + GUNCmetagenome-assembly, contamination-detection
Hi-C reads availableafter contigs+polish: scaffold, curate contact mapscaffolding
Reads not yet QC'dstart at step 1read-qc/fastp-workflow, long-read-sequencing/long-read-qc

Cross-Cutting Gotchas (surface these at every project)

  • Basecaller era must match the assembler flag. --nano-raw on R10/Dorado-SUP reads silently collapses real repeats while RAISING N50; --nano-hq is the R10 default. The platform + basecaller model is an assembly parameter, not metadata.
  • A primary assembly is not a haplotype. The hifiasm primary is a maternal/paternal mosaic that exists in no cell; for any allele-aware downstream use hap1/hap2 phased with trio or Hi-C, and treat HiFi-only hap1/hap2 as only partially phased.
  • N50 is gamed. It rises when an assembly gets WORSE (misjoins, collapsed repeats, retained haplotigs). Report the triad (auN/NG50 + BUSCO + Merqury QV), never N50 alone.
  • A MAG is a population consensus, not a genome. The unit of success is a binned, MIMAG-gated MAG, and "% contamination" conflates foreign-organism mixing, strain mixing, and assembly artifacts.
  • "Contamination" is two disjoint problems. Single-organism cross-kingdom foreign sequence (FCS-GX, blob plot) is a different question from intra-domain MAG contamination/chimerism (CheckM2 + GUNC); do not apply one tool's question to the other's input.
  • Scaffold N50 >> contig N50 because gaps are Ns. Scaffold contiguity is glue, not sequence; every join is a hypothesis a contact map must confirm. Report contig N50 alongside scaffold N50.

Step 0: Profile the Genome (do this first)

Route the full treatment to genome-assembly/genome-profiling. The minimal orchestration step:

bash
# k-mer count from ACCURATE reads (Illumina/HiFi, NEVER noisy ONT), then GenomeScope2 for size / heterozygosity / ploidy
meryl count k=21 output reads.meryl accurate_reads.fq.gz
meryl histogram reads.meryl > reads.hist
genomescope2 -i reads.hist -o gscope_out -k 21
# read off: estimated haploid genome size, heterozygosity %, and (with -p) ploidy.
# These set the NG50 denominator, the expected number of haplotypes, and the purge decision.

Step 1: QC Reads

Short reads route to read-qc/fastp-workflow; long reads to long-read-sequencing/long-read-qc.

bash
fastp -i R1.fq.gz -I R2.fq.gz -o t_R1.fq.gz -O t_R2.fq.gz \
    --detect_adapter_for_pe --qualified_quality_phred 20 --length_required 50 --html qc.html
Show full SKILL.md (458 more words)Show less

Step 2: Assemble (route by data type)

Give the assembler the exact preset for the chemistry; the wrong preset is silent. Detailed options live in the owning skills.

bash
# Illumina-only small/isolate genome -> short-read-assembly
spades.py --isolate -1 t_R1.fq.gz -2 t_R2.fq.gz -o spades_out -t 16
# NOTE: --careful is small-genome-only; do NOT use it on large eukaryote genomes.

# Noisy ONT (R10/Dorado-SUP) -> long-read-assembly. --nano-hq is the modern default.
flye --nano-hq ont.fq.gz --out-dir flye_out --threads 16     # --genome-size optional in recent Flye

# PacBio HiFi -> hifi-assembly (phased by default; verify output filenames per version)
hifiasm -o asm -t 16 hifi.fq.gz                              # add --h1/--h2 (Hi-C) or -1/-2 (trio) to phase

# Community sample -> metagenome-assembly
flye --nano-hq ont.fq.gz --meta --out-dir metaflye_out --threads 16    # --meta is a modifier; still need a read-type selector. Then bin + CheckM2/GUNC

Step 3: Polish IF Needed

Polishing is read-type-matched and conditional. Route to genome-assembly/assembly-polishing.

bash
# Noisy long-read assembly: medaka with the MATCHING model. medaka_consensus does its own
# read-to-assembly alignment from -i/-d (no separate minimap2/BAM step needed); add a Racon
# round upstream only if the assembler did not already polish - see assembly-polishing.
medaka_consensus -i ont.fq.gz -d flye_out/assembly.fasta -o medaka_out -t 16 \
    -m r1041_e82_400bps_sup_v5.0.0   # MUST match the basecaller model used to call the reads
# For a BACTERIAL isolate, prefer the methylation-aware bacterial model (medaka 2.0+):
#   medaka_consensus -i ont.fq.gz -d assembly.fasta -o medaka_out --bacteria

Do NOT reflexively polish a HiFi assembly (already ~Q30+; over-polishing lowers QV). SPAdes output needs no separate long-read polish. The stop signal is a Merqury QV plateau, not a fixed iteration count, and the QV must be measured against reads independent of those used to polish.

Step 4: Decontaminate (route by sample type)

bash
# Single-organism assembly (GenBank-mandatory foreign screen + blob plot)
python3 ./fcs.py screen genome --fasta assembly.fa --out-dir gx_out/ --gx-db "$GXDB/gxdb" --tax-id <taxid>
# acts on EXCLUDE/TRIM/FIX cross-kingdom contigs; keep host-integrated foreign sequence (see contamination-detection)

# MAG (intra-domain contamination + chimerism)
checkm2 predict --input bins/ --output-directory checkm2_out --threads 16
gunc run --input_dir bins/ --out_dir gunc_out                            # chimerism, orthogonal to CheckM2

Step 5: Scaffold IF Hi-C Is Available

YaHS produces a draft; the contact map is the QC, not decoration. Route to genome-assembly/scaffolding.

bash
# Map Hi-C to contigs, then YaHS; inspect the contact map (PretextMap/Juicer) and break misjoins.
yahs assembly.fasta hic_to_contigs.bam -o yahs_out          # output scaffolds + AGP; curate before publishing

Step 6: Three-Axis QC (contiguity + completeness + correctness)

Route the full treatment to genome-assembly/assembly-qc. Report all three axes; lead with the QV.

bash
# Contiguity vs the PROFILED genome size (NG50/auN, not bare N50)
quast.py final.fasta -o quast_out -t 16 --est-ref-size <profiled_size>

# Completeness on the DEEPEST applicable clade (compleasm on good genomes; BUSCO otherwise)
busco -i final.fasta -l <clade>_odb10 -o busco_out -m genome -c 16

# Correctness: Merqury QV from ACCURATE reads (k from best_k.sh, not hardcoded)
K=$(sh $MERQURY/best_k.sh <genome_size_bp> | tail -n1 | awk '{print int($1+0.5)}')   # round float->int
meryl count k=$K output reads.meryl accurate_reads.fq.gz
merqury.sh reads.meryl final.fasta merqury_out          # QV + k-mer completeness + spectra-cn

Common Errors

SymptomCauseFix
Assembly ~1.5-2x profiled size, high BUSCO-Duplicateduncollapsed haplotigs (false duplication)purge_dups; check half-coverage depth peak; do not over-purge real segmental duplications
Contiguous but gene models frameshiftnoisy long-read assembly not polishedRacon -> medaka (matched model); measure QV
QV drops after polishingover-polishing an already-accurate (HiFi) assemblystop polishing; HiFi rarely needs short-read polish
medaka consensus worse than inputwrong basecaller model stringset -m to the model the reads were basecalled with
Fewer contigs than expected but repeats collapsed--nano-raw used on R10 readsre-run Flye with --nano-hq
CheckM2 says clean but bin looks mixedchimera with disjoint markersrun GUNC; CheckM2 marker redundancy cannot see chimerism
Scaffold N50 huge, contig N50 smallscaffolding glue, not sequenceinspect contact map, break off-diagonal misjoins
  • genome-assembly/genome-profiling - Step 0: k-mer spectrum for size, heterozygosity, ploidy; sets expectations before assembling
  • genome-assembly/short-read-assembly - SPAdes path for Illumina-only small/isolate genomes
  • genome-assembly/long-read-assembly - Flye/Canu path for noisy ONT/CLR reads
  • genome-assembly/hifi-assembly - hifiasm phased path for PacBio HiFi
  • genome-assembly/metagenome-assembly - metaFlye/metaSPAdes + binning for community samples
  • genome-assembly/assembly-polishing - Racon/medaka/Pilon, applied only when needed
  • genome-assembly/contamination-detection - FCS-GX/BlobToolKit (single organism) vs CheckM2/GUNC (MAG)
  • genome-assembly/scaffolding - YaHS Hi-C scaffolding and contact-map curation
  • genome-assembly/assembly-qc - Three-axis QC: auN/NG50 + BUSCO + Merqury QV
  • read-qc/fastp-workflow - Short-read QC before assembly
  • long-read-sequencing/long-read-qc - Long-read length/quality QC and basecaller-era awareness
  • workflows/genome-annotation-pipeline - Downstream: only a decontaminated, QC-passed FASTA should hand off to annotation

References

  • Rhie A, Walenz BP, Koren S, Phillippy AM (2020) Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biology 21:245. DOI 10.1186/s13059-020-02134-9. (QV/completeness from k-mers.)
  • Manni M, Berkeley MR, Seppey M, Simao FA, Zdobnov EM (2021) BUSCO update: novel and streamlined workflows. Molecular Biology and Evolution 38:4647-4654. DOI 10.1093/molbev/msab199.
  • Rhie A, McCarthy SA, Fedrigo O, et al (2021) Towards complete and error-free genome assemblies of all vertebrate species. Nature 592:737-746. DOI 10.1038/s41586-021-03451-0. (three-axis / VGP standard.)
  • Ranallo-Benavidez TR, Jaron KS, Schatz MC (2020) GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes. Nature Communications 11:1432. DOI 10.1038/s41467-020-14998-3.

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in workflows/genome-assembly-pipeline of GPTomics/bioSkills.

  • SKILL.md
  • examples/bacterial_assembly.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Workflows Genome Assembly Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Workflows Genome Assembly Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Workflows Genome Assembly Pipeline this skillGPTomics/bioSkills1.2k1 repos~4.1kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Workflows Genome Assembly Pipeline

What does Bio Workflows Genome Assembly Pipeline do?

Orchestrates an end-to-end de novo genome assembly project, routing each step to the right genome-assembly skill rather than restating it. Bio Workflows Genome Assembly Pipeline is an agent skill from GPTomics/bioSkills. Orchestrates an end-to-end de novo genome assembly project, routing each step to the right genome-assembly skill rather than restating it.

When should I use Bio Workflows Genome Assembly Pipeline?

Bio Workflows Genome Assembly Pipeline fits situations like: assembling a genome from raw reads and deciding which assembler; whether to polish; how to prove the result is good.

How do I install Bio Workflows Genome Assembly Pipeline in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-workflows-genome-assembly-pipeline -a claude-code`. Or copy the skill folder (workflows/genome-assembly-pipeline in GPTomics/bioSkills) into .claude/skills/bio-workflows-genome-assembly-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Bio Workflows Genome Assembly Pipeline in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-workflows-genome-assembly-pipeline -a codex`. Or copy the skill folder (workflows/genome-assembly-pipeline in GPTomics/bioSkills) into .agents/skills/bio-workflows-genome-assembly-pipeline in your project. Codex loads it when a task matches its description.

Can I use Bio Workflows Genome Assembly Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-workflows-genome-assembly-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-workflows-genome-assembly-pipeline, .gemini/skills/bio-workflows-genome-assembly-pipeline, .github/skills/bio-workflows-genome-assembly-pipeline and .opencode/skills/bio-workflows-genome-assembly-pipeline in your project.

What does Bio Workflows Genome Assembly Pipeline need to run?

Going by SKILL.md and its folder, Bio Workflows Genome Assembly Pipeline needs a shell for the scripts in its folder and the command-line tools its instructions call (python3 and sh). Our summary lists: Python 3; A Bash shell.

Does Bio Workflows Genome Assembly Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Workflows Genome Assembly Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Workflows Genome Assembly Pipeline use?

Bio Workflows Genome Assembly Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Workflows Genome Assembly Pipeline use?

About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Workflows Genome Assembly Pipeline?

Skills that share tags, products or a category with Bio Workflows Genome Assembly Pipeline: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Workflows Genome Assembly Pipeline?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.