Agent skill

Bio Genome Assembly Metagenome Assembly

by GPTomics in GPTomics/bioSkills

Assembles microbial-community sequencing into metagenome-assembled genomes (MAGs) with metaFlye (ONT), metaSPAdes/MEGAHIT (Illumina), and hifiasm-meta/metaMDBG (PacBio HiFi), then recovers genomes…

MITAuto-check passedResearch & Science

Install Bio Genome Assembly Metagenome Assembly

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-metagenome-assembly -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-assembly-metagenome-assembly --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-assembly/metagenome-assembly .claude/skills/bio-genome-assembly-metagenome-assembly && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-assembly-metagenome-assembly
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5k tokens
SKILL.md length
2,012 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Assembles microbial-community sequencing into metagenome-assembled genomes (MAGs) with metaFlye (ONT), metaSPAdes/MEGAHIT (Illumina), and hifiasm-meta/metaMDBG (PacBio HiFi), then recovers genomes…

  • Works in 3 steps: A MAG is a population consensus, not an… → The deliverable is a community of MAGs,… → Modern practice is multi-binner ->…
  • Reconstructing genomes from a microbiome
  • SKILL.md covers Version Compatibility, The Single Most Important…, Assembler Taxonomy and Binner Taxonomy, plus 12 more sections
  • Runs Shell scripts from its folder; calls pip

What it does

Bio Genome Assembly Metagenome Assembly is an agent skill from GPTomics/bioSkills. Assembles microbial-community sequencing into metagenome-assembled genomes (MAGs) with metaFlye (ONT), metaSPAdes/MEGAHIT (Illumina), and hifiasm-meta/metaMDBG (PacBio HiFi), then recovers genomes via multi-binner consolidation (MetaBAT2, MaxBin2, CONCOCT, SemiBin2, VAMB - DASTool) and QCs them against MIMAG with CheckM2, GUNC, and GTDB-Tk. Covers why a metagenome is not a genome (uneven coverage, micro-diversity, strain collapse to consensus), differential-coverage binning, co-assembly vs per-sample, the…

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/metagenome_workflow.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Reconstructing genomes from a microbiome
  • Recovering MAGs
  • Resolving strain-level variation

Example prompts

  • “Use the bio-genome-assembly-metagenome-assembly skill to assemble microbial-community sequencing into metagenome-assembled genomes (MAGs) with…”
  • “/bio-genome-assembly-metagenome-assembly”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A MAG is a population consensus, not an organism's genome. Co-occurring strains differing by <1% ANI become bubbles the assembler…
  2. The deliverable is a community of MAGs, not one assembly -- a community has no N50. N50 is dominated by whichever few abundant genomes…
  3. Modern practice is multi-binner -> consolidate -> CheckM2 + GUNC -> GTDB-Tk. Never trust one binner; run several (each weights composition…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Assembly Metagenome Assembly loads about 5k tokens when it runs. Until then it costs about 194 tokens; SKILL.md has 2,012 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~194
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,012 words, ~4,990 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-assembly-metagenome-assembly/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-genome-assembly-metagenome-assembly
description
Assembles microbial-community sequencing into metagenome-assembled genomes (MAGs) with metaFlye (ONT), metaSPAdes/MEGAHIT (Illumina), and hifiasm-meta/metaMDBG (PacBio HiFi), then recovers genomes via multi-binner consolidation (MetaBAT2, MaxBin2, CONCOCT, SemiBin2, VAMB -> DAS_Tool) and QCs them against MIMAG with CheckM2, GUNC, and GTDB-Tk. Covers why a metagenome is not a genome (uneven coverage, micro-diversity, strain collapse to consensus), differential-coverage binning, co-assembly vs per-sample, the rRNA-operon collapse that fails short-read MAGs, and strain resolution with inStrain. Use when reconstructing genomes from a microbiome, soil, ocean, or gut community, recovering MAGs, or resolving strain-level variation.
tool_type
cli
primary_tool
metaFlye

Version Compatibility

Reference examples tested with: Flye 2.9+, SPAdes 3.15+ (metaSPAdes), MEGAHIT 1.2+, hifiasm-meta 0.3+, metaMDBG 1.0+, MetaBAT2 2.15+, MaxBin 2.2.7+, CONCOCT 1.1+, SemiBin 2.0+, VAMB 4.1+, DAS_Tool 1.1.6+, CheckM2 1.0+, GUNC 1.0+, GTDB-Tk 2.4+, inStrain 1.7+, minimap2 2.26+, samtools 1.19+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures

GTDB-Tk results track the reference-package RELEASE (e.g. R214 vs R220); the DB release MUST match the GTDB-Tk binary or classification silently fails. CheckM2 and GUNC each download their own DIAMOND DB. SemiBin2's pretrained --environment models are versioned. If code throws an error, introspect the installed tool and adapt rather than retrying.

Metagenome Assembly

"Assemble genomes from my metagenome" -> Co-assemble a community at uneven, strain-mixed coverage, then bin the contigs into a set of consensus population genomes (MAGs) and QC each against MIMAG. The deliverable is MAGs, not a single assembly.

  • CLI: flye --meta --nano-hq reads.fq (ONT), spades.py --meta -1 R1.fq -2 R2.fq or megahit -1 R1.fq -2 R2.fq (Illumina), hifiasm_meta/metaMDBG (HiFi); then binners -> DAS_Tool -> checkm2 predict + gunc run + gtdbtk classify_wf

The Single Most Important Modern Insight -- A Metagenome Is Not a Genome; the Assembler Cannot Assume Uniform Coverage

Every isolate assembler is built on the premise that the true sequence sits at roughly one depth, so a coverage drop or spike signals a repeat or an error. In a community that premise is false by construction: an abundant species at 500x and a rare one at 3x are both real. Running plain SPAdes/Unicycler or single-genome Flye on a community treats the abundance spread and strain bubbles as errors to "fix" and produces garbage. Use --meta modes. Three consequences cascade:

  1. A MAG is a population consensus, not an organism's genome. Co-occurring strains differing by <1% ANI become bubbles the assembler collapses into one consensus path -- a sequence that may match no actual cell in the sample. A 99%-complete circular MAG is still the consensus of the dominant strain; minority-strain accessory genome is averaged away. Treat every per-strain, per-allele, or pangenome claim from a single consensus MAG as suspect until read-level microdiversity (inStrain) or strain-aware assembly confirms it.
  2. The deliverable is a community of MAGs, not one assembly -- a community has no N50. N50 is dominated by whichever few abundant genomes assembled well and says nothing about the community; a "better N50" assembly can have recovered fewer genomes. Report MAG count split by MIMAG tier (HQ/medium/low) and community fraction binned. Bigger total assembly size is not better -- it can mean more chimeras and strain-fragmentation.
  3. Modern practice is multi-binner -> consolidate -> CheckM2 + GUNC -> GTDB-Tk. Never trust one binner; run several (each weights composition vs coverage differently and recovers a partially-different genome set), reconcile with DAS_Tool, then QC every bin with CheckM2 AND GUNC (completeness lies about chimeras) before classifying against the MIMAG 90%/5% bar. HiFi/long reads are the single biggest quality jump: they span the conserved rRNA operons and strain bubbles short reads shred, yielding complete, circular, genuinely-HQ MAGs.

Assembler Taxonomy

ToolCitationMechanism / roleWhen
metaFlyeKolmogorov 2020 Nat Methodsrepeat-graph long-read meta-assembler (--meta mandatory)ONT/CLR community de novo; polish after
metaSPAdesNurk 2017 Genome Resstrain-aware multi-k de Bruijn (spades.py --meta)Illumina, contiguity priority; ONE paired library only
MEGAHITLi 2015 Bioinformaticssuccinct de Bruijn, low-memoryhuge/complex Illumina co-assemblies, soil; lower contiguity
hifiasm-metaFeng 2022 Nat Methodsstrain-resolved HiFi string graphPacBio HiFi communities; keeps strains apart
metaMDBGBenoit 2024 Nat Biotechnolminimizer de Bruijn for HiFiHiFi; often ~2x the HQ circular MAGs, low RAM
OPERA-MSBertrand 2019 Nat Biotechnolshort-read meta scaffolded by long readshybrid short + long

Binner Taxonomy

BinnerCitationSignalWhen
MetaBAT2Kang 2019 PeerJtetranucleotide freq (TNF) + coverage, parameter-freefast default workhorse (-m 1500)
MaxBin2Wu 2016 BioinformaticsEM over TNF + marker genes + coveragesingle/few samples
CONCOCTAlneberg 2014 Nat MethodsGMM on composition+coverage of cut-up contigsmany samples; 4-step pipeline, not one command
SemiBin2Pan 2023 Bioinformaticsself-supervised contrastive deep learningcurrent SOTA; short + long; pretrained env models
VAMBNissen 2021 Nat Biotechnolvariational autoencoder of coabundance + k-mermulti-sample; separates close strains
DAS_ToolSieber 2018 Nat Microbioldereplicate-aggregate-score across binners (consolidation)ALWAYS run; non-redundant set beats any single binner

Decision Tree by Scenario

ScenarioRecommendedWhy
Illumina, complex/huge/soil, low RAMMEGAHIT --presets meta-sensitivesuccinct dBG fits in memory; multi-library
Illumina, contiguity priority, tractable sizemetaSPAdes (spades.py --meta)strain-aware repeat resolution; merge libraries first (one paired lib only)
ONT-only communitymetaFlye --meta --nano-hq -> polishrepeat-graph meta mode; ONT needs polishing -> assembly-polishing
PacBio HiFi communityhifiasm-meta or metaMDBGcomplete circular strain-resolved MAGs; fixes rRNA collapse
Hybrid short + longOPERA-MSshort-read meta scaffolded with long reads
Recover MAGs (any assembly)>=2-3 binners -> DAS_Tool -> CheckM2 + GUNC -> GTDB-Tkensemble beats one binner; chimera + taxonomy gates
Only ONE samplecomposition-only binning, expect weak binsdifferential coverage needs multiple samples; add samples, not tuning
Multiple samples availablemap ALL samples to each assembly for binning depthdifferential-coverage is the strongest binning signal
Strain-level question-> inStrain on reads mapped to MAGsconsensus MAGs blur strains; needs read-level microdiversity
Read-based taxonomy / rare biosphere-> metagenomics/kraken-classificationassembly is blind below the abundance-detection limit
Reads not QC'd / host-contaminated-> long-read-sequencing/long-read-qcremove host reads vs a T2T reference before assembly
MAG contamination forensics-> contamination-detectiondetailed CheckM2/GUNC interpretation

metaFlye (ONT / Long Reads)

bash
flye --meta --nano-hq ont.fastq.gz --out-dir flye_out -t 32
#   --meta        uneven-coverage metagenome mode (REQUIRED for communities)
#   read-type flag (mutually exclusive): --nano-hq (Guppy5+/Q20) | --nano-raw (older) |
#       --pacbio-hifi | --pacbio-raw (CLR)
# outputs: assembly.fasta, assembly_graph.gfa, assembly_info.txt (circularity flag in col 'circ.')

ONT contigs are contiguous but error-prone (indels in homopolymers); polish before downstream use (-> assembly-polishing; medaka needs the matching basecaller model). HiFi usually needs no polishing.

metaSPAdes / MEGAHIT (Illumina)

bash
# metaSPAdes -- contiguity priority; exactly ONE paired library
spades.py --meta -1 R1.fastq.gz -2 R2.fastq.gz -o spades_out -t 32 -m 500
#   -m memory cap in GB (SPAdes aborts if exceeded); -k auto by default
# outputs: contigs.fasta, scaffolds.fasta

# MEGAHIT -- huge/low-RAM; accepts comma-separated multiple libraries
megahit -1 a1.fq.gz,b1.fq.gz -2 a2.fq.gz,b2.fq.gz -o megahit_out -t 32 \
        --presets meta-sensitive --min-contig-len 1000
#   --presets meta-sensitive | meta-large (huge complex); raise --min-contig-len to ~1000 for binning

metaSPAdes --meta supports exactly ONE paired-end library -- a real constraint people miss; concatenate libraries first or use MEGAHIT for many. metaSPAdes is heavier on RAM/time and chokes on soil-scale co-assembly; MEGAHIT assembled a 252 Gbp soil set on one node at the cost of somewhat more fragmentation.

HiFi (the transformative case)

bash
hifiasm_meta -t 32 -o asm hifi.fastq.gz
awk '/^S/{print ">"$2"\n"$3}' asm.p_ctg.gfa > asm.p_ctg.fa   # GFA -> FASTA

metaMDBG asm --out-dir mdbg_out --in-hifi hifi.fastq.gz --threads 32   # often ~2x HQ circular MAGs

Coverage for Binning, then Bin

bash
# Map reads back to the assembly -- per sample for differential coverage
minimap2 -ax map-ont -t 32 contigs.fa reads.fq.gz | samtools sort -@ 32 -o s1.sorted.bam -
samtools index s1.sorted.bam
jgi_summarize_bam_contig_depths --outputDepth depth.txt s1.sorted.bam s2.sorted.bam   # all samples

metabat2 -i contigs.fa -a depth.txt -o metabat/bin -m 1500 -t 32   # -m 1500 = min contig (do not go below ~1000)
run_MaxBin.pl -contig contigs.fa -abund abund1.txt -out maxbin/bin -thread 32 -min_contig_length 1000
SemiBin2 single_easy_bin -i contigs.fa -b s1.sorted.bam -o semibin_out   # add --environment human_gut for a pretrained model

CONCOCT is a 4-step pipeline (cut_up_fasta.py -> concoct_coverage_table.py -> concoct -> merge_cutup_clustering.py -> extract_fasta_bins.py), not one command. Differential-coverage binning needs MULTIPLE samples with abundance variation; with one sample binning collapses to weak composition-only signal -- more samples, not more tuning.

Consolidate (DAS_Tool), then QC

bash
# Convert each binner's output to contig2bin tables, then aggregate-and-score
Fasta_to_Contig2Bin.sh -i metabat/ -e fa    > metabat.tsv
Fasta_to_Contig2Bin.sh -i maxbin/  -e fasta > maxbin.tsv               # MaxBin emits .fasta
gunzip -k semibin_out/output_bins/*.gz 2>/dev/null || true             # SemiBin2 bins are gzipped
Fasta_to_Contig2Bin.sh -i semibin_out/output_bins/ -e fa > semibin.tsv
DAS_Tool -i metabat.tsv,maxbin.tsv,semibin.tsv -l metabat,maxbin,semibin \
         -c contigs.fa -o dastool/DAS --write_bins -t 32
# DAS_Tool assumes the SAME contig set across binners -- feeding bins from different assemblies is a silent error

checkm2 predict --input dastool/DAS_DASTool_bins/ -x fa --output-directory checkm2_out -t 32
gunc run --input_dir dastool/DAS_DASTool_bins/ --file_suffix .fa --out_dir gunc_out --threads 32
gtdbtk classify_wf --genome_dir dastool/DAS_DASTool_bins/ -x fa --out_dir gtdbtk_out --cpus 32

Run CheckM2 AND GUNC: CheckM2 counts marker copy number (completeness/contamination), GUNC tests whether a genome's genes share one lineage (chimerism). A bin made of two half-genomes can score high completeness, low contamination, and still be a chimera -- GUNC is the orthogonal catch. Dereplicate MAGs across samples (dRep ~95% ANI species, ~99% strain) before reporting counts.

Strain Resolution

bash
# Consensus MAGs blur strains; recover read-level microdiversity
inStrain profile sample.sorted.bam mags.fa -o instrain_out -p 16 -g genes.fna
inStrain compare -i instrain_A instrain_B -o instrain_compare   # popANI: shared-strain detection across samples

Per-Method Failure Modes

Isolate assembler on a community

Trigger: plain spades.py/flye (no --meta) on community reads. Mechanism: uniform-coverage assumption deletes rare-taxon contigs and mis-resolves strain bubbles as errors. Symptom: few short contigs, missing abundant taxa. Fix: --meta mode for every meta-assembler.

Single-binner pipeline

Trigger: "we used SemiBin2 because it's SOTA," one binner. Mechanism: each binner recovers a partially-different genome set. Symptom: real MAGs left on the table; lower count than peers. Fix: run >=2-3 binners -> DAS_Tool consolidation.

Single-sample differential-coverage expectation

Trigger: MetaBAT2/CONCOCT on one sample, surprised bins are bad. Mechanism: one coverage value -> composition (TNF) only, which is weak (related genera share TNF). Symptom: poor, split bins. Fix: more samples with abundance variation, then map all back; do not retune.

Show full SKILL.md (827 more words)Show less
Short-read MAG reported HQ on 90/5 alone

Trigger: calling a 95%-complete/2%-contam short-read MAG "high-quality." Mechanism: conserved + multi-copy rRNA operon tangles the short-read graph and stays unbinned. Symptom: MAG fails MIMAG HQ for missing 16S/23S/5S despite great completeness. Fix: check the FULL MIMAG HQ definition (rRNA + tRNA); use HiFi/long reads to span the operon.

High CheckM2 completeness read as quality

Trigger: trusting completeness/contamination alone. Mechanism: non-overlapping markers from two half-genomes score high completeness, low contamination. Symptom: "clean" MAG that is a chimera. Fix: always pair CheckM2 with GUNC -> contamination-detection.

Strain claims from a consensus MAG

Trigger: per-allele/per-strain interpretation of one MAG. Mechanism: the assembler collapsed strains to consensus. Symptom: strain findings that read-level data contradict. Fix: inStrain popANI or strain-aware (hifiasm-meta) assembly.

Reading MAG richness as community richness

Trigger: "200 MAGs cover 60% of reads" treated as the whole community. Mechanism: the rare biosphere never crosses the assembly detection limit. Symptom: species richness massively undercounted. Fix: complement with read-based profiling -> metagenomics/kraken-classification.

Quantitative Thresholds

ThresholdSourceRationale
MIMAG high-quality: completeness >90%, contamination <5%, AND 16S/23S/5S rRNA + >=18 tRNAsBowers 2017 Nat BiotechnolrRNA criterion is the silent killer; short-read MAGs fail it
MIMAG medium-quality: completeness >=50%, contamination <10%Bowers 2017most short-read MAGs land here
Min contig length for binning ~1000-1500 bpMetaBAT2 default 1500shorter contigs have unreliable TNF/coverage -> noise/chimeras
Differential-coverage binning needs N >= ~3-5 samples with abundance variationbinning practiceone coverage value cannot separate co-abundant genomes
Assembly detection limit ~3-5x coveragede Bruijn requirementbelow this the rare biosphere does not assemble at all
MAG dereplication 95% ANI (species), 99% (strain)dRep/ANI conventioncollapse redundant per-sample MAGs before counting
metaSPAdes input: exactly ONE paired library--meta constraintmerge libraries or use MEGAHIT for many
N50 / contiguitynot a community metricreport MAG count + MIMAG tiers instead

Common Errors

Error / symptomCauseSolution
Few short contigs, missing abundant taxaisolate mode on a communityadd --meta
metaSPAdes rejects multiple libraries--meta supports one paired libraryconcatenate, or use MEGAHIT
metaSPAdes aborts / out of memoryRAM cap hit on a complex co-assemblyraise -m, or switch to MEGAHIT
Poor bins from one sampleno differential-coverage signaladd samples; map all back for depth
"HQ MAG" lacks rRNArRNA-operon collapse in short readsfull MIMAG check; HiFi/long reads
Clean CheckM2 but suspect binchimera invisible to marker countsrun GUNC alongside CheckM2
GTDB-Tk classify_wf errorsDB release != binary versionmatch the GTDB reference-package release
Strain finding not reproducibleconsensus MAG blurs strainsinStrain popANI / strain-aware assembly

References

  • Nurk S, Meleshko D, Korobeynikov A, Pevzner PA. 2017. metaSPAdes: a new versatile metagenomic assembler. Genome Res 27:824-834.
  • Li D, Liu CM, Luo R, Sadakane K, Lam TW. 2015. MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics 31:1674-1676.
  • Kolmogorov M, et al. 2020. metaFlye: scalable long-read metagenome assembly using repeat graphs. Nat Methods 17:1103-1110.
  • Feng X, Cheng H, Portik D, Li H. 2022. Metagenome assembly of high-fidelity long reads with hifiasm-meta. Nat Methods 19:671-674.
  • Benoit G, et al. 2024. High-quality metagenome assembly from long accurate reads with metaMDBG. Nat Biotechnol 42:1378-1383.
  • Bertrand D, et al. 2019. Hybrid metagenomic assembly enables high-resolution analysis of resistance determinants and mobile elements in human microbiomes (OPERA-MS). Nat Biotechnol 37:937-944.
  • Kang DD, et al. 2019. MetaBAT 2: an adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies. PeerJ 7:e7359.
  • Wu YW, Simmons BA, Singer SW. 2016. MaxBin 2.0: an automated binning algorithm to recover genomes from multiple metagenomic datasets. Bioinformatics 32:605-607.
  • Alneberg J, et al. 2014. Binning metagenomic contigs by coverage and composition (CONCOCT). Nat Methods 11:1144-1146.
  • Pan S, Zhao XM, Coelho LP. 2023. SemiBin2: self-supervised contrastive learning leads to better MAGs for short- and long-read sequencing. Bioinformatics 39:i21-i29.
  • Nissen JN, et al. 2021. Improved metagenome binning and assembly using deep variational autoencoders (VAMB). Nat Biotechnol 39:555-560.
  • Sieber CMK, et al. 2018. Recovery of genomes from metagenomes via a dereplication, aggregation and scoring strategy (DAS_Tool). Nat Microbiol 3:836-843.
  • Chklovski A, et al. 2023. CheckM2: a rapid, scalable and accurate tool for assessing microbial genome quality using machine learning. Nat Methods 20:1203-1212.
  • Orakov A, et al. 2021. GUNC: detection of chimerism and contamination in prokaryotic genomes. Genome Biol 22:178.
  • Chaumeil PA, Mussig AJ, Hugenholtz P, Parks DH. 2020. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics 36:1925-1927.
  • Bowers RM, et al. 2017. Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea. Nat Biotechnol 35:725-731.
  • Olm MR, et al. 2021. inStrain profiles population microdiversity from metagenomic data and sensitively detects shared microbial strains. Nat Biotechnol 39:727-736.
  • contamination-detection - CheckM2/GUNC interpretation and chimerism forensics for MAGs
  • assembly-qc - Isolate-assembly QC; the uniform-coverage assumption metagenomes abandon
  • assembly-polishing - Polish ONT/CLR meta-contigs before binning (HiFi usually needs none)
  • metagenomics/kraken-classification - Read-based taxonomy; recovers the rare biosphere assembly cannot
  • metagenomics/abundance-estimation - Community abundance downstream of recovered MAGs
  • metagenomics/functional-profiling - Functional potential complementary to genome recovery
  • long-read-sequencing/long-read-qc - Read-level QC and host removal before assembly

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in genome-assembly/metagenome-assembly of GPTomics/bioSkills.

  • SKILL.md
  • examples/metagenome_workflow.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Assembly Metagenome Assembly next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Assembly Metagenome Assembly compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Assembly Metagenome Assembly this skillGPTomics/bioSkills1.2k1 repos~5kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Genome Assembly Metagenome Assembly

What does Bio Genome Assembly Metagenome Assembly do?

Assembles microbial-community sequencing into metagenome-assembled genomes (MAGs) with metaFlye (ONT), metaSPAdes/MEGAHIT (Illumina), and hifiasm-meta/metaMDBG (PacBio HiFi), then recovers genomes…. Bio Genome Assembly Metagenome Assembly is an agent skill from GPTomics/bioSkills. Assembles microbial-community sequencing into metagenome-assembled genomes (MAGs) with metaFlye (ONT), metaSPAdes/MEGAHIT (Illumina), and hifiasm-meta/metaMDBG (PacBio HiFi), then recovers genomes via multi-binner consolidation (MetaBAT2, MaxBin2, CONCOCT, SemiBin2, VAMB - DASTool) and QCs them against MIMAG with CheckM2, GUNC, and GTDB-Tk.

When should I use Bio Genome Assembly Metagenome Assembly?

Bio Genome Assembly Metagenome Assembly fits situations like: reconstructing genomes from a microbiome; recovering MAGs; resolving strain-level variation.

How do I install Bio Genome Assembly Metagenome Assembly in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-metagenome-assembly -a claude-code`. Or copy the skill folder (genome-assembly/metagenome-assembly in GPTomics/bioSkills) into .claude/skills/bio-genome-assembly-metagenome-assembly in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Assembly Metagenome Assembly in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-metagenome-assembly -a codex`. Or copy the skill folder (genome-assembly/metagenome-assembly in GPTomics/bioSkills) into .agents/skills/bio-genome-assembly-metagenome-assembly in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Assembly Metagenome Assembly in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-metagenome-assembly -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-assembly-metagenome-assembly, .gemini/skills/bio-genome-assembly-metagenome-assembly, .github/skills/bio-genome-assembly-metagenome-assembly and .opencode/skills/bio-genome-assembly-metagenome-assembly in your project.

What does Bio Genome Assembly Metagenome Assembly need to run?

Going by SKILL.md and its folder, Bio Genome Assembly Metagenome Assembly needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.

Does Bio Genome Assembly Metagenome Assembly access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Genome Assembly Metagenome Assembly safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Assembly Metagenome Assembly use?

Bio Genome Assembly Metagenome Assembly is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Assembly Metagenome Assembly use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Assembly Metagenome Assembly?

Skills that share tags, products or a category with Bio Genome Assembly Metagenome Assembly: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Assembly Metagenome Assembly?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.