Agent skill

Bio Genome Assembly Long Read Assembly

by GPTomics in GPTomics/bioSkills

Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm…

MITAuto-check passedResearch & Science

Install Bio Genome Assembly Long Read Assembly

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-assembly-long-read-assembly --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-assembly/long-read-assembly .claude/skills/bio-genome-assembly-long-read-assembly && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-assembly-long-read-assembly
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.6k tokens
SKILL.md length
2,070 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm…

  • Works in 2 steps: The basecaller era dictates the input… → The bottleneck flipped from contiguity…
  • Assembling a bacterial
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Shell scripts from its folder

What it does

Bio Genome Assembly Long Read Assembly is an agent skill from GPTomics/bioSkills. Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm, and reconciles bacterial assemblies into a consensus with Trycycler/Autocycler. Covers matching the input flag to the basecaller era (--nano-hq vs --nano-raw), why a raw long-read assembly is contiguous but low-QV and not finished until polished, haplotig false-duplication and purgedups, coverage and read-N50 as…

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/flye_assembly.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Assembling a bacterial
  • Eukaryotic genome from ONT
  • PacBio noisy reads
  • Choosing a long-read assembler

Example prompts

  • “Use the bio-genome-assembly-long-read-assembly skill to assemble genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR)…”
  • “/bio-genome-assembly-long-read-assembly”

Requirements

  • A Bash shell

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The basecaller era dictates the input flag, and a mismatch silently wrecks the assembly while the report looks finished. Noisy long reads…
  2. The bottleneck flipped from contiguity to consensus accuracy. A raw noisy-long-read assembly emits a FASTA with a spectacular N50 that…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Assembly Long Read Assembly loads about 4.6k tokens when it runs. Until then it costs about 208 tokens; SKILL.md has 2,070 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~208
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,070 words, ~4,627 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-assembly-long-read-assembly/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-genome-assembly-long-read-assembly
description
Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm, and reconciles bacterial assemblies into a consensus with Trycycler/Autocycler. Covers matching the input flag to the basecaller era (--nano-hq vs --nano-raw), why a raw long-read assembly is contiguous but low-QV and not finished until polished, haplotig false-duplication and purge_dups, coverage and read-N50 as non-substitutable inputs, and mid-read adapter de-chimerization. Use when assembling a bacterial or eukaryotic genome from ONT or PacBio noisy reads, choosing a long-read assembler, or diagnosing an over-collapsed or duplicated assembly. For PacBio HiFi use hifi-assembly instead.
tool_type
cli
primary_tool
Flye

Version Compatibility

Reference examples tested with: Flye 2.9+, Canu 2.2+, NextDenovo 2.5+, Shasta 0.11+, Raven 1.8+, wtdbg2 2.5+, miniasm 0.3+, minimap2 2.26+, purge_dups 1.2+, Porechop_ABI 0.5+, Trycycler 0.5+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

Two behaviors are version-dependent and load-bearing: Flye --genome-size became optional (auto-estimated) and --scaffold flipped OFF by default in recent releases - confirm against the installed Flye. Shasta --config names are dated and chemistry-specific (e.g. Nanopore-R10-Fast-Nov2022) - list with shasta --command listConfigurations rather than hard-coding one. Canu read-type flag spellings and correctedErrorRate defaults changed across major versions. If a command errors, introspect the installed tool and adapt rather than retrying.

Long-Read Assembly (Noisy Reads)

"Assemble a genome from Nanopore/PacBio-CLR long reads" -> Build a contiguous de novo assembly whose input mode matches the basecaller error regime, then hand off to polishing because the raw consensus is contiguous but not yet accurate.

  • CLI: flye --nano-hq reads.fq.gz --out-dir out -t 16 (modern ONT R10/Dorado), canu -p asm -d out genomeSize=4.6m -nanopore reads.fq.gz (thorough), wtdbg2 -x ont -g 4.6m -i reads.fq.gz -fo asm && wtpoa-cns -i asm.ctg.lay.gz -fo asm.fa (fast draft)

The Single Most Important Modern Insight -- The Basecaller Era Picks the Flag, and the Assembly Is Only Half-Done Until Polished

Two load-bearing facts no assembler README states:

  1. The basecaller era dictates the input flag, and a mismatch silently wrecks the assembly while the report looks finished. Noisy long reads are not one error regime: ONT R9.4.1+Guppy is ~5-10% error, ONT R10.4.1+Dorado SUP is Q20+ (~1-2%), PacBio CLR is ~10-15%. Each assembler has separate modes tuned to each (Flye --nano-raw / --nano-hq / --pacbio-raw). The dangerous direction is telling the assembler the reads are noisier than they are - feeding R10/Dorado data to --nano-raw makes Flye treat real repeat-copy and allele differences as noise and over-collapse them. The result is fewer contigs and a HIGHER N50 as the assembly gets worse - the most dangerous failure mode because the headline metric improved and nothing crashes. The chemistry/kit/basecaller model is an assembly parameter, not optional metadata; if it is unknown, the flag cannot be chosen and the assembly is uninterpretable - find out, do not guess.

  2. The bottleneck flipped from contiguity to consensus accuracy. A raw noisy-long-read assembly emits a FASTA with a spectacular N50 that looks finished, but per-base accuracy is often Q20-Q30 (one error every ~100-1000 bp). The dominant ONT error is the indel, especially in homopolymers, and a single indel frameshifts a protein - so a contiguous assembly can have every gene model broken. Contiguity and correctness are orthogonal axes; N50 is blind to QV. The assembler is the halfway point: the deliverable is a polished assembly with a measured QV (Merqury), not a high-N50 FASTA. Flye does one internal polishing round and stops - that is not "polished." Hand off to assembly-polishing and assembly-qc.

Tool Taxonomy

ToolCitationParadigmWhen
FlyeKolmogorov 2019 Nat Biotechnolrepeat graph (disjointigs -> explicit repeat structure)the fast general-purpose default; bacteria -> eukaryote
CanuKoren 2017 Genome Rescorrect -> trim -> assemble (OLC, MHAP)maximum single-assembler quality; slow, grid-oriented
NextDenovoHu 2024 Genome Biolcorrect (NextCorrect) -> string-graph (NextGraph)large/repetitive plant and animal genomes; contiguity-first
ShastaShafin 2020 Nat Biotechnolrun-length-encoded marker graphONT human-scale, speed/cost critical
RavenVaser & Sikic 2021 Nat Comput SciOLC string graph, near parameter-freefast simple draft; common Trycycler input
wtdbg2Ruan & Li 2020 Nat Methodsfuzzy de Bruijn graph (+ mandatory wtpoa-cns)fastest, lowest RAM; lowest accuracy -> polish hard
miniasmLi 2016 Bioinformaticsstring-graph layout, NO consensuslayout demo / pipeline component only; output = raw read error
Trycycler / AutocyclerWick 2021 Genome Biol / Wick 2025consensus of multiple independent assembliesbacterial reliability; catches single-assembler structural errors
purge_dupsGuan 2020 Bioinformaticsread-depth + self-alignmentremove haplotig false-duplication from a diploid primary

PacBio CLR (--pacbio-raw) is legacy - superseded by HiFi for all new PacBio work; treat CLR support as maintenance for archival data, do not recommend generating new CLR.

Decision Tree by Scenario

ScenarioRecommendedWhy
ONT R10.4.1 / Dorado HAC-SUP, any genomeFlye --nano-hqmodern default mode; matches the Q20+ error regime
ONT R9.4.1 SUP (Guppy5+/Dorado)Flye --nano-hq --read-error 0.05hq mode with the error floor raised to R9-SUP
ONT R9.4.1 legacy fast/HACFlye --nano-rawthe genuinely-noisy mode; do not use on R10
PacBio CLR (archival)Flye --pacbio-raw or Canu -pacbiolegacy noisy mode; polish hard (arrow/GCpp)
Bacterial isolate, want a correct finished genomeTrycycler (interactive) / Autocycler (automated)consensus across assemblers; fixes structural errors polishing can't
Large repetitive plant/animal, contiguity-firstNextDenovomemory-efficient, top contiguity on big genomes
ONT human-scale, speed-criticalShasta --config <era-matched>RLE marker graph; human genomes in days
Quick draft / compute is the bottleneckwtdbg2 or Ravenfastest; accept lower accuracy then polish
PacBio HiFi (Q30+, CCS)-> hifi-assemblyhifiasm phased haplotypes; wrong tool here
After assembling (always)-> assembly-polishing then assembly-qcraw consensus is low-QV; not finished until polished + QV-measured
Reads not yet QC'd / unknown chemistry-> long-read-sequencing/long-read-qc, long-read-sequencing/basecallinggarbage-in caps the assembly; basecaller model sets the flag
Genome size / coverage unknownGenomeScope2 on accurate short reads (not raw ONT)k-mer histograms from noisy reads inflate unique k-mers

Flye (the default)

bash
flye --nano-hq reads.fq.gz --out-dir out -t 16                              # ONT R10 / Dorado SUP
flye --nano-hq reads.fq.gz --read-error 0.05 --out-dir out -t 16            # ONT R9 SUP
flye --nano-raw reads.fq.gz --out-dir out -t 16                             # legacy ONT R9 fast/HAC
flye --pacbio-raw reads.fq.gz --out-dir out -t 16                           # PacBio CLR (legacy)
flye --nano-hq reads.fq.gz --genome-size 3g --asm-coverage 40 -o out -t 32  # large genome: use longest 40x for initial assembly

--genome-size is optional in recent Flye (auto-estimated) but required when paired with --asm-coverage, which downsamples to the longest N-coverage of reads for the initial disjointig step (cuts runtime/RAM on deep large-genome data; the rest are still used). --iterations defaults to 1 polishing round (0 to skip); --keep-haplotypes retains alt bubble paths for diploid awareness; --meta is metaFlye for uneven-coverage communities (-> metagenome-assembly). Output: assembly.fasta, assembly_info.txt (per-contig length/coverage/circularity), assembly_graph.gfa.

Canu (thorough), wtdbg2 / Raven (fast), miniasm (layout only)

bash
canu -p asm -d out genomeSize=4.6m -nanopore reads.fq.gz useGrid=false maxThreads=16   # correct->trim->assemble
wtdbg2 -x ont -g 4.6m -t 16 -i reads.fq.gz -fo asm && wtpoa-cns -t 16 -i asm.ctg.lay.gz -fo asm.ctg.fa  # consensus step is MANDATORY
raven -t 16 reads.fq.gz > asm.fasta                                                     # near parameter-free

Canu's master meta-parameter is correctedErrorRate (max expected difference between two corrected reads): defaults ~0.144 (Nanopore) / ~0.045 (PacBio); raise for heterozygosity/divergence, lower for clean high-coverage data. Read-type flags -nanopore / -pacbio / -pacbio-hifi (there is no -nanopore-hifi; high-accuracy ONT still uses -nanopore); useGrid=false forces a single machine. wtdbg2 needs the separate wtpoa-cns consensus call - the -x preset (ont/sq/rs/ccs) is set FIRST, and -L discards short reads (default 5000 for ont/sq). miniasm does no consensus at all - its output carries the full raw read error rate and is unusable until polished, so use it only as a fast layout inside a polished pipeline.

Haplotig False-Duplication and purge_dups

A diploid assembly from noisy reads either collapses heterozygous loci or emits both haplotypes as separate primary contigs (haplotig duplication). The diagnostic trifecta: (1) assembly size 1.5-2x the expected genome size; (2) a bimodal read-depth histogram with a half-coverage peak (haplotigs split reads between two copies); (3) inflated BUSCO-Duplicated. Fix with purge_dups (read-depth + self-alignment):

bash
minimap2 -xmap-ont asm.fa reads.fq.gz | gzip > aln.paf.gz
pbcstat aln.paf.gz && calcuts PB.stat > cutoffs          # INSPECT the coverage histogram before trusting auto-cuts
split_fa asm.fa > asm.split && minimap2 -xasm5 -DP asm.split asm.split | gzip > self.paf.gz
purge_dups -2 -T cutoffs -c PB.base.cov self.paf.gz > dups.bed
get_seqs -e dups.bed asm.fa                              # -> purged.fa (primary) + hap.fa (haplotigs)

purge_dups cannot tell a haplotig from a real recent segmental duplication/paralog - both look like similar sequence at fractional depth. On a genome with known recent WGD or high SD content (many plants), over-purging deletes real genes; eyeball the histogram and validate against a related assembly. True haplotype-resolved assembly is a HiFi capability - do not promise phasing from noisy reads (-> hifi-assembly).

Pre-Assembly: Mid-Read Adapter De-Chimerization

bash
porechop_abi -abi -i reads.fq.gz -o trimmed.fq.gz -t 16   # ab initio adapter detection; SPLITS internal-adapter chimeras

ONT occasionally sequences two molecules as one read with an internal adapter - an untrimmed chimera becomes a structural mis-join (a layout error polishing cannot fix). Porechop_ABI detects adapters ab initio (no fixed DB, which matters because kit adapter sequences change) and splits chimeric reads. Use it, not the original Porechop (unmaintained since 2018, frozen adapter DB). PacBio CLR handles adapter/scrap removal upstream on the instrument, so this is an ONT-specific concern.

Per-Method Failure Modes

Flag noisier than reads actually are

Trigger: --nano-raw (or low-error-rate omission) on R10/Dorado-SUP data. Mechanism: assembler treats real repeat-copy and allele differences as noise and over-collapses. Symptom: fewer contigs, HIGHER N50, lost repeats/SVs - looks better. Fix: match the flag to the basecaller era (--nano-hq for R10); record pore+kit+model before assembling.

Show full SKILL.md (808 more words)Show less
"It assembled, so it's done"

Trigger: shipping the raw assembler FASTA on its high N50. Mechanism: raw consensus is Q20-Q30; indels frameshift genes. Symptom: broken gene models, failed variant calling, BLAST misses present genes. Fix: polish (long-read Medaka/Racon, model-matched), then measure QV with Merqury; an assembly without a stated QV is a draft.

Trusting polishing to fix a structural error

Trigger: expecting polish to repair a mis-resolved repeat, inversion, collapsed tandem array, or dropped plasmid. Mechanism: polishers correct per-base consensus (QV) only; they never change contig layout. Symptom: a Q50 assembly that is still structurally wrong. Fix: catch layout errors with read-back coverage uniformity, multi-assembler consensus (Trycycler/Autocycler), or Hi-C - not QV.

Haplotig duplication read as completeness

Trigger: "my genome is bigger than expected - more complete!" Mechanism: both haplotypes kept as primary contigs. Symptom: size 1.5-2x expected, half-coverage depth peak, inflated BUSCO-Duplicated. Fix: purge_dups (inspect cutoffs); do not over-purge real SDs.

High coverage from a short-N50 library

Trigger: 200x of reads with a 5 kb N50, expecting contiguity. Mechanism: reads physically cannot span long repeats; depth buys consensus, not spanning. Symptom: fragmented assembly that more depth never fixes. Fix: coverage and read-N50 are non-substitutable; spend effort on read length (extraction, size selection) or ultra-long ONT.

Over-filtering by length

Trigger: aggressive Filtlong/chopper length filter to "keep the best reads." Mechanism: the longest reads are often not the highest quality; the filter discards the long-but-lower-Q reads that span hard repeats. Symptom: lost contiguity. Fix: filter conservatively; spanning reads are precious.

Quantitative Thresholds

ThresholdSourceRationale
Coverage ~30-60xfield convention (approx)below ~20-30x consensus too thin (per-base accuracy collapses); beyond ~60x more of the same reads adds little contiguity and slows overlap
--asm-coverage 40 with --genome-sizeFlye usagedownsample to longest 40x for the initial assembly on deep large genomes
Canu correctedErrorRate ~0.144 ONT / ~0.045 PacBioCanu 2.2 referencemaster knob; raise for heterozygosity, lower for clean high-coverage
Assembly size 1.5-2x expected = red flagdiploid normhaplotig false-duplication; confirm with half-coverage peak + BUSCO-Duplicated -> purge_dups
Merqury QV40 (~1 error/10 kb), Q50 reference-gradeRhie 2020 Genome Biolthe QV stop signal; report it, never an N50 alone
R9 raw ~Q15-17, R10 simplex Q20+, duplex Q30+Wick-blog / ONT (approx)sets the input flag and whether short-read polishing is even needed
wtdbg2 -x ont -L default 5000wtdbg2 presetsilently discards reads shorter than 5 kb

Common Errors

Error / symptomCauseSolution
Fewer contigs, higher N50, but lost variation--nano-raw on R10/Dorado data (over-collapse)use --nano-hq; match the basecaller era
Gene prediction frameshifts everywhereunpolished low-QV assemblypolish then measure QV (Merqury); -> assembly-polishing
Assembly ~2x expected size, BUSCO-Duplicated highuncollapsed haplotigspurge_dups; inspect coverage cutoffs first
wtdbg2 output is empty/shortforgot the wtpoa-cns consensus steprun wtpoa-cns on .ctg.lay.gz
miniasm assembly full of errorsminiasm does no consensuspolish (Racon/Medaka) or use a consensus assembler
Chimeric contigs / structural mis-joinsinternal-adapter chimeric readsde-chimerize with Porechop_ABI before assembly
Canu runs for days, huge RAMnormal for the correction stage on large genomesuseGrid=true on a cluster, or use Flye

References

  • Kolmogorov M, Yuan J, Lin Y, Pevzner PA. 2019. Assembly of long, error-prone reads using repeat graphs (Flye). Nat Biotechnol 37:540-546.
  • Koren S, Walenz BP, Berlin K, Miller JR, Bergman NH, Phillippy AM. 2017. Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Genome Res 27:722-736.
  • Hu J, Wang Z, Sun Z, et al. 2024. NextDenovo: an efficient error correction and accurate assembly tool for noisy long reads. Genome Biol 25:107.
  • Shafin K, Pesout T, Lorig-Roach R, et al. 2020. Nanopore sequencing and the Shasta toolkit enable efficient de novo assembly of eleven human genomes. Nat Biotechnol 38:1044-1053.
  • Vaser R, Sikic M. 2021. Time- and memory-efficient genome assembly with Raven. Nat Comput Sci 1:332-336.
  • Ruan J, Li H. 2020. Fast and accurate long-read assembly with wtdbg2. Nat Methods 17:155-158.
  • Li H. 2016. Minimap and miniasm: fast mapping and de novo assembly for noisy long sequences. Bioinformatics 32:2103-2110.
  • Guan D, McCarthy SA, Wood J, Howe K, Wang Y, Durbin R. 2020. Identifying and removing haplotypic duplication in primary genome assemblies (purge_dups). Bioinformatics 36:2896-2898.
  • Wick RR, Judd LM, Cerdeira LT, et al. 2021. Trycycler: consensus long-read assemblies for bacterial genomes. Genome Biol 22:266.
  • Rhie A, Walenz BP, Koren S, Phillippy AM. 2020. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol 21:245.
  • genome-profiling - k-mer genome-size estimate feeds Flye --genome-size and the haplotig sanity check
  • hifi-assembly - PacBio HiFi (Q30+) phased haplotype-resolved assembly with hifiasm; the right tool for accurate reads
  • assembly-polishing - Polishes the contiguous-but-low-QV contigs this skill produces; assembly is not finished without it
  • assembly-qc - QUAST/BUSCO/Merqury; QV is the deliverable, N50 alone is a vanity metric
  • long-read-sequencing/long-read-qc - Read length/quality QC and conservative filtering before assembly
  • long-read-sequencing/basecalling - Basecaller era and model that determine the assembler input flag
  • read-qc/contamination-screening - Screen reads for host/vector contamination before assembling
  • workflows/genome-assembly-pipeline - End-to-end QC -> assemble -> polish -> scaffold -> QC

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in genome-assembly/long-read-assembly of GPTomics/bioSkills.

  • SKILL.md
  • examples/flye_assembly.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Assembly Long Read Assembly next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Assembly Long Read Assembly compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Assembly Long Read Assembly this skillGPTomics/bioSkills1.2k1 repos~4.6kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Genome Assembly Long Read Assembly

What does Bio Genome Assembly Long Read Assembly do?

Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm…. Bio Genome Assembly Long Read Assembly is an agent skill from GPTomics/bioSkills. Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm, and reconciles bacterial assemblies into a consensus with Trycycler/Autocycler.

When should I use Bio Genome Assembly Long Read Assembly?

Bio Genome Assembly Long Read Assembly fits situations like: assembling a bacterial; eukaryotic genome from ONT; pacBio noisy reads; choosing a long-read assembler.

How do I install Bio Genome Assembly Long Read Assembly in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a claude-code`. Or copy the skill folder (genome-assembly/long-read-assembly in GPTomics/bioSkills) into .claude/skills/bio-genome-assembly-long-read-assembly in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Assembly Long Read Assembly in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a codex`. Or copy the skill folder (genome-assembly/long-read-assembly in GPTomics/bioSkills) into .agents/skills/bio-genome-assembly-long-read-assembly in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Assembly Long Read Assembly in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-assembly-long-read-assembly, .gemini/skills/bio-genome-assembly-long-read-assembly, .github/skills/bio-genome-assembly-long-read-assembly and .opencode/skills/bio-genome-assembly-long-read-assembly in your project.

What does Bio Genome Assembly Long Read Assembly need to run?

Going by SKILL.md and its folder, Bio Genome Assembly Long Read Assembly needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Bio Genome Assembly Long Read Assembly access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Genome Assembly Long Read Assembly safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Assembly Long Read Assembly use?

Bio Genome Assembly Long Read Assembly is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Assembly Long Read Assembly use?

About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Assembly Long Read Assembly?

Skills that share tags, products or a category with Bio Genome Assembly Long Read Assembly: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Assembly Long Read Assembly?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.