Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm…
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-long-read-assembly --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-assembly/long-read-assembly .claude/skills/bio-genome-assembly-long-read-assembly && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-genome-assembly-long-read-assembly" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/long-read-assembly into .claude/skills/bio-genome-assembly-long-read-assembly/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-long-read-assembly", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/long-read-assemblyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-long-read-assembly --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/genome-assembly/long-read-assembly .agents/skills/bio-genome-assembly-long-read-assembly && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-genome-assembly-long-read-assembly" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/long-read-assembly into .agents/skills/bio-genome-assembly-long-read-assembly/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-long-read-assembly", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-long-read-assembly --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/genome-assembly/long-read-assembly .cursor/skills/bio-genome-assembly-long-read-assembly && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-genome-assembly-long-read-assembly" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/long-read-assembly into .cursor/skills/bio-genome-assembly-long-read-assembly/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-long-read-assembly", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path genome-assembly/long-read-assembly--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-long-read-assembly --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/genome-assembly/long-read-assembly .gemini/skills/bio-genome-assembly-long-read-assembly && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-genome-assembly-long-read-assembly" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/long-read-assembly into .gemini/skills/bio-genome-assembly-long-read-assembly/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-long-read-assembly", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-long-read-assemblyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/genome-assembly/long-read-assembly .github/skills/bio-genome-assembly-long-read-assembly && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-genome-assembly-long-read-assembly" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/long-read-assembly into .github/skills/bio-genome-assembly-long-read-assembly/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-long-read-assembly", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-long-read-assembly --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/genome-assembly/long-read-assembly .opencode/skills/bio-genome-assembly-long-read-assembly && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-genome-assembly-long-read-assembly" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/long-read-assembly into .opencode/skills/bio-genome-assembly-long-read-assembly/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-long-read-assembly", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-genome-assembly-long-read-assemblyAssembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm…
Bio Genome Assembly Long Read Assembly is an agent skill from GPTomics/bioSkills. Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm, and reconciles bacterial assemblies into a consensus with Trycycler/Autocycler. Covers matching the input flag to the basecaller era (--nano-hq vs --nano-raw), why a raw long-read assembly is contiguous but low-QV and not finished until polished, haplotig false-duplication and purgedups, coverage and read-N50 as…
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/flye_assembly.sh` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Genome Assembly Long Read Assembly loads about 4.6k tokens when it runs. Until then it costs about 208 tokens; SKILL.md has 2,070 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,070 words, ~4,627 tokens.
.claude/skills/bio-genome-assembly-long-read-assembly/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: Flye 2.9+, Canu 2.2+, NextDenovo 2.5+, Shasta 0.11+, Raven 1.8+, wtdbg2 2.5+, miniasm 0.3+, minimap2 2.26+, purge_dups 1.2+, Porechop_ABI 0.5+, Trycycler 0.5+.
Before using code patterns, verify installed versions match. If versions differ:
<tool> --version then <tool> --help to confirm flagsTwo behaviors are version-dependent and load-bearing: Flye --genome-size became optional (auto-estimated) and --scaffold flipped OFF by default in recent releases - confirm against the installed Flye. Shasta --config names are dated and chemistry-specific (e.g. Nanopore-R10-Fast-Nov2022) - list with shasta --command listConfigurations rather than hard-coding one. Canu read-type flag spellings and correctedErrorRate defaults changed across major versions. If a command errors, introspect the installed tool and adapt rather than retrying.
"Assemble a genome from Nanopore/PacBio-CLR long reads" -> Build a contiguous de novo assembly whose input mode matches the basecaller error regime, then hand off to polishing because the raw consensus is contiguous but not yet accurate.
flye --nano-hq reads.fq.gz --out-dir out -t 16 (modern ONT R10/Dorado), canu -p asm -d out genomeSize=4.6m -nanopore reads.fq.gz (thorough), wtdbg2 -x ont -g 4.6m -i reads.fq.gz -fo asm && wtpoa-cns -i asm.ctg.lay.gz -fo asm.fa (fast draft)Two load-bearing facts no assembler README states:
The basecaller era dictates the input flag, and a mismatch silently wrecks the assembly while the report looks finished. Noisy long reads are not one error regime: ONT R9.4.1+Guppy is ~5-10% error, ONT R10.4.1+Dorado SUP is Q20+ (~1-2%), PacBio CLR is ~10-15%. Each assembler has separate modes tuned to each (Flye --nano-raw / --nano-hq / --pacbio-raw). The dangerous direction is telling the assembler the reads are noisier than they are - feeding R10/Dorado data to --nano-raw makes Flye treat real repeat-copy and allele differences as noise and over-collapse them. The result is fewer contigs and a HIGHER N50 as the assembly gets worse - the most dangerous failure mode because the headline metric improved and nothing crashes. The chemistry/kit/basecaller model is an assembly parameter, not optional metadata; if it is unknown, the flag cannot be chosen and the assembly is uninterpretable - find out, do not guess.
The bottleneck flipped from contiguity to consensus accuracy. A raw noisy-long-read assembly emits a FASTA with a spectacular N50 that looks finished, but per-base accuracy is often Q20-Q30 (one error every ~100-1000 bp). The dominant ONT error is the indel, especially in homopolymers, and a single indel frameshifts a protein - so a contiguous assembly can have every gene model broken. Contiguity and correctness are orthogonal axes; N50 is blind to QV. The assembler is the halfway point: the deliverable is a polished assembly with a measured QV (Merqury), not a high-N50 FASTA. Flye does one internal polishing round and stops - that is not "polished." Hand off to assembly-polishing and assembly-qc.
| Tool | Citation | Paradigm | When |
|---|---|---|---|
| Flye | Kolmogorov 2019 Nat Biotechnol | repeat graph (disjointigs -> explicit repeat structure) | the fast general-purpose default; bacteria -> eukaryote |
| Canu | Koren 2017 Genome Res | correct -> trim -> assemble (OLC, MHAP) | maximum single-assembler quality; slow, grid-oriented |
| NextDenovo | Hu 2024 Genome Biol | correct (NextCorrect) -> string-graph (NextGraph) | large/repetitive plant and animal genomes; contiguity-first |
| Shasta | Shafin 2020 Nat Biotechnol | run-length-encoded marker graph | ONT human-scale, speed/cost critical |
| Raven | Vaser & Sikic 2021 Nat Comput Sci | OLC string graph, near parameter-free | fast simple draft; common Trycycler input |
| wtdbg2 | Ruan & Li 2020 Nat Methods | fuzzy de Bruijn graph (+ mandatory wtpoa-cns) | fastest, lowest RAM; lowest accuracy -> polish hard |
| miniasm | Li 2016 Bioinformatics | string-graph layout, NO consensus | layout demo / pipeline component only; output = raw read error |
| Trycycler / Autocycler | Wick 2021 Genome Biol / Wick 2025 | consensus of multiple independent assemblies | bacterial reliability; catches single-assembler structural errors |
| purge_dups | Guan 2020 Bioinformatics | read-depth + self-alignment | remove haplotig false-duplication from a diploid primary |
PacBio CLR (--pacbio-raw) is legacy - superseded by HiFi for all new PacBio work; treat CLR support as maintenance for archival data, do not recommend generating new CLR.
| Scenario | Recommended | Why |
|---|---|---|
| ONT R10.4.1 / Dorado HAC-SUP, any genome | Flye --nano-hq | modern default mode; matches the Q20+ error regime |
| ONT R9.4.1 SUP (Guppy5+/Dorado) | Flye --nano-hq --read-error 0.05 | hq mode with the error floor raised to R9-SUP |
| ONT R9.4.1 legacy fast/HAC | Flye --nano-raw | the genuinely-noisy mode; do not use on R10 |
| PacBio CLR (archival) | Flye --pacbio-raw or Canu -pacbio | legacy noisy mode; polish hard (arrow/GCpp) |
| Bacterial isolate, want a correct finished genome | Trycycler (interactive) / Autocycler (automated) | consensus across assemblers; fixes structural errors polishing can't |
| Large repetitive plant/animal, contiguity-first | NextDenovo | memory-efficient, top contiguity on big genomes |
| ONT human-scale, speed-critical | Shasta --config <era-matched> | RLE marker graph; human genomes in days |
| Quick draft / compute is the bottleneck | wtdbg2 or Raven | fastest; accept lower accuracy then polish |
| PacBio HiFi (Q30+, CCS) | -> hifi-assembly | hifiasm phased haplotypes; wrong tool here |
| After assembling (always) | -> assembly-polishing then assembly-qc | raw consensus is low-QV; not finished until polished + QV-measured |
| Reads not yet QC'd / unknown chemistry | -> long-read-sequencing/long-read-qc, long-read-sequencing/basecalling | garbage-in caps the assembly; basecaller model sets the flag |
| Genome size / coverage unknown | GenomeScope2 on accurate short reads (not raw ONT) | k-mer histograms from noisy reads inflate unique k-mers |
flye --nano-hq reads.fq.gz --out-dir out -t 16 # ONT R10 / Dorado SUP
flye --nano-hq reads.fq.gz --read-error 0.05 --out-dir out -t 16 # ONT R9 SUP
flye --nano-raw reads.fq.gz --out-dir out -t 16 # legacy ONT R9 fast/HAC
flye --pacbio-raw reads.fq.gz --out-dir out -t 16 # PacBio CLR (legacy)
flye --nano-hq reads.fq.gz --genome-size 3g --asm-coverage 40 -o out -t 32 # large genome: use longest 40x for initial assembly--genome-size is optional in recent Flye (auto-estimated) but required when paired with --asm-coverage, which downsamples to the longest N-coverage of reads for the initial disjointig step (cuts runtime/RAM on deep large-genome data; the rest are still used). --iterations defaults to 1 polishing round (0 to skip); --keep-haplotypes retains alt bubble paths for diploid awareness; --meta is metaFlye for uneven-coverage communities (-> metagenome-assembly). Output: assembly.fasta, assembly_info.txt (per-contig length/coverage/circularity), assembly_graph.gfa.
canu -p asm -d out genomeSize=4.6m -nanopore reads.fq.gz useGrid=false maxThreads=16 # correct->trim->assemble
wtdbg2 -x ont -g 4.6m -t 16 -i reads.fq.gz -fo asm && wtpoa-cns -t 16 -i asm.ctg.lay.gz -fo asm.ctg.fa # consensus step is MANDATORY
raven -t 16 reads.fq.gz > asm.fasta # near parameter-freeCanu's master meta-parameter is correctedErrorRate (max expected difference between two corrected reads): defaults ~0.144 (Nanopore) / ~0.045 (PacBio); raise for heterozygosity/divergence, lower for clean high-coverage data. Read-type flags -nanopore / -pacbio / -pacbio-hifi (there is no -nanopore-hifi; high-accuracy ONT still uses -nanopore); useGrid=false forces a single machine. wtdbg2 needs the separate wtpoa-cns consensus call - the -x preset (ont/sq/rs/ccs) is set FIRST, and -L discards short reads (default 5000 for ont/sq). miniasm does no consensus at all - its output carries the full raw read error rate and is unusable until polished, so use it only as a fast layout inside a polished pipeline.
A diploid assembly from noisy reads either collapses heterozygous loci or emits both haplotypes as separate primary contigs (haplotig duplication). The diagnostic trifecta: (1) assembly size 1.5-2x the expected genome size; (2) a bimodal read-depth histogram with a half-coverage peak (haplotigs split reads between two copies); (3) inflated BUSCO-Duplicated. Fix with purge_dups (read-depth + self-alignment):
minimap2 -xmap-ont asm.fa reads.fq.gz | gzip > aln.paf.gz
pbcstat aln.paf.gz && calcuts PB.stat > cutoffs # INSPECT the coverage histogram before trusting auto-cuts
split_fa asm.fa > asm.split && minimap2 -xasm5 -DP asm.split asm.split | gzip > self.paf.gz
purge_dups -2 -T cutoffs -c PB.base.cov self.paf.gz > dups.bed
get_seqs -e dups.bed asm.fa # -> purged.fa (primary) + hap.fa (haplotigs)purge_dups cannot tell a haplotig from a real recent segmental duplication/paralog - both look like similar sequence at fractional depth. On a genome with known recent WGD or high SD content (many plants), over-purging deletes real genes; eyeball the histogram and validate against a related assembly. True haplotype-resolved assembly is a HiFi capability - do not promise phasing from noisy reads (-> hifi-assembly).
porechop_abi -abi -i reads.fq.gz -o trimmed.fq.gz -t 16 # ab initio adapter detection; SPLITS internal-adapter chimerasONT occasionally sequences two molecules as one read with an internal adapter - an untrimmed chimera becomes a structural mis-join (a layout error polishing cannot fix). Porechop_ABI detects adapters ab initio (no fixed DB, which matters because kit adapter sequences change) and splits chimeric reads. Use it, not the original Porechop (unmaintained since 2018, frozen adapter DB). PacBio CLR handles adapter/scrap removal upstream on the instrument, so this is an ONT-specific concern.
Trigger: --nano-raw (or low-error-rate omission) on R10/Dorado-SUP data. Mechanism: assembler treats real repeat-copy and allele differences as noise and over-collapses. Symptom: fewer contigs, HIGHER N50, lost repeats/SVs - looks better. Fix: match the flag to the basecaller era (--nano-hq for R10); record pore+kit+model before assembling.
Trigger: shipping the raw assembler FASTA on its high N50. Mechanism: raw consensus is Q20-Q30; indels frameshift genes. Symptom: broken gene models, failed variant calling, BLAST misses present genes. Fix: polish (long-read Medaka/Racon, model-matched), then measure QV with Merqury; an assembly without a stated QV is a draft.
Trigger: expecting polish to repair a mis-resolved repeat, inversion, collapsed tandem array, or dropped plasmid. Mechanism: polishers correct per-base consensus (QV) only; they never change contig layout. Symptom: a Q50 assembly that is still structurally wrong. Fix: catch layout errors with read-back coverage uniformity, multi-assembler consensus (Trycycler/Autocycler), or Hi-C - not QV.
Trigger: "my genome is bigger than expected - more complete!" Mechanism: both haplotypes kept as primary contigs. Symptom: size 1.5-2x expected, half-coverage depth peak, inflated BUSCO-Duplicated. Fix: purge_dups (inspect cutoffs); do not over-purge real SDs.
Trigger: 200x of reads with a 5 kb N50, expecting contiguity. Mechanism: reads physically cannot span long repeats; depth buys consensus, not spanning. Symptom: fragmented assembly that more depth never fixes. Fix: coverage and read-N50 are non-substitutable; spend effort on read length (extraction, size selection) or ultra-long ONT.
Trigger: aggressive Filtlong/chopper length filter to "keep the best reads." Mechanism: the longest reads are often not the highest quality; the filter discards the long-but-lower-Q reads that span hard repeats. Symptom: lost contiguity. Fix: filter conservatively; spanning reads are precious.
| Threshold | Source | Rationale |
|---|---|---|
| Coverage ~30-60x | field convention (approx) | below ~20-30x consensus too thin (per-base accuracy collapses); beyond ~60x more of the same reads adds little contiguity and slows overlap |
--asm-coverage 40 with --genome-size | Flye usage | downsample to longest 40x for the initial assembly on deep large genomes |
Canu correctedErrorRate ~0.144 ONT / ~0.045 PacBio | Canu 2.2 reference | master knob; raise for heterozygosity, lower for clean high-coverage |
| Assembly size 1.5-2x expected = red flag | diploid norm | haplotig false-duplication; confirm with half-coverage peak + BUSCO-Duplicated -> purge_dups |
| Merqury QV40 (~1 error/10 kb), Q50 reference-grade | Rhie 2020 Genome Biol | the QV stop signal; report it, never an N50 alone |
| R9 raw ~Q15-17, R10 simplex Q20+, duplex Q30+ | Wick-blog / ONT (approx) | sets the input flag and whether short-read polishing is even needed |
wtdbg2 -x ont -L default 5000 | wtdbg2 preset | silently discards reads shorter than 5 kb |
| Error / symptom | Cause | Solution |
|---|---|---|
| Fewer contigs, higher N50, but lost variation | --nano-raw on R10/Dorado data (over-collapse) | use --nano-hq; match the basecaller era |
| Gene prediction frameshifts everywhere | unpolished low-QV assembly | polish then measure QV (Merqury); -> assembly-polishing |
| Assembly ~2x expected size, BUSCO-Duplicated high | uncollapsed haplotigs | purge_dups; inspect coverage cutoffs first |
| wtdbg2 output is empty/short | forgot the wtpoa-cns consensus step | run wtpoa-cns on .ctg.lay.gz |
| miniasm assembly full of errors | miniasm does no consensus | polish (Racon/Medaka) or use a consensus assembler |
| Chimeric contigs / structural mis-joins | internal-adapter chimeric reads | de-chimerize with Porechop_ABI before assembly |
| Canu runs for days, huge RAM | normal for the correction stage on large genomes | useGrid=true on a cluster, or use Flye |
--genome-size and the haplotig sanity check© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in genome-assembly/long-read-assembly of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Genome Assembly Long Read Assembly next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Genome Assembly Long Read Assembly this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.6k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm…. Bio Genome Assembly Long Read Assembly is an agent skill from GPTomics/bioSkills. Assembles genomes de novo from noisy long reads (Oxford Nanopore R9/R10/Dorado, PacBio CLR) with Flye (repeat graph), Canu (correct-trim-assemble OLC), NextDenovo, Shasta, Raven, wtdbg2, or miniasm, and reconciles bacterial assemblies into a consensus with Trycycler/Autocycler.
Bio Genome Assembly Long Read Assembly fits situations like: assembling a bacterial; eukaryotic genome from ONT; pacBio noisy reads; choosing a long-read assembler.
Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a claude-code`. Or copy the skill folder (genome-assembly/long-read-assembly in GPTomics/bioSkills) into .claude/skills/bio-genome-assembly-long-read-assembly in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a codex`. Or copy the skill folder (genome-assembly/long-read-assembly in GPTomics/bioSkills) into .agents/skills/bio-genome-assembly-long-read-assembly in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-long-read-assembly -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-assembly-long-read-assembly, .gemini/skills/bio-genome-assembly-long-read-assembly, .github/skills/bio-genome-assembly-long-read-assembly and .opencode/skills/bio-genome-assembly-long-read-assembly in your project.
Going by SKILL.md and its folder, Bio Genome Assembly Long Read Assembly needs a shell for the scripts in its folder. Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Genome Assembly Long Read Assembly is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Genome Assembly Long Read Assembly: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.