Agent skill

Bio Genome Assembly Scaffolding

by GPTomics in GPTomics/bioSkills

Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence).

MITAuto-check passedDevelopment

Install Bio Genome Assembly Scaffolding

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-assembly-scaffolding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-assembly/scaffolding .claude/skills/bio-genome-assembly-scaffolding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-assembly-scaffolding
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.6k tokens
SKILL.md length
2,485 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence).

  • Works in 3 steps: The contact map is the QC, not a… → Scaffold N50 is not contig N50, and… → Reordering cannot rescue a bad contig.…
  • Turning contigs into chromosomes with Hi-C
  • SKILL.md covers Version Compatibility, The Single Most Important…, Linking-Data Modality Taxonomy and Decision Tree by Available…, plus 12 more sections
  • Runs Shell scripts from its folder; calls java and python

What it does

Bio Genome Assembly Scaffolding is an agent skill from GPTomics/bioSkills. Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence). Covers Hi-C/Omni-C scaffolding (YaHS, SALSA2, 3D-DNA/Juicer), Hi-C read-mapping prerequisites (map each end separately, no mate rescue, dedup, enzyme-aware), reading the contact map for misjoins/inversions/false-duplications, manual curation in Juicebox/PretextView (the VGP/DToL standard), reference-guided scaffolding (RagTag) and its karyotype-erasure hazard, genetic-map…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/yahs_scaffolding.sh` and `usage-guide.md`).

It sits in Development, covering Project scaffolding and Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Turning contigs into chromosomes with Hi-C
  • Integrating a linkage map
  • Choosing a scaffolder by available linking data
  • Judging whether a chromosome-scale assembly is trustworthy

Example prompts

  • “Use the bio-genome-assembly-scaffolding skill to order and orients assembled contigs into chromosome-scale scaffolds from long-range linking data…”
  • “/bio-genome-assembly-scaffolding”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The contact map is the QC, not a decoration. A correct chromosome shows one bright diagonal with smooth off-diagonal decay; off-diagonal…
  2. Scaffold N50 is not contig N50, and conflating them is the classic reviewer catch. Scaffolding inserts runs of N (estimated, often…
  3. Reordering cannot rescue a bad contig. Scaffolders take contigs as given; a chimeric contig drags the wrong sequence into the wrong…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • java
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Assembly Scaffolding loads about 5.6k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 2,485 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~225
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,485 words, ~5,574 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-assembly-scaffolding/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-genome-assembly-scaffolding
description
Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence). Covers Hi-C/Omni-C scaffolding (YaHS, SALSA2, 3D-DNA/Juicer), Hi-C read-mapping prerequisites (map each end separately, no mate rescue, dedup, enzyme-aware), reading the contact map for misjoins/inversions/false-duplications, manual curation in Juicebox/PretextView (the VGP/DToL standard), reference-guided scaffolding (RagTag) and its karyotype-erasure hazard, genetic-map (ALLMAPS) and Bionano optical-map integration, chimera-breaking before scaffolding, gap-filling, and telomere/contig-vs-scaffold-N50 QC (tidk). Use when turning contigs into chromosomes with Hi-C, integrating a linkage map or optical map, choosing a scaffolder by available linking data, or judging whether a chromosome-scale assembly is trustworthy.
tool_type
cli
primary_tool
YaHS

Version Compatibility

Reference examples tested with: YaHS 1.2+, SALSA2 2.3+, juicer_tools 1.22/2.0+, bwa 0.7.17+, chromap 0.2+, samtools 1.19+, RagTag 2.1+, tidk 0.2.3+, seqkit 2.6+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

hifiasm output filenames and YaHS/SALSA2 enzyme and resolution defaults have changed across releases; confirm output names and -e/-r behaviour against the installed version. The Hi-C-mapping flag set (bwa mem -5SP vs the Arima/VGP per-end pipeline) varies by vendor (Arima, Dovetail/Omni-C, Phase Genomics) - check the current mapping guide before pasting flags. If a command errors, introspect the tool and adapt rather than retrying.

Genome Scaffolding

"Turn my contigs into chromosomes" -> Order and orient contigs by long-range linking signal, insert N-gap spacers, then curate the contact map - scaffolding adds no sequence, so the output is a draft chromosome structure, not finished sequence.

  • CLI: yahs contigs.fa hic_to_contigs.bam (Hi-C default), run_pipeline.py -a contigs.fa -b sorted.bed -m yes (SALSA2), ragtag.py scaffold ref.fa contigs.fa (reference-guided), allmaps path (linkage maps)

The Single Most Important Modern Insight -- Automated Scaffolding Produces a Draft; the Genome Is Trustworthy Only After Manual Curation of the Contact Map

YaHS finishes in minutes and emits a file literally named _scaffolds_final.agp. It is not final. Hi-C scaffolding is a contact-frequency inference - it orders/orients contigs by the polymer-physics signal that contact frequency decays with 1D distance - and it is error-prone exactly at the joins. Every reference-grade pipeline (VGP, Darwin Tree of Life, Earth BioGenome) treats the automated AGP as a first pass that a human curator then corrects in PretextView or Juicebox: breaking misjoins, flipping inversions, removing false duplications, assigning chromosomes by eye. Howe 2021 quantified the hidden labor across 111 VGP/DToL assemblies: on average 221 interventions per Gb (67 breaks, 105 joins, 49 false-duplication removals). Three load-bearing consequences:

  1. The contact map is the QC, not a decoration. A correct chromosome shows one bright diagonal with smooth off-diagonal decay; off-diagonal blocks, anti-diagonal "bowties," and bleeding between chromosomes are errors to inspect and break (see Reading the Contact Map). An assembly paper whose methods say "scaffolded with [tool]" but never mention curation/Juicebox/PretextView is shipping a draft - downgrade any claim that depends on large-scale order (synteny, fusions, structural variants).

  2. Scaffold N50 is not contig N50, and conflating them is the classic reviewer catch. Scaffolding inserts runs of N (estimated, often arbitrary length) and adds zero sequence. A genome can post "scaffold N50 = 60 Mb, chromosome-scale!" while its contig N50 is 200 kb - the contiguity is borne by gaps, not finished sequence. Always report contig N50, scaffold N50, gap count, and total N bases separately.

  3. Reordering cannot rescue a bad contig. Scaffolders take contigs as given; a chimeric contig drags the wrong sequence into the wrong chromosome. Chimeras must be broken before scaffolding (see Chimera-Breaking).

Linking-Data Modality Taxonomy

ModalityHow it linksScaleStatus
Hi-C / Omni-Cproximity ligation; contact frequency ~ 1/(1D distance)chromosome-scaleDominant modern method. Omni-C is enzyme-free/sequence-agnostic; Arima uses fixed enzyme motifs
Bionano optical maps (DLS)labeled motif patterns aligned to in-silico contig mapsmegabaseResolves large SVs and bridges complex repeats; separate instrument + library (cost decision)
10x linked readsbarcoded short reads from one long molecule~10-100 kbDISCONTINUED by 10x (2020). Legacy data only; Tigmint/ARCS still run, no new data
Genetic / linkage mapsmarker recombination order = chromosome orderchromosome-scale, low resolutionALLMAPS integrates multiple maps; orthogonal validation of Hi-C
Reference (homology)align contigs to a related genome, copy its orderas good as syntenyRagTag. Fast, but imposes the reference's karyotype - see hazard below
Long reads as linkersk-mer pairs spanning junctionsread lengthLINKS; mostly superseded by assembling with the long reads directly
Mate-pair / jumpinglarge-insert paired short reads2-40 kbLegacy (SSPACE); chimeric-insert artifacts; never recommend today

Decision Tree by Available Linking Data

Available linking dataUseWhy
Hi-C/Omni-C, want fast contiguous defaultYaHScommunity default; fast, high N90, AGP + Juicebox outputs in one command
Hi-C, want graph-aware conservative joins + input-error correctionSALSA2 -m yes (-g graph.gfa)iterative; breaks chimeric input; uses assembly graph to avoid orientation errors
Hi-C, interactive curation central to workflow3D-DNA + Juicer + Juicebox (JBAT)the Aiden-lab .assembly <-> Juicebox round-trip is built around hand-editing
Hi-C, vertebrate/reference-gradeYaHS or SALSA2 -> PretextView/JBAT curationVGP/DToL standard: automate then curate
Diploid, Hi-C, want both haplotypes as chromosomeshifiasm --h1/--h2 to PHASE first, then YaHS to SCAFFOLD each haplotypesame Hi-C, two jobs - phasing picks the homolog, scaffolding orders along it (see Phasing vs Scaffolding)
Closely related reference, karyotype NOT a questionRagTag correct then scaffoldhomology ordering in minutes - but never if karyotype is the biology
One or more genetic/linkage mapsALLMAPSintegrates multiple maps; robust to marker errors; validates/anchors Hi-C
Bionano optical maps + sequence assemblyBionano Solve hybrid scaffold (before Hi-C)megabase maps bridge complex repeats and large SVs
10x linked reads (legacy data)Tigmint (break) -> ARCS/ARKS (link)break misassemblies first; no new 10x data exists
Contigs not yet QC'd / haplotigs present-> assembly-qc, purge_dupsscaffolding haplotigs strings them up as fake chromosomes
Hi-C contact map for TADs/loops, not chromosomes-> hi-c-analysis/matrix-operationsdifferent use of the same assay

Modern reference-grade recipe (vertebrate/eukaryote): HiFi -> hifiasm contigs -> (optional Bionano hybrid scaffold) -> Hi-C scaffold (YaHS/SALSA2) -> manual curation in PretextView/JBAT -> gap-fill (TGS-GapCloser) -> QC (tidk telomeres, contact-map diagonal, contig-vs-scaffold N50).

Hi-C Read Mapping (the step people botch)

Hi-C reads are not a normal paired-end library: the two ends come from different genomic loci ligated together, so standard PE proper-pair/insert-size logic mis-handles them.

bash
# Map each end SEPARATELY with no mate rescue/pairing; -5 reports the 5' (junction) portion as primary.
bwa index contigs.fa
bwa mem -5SP -T0 -t 16 contigs.fa hic_R1.fq.gz hic_R2.fq.gz | \
    samtools view -@ 8 -b - > aligned.bam
# Mandatory: mark/remove PCR + optical duplicates (Hi-C is duplicate-rich) before scaffolding.
samtools sort -@ 8 -n aligned.bam | samtools fixmate -m - - | \
    samtools sort -@ 8 - | samtools markdup -@ 8 - hic_to_contigs.bam
samtools index hic_to_contigs.bam
  • Each end separately, no mate rescue (-5SP): -5 = 5' portion of a chimeric junction read as primary; -S/-P skip mate rescue and pairing. The Arima/VGP per-end pipeline (bwa mem on R1 and R2 independently, filter_five_end.pl, two_read_bam_combiner.pl) is the higher-fidelity equivalent.
  • MAPQ filter at scaffolding (YaHS -q, SALSA2 reads MAPQ) to drop repeat-ambiguous reads.
  • Enzyme awareness: Arima/Dovetail-DpnII cut at fixed motifs (GATC); the scaffolder uses -e to model legitimate read starts. Omni-C / DNase Hi-C is enzyme-free -> omit -e in YaHS, use -e DNASE in SALSA2.
  • chromap (Zhang 2021 Nat Commun 12:6566) is the fast modern alternative: chromap --preset hic -r contigs.fa -x index -1 R1 -2 R2 ... aligns + dedups Hi-C ~10x faster; YaHS accepts its output.

YaHS (the current default)

bash
samtools faidx contigs.fa
# Enzyme: -e GATC (DpnII/Dovetail), -e GATC,GANTC (Arima 2-enzyme); OMIT -e for Omni-C/enzyme-free.
yahs -e GATC -o out contigs.fa hic_to_contigs.bam
# Outputs: out_scaffolds_final.agp + out_scaffolds_final.fa, intermediate out_rNN.agp, out.bin
# Contig error-correction is ON by default (--no-contig-ec to disable; usually a mistake to disable).

# Juicebox prep for curation (.hic + .assembly to load in JBAT):
# NOTE: `juicer` here is the small utility BUNDLED with YaHS (operates on the .bin), NOT Aiden-lab Juicer;
# `juicer_tools.jar` below IS the separate Aiden-lab jar. Do not confuse the two.
juicer pre -a -o out_JBAT out.bin out_scaffolds_final.agp contigs.fa.fai 2>tmp_assembly.log
java -Xmx48G -jar juicer_tools.jar pre out_JBAT.txt out_JBAT.hic <(cat tmp_assembly.log | grep PRE_C_SIZE | awk '{print $2" "$3}')
# After hand-editing in Juicebox, export out_JBAT.review.assembly, then:
juicer post -o out_curated out_JBAT.review.assembly out_JBAT.liftover.agp contigs.fa

YaHS also accepts a BED (bamToBed) or PA5/pairs input. The _scaffolds_final.agp is the editable, reviewable object - curation edits the AGP, then regenerates FASTA. --telo-motif lets YaHS use telomere signal during scaffolding.

SALSA2 (graph-aware, iterative, breaks chimeras)

bash
bamToBed -i aligned.bam > alignment.bed
sort -k4 alignment.bed > sorted.bed                 # MUST sort by read name (column 4)
python run_pipeline.py -a contigs.fa -l contigs.fa.fai -b sorted.bed \
       -e GATC -o salsa_out -m yes -g contigs_graph.gfa -p yes
# -e DNASE for Omni-C/enzyme-free; -m yes enables input-error (chimera) correction; -p yes writes AGP+FASTA per iteration

The signature -m yes step breaks input contigs where the Hi-C signal contradicts the assembler's join before scaffolding. Assembly headers must not contain :. -i sets iterations (default 3), -c min contig length (default 1000).

Reference-Guided Scaffolding (RagTag) -- Power and Peril

bash
ragtag.py correct ref.fa contigs.fa                 # break query at putative misassemblies vs reference FIRST
ragtag.py scaffold ref.fa ragtag_output/ragtag.correct.fasta -t 16   # order/orient by homology (minimap2 default)
# Outputs ragtag.scaffold.agp + .fasta; inserts 100 bp N gaps by default; -C collapses unplaced into chr0.

The peril (load-bearing): RagTag orders contigs to match the reference, so by construction the output looks like the reference. Every real biological rearrangement that distinguishes the target organism - a chromosome fusion, fission, translocation, or inversion polymorphism - is silently erased and replaced by the reference's structure. Never reference-scaffold a genome whose karyotype or large-scale structure is a biological question - the pipeline "discovers" the reference's karyotype because it was imposed (this has happened in the literature). Acceptable uses: a structurally conserved close relative needing only a coordinate system for gene content; a scaffold hypothesis to then test/correct against Hi-C (ragtag.py merge -b hic.bam lets Hi-C arbitrate). On a genome with interesting structure the honest move is Hi-C + curation, full stop.

Genetic/Linkage Maps (ALLMAPS) and Optical Maps (Bionano)

ALLMAPS (Tang 2015 Genome Biol 16:3) integrates one or more genetic maps to order/orient scaffolds: allmaps merge map1.csv map2.csv -o maps.bed then allmaps path maps.bed contigs.fa. It is the gold standard for validating chromosome assignment and for species where Hi-C is hard; weight multiple maps by quality. Bionano Solve hybrid scaffolding (no journal paper; cite the Bionano Solve Theory of Operation: Hybrid Scaffold technical document) aligns DLS optical maps to in-silico contig maps, resolves conflicts, and merges into hybrid scaffolds - run it before Hi-C on large/repetitive genomes where megabase maps bridge repeat arrays Hi-C fumbles. The decision is economic (separate instrument + library), not purely technical.

Reading the Contact Map (the skill nobody writes down)

The map displays its own errors to a trained eye - render with PretextMap -> PretextView (interactive, emits a curated AGP), PretextSnapshot (static PNG), or Juicebox/JBAT.

  • Correct chromosome: one bright diagonal, contacts decaying smoothly off-diagonal; inter-chromosome space dim and uniform.
  • Misjoin: the diagonal breaks - signal jumps off-diagonal into a separate block at the junction (two diagonals stitched at a corner). Break it.
  • Inversion: an anti-diagonal "bowtie"/butterfly - the gradient runs backwards through the segment. The most common curation fix once recognized; flip it.
  • Translocation / wrong assignment: off-diagonal bright blocks bleeding between what should be separate chromosomes.
  • False duplication (leaked haplotig): a faint off-diagonal stripe parallel to the diagonal contacting the same neighborhood. Remove it (49 of the 221 interventions/Gb).
  • Hardest judgment: the map is noisy near the diagonal even when correct. A real weak join is faint but coherent with the right decay shape; mapping noise is structureless speckle. Treat ambiguous near-diagonal signal as a flag to inspect, not a number to trust.
Show full SKILL.md (991 more words)Show less

Chimera-Breaking, Gap-Filling, and Phasing-vs-Scaffolding

  • Break, then scaffold: SALSA2 -m yes and YaHS default contig-EC break chimeric input where Hi-C contradicts the join; Tigmint (Jackman 2018 BMC Bioinformatics 19:393) breaks before ARCS for linked reads; RagTag correct for the reference-based version. Disabling these to "preserve the assembly" is a common self-inflicted wound.
  • Gap-filling is a SEPARATE downstream step that replaces N-runs with real sequence by bridging long reads across the gap (TGS-GapCloser; LR_Gapcloser). Gaps sit at repeats (that is why the assembler stopped there), so a read can bridge into the wrong repeat copy and insert locally-plausible but globally-wrong sequence - worse than an honest N. Gap-fill after curation, sanity-check filled lengths against the gap estimate, and report which gaps were closed vs left.
  • Hi-C-for-scaffolding vs Hi-C-for-phasing: the same library, two jobs. Scaffolding asks where on the chromosome (orders collapsed contigs); phasing asks which homolog (hifiasm --h1/--h2 during assembly). When handed "Hi-C, make chromosomes," the first question is haploid scaffolding or diploid phasing+scaffolding? - the tools and failure modes differ entirely.

Per-Method Failure Modes

Shipping the automated AGP as final

Trigger: publishing _scaffolds_final without curation. Mechanism: automated joins are inferences error-prone at junctions (~221 edits/Gb in VGP/DToL). Symptom: misjoins/inversions in the contact map; later-discovered wrong synteny/fusions. Fix: curate in PretextView/JBAT before any large-scale-order claim.

Conflating scaffold N50 with contig N50

Trigger: reporting only scaffold N50. Mechanism: scaffolding adds N-gaps, not sequence. Symptom: spectacular scaffold N50 over a small contig N50 - contiguity is nitrogen. Fix: report both N50s, gap count, total N bases separately.

Reference-scaffolding a structurally interesting genome

Trigger: RagTag on a non-model genome whose karyotype is studied. Mechanism: ordering to the reference imposes the reference's structure. Symptom: fusions/translocations/inversions silently erased; circular "discovery" of the reference karyotype. Fix: Hi-C + curation; use RagTag only as a hypothesis to test.

Scaffolding before breaking chimeras

Trigger: scaffolding contigs as given, contig-EC off. Mechanism: a chimeric contig drags region B into region A's chromosome; every join to it is wrong. Symptom: misjoins the contact map blames on the scaffolder. Fix: keep SALSA2 -m yes / YaHS contig-EC on; Tigmint for linked reads.

Treating Hi-C reads as a normal PE library

Trigger: bwa mem without -5SP, no dedup, or proper-pair logic. Mechanism: the two ends are different loci; mate rescue and insert-size filtering corrupt placement. Symptom: sparse/noisy contact map, weak joins. Fix: map each end separately (-5SP or per-end pipeline), dedup, MAPQ filter, enzyme-aware.

Gap-filling across a repeat

Trigger: running a gap-closer over the whole assembly to "finish" it. Mechanism: a long read bridges into a different copy of the repeat the gap sits in. Symptom: filled length far from the gap estimate; locally clean, globally mis-assembled. Fix: gap-fill after curation, sanity-check lengths, report closed-vs-left.

Quantitative Thresholds

ThresholdSourceRationale
~221 curation interventions/Gb (67 breaks, 105 joins, 49 false-dup)Howe 2021 GigaScienceexpect the automated draft to be substantially editable, not finished
Hi-C ~50-100M valid pairs (vertebrate); rule-of-thumb ~1x coverage per ~1 MbVGP practice (~approx)too little -> weak/missing joins, sparse noisy map; depends on genome size/repeats
Scaffold N50 / contig N50 > ~10xcurator reflexcontiguity is gap-borne; flag genome as contiguous-on-paper but unfinished
T2T chromosome = telomere at BOTH ends + zero internal gapsT2T conventiontidk both-end recovery with internal Ns = chromosome-scale but not finished
RagTag default gap = 100 bp NAlonge 2022placeholder length is arbitrary; never a real distance estimate
MAPQ filter on Hi-C reads (e.g. YaHS -q 10)mapping practicedrop repeat-ambiguous reads that create spurious joins

Common Errors

Error / symptomCauseSolution
_scaffolds_final looks chromosome-scale but synteny is wrongshipped uncurated draftcurate the contact map in PretextView/JBAT
Huge scaffold N50, small contig N50contiguity borne by N-gapsreport both; do not conflate
Reference-scaffolded genome "has" the reference's karyotypeRagTag imposed the structurere-scaffold with Hi-C + curation
Sparse/noisy contact map, few joinsHi-C mapped as normal PE / under-sequencedremap with -5SP+dedup; add Hi-C coverage
YaHS makes spurious joinsenzyme/MAPQ misconfigured; chimeric contigsset correct -e (omit for Omni-C), raise -q, keep contig-EC on
SALSA2 errors on input: in FASTA headers or BED not name-sortedrename headers; sort -k4 the BED
Filled gaps far from estimated lengthgap-filler bridged wrong repeat copygap-fill post-curation; sanity-check lengths; report closed-vs-left

References

  • Zhou C, McCarthy SA, Durbin R. 2023. YaHS: yet another Hi-C scaffolding tool. Bioinformatics 39:btac808.
  • Ghurye J, et al. 2019. Integrating Hi-C links with assembly graphs for chromosome-scale assembly (SALSA2). PLoS Comput Biol 15:e1007273.
  • Dudchenko O, et al. 2017. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds (3D-DNA). Science 356:92-95.
  • Durand NC, et al. 2016. Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments. Cell Syst 3:95-98.
  • Alonge M, et al. 2022. Automated assembly scaffolding using RagTag elevates a new tomato system for high-throughput genome editing. Genome Biol 23:258.
  • Tang H, et al. 2015. ALLMAPS: robust scaffold ordering based on multiple maps. Genome Biol 16:3.
  • Yeo S, et al. 2018. ARCS: scaffolding genome drafts with linked reads. Bioinformatics 34:725-731.
  • Jackman SD, et al. 2018. Tigmint: correcting assembly errors using linked reads from large molecules. BMC Bioinformatics 19:393.
  • Zhang H, et al. 2021. Fast alignment and preprocessing of chromatin profiles with chromap. Nat Commun 12:6566.
  • Rhie A, et al. 2021. Towards complete and error-free genome assemblies of all vertebrate species (VGP). Nature 592:737-746.
  • Howe K, et al. 2021. Significantly improving the quality of genome assemblies through curation. GigaScience 10:giaa153.
  • Brown MR, Gonzalez de la Rosa PM, Blaxter M. 2025. tidk: a toolkit to rapidly identify telomeric repeats from genomic datasets. Bioinformatics 41:btaf049.
  • Bionano Genomics. Bionano Solve Theory of Operation: Hybrid Scaffold (technical document) - hybrid scaffolding has no journal paper.
  • long-read-assembly - Produces the contigs this skill orders into chromosomes
  • hifi-assembly - hifiasm phases haplotypes (Hi-C --h1/--h2) before each is scaffolded
  • assembly-polishing - Polish contigs before scaffolding; gaps remain N until gap-filling
  • assembly-qc - Contig-vs-scaffold N50, BUSCO, and Merqury QV on the scaffolded result
  • hi-c-analysis/hic-data-io - Hi-C read/pairs handling feeding the scaffolder
  • comparative-genomics/synteny-analysis - Validate scaffold order/orientation against a related genome
  • workflows/genome-assembly-pipeline - End-to-end QC -> assemble -> polish -> scaffold -> curate -> QC

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in genome-assembly/scaffolding of GPTomics/bioSkills.

  • SKILL.md
  • examples/yahs_scaffolding.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Assembly Scaffolding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Assembly Scaffolding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Assembly Scaffolding this skillGPTomics/bioSkills1.2k1 repos~5.6kAutomated safety check: PassMIT
Repo Genomeruvnet/metaharness696—~772Automated safety check: PassMIT
Jj Flowseandavi/GEOquery118—~981Automated safety check: PassCustom licence
Light Project StructureLight0305/Light-skills640—~3kAutomated safety check: NotesMIT
Harness Similarityruvnet/ruflo74k—~866Automated safety check: NotesMIT
Nx Generatenomcopter/react-mosaic4.8k7 repos~1.9kAutomated safety check: PassCustom licence

Similar skills

  • Repo Genome

    ruvnet/metaharness

    7-section readiness scorecard for a LOCAL repo. An agent skill from ruvnet/metaharness.

    696 GitHub stars~772 tokensUpdated yesterday
    Product & Project ManagementAuto-check passed
  • Jj Flow

    seandavi/GEOquery

    jujutsu (jj) command cheatsheet for this colocated jj+git Bioconductor repo.

    118 GitHub stars~981 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Light Project Structure

    Light0305/Light-skills

    Audits, scaffolds and safely migrates research project folder structures, keeping existing repositories read-only until you approve exact moves from a plan.

    640 GitHub stars~3k tokensUpdated 3 mo ago
    DevelopmentAuto-check: notes
  • Harness Similarity

    ruvnet/ruflo

    ADR-152 — weighted similarity between two harness fingerprints (genome + score JSON).

    74k GitHub stars~866 tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Nx Generate

    nomcopter/react-mosaic

    Generate code using nx generators. An agent skill from nomcopter/react-mosaic.

    4.8k GitHub starsUsed in 7 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Ponytail

    DavidObando/gsharp

    Forces the laziest solution that actually works, simplest, shortest, most minimal.

    565 GitHub starsUsed in 7 repos~1.7k tokens
    DevelopmentAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Genome Assembly Scaffolding

What does Bio Genome Assembly Scaffolding do?

Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence). Bio Genome Assembly Scaffolding is an agent skill from GPTomics/bioSkills. Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence).

When should I use Bio Genome Assembly Scaffolding?

Bio Genome Assembly Scaffolding fits situations like: turning contigs into chromosomes with Hi-C; integrating a linkage map; choosing a scaffolder by available linking data; judging whether a chromosome-scale assembly is trustworthy.

How do I install Bio Genome Assembly Scaffolding in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a claude-code`. Or copy the skill folder (genome-assembly/scaffolding in GPTomics/bioSkills) into .claude/skills/bio-genome-assembly-scaffolding in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Assembly Scaffolding in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a codex`. Or copy the skill folder (genome-assembly/scaffolding in GPTomics/bioSkills) into .agents/skills/bio-genome-assembly-scaffolding in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Assembly Scaffolding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-assembly-scaffolding, .gemini/skills/bio-genome-assembly-scaffolding, .github/skills/bio-genome-assembly-scaffolding and .opencode/skills/bio-genome-assembly-scaffolding in your project.

What does Bio Genome Assembly Scaffolding need to run?

Going by SKILL.md and its folder, Bio Genome Assembly Scaffolding needs a shell for the scripts in its folder and the command-line tools its instructions call (java and python). Our summary lists: Python 3; A Bash shell.

Does Bio Genome Assembly Scaffolding access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Genome Assembly Scaffolding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Assembly Scaffolding use?

Bio Genome Assembly Scaffolding is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Assembly Scaffolding use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Assembly Scaffolding?

Skills that share tags, products or a category with Bio Genome Assembly Scaffolding: Repo Genome (ruvnet/metaharness, 696 stars), Jj Flow (seandavi/GEOquery, 118 stars), Light Project Structure (Light0305/Light-skills, 640 stars) and Harness Similarity (ruvnet/ruflo, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Assembly Scaffolding?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.