Repo Genome
ruvnet/metaharness
7-section readiness scorecard for a LOCAL repo. An agent skill from ruvnet/metaharness.
Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence).
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-scaffolding --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-assembly/scaffolding .claude/skills/bio-genome-assembly-scaffolding && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-genome-assembly-scaffolding" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/scaffolding into .claude/skills/bio-genome-assembly-scaffolding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-scaffolding", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/scaffoldingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-scaffolding --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/genome-assembly/scaffolding .agents/skills/bio-genome-assembly-scaffolding && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-genome-assembly-scaffolding" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/scaffolding into .agents/skills/bio-genome-assembly-scaffolding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-scaffolding", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-scaffolding --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/genome-assembly/scaffolding .cursor/skills/bio-genome-assembly-scaffolding && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-genome-assembly-scaffolding" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/scaffolding into .cursor/skills/bio-genome-assembly-scaffolding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-scaffolding", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path genome-assembly/scaffolding--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-scaffolding --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/genome-assembly/scaffolding .gemini/skills/bio-genome-assembly-scaffolding && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-genome-assembly-scaffolding" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/scaffolding into .gemini/skills/bio-genome-assembly-scaffolding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-scaffolding", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-scaffoldingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/genome-assembly/scaffolding .github/skills/bio-genome-assembly-scaffolding && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-genome-assembly-scaffolding" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/scaffolding into .github/skills/bio-genome-assembly-scaffolding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-scaffolding", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-assembly-scaffolding --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/genome-assembly/scaffolding .opencode/skills/bio-genome-assembly-scaffolding && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-genome-assembly-scaffolding" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-assembly/scaffolding into .opencode/skills/bio-genome-assembly-scaffolding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-assembly-scaffolding", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-genome-assembly-scaffoldingOrders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence).
Bio Genome Assembly Scaffolding is an agent skill from GPTomics/bioSkills. Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence). Covers Hi-C/Omni-C scaffolding (YaHS, SALSA2, 3D-DNA/Juicer), Hi-C read-mapping prerequisites (map each end separately, no mate rescue, dedup, enzyme-aware), reading the contact map for misjoins/inversions/false-duplications, manual curation in Juicebox/PretextView (the VGP/DToL standard), reference-guided scaffolding (RagTag) and its karyotype-erasure hazard, genetic-map…
Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/yahs_scaffolding.sh` and `usage-guide.md`).
It sits in Development, covering Project scaffolding and Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
javapythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Genome Assembly Scaffolding loads about 5.6k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 2,485 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,485 words, ~5,574 tokens.
.claude/skills/bio-genome-assembly-scaffolding/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: YaHS 1.2+, SALSA2 2.3+, juicer_tools 1.22/2.0+, bwa 0.7.17+, chromap 0.2+, samtools 1.19+, RagTag 2.1+, tidk 0.2.3+, seqkit 2.6+.
Before using code patterns, verify installed versions match. If versions differ:
<tool> --version then <tool> --help to confirm flagshifiasm output filenames and YaHS/SALSA2 enzyme and resolution defaults have changed across releases; confirm output names and -e/-r behaviour against the installed version. The Hi-C-mapping flag set (bwa mem -5SP vs the Arima/VGP per-end pipeline) varies by vendor (Arima, Dovetail/Omni-C, Phase Genomics) - check the current mapping guide before pasting flags. If a command errors, introspect the tool and adapt rather than retrying.
"Turn my contigs into chromosomes" -> Order and orient contigs by long-range linking signal, insert N-gap spacers, then curate the contact map - scaffolding adds no sequence, so the output is a draft chromosome structure, not finished sequence.
yahs contigs.fa hic_to_contigs.bam (Hi-C default), run_pipeline.py -a contigs.fa -b sorted.bed -m yes (SALSA2), ragtag.py scaffold ref.fa contigs.fa (reference-guided), allmaps path (linkage maps)YaHS finishes in minutes and emits a file literally named _scaffolds_final.agp. It is not final. Hi-C scaffolding is a contact-frequency inference - it orders/orients contigs by the polymer-physics signal that contact frequency decays with 1D distance - and it is error-prone exactly at the joins. Every reference-grade pipeline (VGP, Darwin Tree of Life, Earth BioGenome) treats the automated AGP as a first pass that a human curator then corrects in PretextView or Juicebox: breaking misjoins, flipping inversions, removing false duplications, assigning chromosomes by eye. Howe 2021 quantified the hidden labor across 111 VGP/DToL assemblies: on average 221 interventions per Gb (67 breaks, 105 joins, 49 false-duplication removals). Three load-bearing consequences:
The contact map is the QC, not a decoration. A correct chromosome shows one bright diagonal with smooth off-diagonal decay; off-diagonal blocks, anti-diagonal "bowties," and bleeding between chromosomes are errors to inspect and break (see Reading the Contact Map). An assembly paper whose methods say "scaffolded with [tool]" but never mention curation/Juicebox/PretextView is shipping a draft - downgrade any claim that depends on large-scale order (synteny, fusions, structural variants).
Scaffold N50 is not contig N50, and conflating them is the classic reviewer catch. Scaffolding inserts runs of N (estimated, often arbitrary length) and adds zero sequence. A genome can post "scaffold N50 = 60 Mb, chromosome-scale!" while its contig N50 is 200 kb - the contiguity is borne by gaps, not finished sequence. Always report contig N50, scaffold N50, gap count, and total N bases separately.
Reordering cannot rescue a bad contig. Scaffolders take contigs as given; a chimeric contig drags the wrong sequence into the wrong chromosome. Chimeras must be broken before scaffolding (see Chimera-Breaking).
| Modality | How it links | Scale | Status |
|---|---|---|---|
| Hi-C / Omni-C | proximity ligation; contact frequency ~ 1/(1D distance) | chromosome-scale | Dominant modern method. Omni-C is enzyme-free/sequence-agnostic; Arima uses fixed enzyme motifs |
| Bionano optical maps (DLS) | labeled motif patterns aligned to in-silico contig maps | megabase | Resolves large SVs and bridges complex repeats; separate instrument + library (cost decision) |
| 10x linked reads | barcoded short reads from one long molecule | ~10-100 kb | DISCONTINUED by 10x (2020). Legacy data only; Tigmint/ARCS still run, no new data |
| Genetic / linkage maps | marker recombination order = chromosome order | chromosome-scale, low resolution | ALLMAPS integrates multiple maps; orthogonal validation of Hi-C |
| Reference (homology) | align contigs to a related genome, copy its order | as good as synteny | RagTag. Fast, but imposes the reference's karyotype - see hazard below |
| Long reads as linkers | k-mer pairs spanning junctions | read length | LINKS; mostly superseded by assembling with the long reads directly |
| Mate-pair / jumping | large-insert paired short reads | 2-40 kb | Legacy (SSPACE); chimeric-insert artifacts; never recommend today |
| Available linking data | Use | Why |
|---|---|---|
| Hi-C/Omni-C, want fast contiguous default | YaHS | community default; fast, high N90, AGP + Juicebox outputs in one command |
| Hi-C, want graph-aware conservative joins + input-error correction | SALSA2 -m yes (-g graph.gfa) | iterative; breaks chimeric input; uses assembly graph to avoid orientation errors |
| Hi-C, interactive curation central to workflow | 3D-DNA + Juicer + Juicebox (JBAT) | the Aiden-lab .assembly <-> Juicebox round-trip is built around hand-editing |
| Hi-C, vertebrate/reference-grade | YaHS or SALSA2 -> PretextView/JBAT curation | VGP/DToL standard: automate then curate |
| Diploid, Hi-C, want both haplotypes as chromosomes | hifiasm --h1/--h2 to PHASE first, then YaHS to SCAFFOLD each haplotype | same Hi-C, two jobs - phasing picks the homolog, scaffolding orders along it (see Phasing vs Scaffolding) |
| Closely related reference, karyotype NOT a question | RagTag correct then scaffold | homology ordering in minutes - but never if karyotype is the biology |
| One or more genetic/linkage maps | ALLMAPS | integrates multiple maps; robust to marker errors; validates/anchors Hi-C |
| Bionano optical maps + sequence assembly | Bionano Solve hybrid scaffold (before Hi-C) | megabase maps bridge complex repeats and large SVs |
| 10x linked reads (legacy data) | Tigmint (break) -> ARCS/ARKS (link) | break misassemblies first; no new 10x data exists |
| Contigs not yet QC'd / haplotigs present | -> assembly-qc, purge_dups | scaffolding haplotigs strings them up as fake chromosomes |
| Hi-C contact map for TADs/loops, not chromosomes | -> hi-c-analysis/matrix-operations | different use of the same assay |
Modern reference-grade recipe (vertebrate/eukaryote): HiFi -> hifiasm contigs -> (optional Bionano hybrid scaffold) -> Hi-C scaffold (YaHS/SALSA2) -> manual curation in PretextView/JBAT -> gap-fill (TGS-GapCloser) -> QC (tidk telomeres, contact-map diagonal, contig-vs-scaffold N50).
Hi-C reads are not a normal paired-end library: the two ends come from different genomic loci ligated together, so standard PE proper-pair/insert-size logic mis-handles them.
# Map each end SEPARATELY with no mate rescue/pairing; -5 reports the 5' (junction) portion as primary.
bwa index contigs.fa
bwa mem -5SP -T0 -t 16 contigs.fa hic_R1.fq.gz hic_R2.fq.gz | \
samtools view -@ 8 -b - > aligned.bam
# Mandatory: mark/remove PCR + optical duplicates (Hi-C is duplicate-rich) before scaffolding.
samtools sort -@ 8 -n aligned.bam | samtools fixmate -m - - | \
samtools sort -@ 8 - | samtools markdup -@ 8 - hic_to_contigs.bam
samtools index hic_to_contigs.bam-5SP): -5 = 5' portion of a chimeric junction read as primary; -S/-P skip mate rescue and pairing. The Arima/VGP per-end pipeline (bwa mem on R1 and R2 independently, filter_five_end.pl, two_read_bam_combiner.pl) is the higher-fidelity equivalent.-q, SALSA2 reads MAPQ) to drop repeat-ambiguous reads.GATC); the scaffolder uses -e to model legitimate read starts. Omni-C / DNase Hi-C is enzyme-free -> omit -e in YaHS, use -e DNASE in SALSA2.chromap --preset hic -r contigs.fa -x index -1 R1 -2 R2 ... aligns + dedups Hi-C ~10x faster; YaHS accepts its output.samtools faidx contigs.fa
# Enzyme: -e GATC (DpnII/Dovetail), -e GATC,GANTC (Arima 2-enzyme); OMIT -e for Omni-C/enzyme-free.
yahs -e GATC -o out contigs.fa hic_to_contigs.bam
# Outputs: out_scaffolds_final.agp + out_scaffolds_final.fa, intermediate out_rNN.agp, out.bin
# Contig error-correction is ON by default (--no-contig-ec to disable; usually a mistake to disable).
# Juicebox prep for curation (.hic + .assembly to load in JBAT):
# NOTE: `juicer` here is the small utility BUNDLED with YaHS (operates on the .bin), NOT Aiden-lab Juicer;
# `juicer_tools.jar` below IS the separate Aiden-lab jar. Do not confuse the two.
juicer pre -a -o out_JBAT out.bin out_scaffolds_final.agp contigs.fa.fai 2>tmp_assembly.log
java -Xmx48G -jar juicer_tools.jar pre out_JBAT.txt out_JBAT.hic <(cat tmp_assembly.log | grep PRE_C_SIZE | awk '{print $2" "$3}')
# After hand-editing in Juicebox, export out_JBAT.review.assembly, then:
juicer post -o out_curated out_JBAT.review.assembly out_JBAT.liftover.agp contigs.faYaHS also accepts a BED (bamToBed) or PA5/pairs input. The _scaffolds_final.agp is the editable, reviewable object - curation edits the AGP, then regenerates FASTA. --telo-motif lets YaHS use telomere signal during scaffolding.
bamToBed -i aligned.bam > alignment.bed
sort -k4 alignment.bed > sorted.bed # MUST sort by read name (column 4)
python run_pipeline.py -a contigs.fa -l contigs.fa.fai -b sorted.bed \
-e GATC -o salsa_out -m yes -g contigs_graph.gfa -p yes
# -e DNASE for Omni-C/enzyme-free; -m yes enables input-error (chimera) correction; -p yes writes AGP+FASTA per iterationThe signature -m yes step breaks input contigs where the Hi-C signal contradicts the assembler's join before scaffolding. Assembly headers must not contain :. -i sets iterations (default 3), -c min contig length (default 1000).
ragtag.py correct ref.fa contigs.fa # break query at putative misassemblies vs reference FIRST
ragtag.py scaffold ref.fa ragtag_output/ragtag.correct.fasta -t 16 # order/orient by homology (minimap2 default)
# Outputs ragtag.scaffold.agp + .fasta; inserts 100 bp N gaps by default; -C collapses unplaced into chr0.The peril (load-bearing): RagTag orders contigs to match the reference, so by construction the output looks like the reference. Every real biological rearrangement that distinguishes the target organism - a chromosome fusion, fission, translocation, or inversion polymorphism - is silently erased and replaced by the reference's structure. Never reference-scaffold a genome whose karyotype or large-scale structure is a biological question - the pipeline "discovers" the reference's karyotype because it was imposed (this has happened in the literature). Acceptable uses: a structurally conserved close relative needing only a coordinate system for gene content; a scaffold hypothesis to then test/correct against Hi-C (ragtag.py merge -b hic.bam lets Hi-C arbitrate). On a genome with interesting structure the honest move is Hi-C + curation, full stop.
ALLMAPS (Tang 2015 Genome Biol 16:3) integrates one or more genetic maps to order/orient scaffolds: allmaps merge map1.csv map2.csv -o maps.bed then allmaps path maps.bed contigs.fa. It is the gold standard for validating chromosome assignment and for species where Hi-C is hard; weight multiple maps by quality. Bionano Solve hybrid scaffolding (no journal paper; cite the Bionano Solve Theory of Operation: Hybrid Scaffold technical document) aligns DLS optical maps to in-silico contig maps, resolves conflicts, and merges into hybrid scaffolds - run it before Hi-C on large/repetitive genomes where megabase maps bridge repeat arrays Hi-C fumbles. The decision is economic (separate instrument + library), not purely technical.
The map displays its own errors to a trained eye - render with PretextMap -> PretextView (interactive, emits a curated AGP), PretextSnapshot (static PNG), or Juicebox/JBAT.
-m yes and YaHS default contig-EC break chimeric input where Hi-C contradicts the join; Tigmint (Jackman 2018 BMC Bioinformatics 19:393) breaks before ARCS for linked reads; RagTag correct for the reference-based version. Disabling these to "preserve the assembly" is a common self-inflicted wound.--h1/--h2 during assembly). When handed "Hi-C, make chromosomes," the first question is haploid scaffolding or diploid phasing+scaffolding? - the tools and failure modes differ entirely.Trigger: publishing _scaffolds_final without curation. Mechanism: automated joins are inferences error-prone at junctions (~221 edits/Gb in VGP/DToL). Symptom: misjoins/inversions in the contact map; later-discovered wrong synteny/fusions. Fix: curate in PretextView/JBAT before any large-scale-order claim.
Trigger: reporting only scaffold N50. Mechanism: scaffolding adds N-gaps, not sequence. Symptom: spectacular scaffold N50 over a small contig N50 - contiguity is nitrogen. Fix: report both N50s, gap count, total N bases separately.
Trigger: RagTag on a non-model genome whose karyotype is studied. Mechanism: ordering to the reference imposes the reference's structure. Symptom: fusions/translocations/inversions silently erased; circular "discovery" of the reference karyotype. Fix: Hi-C + curation; use RagTag only as a hypothesis to test.
Trigger: scaffolding contigs as given, contig-EC off. Mechanism: a chimeric contig drags region B into region A's chromosome; every join to it is wrong. Symptom: misjoins the contact map blames on the scaffolder. Fix: keep SALSA2 -m yes / YaHS contig-EC on; Tigmint for linked reads.
Trigger: bwa mem without -5SP, no dedup, or proper-pair logic. Mechanism: the two ends are different loci; mate rescue and insert-size filtering corrupt placement. Symptom: sparse/noisy contact map, weak joins. Fix: map each end separately (-5SP or per-end pipeline), dedup, MAPQ filter, enzyme-aware.
Trigger: running a gap-closer over the whole assembly to "finish" it. Mechanism: a long read bridges into a different copy of the repeat the gap sits in. Symptom: filled length far from the gap estimate; locally clean, globally mis-assembled. Fix: gap-fill after curation, sanity-check lengths, report closed-vs-left.
| Threshold | Source | Rationale |
|---|---|---|
| ~221 curation interventions/Gb (67 breaks, 105 joins, 49 false-dup) | Howe 2021 GigaScience | expect the automated draft to be substantially editable, not finished |
| Hi-C ~50-100M valid pairs (vertebrate); rule-of-thumb ~1x coverage per ~1 Mb | VGP practice (~approx) | too little -> weak/missing joins, sparse noisy map; depends on genome size/repeats |
| Scaffold N50 / contig N50 > ~10x | curator reflex | contiguity is gap-borne; flag genome as contiguous-on-paper but unfinished |
| T2T chromosome = telomere at BOTH ends + zero internal gaps | T2T convention | tidk both-end recovery with internal Ns = chromosome-scale but not finished |
| RagTag default gap = 100 bp N | Alonge 2022 | placeholder length is arbitrary; never a real distance estimate |
MAPQ filter on Hi-C reads (e.g. YaHS -q 10) | mapping practice | drop repeat-ambiguous reads that create spurious joins |
| Error / symptom | Cause | Solution |
|---|---|---|
_scaffolds_final looks chromosome-scale but synteny is wrong | shipped uncurated draft | curate the contact map in PretextView/JBAT |
| Huge scaffold N50, small contig N50 | contiguity borne by N-gaps | report both; do not conflate |
| Reference-scaffolded genome "has" the reference's karyotype | RagTag imposed the structure | re-scaffold with Hi-C + curation |
| Sparse/noisy contact map, few joins | Hi-C mapped as normal PE / under-sequenced | remap with -5SP+dedup; add Hi-C coverage |
| YaHS makes spurious joins | enzyme/MAPQ misconfigured; chimeric contigs | set correct -e (omit for Omni-C), raise -q, keep contig-EC on |
| SALSA2 errors on input | : in FASTA headers or BED not name-sorted | rename headers; sort -k4 the BED |
| Filled gaps far from estimated length | gap-filler bridged wrong repeat copy | gap-fill post-curation; sanity-check lengths; report closed-vs-left |
--h1/--h2) before each is scaffolded© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in genome-assembly/scaffolding of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Genome Assembly Scaffolding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Genome Assembly Scaffolding this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.6k | Automated safety check: Pass | MIT | |
| Repo Genomeruvnet/metaharness | 696 | — | ~772 | Automated safety check: Pass | MIT | |
| Jj Flowseandavi/GEOquery | 118 | — | ~981 | Automated safety check: Pass | Custom licence | |
| Light Project StructureLight0305/Light-skills | 640 | — | ~3k | Automated safety check: Notes | MIT | |
| Harness Similarityruvnet/ruflo | 74k | — | ~866 | Automated safety check: Notes | MIT | |
| Nx Generatenomcopter/react-mosaic | 4.8k | 7 repos | ~1.9k | Automated safety check: Pass | Custom licence |
ruvnet/metaharness
7-section readiness scorecard for a LOCAL repo. An agent skill from ruvnet/metaharness.
seandavi/GEOquery
jujutsu (jj) command cheatsheet for this colocated jj+git Bioconductor repo.
Light0305/Light-skills
Audits, scaffolds and safely migrates research project folder structures, keeping existing repositories read-only until you approve exact moves from a plan.
ruvnet/ruflo
ADR-152 — weighted similarity between two harness fingerprints (genome + score JSON).
nomcopter/react-mosaic
Generate code using nx generators. An agent skill from nomcopter/react-mosaic.
DavidObando/gsharp
Forces the laziest solution that actually works, simplest, shortest, most minimal.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence). Bio Genome Assembly Scaffolding is an agent skill from GPTomics/bioSkills. Orders and orients assembled contigs into chromosome-scale scaffolds from long-range linking data, inserting N-gap spacers (adds no sequence).
Bio Genome Assembly Scaffolding fits situations like: turning contigs into chromosomes with Hi-C; integrating a linkage map; choosing a scaffolder by available linking data; judging whether a chromosome-scale assembly is trustworthy.
Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a claude-code`. Or copy the skill folder (genome-assembly/scaffolding in GPTomics/bioSkills) into .claude/skills/bio-genome-assembly-scaffolding in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a codex`. Or copy the skill folder (genome-assembly/scaffolding in GPTomics/bioSkills) into .agents/skills/bio-genome-assembly-scaffolding in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-scaffolding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-assembly-scaffolding, .gemini/skills/bio-genome-assembly-scaffolding, .github/skills/bio-genome-assembly-scaffolding and .opencode/skills/bio-genome-assembly-scaffolding in your project.
Going by SKILL.md and its folder, Bio Genome Assembly Scaffolding needs a shell for the scripts in its folder and the command-line tools its instructions call (java and python). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Genome Assembly Scaffolding is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Genome Assembly Scaffolding: Repo Genome (ruvnet/metaharness, 696 stars), Jj Flow (seandavi/GEOquery, 118 stars), Light Project Structure (Light0305/Light-skills, 640 stars) and Harness Similarity (ruvnet/ruflo, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.