Agent skill

Bio Genome Assembly Genome Profiling

by GPTomics in GPTomics/bioSkills

Profiles a genome from raw reads BEFORE assembly with a k-mer spectrum (KMC or Jellyfish histogram), then models it with GenomeScope2 to estimate genome size, heterozygosity, repeat content, and…

MITAuto-check passedResearch & Science

Install Bio Genome Assembly Genome Profiling

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-assembly-genome-profiling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-assembly-genome-profiling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-assembly/genome-profiling .claude/skills/bio-genome-assembly-genome-profiling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-assembly-genome-profiling
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4k tokens
SKILL.md length
1,823 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Profiles a genome from raw reads BEFORE assembly with a k-mer spectrum (KMC or Jellyfish histogram), then models it with GenomeScope2 to estimate genome size, heterozygosity, repeat content, and…

  • Works in 3 steps: The estimate is the denominator AND the… → The spectrum reads as a diagnostic, not… → Count with accurate reads only.…
  • Starting any de novo assembly
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Shell scripts from its folder; calls sh

What it does

Bio Genome Assembly Genome Profiling is an agent skill from GPTomics/bioSkills. Profiles a genome from raw reads BEFORE assembly with a k-mer spectrum (KMC or Jellyfish histogram), then models it with GenomeScope2 to estimate genome size, heterozygosity, repeat content, and ploidy, and Smudgeplot to infer ploidy from heterozygous k-mer pairs (diploid AB vs triploid AAB vs tetraploid AABB). Covers choosing k via Merqury bestk.sh, the k-mer-coverage vs sequencing-coverage confusion, reading het/repeat/contamination/organelle peaks, why noisy ONT must not be used for counting, and how the…

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/profile_genome.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Starting any de novo assembly
  • Deciding whether short reads can work
  • Estimating genome size for an unknown organism
  • Diagnosing ploidy

Example prompts

  • “Use the bio-genome-assembly-genome-profiling skill to profile a genome from raw reads BEFORE assembly with a k-mer spectrum (KMC or Jellyfish…”
  • “/bio-genome-assembly-genome-profiling”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The estimate is the denominator AND the sanity check. NG50 is N50 against the expected genome size, not the assembly size, so without the…
  2. The spectrum reads as a diagnostic, not just a size estimate. A single homozygous peak ~= haploid/inbred; a distinct half-coverage (AB)…
  3. Count with accurate reads only. GenomeScope's negative-binomial mixture model assumes errors are rare and Poisson-like. Noisy ONT…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • sh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Assembly Genome Profiling loads about 4k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 1,823 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~224
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,823 words, ~3,955 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-assembly-genome-profiling/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-genome-assembly-genome-profiling
description
Profiles a genome from raw reads BEFORE assembly with a k-mer spectrum (KMC or Jellyfish histogram), then models it with GenomeScope2 to estimate genome size, heterozygosity, repeat content, and ploidy, and Smudgeplot to infer ploidy from heterozygous k-mer pairs (diploid AB vs triploid AAB vs tetraploid AABB). Covers choosing k via Merqury best_k.sh, the k-mer-coverage vs sequencing-coverage confusion, reading het/repeat/contamination/organelle peaks, why noisy ONT must not be used for counting, and how the estimate becomes the NG50 denominator, the Flye -g value, the hifiasm --hom-cov/purge setting, and the 1.5-2x-too-big haplotig sanity check. Use when starting any de novo assembly, deciding whether short reads can work, estimating genome size for an unknown organism, diagnosing ploidy, or sanity-checking an assembly's size against expectation.
tool_type
cli
primary_tool
GenomeScope2

Version Compatibility

Reference examples tested with: GenomeScope2 2.0+, KMC 3.2+, Jellyfish 2.3+, meryl 1.4+ (Merqury 1.3+), Smudgeplot 0.2.5+, KAT 2.4+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

Smudgeplot changed its backend: classic releases (0.2.x) run smudgeplot.py hetkmers/smudgeplot.py plot on a KMC dump; newer releases (0.3+) run smudgeplot hetmers/smudgeplot all on a FastK database. Confirm which interface is installed (smudgeplot.py --version or smudgeplot --version) before scripting. GenomeScope2 ships as genomescope2 and as genomescope.R; both take the same flags. meryl best_k.sh lives in the Merqury install. If code throws an error, introspect the installed tool and adapt rather than retrying.

Genome Profiling

"What am I about to assemble, and what should I expect?" -> Build a k-mer spectrum from raw accurate reads and model it to estimate genome size, heterozygosity, repeat content, and ploidy, which set every downstream assembly expectation and parameter.

  • CLI: kmc -k21 ... reads kmc_db tmp/ && kmc_tools transform kmc_db histogram reads.histo (count) then genomescope2 -i reads.histo -o gs_out -k 21 -p 2 (model); smudgeplot.py hetkmers / smudgeplot.py plot (ploidy)

The Single Most Important Modern Insight -- Profile the Genome Before Assembling, or the Assembler Guesses For Itself

A k-mer spectrum built from the raw reads -- reference-free, before a single contig exists -- estimates genome size, heterozygosity, repeat content, and ploidy, and those four numbers set the rest of the project: the NG50 denominator (assembly-qc), Flye's -g/--genome-size, hifiasm's --hom-cov/purge level, and whether short reads can produce the assembly being asked for at all. Skipping it leaves the assembler to infer the homozygous-coverage peak itself; when it mis-estimates (heterozygosity, odd ploidy, contamination, a bimodal coverage spectrum) it over- or under-purges, and that is exactly why people publish genomes inflated 1.5-2x by uncollapsed haplotigs. Three load-bearing moves:

  1. The estimate is the denominator AND the sanity check. NG50 is N50 against the expected genome size, not the assembly size, so without the estimate NG50 cannot be reported honestly. After assembly, the same number is the haplotig test: an assembly 1.5-2x the GenomeScope size with high BUSCO-Duplicated is uncollapsed haplotypes, not a big genome -- purge before believing the size (see hifi-assembly, assembly-qc).
  2. The spectrum reads as a diagnostic, not just a size estimate. A single homozygous peak ~= haploid/inbred; a distinct half-coverage (AB) peak left of the homozygous (AA) peak is heterozygous diploid, and the het peak's area gives the heterozygosity rate. A heavy high-multiplicity tail is repeat content; a spike at very high multiplicity is organelle or high-copy repeat; a left-shoulder near multiplicity 1 is sequencing error or a low-coverage contaminant. Read it before trusting any number from it.
  3. Count with accurate reads only. GenomeScope's negative-binomial mixture model assumes errors are rare and Poisson-like. Noisy ONT (raw/HAC) injects so many unique error k-mers that the error shoulder swamps the real peaks and the fit fails. Count from Illumina or PacBio HiFi; use the ONT reads to assemble, never to profile.

Tool Taxonomy

ToolCitationRoleWhen
KMCKokot 2017 Bioinformaticsdisk-based k-mer counter -> histogramdefault counter; frugal on RAM, fast on large genomes
JellyfishMarcais & Kingsford 2011 Bioinformaticsin-memory k-mer counter -> histogramalternative counter; classic GenomeScope input
merylRhie 2020 Genome Biolk-mer counter + best_k.shderives k from genome size; feeds Merqury QV downstream
GenomeScope2Ranallo-Benavidez 2020 Nat Communmodel: size, het, repeat, ploidy from the histogramthe profiling model for diploids and polyploids
SmudgeplotRanallo-Benavidez 2020 Nat Communploidy from het k-mer-pair coverage ratiosunknown ploidy; cross-check GenomeScope's -p
KATMapleson 2017 Bioinformaticsspectra plots, reads-vs-assembly k-mer comparisoncontamination triage; post-assembly completeness/spectra-cn

Decision Tree by Scenario

ScenarioWhat the profile indicatesPath
Single sharp peak, size ~ expected, low hethaploid/inbred or clonal; cleanproceed -> short-read-assembly or hifi-assembly with default purge
Two peaks (AB at ~half AA)heterozygous diploid; het rate from AB areaHiFi+phasing best; if short reads only, expect haplotigs -> short-read-assembly (Platanus)
High het + Illumina onlyshort-read DBG will fragment and inflate 1.5-2xget long reads, or plan purge_dups; do not report inflated size
GenomeScope -p 2 fits poorly; Smudgeplot shows AAB/AABBtriploid/tetraploid; ploidy not 2re-run GenomeScope with correct -p; -> hifi-assembly haplotype expectations
Coverage peak < ~15-20xtoo shallow for a stable model fitsequence more, or treat size/het as lower-confidence; -> read-qc/quality-reports
Extra peak at odd multiplicity, or bimodal spectrumcontamination / organelle / mixed sampleKAT spectra triage; screen reads -> read-qc/quality-reports before assembling
Only noisy ONT availablecannot profile reliably from error-dominated spectrumassemble first, then estimate size from the assembly + Merqury -> assembly-qc
Need the genome-size denominator for QCGenomeScope haploid lengthfeed as NG50 expected size and Flye -g -> assembly-qc, long-read-assembly

Choosing k

Goal: Pick a k that is large enough that most k-mers are genomically unique but small enough to keep per-k-mer depth high.

Approach: Derive k from the expected genome size with Merqury's best_k.sh (formula k = log4(G(1-p)/p), default tolerable collision rate p=0.001); a vertebrate-scale ~3 Gb genome returns k=21 (the de facto GenomeScope default), ~1 Gb returns k=20, and a ~12 Mb yeast returns k=17.

bash
sh $MERQURY/best_k.sh 3100000000          # ~3.1 Gb vertebrate -> k=21 (the common GenomeScope default)
sh $MERQURY/best_k.sh 1000000000          # ~1 Gb -> k=20
sh $MERQURY/best_k.sh 12000000            # ~12 Mb yeast -> k=17

Too small a k saturates: nearly every k-mer recurs across the genome by chance, the unique/repeat peaks merge, and size is overestimated. Too large a k loses depth (per-k-mer coverage is c*(L-k+1)/L, so it drops as k rises) and the peaks blur into the error shoulder. k=21 is the long-standing default at vertebrate (~3 Gb) scale because it sits in this window; smaller genomes want smaller k (best_k.sh returns ~20 at 1 Gb, ~17 at 12 Mb). Use the SAME k for counting and for the -k passed to GenomeScope2.

Counting K-mers and Running GenomeScope2

bash
# KMC: -ci1 keeps singletons (the error shoulder GenomeScope models), -cs10000 caps the histogram tail
kmc -k21 -t16 -m64 -ci1 -cs10000 @fastq_list.txt kmc_db tmp/
kmc_tools transform kmc_db histogram reads.histo -cx10000

# GenomeScope2: -p 2 = diploid; raise for known/Smudgeplot-suggested polyploidy
genomescope2 -i reads.histo -o gs_out -k 21 -p 2

# Jellyfish alternative (-C canonical k-mers, mandatory for unstranded WGS)
jellyfish count -C -m 21 -s 4G -t 16 reads_*.fastq -o reads.jf
jellyfish histo -t 16 reads.jf > reads_jf.histo

GenomeScope2 writes a model fit (model.txt), the linear/log spectrum plots, and a summary with haploid genome length, heterozygosity, and a "% unique" (inverse of repeat content). -cs10000/-cx10000 cap the histogram so a single organelle/repeat spike at multiplicity 100000+ does not dominate the file; raise the cap only for very high-coverage data.

K-mer Coverage vs Sequencing Coverage (the confusion that breaks the fit)

GenomeScope reports kmercov (lambda) -- the mean coverage per k-mer, the x-position of the homozygous peak. This is NOT per-base sequencing coverage. They differ by the (L-k+1)/L factor: at read length L=150 and k=21, a k-mer is covered 130/150 ~= 0.87x as often as a base, so a 50x-sequenced genome shows a homozygous peak near multiplicity 43, not 50. Passing the sequencing coverage where GenomeScope expects the k-mer-coverage peak (e.g. as an -l initial guess), or reading lambda back as sequencing depth, mis-scales the model and the size estimate. Read the peak off the plot; let GenomeScope estimate lambda unless the fit fails, then seed -l with the observed peak position.

Show full SKILL.md (732 more words)Show less

Ploidy with Smudgeplot

bash
# Classic (0.2.x, KMC backend): pick L/U coverage cutoffs, dump het k-mer pairs, plot
L=$(smudgeplot.py cutoff kmc_db.histo L)     # ~0.5x the haploid peak (errors below)
U=$(smudgeplot.py cutoff kmc_db.histo U)     # ~8.5x the haploid peak (repeats above)
kmc_tools transform kmc_db -ci"$L" -cx"$U" dump -s kmc_L"$L"_U"$U".dump
smudgeplot.py hetkmers -o kmer_pairs < kmc_L"$L"_U"$U".dump
smudgeplot.py plot kmer_pairs_coverages.tsv -o sample

Smudgeplot infers ploidy from the coverage RATIO of heterozygous k-mer pairs, independent of GenomeScope's model: a diploid AB pair sits at CovB/(CovA+CovB) ~= 0.5, a triploid AAB near 0.33, a tetraploid AABB shows smudges at 0.25 and 0.5. When Smudgeplot's inferred ploidy disagrees with the -p that fit GenomeScope best, that disagreement is itself the finding -- re-examine for polyploidy, aneuploidy, or contamination rather than forcing one answer.

Per-Method Failure Modes

Profiling from noisy ONT

Trigger: counting k-mers from raw/HAC Nanopore reads. Mechanism: ~5-10% per-base error makes nearly every error k-mer unique, burying the real peaks under the multiplicity-1 shoulder. Symptom: GenomeScope fit fails or returns absurd size/het. Fix: count from Illumina or HiFi; profile from the assembly + Merqury if only ONT exists.

Too-low coverage

Trigger: homozygous peak below ~15-20x. Mechanism: the error shoulder and the real peak overlap; the negative-binomial mixture cannot separate them. Symptom: wide confidence intervals, unstable size, no clean peaks. Fix: sequence more depth, or report size/het as low-confidence.

Reading lambda as sequencing depth

Trigger: treating GenomeScope kmercov as per-base coverage, or seeding -l with sequencing depth. Mechanism: k-mer coverage = depth * (L-k+1)/L < depth. Symptom: mis-scaled model, size off by the (L-k+1)/L factor. Fix: read lambda off the peak; do not substitute sequencing coverage.

k too small (saturation)

Trigger: k=15-17 on a large/repeat-rich genome. Mechanism: most k-mers recur by chance; unique and repeat components merge. Symptom: inflated size, blurred peaks. Fix: use best_k.sh for the expected size (k=21 at vertebrate ~3 Gb scale; smaller for smaller genomes).

Contamination/organelle distorting the spectrum

Trigger: an unexplained extra peak or a spike at very high multiplicity. Mechanism: a second organism's k-mers add their own peak; organelle/high-copy DNA spikes the tail. Symptom: GenomeScope size too large or a multi-modal spectrum. Fix: KAT spectra triage; screen reads (-> read-qc/quality-reports) before assembling.

Quantitative Thresholds

ThresholdSourceRationale
k from best_k.sh (21 at ~3 Gb, 20 at ~1 Gb, 17 at ~12 Mb)best_k.sh k=log4(G(1-p)/p), p=0.001balances uniqueness vs per-k-mer depth; k=21 is the common GenomeScope default at vertebrate scale
Homozygous peak >= ~15-20xGenomeScope model stabilitybelow this the error shoulder and real peak cannot be separated
AB peak at ~0.5x the AA peakdiploid k-mer theoryheterozygous k-mers are half-covered (one haplotype); het rate from AB area
Smudgeplot CovB/(CovA+CovB): 0.5 / 0.33 / 0.25k-mer-pair coverage ratioAB diploid / AAB triploid / AABB tetraploid
Assembly size 1.5-2x GenomeScope estimatehaplotig normuncollapsed heterozygous haplotypes; purge before reporting size
KMC -cs/GenomeScope -cx cap ~10000histogram tail controlprevents an organelle/repeat spike from dominating the file
Smudgeplot L ~0.5x, U ~8.5x haploid peakSmudgeplot guidanceexcludes errors (below) and high-copy repeats (above) from pairing

Common Errors

Error / symptomCauseSolution
GenomeScope size far larger than expectedk too small, contamination, or organelle spikeraise k via best_k.sh; cap the tail; screen reads
Model fit fails / "cannot fit"too-low coverage or ONT error k-mersmore depth; count from Illumina/HiFi, not noisy ONT
Histogram peak at multiplicity ~43 for "50x" datak-mer coverage = depth * (L-k+1)/L, not depthexpected; do not pass sequencing depth as -l
GenomeScope -p 2 fits poorlysample is not diploidrun Smudgeplot; re-run with the correct -p
Jellyfish histogram looks halvedcounted without -C (canonical)recount with -C; WGS is unstranded
Assembly 1.8x the profiled sizeuncollapsed haplotigspurge_dups / hifiasm purge; report the haploid size

References

  • Ranallo-Benavidez TR, Jaron KS, Schatz MC. 2020. GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes. Nat Commun 11:1432.
  • Vurture GW, et al. 2017. GenomeScope: fast reference-free genome profiling from short reads. Bioinformatics 33:2202-2204.
  • Kokot M, Dlugosz M, Deorowicz S. 2017. KMC 3: counting and manipulating k-mer statistics. Bioinformatics 33:2759-2761.
  • Marcais G, Kingsford C. 2011. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers (Jellyfish). Bioinformatics 27:764-770.
  • Rhie A, Walenz BP, Koren S, Phillippy AM. 2020. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol 21:245.
  • Mapleson D, et al. 2017. KAT: a K-mer analysis toolkit to quality control NGS datasets and genome assemblies. Bioinformatics 33:574-576.
  • short-read-assembly - High het from the profile predicts short-read fragmentation and haplotig inflation
  • hifi-assembly - The profiled size/het sets hifiasm --hom-cov and the purge level
  • long-read-assembly - The profiled genome size feeds Flye -g and ONT/PacBio expectations
  • assembly-qc - The estimate is the NG50 denominator and the 1.5-2x haplotig sanity check
  • read-qc/quality-reports - QC and contamination-screen reads before profiling and assembling
  • workflows/genome-assembly-pipeline - Profiling is the first step of the end-to-end assembly workflow

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in genome-assembly/genome-profiling of GPTomics/bioSkills.

  • SKILL.md
  • examples/profile_genome.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Assembly Genome Profiling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Assembly Genome Profiling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Assembly Genome Profiling this skillGPTomics/bioSkills1.2k1 repos~4kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Genome Assembly Genome Profiling

What does Bio Genome Assembly Genome Profiling do?

Profiles a genome from raw reads BEFORE assembly with a k-mer spectrum (KMC or Jellyfish histogram), then models it with GenomeScope2 to estimate genome size, heterozygosity, repeat content, and…. Bio Genome Assembly Genome Profiling is an agent skill from GPTomics/bioSkills. Profiles a genome from raw reads BEFORE assembly with a k-mer spectrum (KMC or Jellyfish histogram), then models it with GenomeScope2 to estimate genome size, heterozygosity, repeat content, and ploidy, and Smudgeplot to infer ploidy from heterozygous k-mer pairs (diploid AB vs triploid AAB vs tetraploid AABB).

When should I use Bio Genome Assembly Genome Profiling?

Bio Genome Assembly Genome Profiling fits situations like: starting any de novo assembly; deciding whether short reads can work; estimating genome size for an unknown organism; diagnosing ploidy.

How do I install Bio Genome Assembly Genome Profiling in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-genome-profiling -a claude-code`. Or copy the skill folder (genome-assembly/genome-profiling in GPTomics/bioSkills) into .claude/skills/bio-genome-assembly-genome-profiling in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Assembly Genome Profiling in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-genome-profiling -a codex`. Or copy the skill folder (genome-assembly/genome-profiling in GPTomics/bioSkills) into .agents/skills/bio-genome-assembly-genome-profiling in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Assembly Genome Profiling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-assembly-genome-profiling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-assembly-genome-profiling, .gemini/skills/bio-genome-assembly-genome-profiling, .github/skills/bio-genome-assembly-genome-profiling and .opencode/skills/bio-genome-assembly-genome-profiling in your project.

What does Bio Genome Assembly Genome Profiling need to run?

Going by SKILL.md and its folder, Bio Genome Assembly Genome Profiling needs a shell for the scripts in its folder and the command-line tools its instructions call (sh). Our summary lists: A Bash shell.

Does Bio Genome Assembly Genome Profiling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Genome Assembly Genome Profiling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Assembly Genome Profiling use?

Bio Genome Assembly Genome Profiling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Assembly Genome Profiling use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Assembly Genome Profiling?

Skills that share tags, products or a category with Bio Genome Assembly Genome Profiling: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Assembly Genome Profiling?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.