Agent skill

Bio Genome Annotation Annotation Qc

by GPTomics in GPTomics/bioSkills

Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic…

MITAuto-check passedResearch & Science

Install Bio Genome Annotation Annotation Qc

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-annotation-annotation-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-annotation/annotation-qc .claude/skills/bio-genome-annotation-annotation-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-annotation-annotation-qc
GitHub stars
1.2k
Used in
1 other repo
Token cost
~3.6k tokens
SKILL.md length
1,519 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic…

  • Works in 3 steps: Run assembly-BUSCO vs proteome-BUSCO on… → Gene count is a vanity metric. The same… → Complement BUSCO with OMArk. BUSCO's…
  • Judging whether an annotation is good enough to publish
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Shell and Python scripts from its folder; calls pip

What it does

Bio Genome Annotation Annotation Qc is an agent skill from GPTomics/bioSkills. Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic completeness/contamination), and a gene-set sanity panel (gene count, mono-exonic fraction, protein-length distribution, mRNA:gene ratio, coding density). Covers the assembly-BUSCO-vs-proteome-BUSCO diagnostic, what BUSCO-Duplicated really means, why gene count is a vanity metric, and the QC of transferred…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/busco_genome_vs_proteome.sh`, `examples/gene_set_sanity.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Judging whether an annotation is good enough to publish
  • Diagnosing a suspect annotation
  • Comparing annotation completeness across pipelines

Example prompts

  • “Use the bio-genome-annotation-annotation-qc skill to assess the quality and completeness of a genome annotation with BUSCO (conserved single-copy…”
  • “/bio-genome-annotation-annotation-qc”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Run assembly-BUSCO vs proteome-BUSCO on the same assembly - the diagnostic almost nobody runs. -m genome does its own gene-finding; -m…
  2. Gene count is a vanity metric. The same assembly annotated by two pipelines routinely differs 20-40% in gene count, and a higher count is…
  3. Complement BUSCO with OMArk. BUSCO's single-copy-ortholog lens is blind to over-prediction, chimeras, and contamination; OMArk assesses…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Annotation Annotation Qc loads about 3.6k tokens when it runs. Until then it costs about 180 tokens; SKILL.md has 1,519 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~180
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,519 words, ~3,626 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-annotation-annotation-qc/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-genome-annotation-annotation-qc
description
Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic completeness/contamination), and a gene-set sanity panel (gene count, mono-exonic fraction, protein-length distribution, mRNA:gene ratio, coding density). Covers the assembly-BUSCO-vs-proteome-BUSCO diagnostic, what BUSCO-Duplicated really means, why gene count is a vanity metric, and the QC of transferred annotations. Use when judging whether an annotation is good enough to publish or submit, diagnosing a suspect annotation, or comparing annotation completeness across pipelines.
tool_type
cli
primary_tool
BUSCO

Version Compatibility

Reference examples tested with: BUSCO 5.5+, OMArk 0.3+, CheckM2 1.0+, compleasm 0.2.6+, gffutils 0.12+, matplotlib 3.8+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures

BUSCO results depend on the lineage dataset (e.g. vertebrata_odb10 vs the shallow eukaryota_odb10) and the OrthoDB version - record them; a 99% on a shallow ~255-gene set and a 99% on a deep ~5,500-gene clade set are very different claims. If code throws an error, introspect the installed tool and adapt rather than retrying.

Annotation QC

"Is my genome annotation any good?" -> Measure conserved-gene completeness, proteome consistency/contamination, and gene-set sanity, and decide whether the limiting factor is the assembly or the predictor.

  • CLI: busco -i proteins.faa -m proteins -l <lineage>_odb10 (proteome) and busco -i genome.fa -m genome -l <lineage>_odb10 (assembly), omark, checkm2 predict (prokaryotes)

The Single Most Important Modern Insight -- BUSCO Measures the Easy 10%; the Diagnostic Is Genome-vs-Proteome

BUSCO completeness measures recovery of the ~1,000-5,500 most conserved single-copy orthologs - the most conserved, highly expressed, intron-stable genes any half-broken pipeline will find. A 98%-complete BUSCO confirms the housekeeping core is intact; it says nothing about whether the other ~20,000 models are chimeric, fragmented, frame-shifted, fused, or hallucinated from TEs. BUSCO is a smoke detector in one room of a burning house. Three load-bearing moves:

  1. Run assembly-BUSCO vs proteome-BUSCO on the same assembly - the diagnostic almost nobody runs. -m genome does its own gene-finding; -m proteins scores the delivered proteome. If assembly is 98% and proteome is 85%, the genes are physically present and the predictor missed them - a training/evidence/masking problem, fixable without touching the assembly. If both are low, the genes aren't in the assembly - stop annotating and fix the assembly. This single fork redirects more wasted effort than any other check. (Always run BUSCO on the delivered proteome, -m proteins, not genome mode reported as if it described the annotation.)
  2. Gene count is a vanity metric. The same assembly annotated by two pipelines routinely differs 20-40% in gene count, and a higher count is as likely to mean spurious TE ORFs and split models as more real genes. The diagnostic signal is in the mono-exonic fraction, protein-length distribution, mean exons/gene, and mRNA:gene ratio - not the headline count.
  3. Complement BUSCO with OMArk. BUSCO's single-copy-ortholog lens is blind to over-prediction, chimeras, and contamination; OMArk assesses the whole proteome for completeness AND consistency (are genes consistent with their homologs?) and detects contaminant species. Use both.

Tool Taxonomy

ToolCitationMeasuresDomain
BUSCOManni 2021 Mol Biol Evolconserved single-copy ortholog recovery (C/D/F/M)euk + prok + viral; proteome/genome/transcriptome modes
OMArkNevers 2025 Nat Biotechnolproteome completeness + consistency + contaminationeukaryotic proteomes; catches over-prediction BUSCO misses
compleasmHuang 2023faster miniprot-based BUSCO-compatible completenesseuk; genome mode
CheckM2Chklovski 2023 Nat Methodscompleteness + contamination (ML, lineage-agnostic)bacteria/archaea, isolates + MAGs
Gene-set sanity panel(gffutils)gene count, mono-exonic %, protein length, mRNA:gene, coding densityany; the panel BUSCO can't provide
LAIOu 2018 NARLTR-RT assembly resolution (repeat-space contiguity)LTR-rich genomes; assembly-side QC

Decision Tree by Scenario

ScenarioQC to runWhy
Prokaryotic genome/MAGCheckM2 (completeness/contamination) + coding density + tRNA/rRNA countsCheckM2 is the field-standard completeness call; BUSCO is conservative on prokaryotes
Eukaryotic annotationBUSCO -m proteins AND -m genome + OMArk + gene-set sanity panelthe genome-vs-proteome fork + over-prediction/contamination
Suspect high gene countmono-exonic fraction + protein-length + BUSCO-Duplicateddistinguish haplotigs / unmasked TEs / split models
High BUSCO-Duplicatedcheck synteny + clade ploidyuncollapsed haplotigs (purge) vs real WGD (keep)
Transferred annotationBUSCO on the lifted set vs reference + ORF integrityquantify silently lost conserved genes
Low proteome-BUSCO, high genome-BUSCOfix evidence/masking/trainingthe genes are present; the predictor missed them
Both BUSCO low-> genome-assembly/assembly-qc (purge_dups, more data)the limit is the assembly

BUSCO and How to Read It

bash
busco -i proteins.faa -m proteins -l vertebrata_odb10 -o busco_prot -c 16   # the delivered proteome
busco -i genome.fa    -m genome   -l vertebrata_odb10 -o busco_genome -c 16 # the assembly

Reported as C:[S,D],F,M: Complete (full-length match), Duplicated (subset of Complete, found ≥2x), Fragmented (partial), Missing. Use the deepest applicable clade dataset, not the shallow eukaryota_odb10 (a 99% on a 255-gene set is trivially easy and not comparable to a 99% on a 5,500-gene clade set). High Duplicated is the most-misread signal - three causes, opposite responses: uncollapsed haplotigs (genome-wide, heterozygosity-scaled, non-syntenic -> purge_dups before annotating), real recent WGD (syntenic, ploidy-consistent, Ks duplication peak -> keep), or split models from a fragmented assembly. A clean haploid annotation runs ~1-3% Duplicated; >5-8% with no known WGD screams purge_dups first. compleasm is a faster miniprot-based alternative that often reports higher completeness than BUSCO's metaeuk path.

OMArk (Proteome Consistency and Contamination)

bash
omamer search --db LUCA.h5 --query proteins.faa --out proteins.omamer
omark -f proteins.omamer -d LUCA.h5 -o omark_out

OMArk catches what BUSCO's single-copy lens misses: it classifies the whole proteome as consistent / inconsistent (a gene whose structure conflicts with its homologs - a chimera or fragment) / unknown, and detects contaminant species mixed into the proteome. A high "inconsistent" fraction flags over-prediction or fused/split models even when BUSCO is green.

CheckM2 (Prokaryotic Completeness/Contamination)

bash
checkm2 predict --input genome.fna --output-directory checkm2_out --threads 16

For bacteria/archaea, CheckM2's lineage-agnostic ML model is the standard completeness/contamination call (handles reduced/novel lineages where marker sets fail). Gate annotation on it: contamination >5% mixes two organisms' genes into a chimeric set; completeness <90% with contamination >5-10% makes gene count, coding density, and hypothetical fraction uninterpretable - fix the assembly/binning first.

Gene-Set Sanity Panel with Python

Goal: Compute the triage panel that reveals annotation health where gene count and BUSCO cannot.

Approach: Load the GFF3 into gffutils; compute the mRNA:gene ratio (1.00 = isoform/UTR-naive), the mono-exonic fraction, and the protein-length distribution; flag clade-anomalous values.

python
import gffutils

MONOEXONIC_FLAG = 0.30   # >30% single-exon in a vertebrate suggests unmasked TEs/pseudogenes/fragments (calibrate per clade; fungi run higher)

def sanity_panel(gff_file):
    db = gffutils.create_db(gff_file, ':memory:', merge_strategy='merge')
    genes = list(db.features_of_type('gene'))
    mrnas = list(db.features_of_type(['mRNA', 'transcript']))
    exon_counts = [len(list(db.children(tx, featuretype='exon'))) for tx in mrnas]
    mono_frac = sum(1 for e in exon_counts if e == 1) / len(exon_counts) if exon_counts else 0
    mrna_per_gene = len(mrnas) / len(genes) if genes else 0
    if mrna_per_gene <= 1.001:
        print('WARNING: mRNA:gene == 1.00 -- isoform/UTR-naive; AS/3-prime-tag scRNA-seq analyses untrustworthy')
    if mono_frac > MONOEXONIC_FLAG:
        print(f'WARNING: mono-exonic fraction {mono_frac:.1%} high -- check masking/contamination')
    return {'genes': len(genes), 'mrna_per_gene': mrna_per_gene, 'mono_exonic_fraction': mono_frac}

Read gene count against the nearest well-annotated relative and the species' ploidy, never in isolation (1.5-2x with no WGD = haplotigs/over-prediction; ~0.5x = over-masking/under-training). A healthy protein-length distribution is unimodal near the clade-typical ~300-450 aa; a sub-100-aa spike = spurious/fragmented calls; a fat left tail = partials from a fragmented assembly.

Show full SKILL.md (582 more words)Show less

Per-Method Failure Modes

Reporting genome-mode BUSCO as the annotation's score

Trigger: running BUSCO -m genome and citing it as annotation quality. Mechanism: genome mode does its own gene-finding, often better than the author's pipeline on the conserved core. Symptom: reported BUSCO higher than the real proteome BUSCO. Fix: run -m proteins on the delivered proteome; report both and compare.

High Duplicated read as success

Trigger: treating high BUSCO-D as good. Mechanism: uncollapsed haplotigs vs real WGD vs split models. Symptom: inflated gene count. Fix: plot duplicated pairs against synteny + clade ploidy; if no WGD and D>5-8%, purge_dups the assembly first.

Shallow lineage dataset

Trigger: eukaryota_odb10 (~255 genes) instead of the deep clade set. Mechanism: a small, conserved set is trivially easy to hit 99%. Symptom: misleadingly high completeness. Fix: use the deepest applicable clade dataset; record it.

BUSCO alone, no OMArk

Trigger: judging a proteome by BUSCO completeness only. Mechanism: the single-copy-ortholog lens misses over-prediction, chimeras, and contamination. Symptom: green BUSCO on a proteome full of fused/fragmented or contaminant models. Fix: add OMArk for consistency + contamination.

Annotating before the QC gate (prokaryote)

Trigger: Bakta/Prokka before CheckM2. Mechanism: contamination mixes organisms; low completeness truncates. Symptom: chimeric/inflated gene set, uninterpretable coding density. Fix: CheckM2 first; contamination >5% -> decontaminate.

Quantitative Thresholds

ThresholdSourceRationale
Assembly-BUSCO vs proteome-BUSCO gapthe diagnostic forklarge gap = predictor missed present genes; both low = fix assembly
BUSCO-Duplicated ~1-3% (clean haploid); >5-8% no WGDassembly normpurge_dups before annotating
Use deepest applicable clade _odb10BUSCO guidanceshallow sets trivially hit 99%
Mono-exonic ~10-20% (vertebrate); >25-30% flagclade normunmasked TEs/pseudogenes/fragments; fungi legitimately higher
Protein length unimodal ~300-450 aaeukaryote normsub-100-aa spike = spurious; fat left tail = partials
mRNA:gene ratio == 1.00annotation structureisoform/UTR-naive
CheckM2 contamination ≤5%, completeness ≥90%MIMAG-alignedabove/below -> prokaryotic QC numbers uninterpretable
Prokaryotic coding density ~88-90%bacterial norm<85% wrong table/fragmentation; >93% over-calling

Common Errors

Error / symptomCauseSolution
BUSCO high, annotation still badBUSCO measures only the conserved coreadd OMArk + sanity panel; inspect mono-exonic models
Proteome-BUSCO < genome-BUSCOpredictor missed present genesfix evidence/masking/training, not the assembly
High Duplicatedhaplotigs vs WGD vs split modelssynteny + ploidy check; purge_dups if no WGD
99% completeness looks too goodshallow lineage datasetrerun with the deep clade set
Cross-pipeline gene counts disagreegene count is not a quality metriccompare sanity panels, not counts
CheckM2 high contaminationmixed organisms / poor binningdecontaminate before annotation

References

  • Manni M, et al. 2021. BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes. Mol Biol Evol 38:4647-4654.
  • Simão FA, et al. 2015. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31:3210-3212.
  • Nevers Y, et al. 2025. Quality assessment of gene repertoire annotations with OMArk. Nat Biotechnol 43:124-133.
  • Chklovski A, et al. 2023. CheckM2: a rapid, scalable and accurate tool for assessing microbial genome quality using machine learning. Nat Methods 20:1203-1212.
  • Huang N, Li H. 2023. compleasm: a faster and more accurate reimplementation of BUSCO. Bioinformatics 39:btad595.
  • Guan D, et al. 2020. Identifying and removing haplotypic duplication in primary genome assemblies (purge_dups). Bioinformatics 36:2896-2898.
  • Ou S, Chen J, Jiang N. 2018. Assessing genome assembly quality using the LTR Assembly Index (LAI). Nucleic Acids Res 46:e126.
  • eukaryotic-gene-prediction - The annotation whose proteome/gene-set this QC evaluates
  • prokaryotic-annotation - CheckM2 gate and coding-density sanity for prokaryotes
  • annotation-transfer - BUSCO on a lifted set quantifies silently lost conserved genes
  • repeat-annotation - LAI and over-masking diagnosis in the assembly-to-annotation handoff
  • genome-assembly/assembly-qc - Assembly-side completeness; purge haplotigs before annotating

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in genome-annotation/annotation-qc of GPTomics/bioSkills.

  • SKILL.md
  • examples/busco_genome_vs_proteome.sh
  • examples/gene_set_sanity.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Annotation Annotation Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Annotation Annotation Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Annotation Annotation Qc this skillGPTomics/bioSkills1.2k1 repos~3.6kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Genome Annotation Annotation Qc

What does Bio Genome Annotation Annotation Qc do?

Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic…. Bio Genome Annotation Annotation Qc is an agent skill from GPTomics/bioSkills. Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic completeness/contamination), and a gene-set sanity panel (gene count, mono-exonic fraction, protein-length distribution, mRNA:gene ratio, coding density).

When should I use Bio Genome Annotation Annotation Qc?

Bio Genome Annotation Annotation Qc fits situations like: judging whether an annotation is good enough to publish; diagnosing a suspect annotation; comparing annotation completeness across pipelines.

How do I install Bio Genome Annotation Annotation Qc in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a claude-code`. Or copy the skill folder (genome-annotation/annotation-qc in GPTomics/bioSkills) into .claude/skills/bio-genome-annotation-annotation-qc in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Annotation Annotation Qc in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a codex`. Or copy the skill folder (genome-annotation/annotation-qc in GPTomics/bioSkills) into .agents/skills/bio-genome-annotation-annotation-qc in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Annotation Annotation Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-annotation-annotation-qc, .gemini/skills/bio-genome-annotation-annotation-qc, .github/skills/bio-genome-annotation-annotation-qc and .opencode/skills/bio-genome-annotation-annotation-qc in your project.

What does Bio Genome Annotation Annotation Qc need to run?

Going by SKILL.md and its folder, Bio Genome Annotation Annotation Qc needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Genome Annotation Annotation Qc access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Genome Annotation Annotation Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Annotation Annotation Qc use?

Bio Genome Annotation Annotation Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Annotation Annotation Qc use?

About 3.6k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Annotation Annotation Qc?

Skills that share tags, products or a category with Bio Genome Annotation Annotation Qc: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Annotation Annotation Qc?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.