Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic…
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-annotation-qc --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-annotation/annotation-qc .claude/skills/bio-genome-annotation-annotation-qc && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-genome-annotation-annotation-qc" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/annotation-qc into .claude/skills/bio-genome-annotation-annotation-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-annotation-qc", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/annotation-qcType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-annotation-qc --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/genome-annotation/annotation-qc .agents/skills/bio-genome-annotation-annotation-qc && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-genome-annotation-annotation-qc" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/annotation-qc into .agents/skills/bio-genome-annotation-annotation-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-annotation-qc", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-annotation-qc --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/genome-annotation/annotation-qc .cursor/skills/bio-genome-annotation-annotation-qc && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-genome-annotation-annotation-qc" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/annotation-qc into .cursor/skills/bio-genome-annotation-annotation-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-annotation-qc", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path genome-annotation/annotation-qc--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-annotation-qc --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/genome-annotation/annotation-qc .gemini/skills/bio-genome-annotation-annotation-qc && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-genome-annotation-annotation-qc" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/annotation-qc into .gemini/skills/bio-genome-annotation-annotation-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-annotation-qc", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-annotation-qcInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/genome-annotation/annotation-qc .github/skills/bio-genome-annotation-annotation-qc && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-genome-annotation-annotation-qc" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/annotation-qc into .github/skills/bio-genome-annotation-annotation-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-annotation-qc", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-genome-annotation-annotation-qc --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/genome-annotation/annotation-qc .opencode/skills/bio-genome-annotation-annotation-qc && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-genome-annotation-annotation-qc" agent skill from https://github.com/GPTomics/bioSkills/tree/main/genome-annotation/annotation-qc into .opencode/skills/bio-genome-annotation-annotation-qc/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-genome-annotation-annotation-qc", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-genome-annotation-annotation-qcAssesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic…
Bio Genome Annotation Annotation Qc is an agent skill from GPTomics/bioSkills. Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic completeness/contamination), and a gene-set sanity panel (gene count, mono-exonic fraction, protein-length distribution, mRNA:gene ratio, coding density). Covers the assembly-BUSCO-vs-proteome-BUSCO diagnostic, what BUSCO-Duplicated really means, why gene count is a vanity metric, and the QC of transferred…
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/busco_genome_vs_proteome.sh`, `examples/gene_set_sanity.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell and Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Genome Annotation Annotation Qc loads about 3.6k tokens when it runs. Until then it costs about 180 tokens; SKILL.md has 1,519 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,519 words, ~3,626 tokens.
.claude/skills/bio-genome-annotation-annotation-qc/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: BUSCO 5.5+, OMArk 0.3+, CheckM2 1.0+, compleasm 0.2.6+, gffutils 0.12+, matplotlib 3.8+.
Before using code patterns, verify installed versions match. If versions differ:
<tool> --version then <tool> --help to confirm flagspip show <package> then help(module.function) to check signaturesBUSCO results depend on the lineage dataset (e.g. vertebrata_odb10 vs the shallow eukaryota_odb10) and the OrthoDB version - record them; a 99% on a shallow ~255-gene set and a 99% on a deep ~5,500-gene clade set are very different claims. If code throws an error, introspect the installed tool and adapt rather than retrying.
"Is my genome annotation any good?" -> Measure conserved-gene completeness, proteome consistency/contamination, and gene-set sanity, and decide whether the limiting factor is the assembly or the predictor.
busco -i proteins.faa -m proteins -l <lineage>_odb10 (proteome) and busco -i genome.fa -m genome -l <lineage>_odb10 (assembly), omark, checkm2 predict (prokaryotes)BUSCO completeness measures recovery of the ~1,000-5,500 most conserved single-copy orthologs - the most conserved, highly expressed, intron-stable genes any half-broken pipeline will find. A 98%-complete BUSCO confirms the housekeeping core is intact; it says nothing about whether the other ~20,000 models are chimeric, fragmented, frame-shifted, fused, or hallucinated from TEs. BUSCO is a smoke detector in one room of a burning house. Three load-bearing moves:
-m genome does its own gene-finding; -m proteins scores the delivered proteome. If assembly is 98% and proteome is 85%, the genes are physically present and the predictor missed them - a training/evidence/masking problem, fixable without touching the assembly. If both are low, the genes aren't in the assembly - stop annotating and fix the assembly. This single fork redirects more wasted effort than any other check. (Always run BUSCO on the delivered proteome, -m proteins, not genome mode reported as if it described the annotation.)| Tool | Citation | Measures | Domain |
|---|---|---|---|
| BUSCO | Manni 2021 Mol Biol Evol | conserved single-copy ortholog recovery (C/D/F/M) | euk + prok + viral; proteome/genome/transcriptome modes |
| OMArk | Nevers 2025 Nat Biotechnol | proteome completeness + consistency + contamination | eukaryotic proteomes; catches over-prediction BUSCO misses |
| compleasm | Huang 2023 | faster miniprot-based BUSCO-compatible completeness | euk; genome mode |
| CheckM2 | Chklovski 2023 Nat Methods | completeness + contamination (ML, lineage-agnostic) | bacteria/archaea, isolates + MAGs |
| Gene-set sanity panel | (gffutils) | gene count, mono-exonic %, protein length, mRNA:gene, coding density | any; the panel BUSCO can't provide |
| LAI | Ou 2018 NAR | LTR-RT assembly resolution (repeat-space contiguity) | LTR-rich genomes; assembly-side QC |
| Scenario | QC to run | Why |
|---|---|---|
| Prokaryotic genome/MAG | CheckM2 (completeness/contamination) + coding density + tRNA/rRNA counts | CheckM2 is the field-standard completeness call; BUSCO is conservative on prokaryotes |
| Eukaryotic annotation | BUSCO -m proteins AND -m genome + OMArk + gene-set sanity panel | the genome-vs-proteome fork + over-prediction/contamination |
| Suspect high gene count | mono-exonic fraction + protein-length + BUSCO-Duplicated | distinguish haplotigs / unmasked TEs / split models |
| High BUSCO-Duplicated | check synteny + clade ploidy | uncollapsed haplotigs (purge) vs real WGD (keep) |
| Transferred annotation | BUSCO on the lifted set vs reference + ORF integrity | quantify silently lost conserved genes |
| Low proteome-BUSCO, high genome-BUSCO | fix evidence/masking/training | the genes are present; the predictor missed them |
| Both BUSCO low | -> genome-assembly/assembly-qc (purge_dups, more data) | the limit is the assembly |
busco -i proteins.faa -m proteins -l vertebrata_odb10 -o busco_prot -c 16 # the delivered proteome
busco -i genome.fa -m genome -l vertebrata_odb10 -o busco_genome -c 16 # the assemblyReported as C:[S,D],F,M: Complete (full-length match), Duplicated (subset of Complete, found ≥2x), Fragmented (partial), Missing. Use the deepest applicable clade dataset, not the shallow eukaryota_odb10 (a 99% on a 255-gene set is trivially easy and not comparable to a 99% on a 5,500-gene clade set). High Duplicated is the most-misread signal - three causes, opposite responses: uncollapsed haplotigs (genome-wide, heterozygosity-scaled, non-syntenic -> purge_dups before annotating), real recent WGD (syntenic, ploidy-consistent, Ks duplication peak -> keep), or split models from a fragmented assembly. A clean haploid annotation runs ~1-3% Duplicated; >5-8% with no known WGD screams purge_dups first. compleasm is a faster miniprot-based alternative that often reports higher completeness than BUSCO's metaeuk path.
omamer search --db LUCA.h5 --query proteins.faa --out proteins.omamer
omark -f proteins.omamer -d LUCA.h5 -o omark_outOMArk catches what BUSCO's single-copy lens misses: it classifies the whole proteome as consistent / inconsistent (a gene whose structure conflicts with its homologs - a chimera or fragment) / unknown, and detects contaminant species mixed into the proteome. A high "inconsistent" fraction flags over-prediction or fused/split models even when BUSCO is green.
checkm2 predict --input genome.fna --output-directory checkm2_out --threads 16For bacteria/archaea, CheckM2's lineage-agnostic ML model is the standard completeness/contamination call (handles reduced/novel lineages where marker sets fail). Gate annotation on it: contamination >5% mixes two organisms' genes into a chimeric set; completeness <90% with contamination >5-10% makes gene count, coding density, and hypothetical fraction uninterpretable - fix the assembly/binning first.
Goal: Compute the triage panel that reveals annotation health where gene count and BUSCO cannot.
Approach: Load the GFF3 into gffutils; compute the mRNA:gene ratio (1.00 = isoform/UTR-naive), the mono-exonic fraction, and the protein-length distribution; flag clade-anomalous values.
import gffutils
MONOEXONIC_FLAG = 0.30 # >30% single-exon in a vertebrate suggests unmasked TEs/pseudogenes/fragments (calibrate per clade; fungi run higher)
def sanity_panel(gff_file):
db = gffutils.create_db(gff_file, ':memory:', merge_strategy='merge')
genes = list(db.features_of_type('gene'))
mrnas = list(db.features_of_type(['mRNA', 'transcript']))
exon_counts = [len(list(db.children(tx, featuretype='exon'))) for tx in mrnas]
mono_frac = sum(1 for e in exon_counts if e == 1) / len(exon_counts) if exon_counts else 0
mrna_per_gene = len(mrnas) / len(genes) if genes else 0
if mrna_per_gene <= 1.001:
print('WARNING: mRNA:gene == 1.00 -- isoform/UTR-naive; AS/3-prime-tag scRNA-seq analyses untrustworthy')
if mono_frac > MONOEXONIC_FLAG:
print(f'WARNING: mono-exonic fraction {mono_frac:.1%} high -- check masking/contamination')
return {'genes': len(genes), 'mrna_per_gene': mrna_per_gene, 'mono_exonic_fraction': mono_frac}Read gene count against the nearest well-annotated relative and the species' ploidy, never in isolation (1.5-2x with no WGD = haplotigs/over-prediction; ~0.5x = over-masking/under-training). A healthy protein-length distribution is unimodal near the clade-typical ~300-450 aa; a sub-100-aa spike = spurious/fragmented calls; a fat left tail = partials from a fragmented assembly.
Trigger: running BUSCO -m genome and citing it as annotation quality. Mechanism: genome mode does its own gene-finding, often better than the author's pipeline on the conserved core. Symptom: reported BUSCO higher than the real proteome BUSCO. Fix: run -m proteins on the delivered proteome; report both and compare.
Trigger: treating high BUSCO-D as good. Mechanism: uncollapsed haplotigs vs real WGD vs split models. Symptom: inflated gene count. Fix: plot duplicated pairs against synteny + clade ploidy; if no WGD and D>5-8%, purge_dups the assembly first.
Trigger: eukaryota_odb10 (~255 genes) instead of the deep clade set. Mechanism: a small, conserved set is trivially easy to hit 99%. Symptom: misleadingly high completeness. Fix: use the deepest applicable clade dataset; record it.
Trigger: judging a proteome by BUSCO completeness only. Mechanism: the single-copy-ortholog lens misses over-prediction, chimeras, and contamination. Symptom: green BUSCO on a proteome full of fused/fragmented or contaminant models. Fix: add OMArk for consistency + contamination.
Trigger: Bakta/Prokka before CheckM2. Mechanism: contamination mixes organisms; low completeness truncates. Symptom: chimeric/inflated gene set, uninterpretable coding density. Fix: CheckM2 first; contamination >5% -> decontaminate.
| Threshold | Source | Rationale |
|---|---|---|
| Assembly-BUSCO vs proteome-BUSCO gap | the diagnostic fork | large gap = predictor missed present genes; both low = fix assembly |
| BUSCO-Duplicated ~1-3% (clean haploid); >5-8% no WGD | assembly norm | purge_dups before annotating |
Use deepest applicable clade _odb10 | BUSCO guidance | shallow sets trivially hit 99% |
| Mono-exonic ~10-20% (vertebrate); >25-30% flag | clade norm | unmasked TEs/pseudogenes/fragments; fungi legitimately higher |
| Protein length unimodal ~300-450 aa | eukaryote norm | sub-100-aa spike = spurious; fat left tail = partials |
| mRNA:gene ratio == 1.00 | annotation structure | isoform/UTR-naive |
| CheckM2 contamination ≤5%, completeness ≥90% | MIMAG-aligned | above/below -> prokaryotic QC numbers uninterpretable |
| Prokaryotic coding density ~88-90% | bacterial norm | <85% wrong table/fragmentation; >93% over-calling |
| Error / symptom | Cause | Solution |
|---|---|---|
| BUSCO high, annotation still bad | BUSCO measures only the conserved core | add OMArk + sanity panel; inspect mono-exonic models |
| Proteome-BUSCO < genome-BUSCO | predictor missed present genes | fix evidence/masking/training, not the assembly |
| High Duplicated | haplotigs vs WGD vs split models | synteny + ploidy check; purge_dups if no WGD |
| 99% completeness looks too good | shallow lineage dataset | rerun with the deep clade set |
| Cross-pipeline gene counts disagree | gene count is not a quality metric | compare sanity panels, not counts |
| CheckM2 high contamination | mixed organisms / poor binning | decontaminate before annotation |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in genome-annotation/annotation-qc of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Genome Annotation Annotation Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Genome Annotation Annotation Qc this skillGPTomics/bioSkills | 1.2k | 1 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic…. Bio Genome Annotation Annotation Qc is an agent skill from GPTomics/bioSkills. Assesses the quality and completeness of a genome annotation with BUSCO (conserved single-copy ortholog recovery), OMArk (proteome completeness, consistency, and contamination), CheckM2 (prokaryotic completeness/contamination), and a gene-set sanity panel (gene count, mono-exonic fraction, protein-length distribution, mRNA:gene ratio, coding density).
Bio Genome Annotation Annotation Qc fits situations like: judging whether an annotation is good enough to publish; diagnosing a suspect annotation; comparing annotation completeness across pipelines.
Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a claude-code`. Or copy the skill folder (genome-annotation/annotation-qc in GPTomics/bioSkills) into .claude/skills/bio-genome-annotation-annotation-qc in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a codex`. Or copy the skill folder (genome-annotation/annotation-qc in GPTomics/bioSkills) into .agents/skills/bio-genome-annotation-annotation-qc in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-annotation-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-annotation-annotation-qc, .gemini/skills/bio-genome-annotation-annotation-qc, .github/skills/bio-genome-annotation-annotation-qc and .opencode/skills/bio-genome-annotation-annotation-qc in your project.
Going by SKILL.md and its folder, Bio Genome Annotation Annotation Qc needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Genome Annotation Annotation Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Genome Annotation Annotation Qc: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.