Agent skill

Bio Genome Annotation Eukaryotic Gene Prediction

by GPTomics in GPTomics/bioSkills

Predicts protein-coding gene structures (exons, introns, UTRs) in eukaryotic genomes with BRAKER3 (RNA-seq + protein evidence), BRAKER1/BRAKER2, GALBA (protein-only), Funannotate (fungi), GeMoMa…

MITAuto-check passedResearch & Science

Install Bio Genome Annotation Eukaryotic Gene Prediction

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-eukaryotic-gene-prediction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-annotation-eukaryotic-gene-prediction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-annotation/eukaryotic-gene-prediction .claude/skills/bio-genome-annotation-eukaryotic-gene-prediction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-annotation-eukaryotic-gene-prediction
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.3k tokens
SKILL.md length
1,780 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Predicts protein-coding gene structures (exons, introns, UTRs) in eukaryotic genomes with BRAKER3 (RNA-seq + protein evidence), BRAKER1/BRAKER2, GALBA (protein-only), Funannotate (fungi), GeMoMa…

  • Annotating a newly assembled eukaryotic genome
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 11 more sections
  • Runs Shell scripts from its folder; calls pip
  • Choosing a gene-prediction pipeline based on available evidence

What it does

Bio Genome Annotation Eukaryotic Gene Prediction is an agent skill from GPTomics/bioSkills. Predicts protein-coding gene structures (exons, introns, UTRs) in eukaryotic genomes with BRAKER3 (RNA-seq + protein evidence), BRAKER1/BRAKER2, GALBA (protein-only), Funannotate (fungi), GeMoMa (homology projection), or Helixer/Tiberius (deep-learning ab initio). Covers the evidence-first tool decision, mandatory soft-masking, the training-set-quality-dominates principle, OrthoDB clade-partition selection, the one-isoform-per-locus and missing-UTR traps, merge/split errors, and reference bias against orphan…

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/braker3_pipeline.sh`, `examples/galba_protein_only.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Annotating a newly assembled eukaryotic genome
  • Choosing a gene-prediction pipeline based on available evidence
  • Diagnosing a poor annotation

Example prompts

  • “Use the bio-genome-annotation-eukaryotic-gene-prediction skill to predict protein-coding gene structures (exons, introns, UTRs) in eukaryotic…”
  • “/bio-genome-annotation-eukaryotic-gene-prediction”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Annotation Eukaryotic Gene Prediction loads about 4.3k tokens when it runs. Until then it costs about 181 tokens; SKILL.md has 1,780 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~181
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,780 words, ~4,338 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-annotation-eukaryotic-gene-prediction/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-genome-annotation-eukaryotic-gene-prediction
description
Predicts protein-coding gene structures (exons, introns, UTRs) in eukaryotic genomes with BRAKER3 (RNA-seq + protein evidence), BRAKER1/BRAKER2, GALBA (protein-only), Funannotate (fungi), GeMoMa (homology projection), or Helixer/Tiberius (deep-learning ab initio). Covers the evidence-first tool decision, mandatory soft-masking, the training-set-quality-dominates principle, OrthoDB clade-partition selection, the one-isoform-per-locus and missing-UTR traps, merge/split errors, and reference bias against orphan genes. Use when annotating a newly assembled eukaryotic genome, choosing a gene-prediction pipeline based on available evidence, or diagnosing a poor annotation.
tool_type
cli
primary_tool
BRAKER3

Version Compatibility

Reference examples tested with: BRAKER 3.0+, GALBA 1.0.11+, AUGUSTUS 3.5+, Funannotate 1.8+, HISAT2 2.2.1+, STAR 2.7.11+, BUSCO 5.5+, samtools 1.19+, gffutils 0.12+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures

Reproducibility is engineered, not assumed: run BRAKER/GALBA/Funannotate from their official containers, and pin the OrthoDB partition version (v11 vs v12 give different hints -> different models), the repeat library, and the Dfam release. GeneMark (inside BRAKER) historically required an expiring academic .gm_key; the requirement has relaxed in recent versions but check the installed version's docs. If code throws an error, introspect the installed tool and adapt rather than retrying.

Eukaryotic Gene Prediction

"Predict genes in my eukaryotic genome" -> Identify protein-coding gene structures from a soft-masked assembly using RNA-seq and/or protein evidence to train and guide a gene finder.

  • CLI: braker.pl --genome=masked.fa --prot_seq=orthodb_clade.fa --rnaseq_sets_ids=SRR... --softmasking (BRAKER3)

The Single Most Important Modern Insight -- Training-Set Quality and Evidence Dominate, Not the Algorithm

Gene-prediction accuracy is governed by the quality of the training set and the extrinsic evidence, not by which gene-finder is chosen. A 2024-era HMM (AUGUSTUS) trained on a clean, evidence-validated gene set beats a fancier model trained on garbage. This reframes every decision:

  • BRAKER3's accuracy jump is not a better HMM and not the TSEBRA combiner (the common misattribution). It is GeneMark-ETP's method of mining a high-confidence training set from loci where RNA-seq-assembled transcripts AND protein homology independently agree, then training AUGUSTUS on that (Brůna 2024 Genome Res 34:757; Gabriel 2024 Genome Res 34:769). BRAKER3 beats BRAKER1/2 because it learns from loci it has two reasons to trust.
  • Therefore the first question is "what evidence do I have, and is it from the right clade?" - not "which tool?" Soft-masking and OrthoDB-partition choice are upstream of, and more consequential than, the predictor.
  • The #1 silent killer is training on a bad assembly. A finder trained on a contaminated, fragmented, or repeat-polluted assembly produces confidently wrong models genome-wide that are syntactically valid, BUSCO-complete, and undetectable from the GFF3. Decontaminate and check assembly contiguity/BUSCO before any self-trainer touches the genome.
  • The 2025-2026 deep-learning twist: Tiberius/Helixer can match BRAKER3's evidence-based accuracy with no evidence at all on well-represented clades (vertebrates) - but they degrade on under-represented lineages and are isoform/UTR-naive. The field is mid-transition; do not present DL as universally superior.

Tool Taxonomy

PipelineCitationEvidence usedWhen
BRAKER3Gabriel 2024 Genome ResRNA-seq + protein (OrthoDB)Default when both exist; HC-training-set method
BRAKER1Hoff 2016 BioinformaticsRNA-seq onlyspliced-read intron hints train GeneMark-ET
BRAKER2Brůna 2021 NAR Genom Bioinformprotein only (broad OrthoDB)ProtHint mines hints from a remote protein DB
GALBABrůna 2023 BMC Bioinformaticsprotein only (close relatives)beats BRAKER2 on large vertebrate genomes with good close proteomes
GeMoMaKeilwagen 2018 BMC Bioinformaticsreference annotation + intron-position conservationproject an existing close-relative annotation
FunannotatePalmer & Stajich 2020 (Zenodo)anyfungal de facto standard; train->predict(EVM)->update(PASA)
MAKER2 / EVMHolt 2011; Haas 2008many trackscombiner / build-it-yourself route; transparent evidence weighting
Helixer / TiberiusStiehler 2020; Gabriel 2024none (ab initio, DL)evidence-poor genomes in well-represented clades
AUGUSTUS / GeneMark-ESStanke 2006; Lomsadzehints / self-traincomponents; standalone ab initio is the last resort

OrthoDB rule: BRAKER2/3 want a clade partition (pick the smallest partition that still contains the target clade - Vertebrata for a fish, not all Metazoa). GALBA wants a few close-relative proteomes, not a broad clade.

Decision Tree by Scenario

Evidence / scenarioRecommendedWhy
RNA-seq + protein DBBRAKER3state-of-the-art at all genome sizes
RNA-seq onlyBRAKER1intron hints from spliced reads
Protein only, close relativesGALBAminiprot-aligned close proteomes train AUGUSTUS
Protein only, broad/distantBRAKER2ProtHint mines a remote OrthoDB clade
No evidence, represented cladeTiberius or HelixerDL ab initio now matches evidence-based on vertebrates
Reference annotation of a close relativeGeMoMahomology + intron-position projection
Fungus (any evidence)Funannotatetiny-intron-aware; bundled EVM + PASA update
Need isoforms + UTRsadd PASA / Iso-Seq update steppredictors emit one CDS-only model per locus
Genome not yet masked-> repeat-annotation (soft-mask first)mandatory prerequisite
Want to assess the result-> annotation-qcBUSCO genome-vs-proteome, OMArk, sanity metrics

Soft-Masking Is Mandatory (Prerequisite)

Run repeat-annotation to soft-mask (repeats -> lowercase) before prediction. Unmasked TEs contain ORFs and pseudo-splice-sites; the predictor calls thousands of spurious genes inside repeats (catastrophic in plants where >80% of the genome can be TE) and TE domains pollute training. Pass --softmasking so BRAKER honors lowercase as a soft penalty (a real gene can still span a repeat). Hard-masking (repeats -> N) destroys sequence and truncates real repeat-overlapping genes - avoid it for prediction. But over-aggressive masking with an uncurated library deletes real multi-copy families (NLR/R-genes, zinc-fingers): filter the repeat library against a protein DB and confirm conserved families survive.

BRAKER3 (RNA-seq + Protein)

bash
# Align RNA-seq with a splice-aware aligner; output sorted BAM
hisat2-build masked.fasta idx
hisat2 -x idx -1 R1.fq.gz -2 R2.fq.gz --dta -p 16 | samtools sort -@4 -o rnaseq.bam
samtools index rnaseq.bam

# BRAKER3: protein = an OrthoDB clade partition (smallest that contains the clade)
braker.pl --genome=masked.fasta --prot_seq=Vertebrata.fa \
    --bam=rnaseq.bam --softmasking --threads=16 --species=my_species \
    --gff3 --workingdir=braker3_out

--rnaseq_sets_ids=SRR...,SRR... --rnaseq_sets_dirs=/fastq/ auto-downloads/aligns reads in place of --bam. Outputs: braker.gtf/braker.gff3 (TSEBRA-combined), braker.codingseq, braker.aa.

GALBA (Protein-Only, Close Relatives)

bash
galba.pl --genome=masked.fasta --prot_seq=close_relatives.faa \
    --species=my_species --threads=16

Prefer BRAKER3 whenever RNA-seq exists - intron evidence substantially improves splice-site accuracy.

Funannotate (Fungi)

bash
funannotate mask -i assembly.fa -o masked.fa --cpus 16
funannotate train -i masked.fa -o out -l R1.fq -r R2.fq --species "Genus species"
funannotate predict -i masked.fa -o out -s "Genus species" \
    --transcript_evidence transcripts.fa --protein_evidence proteins.fa
funannotate update -i out --cpus 16     # PASA adds UTRs and isoforms

Gene-Model Sanity Statistics with Python

Goal: Compute the triage panel that reveals annotation health where gene count and BUSCO cannot - the isoform ratio, mono-exonic fraction, and protein-length distribution.

Approach: Load the GFF3 into gffutils; compute the mRNA:gene ratio (1.00 = isoform-naive), the single-exon fraction, and CDS-length stats; flag clade-anomalous values.

python
import gffutils

MONOEXONIC_FLAG = 0.30   # >30% single-exon in a vertebrate suggests unmasked TEs/pseudogenes/fragments (calibrate per clade)

def gene_model_stats(gff_file):
    db = gffutils.create_db(gff_file, ':memory:', merge_strategy='merge')
    genes = list(db.features_of_type('gene'))
    mrnas = list(db.features_of_type(['mRNA', 'transcript']))
    exon_counts = [len(list(db.children(tx, featuretype='exon'))) for tx in mrnas]
    mono_frac = sum(1 for e in exon_counts if e == 1) / len(exon_counts) if exon_counts else 0
    mrna_per_gene = len(mrnas) / len(genes) if genes else 0
    if mrna_per_gene <= 1.001:
        print('WARNING: one isoform per locus (mRNA:gene == 1.00) -- isoform/UTR-naive; AS analyses untrustworthy')
    if mono_frac > MONOEXONIC_FLAG:
        print(f'WARNING: mono-exonic fraction {mono_frac:.1%} high -- check masking/contamination')
    return {'genes': len(genes), 'mrna_per_gene': mrna_per_gene, 'mono_exonic_fraction': mono_frac}

Hard Biology the Pipeline Gets Wrong

  • One isoform per locus, no UTRs. Almost every de novo annotation ships a single CDS-only model per gene (mRNA:gene == 1.00). This silently breaks downstream alternative-splicing/isoform-switching analysis (a switch the reference doesn't contain cannot be detected), and missing 3' UTRs break 3'-tag scRNA-seq (10x reads land "intergenic" and are discarded), APA, and miRNA-target work. The only fix is a transcript-evidence update (PASA, or Iso-Seq via funannotate update). Human GENCODE has ~4-5 isoforms/gene; a fresh annotation has one.
  • Merge/split errors are invisible to automated QC. Tandem arrays (NLR clusters, immune loci) fuse into one elongated model; genes split across contig breaks become two partials; read-through transcription fuses two genes (evidence-supported, so especially nasty); a long intron read as intergenic splits one gene in two. Long-read Iso-Seq + a contiguous assembly prevent these; they hide in the length/exon-count tails otherwise.
  • Reference bias against orphan genes. Protein-evidence pulls models toward known genes and away from lineage-specific/fast-evolving/orphan genes - exactly the novel biology. Apparent "lineage-specific" genes are often just homology-detection failure (Weisman 2020 PLoS Biol 18:e3000862). Keep well-supported ab initio/DL calls in repeat-free RNA-seq-supported regions rather than filtering to "evidence-supported only," which amputates the orphan set.
  • Protists break the spliceosome. Alternative genetic codes (ciliate UAA/UAG -> Gln), trans-splicing (kinetoplastids, nematodes), and polycistronic transcription mean generic eukaryote models truncate or mis-call. Set the correct translation table and know the RNA-processing biology before trusting any predictor.
Show full SKILL.md (682 more words)Show less

Per-Method Failure Modes

Unmasked or hard-masked genome

Trigger: running BRAKER without soft-masking, or hard-masking. Mechanism: TE ORFs become genes / masked sequence truncates real genes. Symptom: 2x inflated gene count, high mono-exonic fraction, or fragmented models. Fix: soft-mask with a curated library; pass --softmasking.

Training on a bad assembly

Trigger: annotating a contaminated/fragmented assembly. Mechanism: self-trainer learns contaminant/truncated gene structure and applies it genome-wide. Symptom: confidently wrong, BUSCO-green models. Fix: decontaminate (FCS-GX/BlobTools) and check assembly BUSCO/N50 first.

Wrong AUGUSTUS species / OrthoDB partition

Trigger: a "close enough" pre-trained species or wrong clade partition. Mechanism: splice-site/intron-length params or protein hints mismatched. Symptom: systematically mis-placed exon boundaries; clean-looking GFF3. Fix: train on the target (BRAKER does this); pick the smallest correct OrthoDB clade.

Expecting isoforms/UTRs from a one-model pipeline

Trigger: AS/3'-tag/APA analysis against a de novo annotation. Mechanism: one CDS-only model per locus. Symptom: discarded scRNA-seq reads; empty AS results. Fix: add a PASA/Iso-Seq update step; check mRNA:gene ratio.

High BUSCO-Duplicated read as success

Trigger: treating high D as good. Mechanism: uncollapsed haplotigs vs real WGD vs split models. Symptom: inflated gene count. Fix: if no known WGD and D>5-8%, purge_dups the assembly first; if known polyploid, confirm via synteny/Ks and keep.

Quantitative Thresholds

ThresholdSourceRationale
Soft-mask before predictionuniversalunmasked TEs inflate spurious genes
Gene count vs nearest relative (±, reconcile with ploidy)clade norm1.5-2x with no WGD = haplotigs/over-prediction; ~0.5x = over-masking/under-training
Mono-exonic fraction ~10-20% (vertebrate)clade norm>25-30% = unmasked TEs/pseudogenes/fragments; fungi legitimately higher
Protein length unimodal ~300-450 aaeukaryote normsub-100-aa spike = spurious/fragmented; fat left tail = partials
BUSCO-Duplicated ~1-3% (clean haploid)assembly norm>5-8% with no WGD -> purge_dups before annotating
mRNA:gene ratioannotation structure== 1.00 means isoform/UTR-naive
Pick smallest OrthoDB partition containing the cladeBRAKER guidancebroader = noisier hints, slower

Common Errors

Error / symptomCauseSolution
Thousands of extra short genesunmasked repeatssoft-mask; --softmasking
BRAKER fails mid-runexpired GeneMark key / special chars in FASTA headersuse the container; sed 's/ .*//' genome.fa
Many single-exon genesunmasked TEs / contamination / fragmented assemblyverify masking; decontaminate; check N50
Low protein-BUSCO, high genome-BUSCOpredictor missed present genes (training/evidence)fix evidence/masking, not the assembly
AS analysis returns nothingone-isoform annotationrun PASA/Iso-Seq update
Suspiciously few NLR/ZNF genesover-masked with uncurated libraryfilter repeat library against a protein DB

References

  • Stanke M, et al. 2006. Gene prediction with a hidden Markov model and a new intron submodel (AUGUSTUS). BMC Bioinformatics 7:62.
  • Hoff KJ, et al. 2016. BRAKER1: unsupervised RNA-Seq-based genome annotation with GeneMark-ET and AUGUSTUS. Bioinformatics 32:767-769.
  • Brůna T, et al. 2021. BRAKER2: automatic eukaryotic genome annotation with GeneMark-EP+ and AUGUSTUS supported by a protein database. NAR Genom Bioinform 3:lqaa108.
  • Gabriel L, et al. 2024. BRAKER3: fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA. Genome Res 34:769-777.
  • Brůna T, et al. 2024. GeneMark-ETP significantly improves the accuracy of automatic annotation of large eukaryotic genomes. Genome Res 34:757-768.
  • Gabriel L, et al. 2021. TSEBRA: transcript selector for BRAKER. BMC Bioinformatics 22:566.
  • Brůna T, et al. 2023. GALBA: genome annotation with miniprot and AUGUSTUS. BMC Bioinformatics 24:327.
  • Keilwagen J, et al. 2018. Combining RNA-seq data and homology-based gene prediction for plants, animals and fungi (GeMoMa). BMC Bioinformatics 19:189.
  • Haas BJ, et al. 2008. Automated eukaryotic gene structure annotation using EVidenceModeler. Genome Biol 9:R7.
  • Haas BJ, et al. 2003. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies (PASA). Nucleic Acids Res 31:5654-5666.
  • Gabriel L, et al. 2024. Tiberius: end-to-end deep learning with an HMM for gene prediction. Bioinformatics 40:btae685.
  • Stiehler F, et al. 2020. Helixer: cross-species gene annotation of large eukaryotic genomes using deep learning. Bioinformatics 36:5291-5298.
  • Manni M, et al. 2021. BUSCO update. Mol Biol Evol 38:4647-4654.
  • Weisman CM, et al. 2020. Many, but not all, lineage-specific genes can be explained by homology detection failure. PLoS Biol 18:e3000862.
  • Palmer JM, Stajich J. 2020. Funannotate v1.8: eukaryotic genome annotation. Zenodo. doi:10.5281/zenodo.4054262.
  • repeat-annotation - PREREQUISITE: soft-mask repeats before prediction
  • functional-annotation - Add GO/KEGG/Pfam to predicted proteins
  • annotation-qc - BUSCO genome-vs-proteome, OMArk, gene-set sanity metrics
  • ncrna-annotation - ncRNAs are not found by protein-coding prediction
  • read-alignment/star-alignment - Splice-aware RNA-seq alignment for evidence
  • genome-assembly/assembly-qc - Verify assembly quality and purge haplotigs before prediction

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in genome-annotation/eukaryotic-gene-prediction of GPTomics/bioSkills.

  • SKILL.md
  • examples/braker3_pipeline.sh
  • examples/galba_protein_only.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Annotation Eukaryotic Gene Prediction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Annotation Eukaryotic Gene Prediction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Annotation Eukaryotic Gene Prediction this skillGPTomics/bioSkills1.2k1 repos~4.3kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Genome Annotation Eukaryotic Gene Prediction

What does Bio Genome Annotation Eukaryotic Gene Prediction do?

Predicts protein-coding gene structures (exons, introns, UTRs) in eukaryotic genomes with BRAKER3 (RNA-seq + protein evidence), BRAKER1/BRAKER2, GALBA (protein-only), Funannotate (fungi), GeMoMa…. Bio Genome Annotation Eukaryotic Gene Prediction is an agent skill from GPTomics/bioSkills. Predicts protein-coding gene structures (exons, introns, UTRs) in eukaryotic genomes with BRAKER3 (RNA-seq + protein evidence), BRAKER1/BRAKER2, GALBA (protein-only), Funannotate (fungi), GeMoMa (homology projection), or Helixer/Tiberius (deep-learning ab initio).

When should I use Bio Genome Annotation Eukaryotic Gene Prediction?

Bio Genome Annotation Eukaryotic Gene Prediction fits situations like: annotating a newly assembled eukaryotic genome; choosing a gene-prediction pipeline based on available evidence; diagnosing a poor annotation.

How do I install Bio Genome Annotation Eukaryotic Gene Prediction in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-eukaryotic-gene-prediction -a claude-code`. Or copy the skill folder (genome-annotation/eukaryotic-gene-prediction in GPTomics/bioSkills) into .claude/skills/bio-genome-annotation-eukaryotic-gene-prediction in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Annotation Eukaryotic Gene Prediction in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-eukaryotic-gene-prediction -a codex`. Or copy the skill folder (genome-annotation/eukaryotic-gene-prediction in GPTomics/bioSkills) into .agents/skills/bio-genome-annotation-eukaryotic-gene-prediction in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Annotation Eukaryotic Gene Prediction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-eukaryotic-gene-prediction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-annotation-eukaryotic-gene-prediction, .gemini/skills/bio-genome-annotation-eukaryotic-gene-prediction, .github/skills/bio-genome-annotation-eukaryotic-gene-prediction and .opencode/skills/bio-genome-annotation-eukaryotic-gene-prediction in your project.

What does Bio Genome Annotation Eukaryotic Gene Prediction need to run?

Going by SKILL.md and its folder, Bio Genome Annotation Eukaryotic Gene Prediction needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Genome Annotation Eukaryotic Gene Prediction access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Genome Annotation Eukaryotic Gene Prediction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Annotation Eukaryotic Gene Prediction use?

Bio Genome Annotation Eukaryotic Gene Prediction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Annotation Eukaryotic Gene Prediction use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Annotation Eukaryotic Gene Prediction?

Skills that share tags, products or a category with Bio Genome Annotation Eukaryotic Gene Prediction: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Annotation Eukaryotic Gene Prediction?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.