Agent skill

Bio Genome Annotation Functional Annotation

by GPTomics in GPTomics/bioSkills

Assigns GO terms, Pfam/InterPro domains, KEGG orthologs, EC numbers, and product names to predicted proteins using eggNOG-mapper (orthology), InterProScan (domain signatures), and KofamScan (KEGG)…

MITAuto-check passedResearch & Science

Install Bio Genome Annotation Functional Annotation

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-annotation-functional-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-annotation-functional-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-annotation/functional-annotation .claude/skills/bio-genome-annotation-functional-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-annotation-functional-annotation
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.2k tokens
SKILL.md length
1,852 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Assigns GO terms, Pfam/InterPro domains, KEGG orthologs, EC numbers, and product names to predicted proteins using eggNOG-mapper (orthology), InterProScan (domain signatures), and KofamScan (KEGG)…

  • Works in 3 steps: The goal is not "maximally annotated" -… → "Domain present" and "function known"… → Record provenance on every label…
  • Adding functional annotation to predicted genes
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Shell and Python scripts from its folder; calls pip

What it does

Bio Genome Annotation Functional Annotation is an agent skill from GPTomics/bioSkills. Assigns GO terms, Pfam/InterPro domains, KEGG orthologs, EC numbers, and product names to predicted proteins using eggNOG-mapper (orthology), InterProScan (domain signatures), and KofamScan (KEGG), routing specialized functions to dbCAN/antiSMASH/AMRFinderPlus/SignalP. Covers the orthology-vs-domain-vs-homology paradigms, the annotation-error percolation cascade, domain-presence-is-not-function, GO IEA circularity in enrichment, evidence tiering, and bit-score/coverage thresholds. Use when adding functional…

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/functional_annotation.sh`, `examples/merge_annotations.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Adding functional annotation to predicted genes
  • Choosing between eggNOG-mapper and InterProScan
  • Judging how much to trust a functional label

Example prompts

  • “Use the bio-genome-annotation-functional-annotation skill to assign GO terms, Pfam/InterPro domains, KEGG orthologs, EC numbers, and product names…”
  • “/bio-genome-annotation-functional-annotation”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The goal is not "maximally annotated" - it is "honestly tiered." A genome that is 40% "hypothetical protein" with the rest correctly…
  2. "Domain present" and "function known" are different claims. A Pfam hit reports architecture, not activity - ~10% of the human kinome are…
  3. Record provenance on every label (method, donor, donor evidence code, identity/coverage/bitscore, DB version). That is the only thing that…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Annotation Functional Annotation loads about 4.2k tokens when it runs. Until then it costs about 171 tokens; SKILL.md has 1,852 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~171
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,852 words, ~4,193 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-annotation-functional-annotation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-genome-annotation-functional-annotation
description
Assigns GO terms, Pfam/InterPro domains, KEGG orthologs, EC numbers, and product names to predicted proteins using eggNOG-mapper (orthology), InterProScan (domain signatures), and KofamScan (KEGG), routing specialized functions to dbCAN/antiSMASH/AMRFinderPlus/SignalP. Covers the orthology-vs-domain-vs-homology paradigms, the annotation-error percolation cascade, domain-presence-is-not-function, GO IEA circularity in enrichment, evidence tiering, and bit-score/coverage thresholds. Use when adding functional annotation to predicted genes, choosing between eggNOG-mapper and InterProScan, or judging how much to trust a functional label.
tool_type
cli
primary_tool
eggNOG-mapper

Version Compatibility

Reference examples tested with: eggNOG-mapper 2.1.15 (pin for reproducibility), InterProScan 5.66+, KofamScan 1.3+, pandas 2.2+, AGAT 1.4+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures

Annotation content tracks database release: record the eggNOG DB version, InterPro/Pfam release, and KEGG/KofamScan profile date, and note whether InterProScan used the EBI precalculated lookup service. eggNOG-mapper v3 is under testing (not production) - pin v2.1.15. If code throws an error, introspect the installed tool and adapt rather than retrying.

Functional Annotation

"Functionally annotate my predicted proteins" -> Transfer GO/KEGG/Pfam/EC/product labels from characterized proteins by orthology and domain signatures, attaching a confidence tier and provenance to each.

  • CLI: emapper.py -i proteins.faa --itype proteins -m diamond (eggNOG-mapper), interproscan.sh -i proteins.faa -f TSV,GFF3 -goterms -pa (InterProScan)

The Single Most Important Modern Insight -- Annotation Is a Propagated Hypothesis, Not a Measurement

Almost every label on a new genome is transferred by homology/orthology/ML from a small island of experimentally characterized proteins. The transfer chain is lossy and self-reinforcing - it behaves like a percolation cascade (Gilks 2002 Bioinformatics 18:1641): an over-specific name assigned in year 0, deposited with no record that it was transferred, becomes the nearest hit for the next genome, whose label becomes evidence for the next. By the time a query reaches NR, "number of hits agreeing" measures how far an error spread, not correctness. Schnoes 2009 (PLoS Comput Biol 5:e1000605) found misannotation reaching ~80% in bulk databases (TrEMBL/NR) and near-zero in curated Swiss-Prot - the gap is the curation. Three load-bearing consequences:

  1. The goal is not "maximally annotated" - it is "honestly tiered." A genome that is 40% "hypothetical protein" with the rest correctly tiered by evidence is a better scientific object than one 95% named with half the names wrong. Prefer curated/orthology donors (Swiss-Prot, eggNOG OG consensus) over best-hits, and demote specificity as identity/coverage fall (full EC -> partial 1.1.1.-; specific name -> superfamily; whole-protein -> per-domain). PI/reviewer pressure to "annotate everything" manufactures the next genome's percolating error.
  2. "Domain present" and "function known" are different claims. A Pfam hit reports architecture, not activity - ~10% of the human kinome are catalytically dead pseudokinases that carry a confident "protein kinase" domain. Moonlighting (GAPDH), promiscuity, and mechanistically-diverse superfamilies (enolase, amidohydrolase, HAD, TIM-barrel: shared fold, divergent substrate) make this a first-order effect. Fold conservation != function conservation any more than sequence does - so structure-based transfer (Foldseek) inherits the same trap with higher false confidence.
  3. Record provenance on every label (method, donor, donor evidence code, identity/coverage/bitscore, DB version). That is the only thing that stops the provenance-amnesia step that turns a transfer into a "fact."

Tool Taxonomy

ParadigmToolMechanismFailure mode
OrthologyeggNOG-mapperseed-ortholog -> orthologous group -> consensus transfertax-scope sensitive; HGT/xenologs break the orthology assumption
Domain/signatureInterProScanprofile HMMs/matrices -> integrated InterPro entriesa domain implies a capability, not the substrate; broad families uninformative
KEGG orthologKofamScanper-KO HMMs + adaptive thresholdsKO assignment, not pathway proof
Homology best-hitDIAMOND vs Swiss-Prottop-hit similarity, transfer labelbest-hit != ortholog; transitive error propagation
ML / structureDeepGO, DeepFRI, Foldseeklearned sequence/structure -> GOlow precision; ontology terms not products; reaches twilight zone only

Default workhorse pair: eggNOG-mapper + InterProScan (orthogonal evidence: orthology vs signatures), reconciled afterward. Add KofamScan if KEGG pathway reconstruction is the goal (its adaptive per-KO thresholds are stricter than eggNOG's KEGG_ko). DIAMOND-vs-Swiss-Prot is the cheap product-name layer; never use it alone for GO.

Decision Tree by Scenario

ScenarioRecommendedWhy
Bacterial isolateBakta/PGAP product names + eggNOG-mapper + InterProScanstructural pipeline first, then orthology + domains
Eukaryotic proteomeInterProScan (domains+GO+pathways) + eggNOG-mapperorthogonal evidence, reconcile
Metagenome / MAGeggNOG-mapper --itype metagenome (+ KofamScan, dbCAN)built-in gene calling; KEGG modules
Twilight-zone / ORFan (no homolog)ML (DeepGOPlus) or structure (ESMFold -> Foldseek -> DeepFRI)only handle on the homology-free fraction; low-confidence leads
CAZymes / BGCs / AMR / signal peptides-> dbCAN / antiSMASH / AMRFinderPlus / SignalP6a generic Pfam hit gives no substrate/phenotype/cluster
GO enrichment downstream-> pathway-analysis/go-enrichment (mind IEA circularity)enrichment on IEA partly tests the pipeline against itself

eggNOG-mapper

bash
download_eggnog_data.py --data_dir db/ -y               # ~44 GB (DIAMOND DB installed by default; -D skips it)
emapper.py -i proteins.faa --itype proteins -m diamond \
    --tax_scope auto --data_dir db/ --cpu 16 -o annot --output_dir out/

Three stages: (1) seed-ortholog search (DIAMOND/MMseqs2/HMMER) anchors the query - this is a best-hit and is not the annotation; (2) orthology assignment retrieves the seed's fine-grained orthologs within the chosen taxonomic scope; (3) functional transfer pools terms across the set of orthologs (which damps single-entry misannotation - this is why eggNOG-mapper beats raw DIAMOND-vs-NR). --tax_scope is the single most consequential parameter: too broad gathers distant orthologs and over-generalizes function; auto lets each seed take its most-informative phylogenetic ceiling. --itype {proteins,CDS,genome,metagenome} (genome/metagenome runs Prodigal first). Output .emapper.annotations columns include seed_ortholog, eggNOG_OGs, COG_category, Description, Preferred_name, GOs, EC, KEGG_ko, PFAMs (read the actual header; - = empty).

InterProScan

bash
interproscan.sh -i proteins.faa -f TSV,GFF3 -goterms -pa -cpu 16

Runs member-database scanners (Pfam, PANTHER, NCBIfam, SUPERFAMILY, CDD, SMART, Gene3D, Hamap, PROSITE, ...) and integrates overlapping signatures into InterPro entries (stable IPRxxxxxx, with a type: Family/Domain/Repeat/Site/Homologous Superfamily). Report at the InterPro-entry level - it is the consensus that survives one member DB being wrong. -goterms adds the interpro2go mapping (these GO are IEA/electronic); -pa maps Reactome/MetaCyc. By default it queries the EBI precalculated lookup service (fast, MD5-keyed); -dp forces local compute (novel/confidential sequences, reproducibility). Java 11+ and a tens-of-GB data bundle required; for millions of proteins, chunk the FASTA into array jobs.

Reconciling Multi-Tool Output with Python

Goal: Merge eggNOG and InterProScan per protein while preserving provenance, so a curated name is never silently overwritten by a generic domain.

Approach: Parse each tool's table, keep source namespaces separate, union GO with source tags, and prefer the orthology Preferred_name/Description for the human-readable product.

python
import pandas as pd

def parse_eggnog(path):
    df = pd.read_csv(path, sep='\t', comment='#', header=None)
    cols = ['query', 'seed_ortholog', 'evalue', 'score', 'eggNOG_OGs', 'max_annot_lvl',
            'COG_category', 'Description', 'Preferred_name', 'GOs', 'EC', 'KEGG_ko']
    df.columns = (cols + [f'c{i}' for i in range(len(df.columns) - len(cols))])[:len(df.columns)]
    return df

def best_product_name(row):
    name = row.get('Preferred_name', '-')
    return name if name not in ('-', '', None) else 'hypothetical protein'   # honest default, not a forced guess

Use AGAT (agat_sp_manage_functional_annotation.pl) to graft BLAST/InterProScan results onto a GFF3 (it handles the spec edge cases). For GO deliverables use GAF (carries the evidence code); keep each tool in its own Dbxref namespace.

Ontology Rigor and the IEA Circularity

  • GO MF vs BP transfer with different reliability. Molecular Function ("DNA helicase activity") is local/chemical and transfers with the fold; Biological Process ("DNA replication") is systemic context and does not transfer reliably - CAFA confirmed BLAST beats naive baselines for MF but not BP (Radivojac 2013 Nat Methods 10:221). Weight MF over BP when evidence is limited; never let a transferred BP term drive a conclusion alone.
  • Essentially all genome-derived GO is IEA (Inferred from Electronic Annotation, never curator-reviewed). Running GO enrichment on IEA against an IEA background partly tests the pipeline against itself - if interpro2go maps a common domain to a term, every genome with that domain looks "enriched." Compounded by annotation bias (58% of human GO covers 16% of genes; Haynes 2018 Sci Rep 8:1362) and True-Path-Rule inflation of shallow terms. Pin versions, match background to foreground pipeline, prefer non-IEA where it exists, and state that the result is annotation-derived.
  • EC numbers are not stable. Deleted/transferred numbers are tombstones (never reused); a partial EC 1.1.1.- is a valid statement of ignorance (the EC equivalent of "hypothetical"). Demand orthology or a curated rule before asserting a full four-level EC.
  • KEGG bulk access is paywalled. Free routes: KofamScan/KofamKOALA (local HMMs + adaptive per-KO thresholds; an * marks above-threshold hits) or eggNOG's KEGG_ko. A "complete module" is a reconstruction (a gap can be non-orthologous gene displacement; a filled step can be a paralog doing something else), not proof of flux.
Show full SKILL.md (676 more words)Show less

Per-Method Failure Modes

Best-hit-as-ortholog

Trigger: transferring a specific function from one DIAMOND/BLAST top hit (esp. TrEMBL/NR). Mechanism: best-hit != ortholog; the bulk-DB hit is likely itself an auto-annotation. Symptom: confident specific names with no provenance. Fix: orthology consensus + Swiss-Prot donors.

Over-specific transfer

Trigger: copying the exact substrate/EC of a characterized homolog onto a distant relative. Mechanism: mechanistically-diverse superfamilies share fold, not substrate. Symptom: a "muconate cycloisomerase" that does something else. Fix: demote to superfamily / partial EC as identity and coverage fall.

Wrong eggNOG tax_scope

Trigger: leaving scope too broad/narrow or unpinned. Mechanism: distant orthologs over-generalize, or no informative orthologs. Symptom: vague or missing function. Fix: auto, or pin the known clade.

Circular GO enrichment

Trigger: enriching IEA annotations against a mismatched background. Mechanism: measures the mapping table and study popularity, not biology. Symptom: "enriched" for whatever well-studied genes are annotated for. Fix: non-IEA where possible; matched background; pin versions; caveat the result.

Reading specialized function from a generic hit

Trigger: inferring CAZyme substrate / AMR phenotype / BGC product from a plain Pfam domain. Mechanism: the substrate/phenotype/cluster signal is not in a generic domain. Symptom: wrong substrate or phenotype call. Fix: route to dbCAN / AMRFinderPlus / antiSMASH.

Quantitative Thresholds

ThresholdSourceRationale
Reason in bits-per-residue, not raw e-valuealignment statisticse-value scales with DB size (a database-size artifact); bits/residue is density
Bidirectional coverage ≥50-70% query and subjecttransfer practiceone-domain coverage justifies only a domain-level claim
~40% identity over full length (well-behaved families only)soft floorno safe identity in mechanistically-diverse superfamilies; demote specificity instead
Named fraction "too high for the taxon" (>90% on a novel isolate)over-annotation smell testloose thresholds manufacturing names; expect 20-50% hypothetical
eggNOG --tax_scope autoeggNOG-mapperper-seed informative ceiling
KofamScan adaptive per-KO threshold (*)Aramaki 2020a single global e-value misfires across KO families

Common Errors

Error / symptomCauseSolution
Low annotation ratefragmented ORFs / narrow scopecheck protein quality; --tax_scope auto; run both tools and merge
Specific name on a distant homologover-specific transferdemote to superfamily / partial EC; record identity
eggNOG DB errorsDB/version mismatchre-download; pin emapper 2.1.15
InterProScan memory/timefull proteome at oncechunk FASTA; keep lookup service on; drop PANTHER/Gene3D if not needed
Enrichment "too clean"IEA circularity / study biasmatched background; pin GO release; caveat
Multidomain protein mislabelednamed by first/best domainreport all domains with coordinates

References

  • Cantalapiedra CP, et al. 2021. eggNOG-mapper v2: functional annotation, orthology assignments, and domain prediction at the metagenomic scale. Mol Biol Evol 38:5825-5829.
  • Huerta-Cepas J, et al. 2019. eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource. Nucleic Acids Res 47:D309-D314.
  • Jones P, et al. 2014. InterProScan 5: genome-scale protein function classification. Bioinformatics 30:1236-1240.
  • Blum M, et al. 2025. InterPro: the protein sequence classification resource in 2025. Nucleic Acids Res 53:D444-D456.
  • Schnoes AM, et al. 2009. Annotation error in public databases: misannotation of molecular function in enzyme superfamilies. PLoS Comput Biol 5:e1000605.
  • Gilks WR, et al. 2002. Modeling the percolation of annotation errors in a database of protein sequences. Bioinformatics 18:1641-1649.
  • Aramaki T, et al. 2020. KofamKOALA: KEGG ortholog assignment based on profile HMM and adaptive score threshold. Bioinformatics 36:2251-2252.
  • Zheng J, et al. 2023. dbCAN3: automated carbohydrate-active enzyme and substrate annotation. Nucleic Acids Res 51:W115-W121.
  • Teufel F, et al. 2022. SignalP 6.0 predicts all five types of signal peptides using protein language models. Nat Biotechnol 40:1023-1025.
  • Feldgarden M, et al. 2021. AMRFinderPlus and the Reference Gene Catalog facilitate examination of the genomic links among antimicrobial resistance, stress response, and virulence. Sci Rep 11:12728.
  • Blin K, et al. 2023. antiSMASH 7.0: new and improved predictions for detection, regulation, chemical structures and visualisation. Nucleic Acids Res 51:W46-W50.
  • Radivojac P, et al. 2013. A large-scale evaluation of computational protein function prediction (CAFA). Nat Methods 10:221-227.
  • Haynes WA, et al. 2018. Gene annotation bias impedes biomedical research. Sci Rep 8:1362.
  • prokaryotic-annotation - Bakta/PGAP product names + locus tags before functional layers
  • eukaryotic-gene-prediction - Produces the protein FASTA to annotate
  • annotation-qc - Annotation-coverage and hypothetical-fraction sanity
  • pathway-analysis/go-enrichment - Enrichment using GO annotations (mind IEA circularity)
  • pathway-analysis/kegg-pathways - Pathway mapping with KEGG orthologs
  • epidemiological-genomics/amr-surveillance - AMRFinderPlus/CARD for resistance genes and point mutations

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in genome-annotation/functional-annotation of GPTomics/bioSkills.

  • SKILL.md
  • examples/functional_annotation.sh
  • examples/merge_annotations.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Annotation Functional Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Annotation Functional Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Annotation Functional Annotation this skillGPTomics/bioSkills1.2k1 repos~4.2kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Genome Annotation Functional Annotation

What does Bio Genome Annotation Functional Annotation do?

Assigns GO terms, Pfam/InterPro domains, KEGG orthologs, EC numbers, and product names to predicted proteins using eggNOG-mapper (orthology), InterProScan (domain signatures), and KofamScan (KEGG)…. Bio Genome Annotation Functional Annotation is an agent skill from GPTomics/bioSkills. Assigns GO terms, Pfam/InterPro domains, KEGG orthologs, EC numbers, and product names to predicted proteins using eggNOG-mapper (orthology), InterProScan (domain signatures), and KofamScan (KEGG), routing specialized functions to dbCAN/antiSMASH/AMRFinderPlus/SignalP.

When should I use Bio Genome Annotation Functional Annotation?

Bio Genome Annotation Functional Annotation fits situations like: adding functional annotation to predicted genes; choosing between eggNOG-mapper and InterProScan; judging how much to trust a functional label.

How do I install Bio Genome Annotation Functional Annotation in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-functional-annotation -a claude-code`. Or copy the skill folder (genome-annotation/functional-annotation in GPTomics/bioSkills) into .claude/skills/bio-genome-annotation-functional-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Annotation Functional Annotation in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-functional-annotation -a codex`. Or copy the skill folder (genome-annotation/functional-annotation in GPTomics/bioSkills) into .agents/skills/bio-genome-annotation-functional-annotation in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Annotation Functional Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-annotation-functional-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-annotation-functional-annotation, .gemini/skills/bio-genome-annotation-functional-annotation, .github/skills/bio-genome-annotation-functional-annotation and .opencode/skills/bio-genome-annotation-functional-annotation in your project.

What does Bio Genome Annotation Functional Annotation need to run?

Going by SKILL.md and its folder, Bio Genome Annotation Functional Annotation needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Genome Annotation Functional Annotation access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Genome Annotation Functional Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Annotation Functional Annotation use?

Bio Genome Annotation Functional Annotation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Annotation Functional Annotation use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Annotation Functional Annotation?

Skills that share tags, products or a category with Bio Genome Annotation Functional Annotation: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Annotation Functional Annotation?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.