Agent skill

Bio Pathway Go Enrichment

by GPTomics in GPTomics/bioSkills

Runs Gene Ontology over-representation analysis (ORA) on a gene LIST with clusterProfiler enrichGO, the one-sided hypergeometric/Fisher 2x2 test phyper(k-1, M, N-M, n, lower.tail=FALSE).

MITAuto-check passedResearch & Science

Install Bio Pathway Go Enrichment

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-pathway-go-enrichment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-pathway-go-enrichment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pathway-analysis/go-enrichment .claude/skills/bio-pathway-go-enrichment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-pathway-go-enrichment
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.1k tokens
SKILL.md length
2,266 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Runs Gene Ontology over-representation analysis (ORA) on a gene LIST with clusterProfiler enrichGO, the one-sided hypergeometric/Fisher 2x2 test phyper(k-1, M, N-M, n, lower.tail=FALSE).

  • Works in 3 steps: The universe is the null, not a setting.… → ORA throws away magnitude and inherits… → The deliverable is never "the enriched…
  • A pre-selected gene list (DE hits
  • SKILL.md covers Version Compatibility, The Single Most Important…, ORA vs GSEA (the central fork) and Tool Taxonomy, plus 12 more sections
  • Runs R scripts from its folder

What it does

Bio Pathway Go Enrichment is an agent skill from GPTomics/bioSkills. Runs Gene Ontology over-representation analysis (ORA) on a gene LIST with clusterProfiler enrichGO, the one-sided hypergeometric/Fisher 2x2 test phyper(k-1, M, N-M, n, lower.tail=FALSE). Covers why the BACKGROUND universe (not the gene list) is the null and decides significance, why omitting universe= is a bug, why enrichGO defaults to ont='MF' not 'BP', why pvalueCutoff filters p.adjust not raw p, why ORA discards effect magnitude and inherits GO-DAG true-path redundancy (simplify, topGO), why RNA-seq…

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • A pre-selected gene list (DE hits
  • Co-expression module
  • GWAS-mapped) needs GO annotation

Example prompts

  • “/bio-pathway-go-enrichment”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The universe is the null, not a setting. Change N or M and every p-value changes. Omitting universe= defaults N to ALL annotated genes…
  2. ORA throws away magnitude and inherits the GO DAG. A gene is in or out at one threshold (no effect size), so a 2000-gene term at 1.2x fold…
  3. The deliverable is never "the enriched pathways." It is a correctly-backgrounded, effect-sized, redundancy-resolved short list of…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Pathway Go Enrichment loads about 5.1k tokens when it runs. Until then it costs about 243 tokens; SKILL.md has 2,266 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~243
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,266 words, ~5,137 tokens.

Download SKILL.mdSave it as .claude/skills/bio-pathway-go-enrichment/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-pathway-go-enrichment
description
Runs Gene Ontology over-representation analysis (ORA) on a gene LIST with clusterProfiler enrichGO, the one-sided hypergeometric/Fisher 2x2 test phyper(k-1, M, N-M, n, lower.tail=FALSE). Covers why the BACKGROUND universe (not the gene list) is the null and decides significance, why omitting universe= is a bug, why enrichGO defaults to ont='MF' not 'BP', why pvalueCutoff filters p.adjust not raw p, why ORA discards effect magnitude and inherits GO-DAG true-path redundancy (simplify, topGO), why RNA-seq gene-length bias inflates long-gene terms (GOseq Wallenius), plus GeneRatio/BgRatio, bitr ID mapping, minGSSize/maxGSSize, groupGO. Use when a pre-selected gene list (DE hits, co-expression module, screen, GWAS-mapped) needs GO annotation. For a ranked no-cutoff analysis see gsea; for other databases see kegg-pathways, reactome-pathways, wikipathways; DE source is differential-expression/de-results; plots in enrichment-visualization.
tool_type
r
primary_tool
clusterProfiler

Version Compatibility

Reference examples tested with: clusterProfiler 4.18.4+, org.Hs.eg.db 3.22+ (goseq 1.54+ for the length-bias snippet).

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

GO annotation lives in the local org.*.eg.db OrgDb and GO.db, both pinned to the Bioconductor release, so a GO ORA is reproducible given the package versions - record packageVersion('org.Hs.eg.db') and packageVersion('GO.db') with results. enrichGO moved no core arguments recently, but several plot helpers migrated to enrichplot in clusterProfiler 4.x (those live in enrichment-visualization).

GO Over-Representation Analysis

"Which biological processes are enriched in my gene list?" -> Test each GO term for over-representation of the query genes against a defined background with the one-sided hypergeometric test - because the BACKGROUND universe, not the gene list, is what decides which terms look significant.

  • R: enrichGO(gene, universe, OrgDb, keyType='ENTREZID', ont='BP')

Scope: hypergeometric ORA of a gene LIST against GO terms, with background-universe selection, ID conversion, GO-DAG redundancy reduction, RNA-seq length-bias correction, and the generic enricher test for custom gene sets. A ranked-list / no-cutoff analysis -> gsea. KEGG/Reactome/WikiPathways gene sets -> kegg-pathways, reactome-pathways, wikipathways. The DE list source -> differential-expression/de-results. Plots -> enrichment-visualization.

The Single Most Important Modern Insight -- ORA Is a Competitive 2x2 Hypergeometric Test Whose Null IS the Chosen Universe

ORA does not answer "which pathways are in my gene list". It is a competitive gene-sampling test (Goeman & Buhlmann 2007 Bioinformatics 23:980): of the genes flagged (the foreground), are more annotated to term T than expected when drawing the same number at random from the universe? The p-value is the upper tail of the hypergeometric, computed verbatim by DOSE/clusterProfiler as phyper(k-1, M, N-M, n, lower.tail=FALSE) = P(X>=k), the one-sided Fisher exact test on the 2x2 table. Here N = universe genes carrying any GO annotation, M = universe genes in T, n = foreground genes annotated, k = the overlap (the Count column). The report columns are GeneRatio = k/n and BgRatio = M/N - both denominators restricted to ANNOTATED genes - and fold enrichment = GeneRatio/BgRatio.

Three consequences drive every misuse:

  1. The universe is the null, not a setting. Change N or M and every p-value changes. Omitting universe= defaults N to ALL annotated genes (~18k for human BP); if the assay only measured ~12k genes, terms for tissue-restricted and lowly-expressed genes go spuriously significant. Omitting universe= is a bug, not a default - set it to the genes that COULD have entered the foreground (the tested-gene set), map foreground and universe identically, and report N. The whole-genome background is defensible only when every gene truly could have been detected (Wijesooriya 2022; Timmons 2015).
  2. ORA throws away magnitude and inherits the GO DAG. A gene is in or out at one threshold (no effect size), so a 2000-gene term at 1.2x fold can beat a 12-gene term at 4x on p-value alone - always read fold enrichment alongside p.adjust. The true-path rule propagates each annotation to all ancestors, so one real signal lights up a whole lineage ("cell cycle", "cell cycle process", "mitotic cell cycle" together); those tests are positively correlated, BH still valid but the term list over-reports. Resolve with simplify() (semantic collapse, per ontology) or topGO elim/weight (decorrelation in the test).
  3. The deliverable is never "the enriched pathways." It is a correctly-backgrounded, effect-sized, redundancy-resolved short list of HYPOTHESES - and it cannot be used to validate the DE list that produced it (circular: the terms are a deterministic function of the same genes; Timmons 2015).

ORA vs GSEA (the central fork)

ORA needs a pre-selected LIST plus a BACKGROUND and binarizes significant/not; GSEA needs a RANKED vector of ALL genes and no cutoff. Pick by whether a ranking exists and whether the cutoff would be arbitrary. The full three-generations taxonomy (ORA vs FCS vs topology) and competitive-vs-self-contained null theory live in the category README - this skill owns the ORA/GO slice.

ScenarioMethodWhy
All genes carry a DE statistic, cutoff would be arbitraryGSEA (gseGO) -> gseauses the full ranking; no threshold
Pre-selected list (co-expression module, GWAS-mapped, screen hits, markers)ORA (enrichGO)no ranking available; ORA is appropriate
Very small list (< ~15-20 genes)low ORA power; report fold enrichment + counts, consider GSEAhypergeometric power collapses on tiny lists
RNA-seq DE list with length/selection biasGOseq (Wallenius)length-corrected ORA; standard ORA inflates long-gene terms

Tool Taxonomy

Source / methodCitationMechanism / roleWhen
enrichGO (clusterProfiler)Yu 2012 OMICS 16:284; Wu 2021 Innovation 2:100141one-sided hypergeometric per GO term; local OrgDbthe default ORA workhorse for a gene list
GO DAG (BP/MF/CC)Ashburner 2000 Nat Genet 25:25three DAGs; true-path propagation to ancestorsthe annotation structure being tested
simplify (GOSemSim)Wang 2007 Bioinformatics 23:1274semantic-similarity de-redundancy, per ontologycollapse redundant ancestor lineages, keep calibrated p/FDR
topGO elim/weight/weight01Alexa 2006 Bioinformatics 22:1600decorrelates the GO graph inside the testspecificity-resolved short list (treat scores as ranking, not FDR)
GOseqYoung 2010 Genome Biol 11:R14Wallenius noncentral hypergeometric weighted by a length PWFRNA-seq DE with gene-length/selection bias
enricher (clusterProfiler)Yu 2012 OMICS 16:284same hypergeometric engine on a custom TERM2GENEany gene set (MSigDB, in-house) not in a DB function
gseGO / GSEA(route -> gsea)rank-based running-sum, permutation nulla ranking exists; no arbitrary cutoff

Run the GO ORA

Goal: Find GO terms over-represented in a gene list relative to the genes that could have been selected.

Approach: Build the foreground and the universe with the SAME ID mapping, set ont explicitly (the source default is 'MF'), pass universe= (omitting it is a bug), and read fold enrichment alongside p.adjust.

r
library(clusterProfiler)
library(org.Hs.eg.db)

ego <- enrichGO(gene          = gene_list,        # foreground ENTREZ IDs
                universe      = universe_ids,     # tested-gene set, mapped identically -- NOT the genome
                OrgDb         = org.Hs.eg.db,
                keyType       = 'ENTREZID',
                ont           = 'BP',             # SET explicitly: source default is 'MF', not 'BP'
                pAdjustMethod = 'BH',
                pvalueCutoff  = 0.05,             # filters p.adjust (despite the name), not raw pvalue
                qvalueCutoff  = 0.2,
                minGSSize     = 10,
                maxGSSize     = 500,
                readable      = TRUE)             # map ENTREZ -> SYMBOL in the output

The returned enrichResult has columns ID, Description, GeneRatio, BgRatio, pvalue, p.adjust, qvalue, geneID, Count (plus ONTOLOGY when ont='ALL'). pvalueCutoff filters the ADJUSTED p, so an empty table usually means the cutoff or the universe, not biology - inspect everything with pvalueCutoff=1, qvalueCutoff=1.

Build the Foreground and Universe from DE Results

Goal: Turn a DE table into the foreground gene vector and the matched background universe.

Approach: Filter the DE table to the hits for the foreground; take the genes that were actually TESTED for the universe (DESeq2: rows with non-NA pvalue survive independent filtering); map both with the same bitr call. The DE mechanics and the $padj/$adj.P.Val column choice live in differential-expression/de-results.

r
de <- read.csv('de_results.csv')

sig_genes  <- de$gene_id[de$padj < 0.05 & abs(de$log2FoldChange) > 1]   # foreground = hits
all_tested <- de$gene_id[!is.na(de$pvalue)]                            # universe = tested genes, NOT all rows, NOT the genome

fg_map <- bitr(sig_genes,  fromType = 'SYMBOL', toType = 'ENTREZID', OrgDb = org.Hs.eg.db)
bg_map <- bitr(all_tested, fromType = 'SYMBOL', toType = 'ENTREZID', OrgDb = org.Hs.eg.db)

gene_list    <- unique(fg_map$ENTREZID)   # deduplicate one-to-many maps before counting
universe_ids <- unique(bg_map$ENTREZID)

bitr one-to-many maps produce duplicate rows that inflate Count; deduplicate. If more than ~15% of genes fail to convert the result is unreliable - report the conversion rate. Mixed up- and down-regulated genes cancel in one list: run ORA separately per direction when direction matters.

Reduce GO-DAG Redundancy with simplify

Goal: Collapse the redundant ancestor lineage so one biological signal is one entry, not a dozen.

Approach: simplify() removes terms whose semantic similarity to a kept term exceeds the cutoff. It operates on ONE ontology (GOSemSim defines similarity within a single DAG), so run BP/MF/CC separately and simplify each - it does NOT de-redundify an ont='ALL' object.

r
ego_bp <- enrichGO(gene_list, universe = universe_ids, OrgDb = org.Hs.eg.db, keyType = 'ENTREZID', ont = 'BP', readable = TRUE)
ego_bp <- simplify(ego_bp, cutoff = 0.7, by = 'p.adjust', select_fun = min, measure = 'Wang')

measure='Wang' (the default) is graph-topology-based and stable across annotation releases; IC-based measures ('Resnik', 'Lin', 'Jiang', 'Rel') shift with the annotation corpus. topGO elim/weight01 is the alternative that decorrelates inside the test, returning a specificity-resolved list directly - but its conditioned p-values are best treated as a ranking, not calibrated FDR (Alexa 2006).

Correct RNA-seq Length Bias with GOseq

Goal: Stop long, highly-expressed genes from looking enriched for a purely technical reason.

Approach: DE-detection power scales with read count, which scales with transcript length and expression, so the foreground is enriched for long genes - and RPKM/TMM normalization does NOT fix it (it corrects abundance, not detection power). GOseq fits a probability weighting function (PWF) over the bias variable and tests with the Wallenius noncentral hypergeometric (Young 2010). The input is a NAMED 0/1 vector over ALL tested genes; goseq returns UNADJUSTED p-values, so apply BH afterward.

r
library(goseq)

all_genes <- de$gene_id[!is.na(de$pvalue)]
de_genes  <- as.integer(all_genes %in% sig_genes)   # named binary vector over the tested set
names(de_genes) <- all_genes

pwf <- nullp(de_genes, 'hg38', 'ensGene')           # fits the length PWF; inspect the fit plot
go  <- goseq(pwf, 'hg38', 'ensGene', method = 'Wallenius')   # default; 'Hypergeometric' ignores bias (= standard ORA)
go$padj <- p.adjust(go$over_represented_pvalue, method = 'BH')   # goseq does NOT BH-correct internally

GSEA on a length-neutral ranking statistic (the moderated t / Wald z) is largely immune to this bias - one more reason to consider gsea for RNA-seq.

All Three Ontologies and a Descriptive Breakdown

ont='ALL' runs BP/MF/CC separately and rbinds them with an ONTOLOGY column (pool=FALSE default; pool=TRUE treats the three as one set). groupGO is NOT a test - it classifies genes at a fixed DAG level for a GO-slim overview (counts, no p-values); never read its counts as significance.

r
ego_all <- enrichGO(gene_list, universe = universe_ids, OrgDb = org.Hs.eg.db, keyType = 'ENTREZID', ont = 'ALL', readable = TRUE)
ggo     <- groupGO(gene_list, OrgDb = org.Hs.eg.db, keyType = 'ENTREZID', ont = 'BP', level = 3, readable = TRUE)
Show full SKILL.md (902 more words)Show less

Custom Gene Sets with enricher

For gene sets not covered by a DB function (MSigDB collections, in-house sets), enricher runs the SAME hypergeometric engine against a two-column TERM2GENE table; pass the same explicit universe.

r
ego_custom <- enricher(gene_list, TERM2GENE = t2g, universe = universe_ids,
                       pvalueCutoff = 0.05, pAdjustMethod = 'BH', minGSSize = 10, maxGSSize = 500, qvalueCutoff = 0.2)

Other Organisms

Swap the OrgDb: org.Mm.eg.db (mouse), org.Dr.eg.db (zebrafish), org.Sc.sgd.db (yeast, keyType='ORF'). Check usable key types with keytypes(OrgDb).

Per-Method Failure Modes

Whole-genome or default universe

Trigger: omitting universe=, or passing the genome when the assay measured fewer genes. Mechanism: N defaults to all annotated genes, inflating the denominator with genes that never could have been selected. Symptom: a confident table where tissue-restricted / lowly-expressed-gene terms dominate. Fix: set universe= to the tested-gene set, map foreground and universe identically, report N.

p read without fold enrichment (term-size trap)

Trigger: ranking results by p.adjust alone. Mechanism: a 2000-gene term has enormous power at tiny fold enrichment; p scales with term size. Symptom: vague broad terms ("cellular process") top the list, specific terms buried. Fix: read fold enrichment = (k/n)/(M/N) alongside p.adjust; trim extremes with minGSSize=10, maxGSSize=500.

Redundant ancestor lineage counted as findings

Trigger: reporting "cell cycle", "cell cycle process", "mitotic cell cycle" as separate discoveries. Mechanism: true-path propagation lights up a whole lineage from one signal; the tests are positively correlated. Symptom: the top 20 is one biological theme repeated. Fix: simplify() per ontology, or topGO weight01; never count lineage members as independent hits.

RNA-seq length / selection bias

Trigger: standard ORA on an RNA-seq DE list without length correction. Mechanism: detection power scales with count ~ length/expression; TMM/RPKM fixes abundance, not power. Symptom: long-gene categories (ECM, adhesion) enriched, short-gene (ribosomal) depleted - and it survives FDR. Fix: GOseq with a length PWF + method='Wallenius', then BH; or GSEA on a bias-neutral statistic.

Wrong ID type or silent gene loss

Trigger: passing ENSEMBL/SYMBOL with a mismatched keyType, or not checking the bitr conversion rate. Mechanism: unmapped IDs are dropped, shrinking the foreground; one-to-many maps inflate Count. Symptom: "no gene can be mapped", or a suspiciously small/large Count. Fix: match keyType to one of keytypes(OrgDb), deduplicate after bitr, report conversion rate (flag >15% loss).

pvalueCutoff misread as raw-p filter

Trigger: concluding "no significant terms" when strong raw p exists. Mechanism: pvalueCutoff filters p.adjust, not pvalue. Symptom: an empty table despite plausible signal. Fix: inspect with pvalueCutoff=1, qvalueCutoff=1, then judge on p.adjust.

simplify on ont='ALL'

Trigger: calling simplify() on an ont='ALL' object. Mechanism: semantic similarity is defined within ONE ontology, not across BP/MF/CC. Symptom: redundancy not removed, or an error. Fix: run BP/MF/CC separately and simplify each.

Quantitative Thresholds

ThresholdSourceRationale
pvalueCutoff = 0.05clusterProfiler defaultfilters on p.adjust (NOT raw pvalue); standard FDR gate
qvalueCutoff = 0.2clusterProfiler defaultsecondary q-value gate; loosen to 1 to inspect all terms
pAdjustMethod = 'BH'Benjamini-Hochbergcontrols FDR; valid under the positive dependence of true-path-correlated terms (Bonferroni is needlessly strict here)
minGSSize = 10enrichGO defaultdrop tiny sets that overfit and are noisy
maxGSSize = 500enrichGO defaultdrop huge general sets that always "enrich" with trivial fold
simplify(cutoff = 0.7)GOSemSim/Wangsemantic-similarity redundancy cutoff; lower keeps more terms, higher is more aggressive
fold enrichment > 2heuristic(k/n)/(M/N); a rough "strong" flag, never a substitute for p.adjust
ID-conversion loss > 15%heuristicabove this the foreground is too eroded to trust; report the rate

Common Errors

Error / symptomCauseSolution
--> No gene can be mappedwrong keyType / OrgDb, or IDs not in the OrgDbmatch keyType to keytypes(OrgDb); bitr to ENTREZID first
Empty result tablepvalueCutoff filters p.adjust; or universe too large; or IDs lostset cutoffs to 1 to inspect; fix the universe; check conversion rate
Vague broad terms dominateranking by p alone (term-size trap)read fold enrichment; trim with minGSSize/maxGSSize
Many redundant ancestor termsGO-DAG true-path propagationsimplify() per ontology, or topGO weight01
simplify does nothing / errors on ALLsimilarity is per-ontologyrun BP/MF/CC separately
Description column shows IDs not namesnot readablereadable=TRUE or setReadable(ego, OrgDb, 'ENTREZID')
Tested MF when expecting BPenrichGO default ont='MF'set ont explicitly every call

References

  • Ashburner M, Ball CA, Blake JA, et al. 2000. Gene Ontology: tool for the unification of biology. Nat Genet 25:25-29.
  • Yu G, Wang LG, Han Y, He QY. 2012. clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS 16:284-287.
  • Wu T, Hu E, Xu S, et al. 2021. clusterProfiler 4.0: a universal enrichment tool for interpreting omics data. The Innovation 2(3):100141.
  • Goeman JJ, Buhlmann P. 2007. Analyzing gene expression data in terms of gene sets: methodological issues. Bioinformatics 23:980-987.
  • Alexa A, Rahnenfuhrer J, Lengauer T. 2006. Improved scoring of functional groups from gene expression data by decorrelating GO graph structure. Bioinformatics 22(13):1600-1607.
  • Young MD, Wakefield MJ, Smyth GK, Oshlack A. 2010. Gene ontology analysis for RNA-seq: accounting for selection bias. Genome Biol 11(2):R14.
  • Wang JZ, Du Z, Payattakool R, et al. 2007. A new method to measure the semantic similarity of GO terms. Bioinformatics 23(10):1274-1281.
  • Timmons JA, Szkop KJ, Gallagher IJ. 2015. Multiple sources of bias confound functional enrichment analysis of global -omics data. Genome Biol 16:186.
  • Wijesooriya K, Jadaan SA, Perera KL, et al. 2022. Urgent need for consistent standards in functional enrichment analysis. PLoS Comput Biol 18(3):e1009935.
  • gsea - Ranked-list GSEA alternative when a full ranking exists and a cutoff is arbitrary
  • kegg-pathways - KEGG pathway and module enrichment
  • reactome-pathways - Reactome curated-pathway enrichment
  • wikipathways - WikiPathways community-pathway enrichment
  • enrichment-visualization - Dot/bar/cnet/emap/tree plots of enrichment results
  • differential-expression/de-results - Source of the gene list and the tested-gene universe
  • database-access/entrez-fetch - Fetch gene annotations / ID maps from NCBI
  • workflows/expression-to-pathways - End-to-end DE-to-enrichment pipeline

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in pathway-analysis/go-enrichment of GPTomics/bioSkills.

  • SKILL.md
  • examples/go_all_ontologies.R
  • examples/go_enrichment_basic.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Pathway Go Enrichment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Pathway Go Enrichment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Pathway Go Enrichment this skillGPTomics/bioSkills1.2k1 repos~5.1kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Pathway Go Enrichment

What does Bio Pathway Go Enrichment do?

Runs Gene Ontology over-representation analysis (ORA) on a gene LIST with clusterProfiler enrichGO, the one-sided hypergeometric/Fisher 2x2 test phyper(k-1, M, N-M, n, lower.tail=FALSE). Bio Pathway Go Enrichment is an agent skill from GPTomics/bioSkills.tail=FALSE).

When should I use Bio Pathway Go Enrichment?

Bio Pathway Go Enrichment fits situations like: A pre-selected gene list (DE hits; co-expression module; GWAS-mapped) needs GO annotation.

How do I install Bio Pathway Go Enrichment in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-pathway-go-enrichment -a claude-code`. Or copy the skill folder (pathway-analysis/go-enrichment in GPTomics/bioSkills) into .claude/skills/bio-pathway-go-enrichment in your project. Claude Code loads it when a task matches its description.

How do I install Bio Pathway Go Enrichment in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-pathway-go-enrichment -a codex`. Or copy the skill folder (pathway-analysis/go-enrichment in GPTomics/bioSkills) into .agents/skills/bio-pathway-go-enrichment in your project. Codex loads it when a task matches its description.

Can I use Bio Pathway Go Enrichment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-pathway-go-enrichment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-pathway-go-enrichment, .gemini/skills/bio-pathway-go-enrichment, .github/skills/bio-pathway-go-enrichment and .opencode/skills/bio-pathway-go-enrichment in your project.

What does Bio Pathway Go Enrichment need to run?

Going by SKILL.md and its folder, Bio Pathway Go Enrichment needs R for the scripts in its folder.

Does Bio Pathway Go Enrichment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Pathway Go Enrichment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Pathway Go Enrichment use?

Bio Pathway Go Enrichment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Pathway Go Enrichment use?

About 5.1k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Pathway Go Enrichment?

Skills that share tags, products or a category with Bio Pathway Go Enrichment: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Pathway Go Enrichment?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.