Agent skill

Bio Workflows Expression To Pathways

by GPTomics in GPTomics/bioSkills

Orchestrates the full path from differential expression results to redundancy-collapsed functional enrichment: choose ORA vs GSEA, convert gene IDs per method, run…

MITAuto-check passedData & Analytics

Install Bio Workflows Expression To Pathways

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-workflows-expression-to-pathways -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-workflows-expression-to-pathways --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workflows/expression-to-pathways .claude/skills/bio-workflows-expression-to-pathways && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-workflows-expression-to-pathways
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.5k tokens
SKILL.md length
1,973 words
Files
3
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Orchestrates the full path from differential expression results to redundancy-collapsed functional enrichment: choose ORA vs GSEA, convert gene IDs per method, run…

  • Works in 3 steps: Prepare the Input (list, ranked vector,… → Convert Gene IDs to What Each Method Needs → Collapse Redundancy, Then Visualize
  • A DESeq2/edgeR/limma result must become enriched GO terms
  • SKILL.md covers Version Compatibility, The Single Most Important…, Pipeline Overview and Stage Map, plus 12 more sections
  • Runs R scripts from its folder

What it does

Bio Workflows Expression To Pathways is an agent skill from GPTomics/bioSkills. Orchestrates the full path from differential expression results to redundancy-collapsed functional enrichment: choose ORA vs GSEA, convert gene IDs per method, run enrichGO/enrichKEGG/enrichPathway/enrichWP or gseGO/gseKEGG (clusterProfiler, ReactomePA, rWikiPathways), and visualize. Use when a DESeq2/edgeR/limma result must become enriched GO terms, KEGG/Reactome/WikiPathways pathways, or a GSEA leading edge; when the input is a full ranking for all genes (GSEA, named decreasing vector) or only a pre-selected…

Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Data & Analytics, covering Statistics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • A DESeq2/edgeR/limma result must become enriched GO terms
  • KEGG/Reactome/WikiPathways pathways
  • A GSEA leading edge
  • The input is a full ranking for all genes (GSEA

Example prompts

  • “Use the bio-workflows-expression-to-pathways skill to orchestrate the full path from differential expression results to redundancy-collapsed…”
  • “/bio-workflows-expression-to-pathways”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Prepare the Input (list, ranked vector, and the universe)
  2. Convert Gene IDs to What Each Method Needs
  3. Collapse Redundancy, Then Visualize

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Workflows Expression To Pathways loads about 5.5k tokens when it runs. Until then it costs about 196 tokens; SKILL.md has 1,973 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~196
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,973 words, ~5,492 tokens.

Download SKILL.mdSave it as .claude/skills/bio-workflows-expression-to-pathways/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-workflows-expression-to-pathways
description
Orchestrates the full path from differential expression results to redundancy-collapsed functional enrichment: choose ORA vs GSEA, convert gene IDs per method, run enrichGO/enrichKEGG/enrichPathway/enrichWP or gseGO/gseKEGG (clusterProfiler, ReactomePA, rWikiPathways), and visualize. Use when a DESeq2/edgeR/limma result must become enriched GO terms, KEGG/Reactome/WikiPathways pathways, or a GSEA leading edge; when the input is a full ranking for all genes (GSEA, named decreasing vector) or only a pre-selected list (ORA plus a defensible background universe); or when assembling DE-to-pathway end to end. The DE list and ranking statistic come from differential-expression/de-results; per-method nuance lives in the pathway-analysis skills.
tool_type
r
primary_tool
clusterProfiler
workflow
true
depends_on
pathway-analysis/go-enrichment, pathway-analysis/gsea, pathway-analysis/kegg-pathways, pathway-analysis/reactome-pathways, pathway-analysis/wikipathways…

Version Compatibility

Reference examples tested with: clusterProfiler 4.10+, org.Hs.eg.db 3.18+, ReactomePA 1.46+, enrichplot 1.22+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Expression to Pathways Workflow

"Find enriched pathways from my differential expression results" -> Decide the generation (ORA vs GSEA) FIRST, then convert IDs to the form each method needs, run enrichment against the chosen database, and collapse redundancy before interpreting - because the enrichment result is a claim conditioned on the method, the background universe, and the database version, not a discovery the algorithm hands back.

  • R: enrichGO(...) / gseGO(...) / enrichKEGG(...) / enrichPathway(...) (clusterProfiler, ReactomePA)

Scope: the ORCHESTRATION of a DE-to-enrichment pipeline - the generation fork, per-method ID conversion, the universe decision, the live-vs-local database caveat, and the handoff to redundancy-collapsed visualization. This workflow does NOT re-teach each method. The null/universe/reproducibility theory and the master method-selection tree -> the pathway-analysis README and the per-method skills (go-enrichment, gsea). ORA mechanics -> go-enrichment; GSEA mechanics -> gsea; per-database IDs and live-DB behavior -> kegg-pathways, reactome-pathways, wikipathways; the DE list and ranking statistic -> differential-expression/de-results; plotting -> enrichment-visualization.

The Single Most Important Modern Insight -- The First Decision Is the Generation, and It Is Set by the Input, Not by Preference

Pathway analysis has three generations (Khatri 2012 PLoS Comput Biol 8:e1002375): over-representation analysis (ORA), functional class scoring / GSEA (FCS), and pathway topology. A workflow that "runs enrichment" without first deciding which generation applies has already made the choice silently - usually ORA, the worst-calibrated corner of the space. The fork is mechanical:

  1. Is there a meaningful per-gene ranking for (nearly) ALL measured genes? A signed test statistic, the DESeq2 Wald stat, or -sign(log2FC)*log10(p) for every tested gene -> GSEA (a NAMED vector sorted in DECREASING order; the ranking metric IS the experiment). No arbitrary cutoff; detects coordinated weak shifts that ORA misses.
  2. Only a pre-selected LIST (DE hits past a cutoff, a co-expression module, GWAS loci, screen hits)? -> ORA, and the deliverable hinges on a defensible background universe - the genes that were testable, not the whole genome. ORA's p-value is whatever the denominator says it is.

The dangerous default is running ORA on data that has a full ranking (binarizing away the signal) or running ORA against the genome (measuring expression, not enrichment). Decide the fork out loud, record it, and record the why (see the pathway-analysis README) - this workflow owns the routing, not the derivation.

Pipeline Overview

DE results (differential-expression/de-results)
    |
    v
[0. Decide the generation: ranking for all genes? -> GSEA | pre-selected list? -> ORA]
    |
    +--> ORA branch: define the TESTABLE-gene universe, convert IDs per method
    |        +--> enrichGO     (OrgDb keyType)        -> go-enrichment
    |        +--> enrichKEGG   ('kegg' / 'ncbi-geneid', LIVE DB)  -> kegg-pathways
    |        +--> enrichPathway (ENTREZ, local DB)    -> reactome-pathways
    |        +--> enrichWP      (ENTREZ, LIVE GMT)    -> wikipathways
    |
    +--> GSEA branch: build a NAMED decreasing vector of ALL genes, set.seed
    |        +--> gseGO / gseKEGG / GSEA(+msigdbr)    -> gsea
    |
    v
[Redundancy collapse + visualization: simplify, pairwise_termsim, dotplot/emapplot/gseaplot2]   (enrichment-visualization)
    |
    v
A claim conditioned on universe + method + database version (record provenance)

Stage Map

StageGoalOwns the nuance
0. Decide generationORA vs GSEA from the available inputpathway-analysis README (method selection)
1. Prepare inputBuild the gene list AND/OR the named ranked vector; define the universedifferential-expression/de-results (the stat); pathway-analysis/go-enrichment (the universe rule)
2. Convert IDsMap to the form each method needs (OrgDb keyType / kegg-id / ENTREZ)go-enrichment, kegg-pathways, reactome-pathways
3a. ORAHypergeometric test of the list vs backgroundgo-enrichment, kegg-pathways, reactome-pathways, wikipathways
3b. GSEARunning-sum over the full rankinggsea
4. Collapse + visualizeReduce redundancy, then plotenrichment-visualization

Decision Tree by Scenario

ScenarioRouteWhy
All genes carry a DE statistic, a cutoff would be arbitraryGSEA (gseGO/gseKEGG) -> gseauses the full ranking; no cutoff; named decreasing vector
Pre-selected list (module, GWAS loci, screen hits), no rankingORA (enrichGO/enrichKEGG) -> go-enrichmentno ranking available; define the universe
Broad function annotationenrichGO / gseGO -> go-enrichment, gseaGO is the broadest LOCAL resource (reproducible)
Metabolic / signaling pathwaysenrichKEGG / gseKEGG -> kegg-pathwaysKEGG maps query a LIVE DB (pin the date)
Reaction-level, peer-reviewed, reproducible offlineenrichPathway -> reactome-pathwayslocal reactome.db, version-pinned
Disease/drug sets the others miss, broad speciesenrichWP -> wikipathwayscommunity-curated; LIVE versioned GMT
Bacterial / prokaryotic dataenrichKEGG with locus tags + KEGG organism code -> kegg-pathwaysKEGG covers prokaryotes; OrgDb usually does not
RNA-seq with strong gene-length biasGOseq -> go-enrichmentlength-aware ORA null
Multiple conditions/clusters side by sidecompareCluster -> any DBone model, faceted dotplot; never compare p across separate runs
The DE list / ranking statistic itself-> differential-expression/de-resultsthat is upstream, not enrichment
Why this null, which universe, version reporting-> go-enrichment (universe), gsea (null)per-method theory owned by each skill

Stage 1: Prepare the Input (list, ranked vector, and the universe)

Goal: Turn a DE table into the two possible inputs - a gene LIST for ORA and a NAMED decreasing vector for GSEA - and define the background universe as the testable genes.

Approach: Read the DE result, derive the significant list, build the ranked vector from a signed statistic (not a bare log2FC), and set the universe to exactly the genes that entered the DE test. The DE mechanics (the $padj vs $adj.P.Val column, shrinkage) live at differential-expression/de-results - this is only input shaping.

r
library(clusterProfiler)
library(org.Hs.eg.db)

res <- read.csv('deseq2_results.csv', row.names = 1)

# ORA input: a pre-selected list (DESeq2 padj column; limma/edgeR name it differently)
sig_genes <- rownames(subset(res, padj < 0.05 & abs(log2FoldChange) > 1))

# Background universe = genes that were TESTABLE (entered the DE test), NOT the genome.
# Using the genome measures expression bias, not enrichment (the universe rule; see pathway-analysis/go-enrichment).
universe_genes <- rownames(res[!is.na(res$pvalue), ])

# GSEA input: a NAMED vector of ALL genes, sorted DECREASING by a signed metric.
# Prefer the Wald stat (magnitude + precision); a bare log2FC over-weights noisy low-count genes.
# TRAP: a table from lfcShrink(type='apeglm'/'ashr') has NO `stat` column -- shrinkage drops it.
# Rank from the UNSHRUNK results(dds)$stat; use edgeR `sign(logFC)*-log10(PValue)` or limma `t`.
ranked <- res$stat
names(ranked) <- rownames(res)
ranked <- sort(ranked[!is.na(ranked)], decreasing = TRUE)

Stage 2: Convert Gene IDs to What Each Method Needs

Goal: Map identifiers to the exact ID type each enrichment function expects, because a mismatch returns zero hits silently.

Approach: Use bitr (OrgDb) for SYMBOL/ENSEMBL -> ENTREZ, keep both list and ranked vector in the same ID space, deduplicate, and track the conversion rate. Per-method ID rules are owned by each DB skill; the table below is the routing summary.

r
# enrichGO accepts ENSEMBL/SYMBOL/ENTREZ via keyType=; ENTREZ is the safe lingua franca downstream
sig_entrez <- bitr(sig_genes, fromType = 'SYMBOL', toType = 'ENTREZID', OrgDb = org.Hs.eg.db)
bg_entrez <- bitr(universe_genes, fromType = 'SYMBOL', toType = 'ENTREZID', OrgDb = org.Hs.eg.db)

# Carry the ranking through conversion: name the kept stat by its ENTREZ id
ranked_map <- bitr(names(ranked), fromType = 'SYMBOL', toType = 'ENTREZID', OrgDb = org.Hs.eg.db)
ranked_list <- ranked[ranked_map$SYMBOL]
names(ranked_list) <- ranked_map$ENTREZID
ranked_list <- ranked_list[!duplicated(names(ranked_list))]   # dedup or GSEA biases the score
ranked_list <- sort(ranked_list, decreasing = TRUE)           # re-sort: bitr remap can reorder rows

conv_rate <- nrow(sig_entrez) / length(sig_genes)   # report it; <0.85 -> wrong ID type/organism
MethodkeyType / ID requiredConvert with
enrichGO / gseGOOrgDb keyType ('ENSEMBL', 'SYMBOL', 'ENTREZID')bitr
enrichKEGG / gseKEGG'kegg' or 'ncbi-geneid' (NOT ENSEMBL/OrgDb)bitr to ENTREZID, pass keyType='ncbi-geneid' (bitr_kegg only converts among KEGG ID flavors)
enrichPathway / gsePathway (ReactomePA)ENTREZbitr
enrichWP / gseWP (WikiPathways)ENTREZ + organism stringbitr

Stage 3a: ORA Branch (pre-selected list + universe)

Goal: Test each gene set for over-representation of the list against the testable-gene background.

Approach: Always pass universe=; run GO ontologies separately; KEGG/Reactome/WikiPathways each need their own ID form. KEGG and WikiPathways query a LIVE database (internet-dependent, not reproducible across releases - pin the run date); GO and Reactome read local annotation (reproducible given the Bioconductor release).

r
# GO ORA - universe is the decision; simplify() collapses DAG redundancy (BP/MF/CC separately, not 'ALL')
go_bp <- enrichGO(sig_entrez$ENTREZID, universe = bg_entrez$ENTREZID, OrgDb = org.Hs.eg.db,
                  ont = 'BP', pAdjustMethod = 'BH', pvalueCutoff = 0.05, readable = TRUE)
go_bp <- simplify(go_bp, cutoff = 0.7, by = 'p.adjust')

# KEGG ORA - LIVE KEGG REST API; needs internet; record the access date for reproducibility
kegg <- enrichKEGG(sig_entrez$ENTREZID, universe = bg_entrez$ENTREZID, organism = 'hsa', keyType = 'ncbi-geneid', pvalueCutoff = 0.05)   # Entrez input; keyType='kegg' = Entrez for eukaryotes / locus tags for prokaryotes, 'ncbi-geneid' is explicit
kegg <- setReadable(kegg, OrgDb = org.Hs.eg.db, keyType = 'ENTREZID')

# Reactome ORA - ENTREZ required; LOCAL reactome.db so reproducible given the release
library(ReactomePA)
reactome <- enrichPathway(sig_entrez$ENTREZID, universe = bg_entrez$ENTREZID, organism = 'human', pvalueCutoff = 0.05, readable = TRUE)

Stage 3b: GSEA Branch (named decreasing vector of all genes)

Goal: Find gene sets whose genes shift coordinately across the full ranking, without a significance cutoff.

Approach: Run on the named decreasing ranked_list, fix the permutation seed so p-values are reproducible, then read the leading edge as the interpretable core. clusterProfiler GSEA is preranked / gene-permutation (the inter-gene-correlation-UNcorrected null) - a discovery screen; see pathway-analysis/gsea for the calibration caveat (CAMERA/ROAST).

r
set.seed(123)   # permutation reproducibility; without it p-values drift across runs

gsea_go <- gseGO(ranked_list, OrgDb = org.Hs.eg.db, ont = 'BP',
                 minGSSize = 10, maxGSSize = 500, pvalueCutoff = 0.05, verbose = FALSE)

gsea_kegg <- gseKEGG(ranked_list, organism = 'hsa',
                     minGSSize = 10, maxGSSize = 500, pvalueCutoff = 0.05, verbose = FALSE)

Stage 4: Collapse Redundancy, Then Visualize

Goal: Reduce overlapping terms to distinct findings before drawing conclusions, then plot deliberately.

Approach: A list of 40 significant GO terms is often a few biological stories told many times (shared genes via the GO true-path rule). Collapse with simplify/pairwise_termsim, then plot - emapplot/treeplot require pairwise_termsim() first (cnetplot does NOT; it draws the gene-concept network from the geneID column directly), and gseaplot2 is for a gseaResult not an enrichResult. Encoding choice and required pre-steps are owned by pathway-analysis/enrichment-visualization.

r
library(enrichplot)

go_bp <- pairwise_termsim(go_bp)              # required before emapplot/treeplot
dotplot(go_bp, showCategory = 20)             # GeneRatio vs Count: pick the encoding deliberately
emapplot(go_bp, showCategory = 30)            # redundancy-collapsed term-similarity map
gseaplot2(gsea_go, geneSetID = 1:3)           # gseaResult only, not enrichResult

Multi-Condition Comparison

Goal: Compare enrichment across conditions in one model instead of comparing p-values from separate runs.

Approach: compareCluster fits all gene lists together and facets the dotplot; never compare raw -log10(p) across separate enrichments (it scales with set size and sample size). For GSEA, compare NES, not p.

r
gene_clusters <- list(A = sig_A, B = sig_B, C = sig_C)
cc <- compareCluster(gene_clusters, fun = 'enrichKEGG', organism = 'hsa',
                     universe = bg_entrez$ENTREZID)   # compareCluster forwards ... to fun; omitting universe silently reverts to the whole-genome background
dotplot(cc, showCategory = 10)

Per-Method Failure Modes

Whole-genome background

Trigger: universe= left at default while only ~12k genes were expressed. Mechanism: the hypergeometric p-value is fully determined by the denominator; the genome inflates any set whose members are expressed in the tissue. Symptom: many tissue-specific terms enrich with tiny p. Fix: set universe to the genes that entered the DE test (the testable set).

Show full SKILL.md (761 more words)Show less
ORA on a ranked dataset

Trigger: filtering all-gene DE results to a list and running ORA. Mechanism: binarizing at an arbitrary cutoff discards magnitude and the coordinated-weak signal. Symptom: GSEA finds sets ORA missed. Fix: if a ranking exists for all genes, run GSEA; reserve ORA for genuinely unranked lists.

Wrong ID type for the database

Trigger: ENSEMBL/SYMBOL passed to enrichKEGG/enrichPathway/enrichWP. Mechanism: those expect kegg-id/ENTREZ; unmatched IDs are dropped. Symptom: zero terms, no error. Fix: bitr/bitr_kegg to the required ID; check conv_rate.

GSEA without set.seed or with an unsorted vector

Trigger: no seed, or a list that is not named and decreasing. Mechanism: permutation p-values drift run to run; an unsorted/unnamed vector errors or mis-ranks. Symptom: different leading edges each run, or a names error. Fix: build the named decreasing vector and set.seed.

Ranking a GSEA vector off a shrunken-LFC object

Trigger: building the ranked vector from a lfcShrink(type='apeglm'/'ashr') table, or ranking by bare log2FoldChange. Mechanism: apeglm/ashr DROP the stat column, so the vector silently falls back to shrunken LFC; low-count genes with unstable large FC then hijack the leading edge. Symptom: the leading edge is dominated by low-baseMean genes, or res$stat is NULL. Fix: rank from the unshrunken results(dds)$stat (DESeq2), limma topTable$t, or edgeR sign(logFC)*-log10(PValue); reserve shrunken LFC for visualization.

Live-DB result reported as reproducible

Trigger: KEGG/WikiPathways result with no recorded date. Mechanism: those query the current data release; the same code returns different pathways later. Symptom: a collaborator cannot reproduce the figure. Fix: record the access date and data version; prefer local GO/Reactome when reproducibility is paramount.

Redundancy read as replication

Trigger: interpreting 40 overlapping GO terms as 40 findings. Mechanism: the true-path rule and pathway overlap mean shared genes drive many sets. Symptom: the same 3-5 genes explain the top 20 terms. Fix: simplify/pairwise_termsim, inspect the leading-edge/geneID core, report clusters of terms.

Quantitative Thresholds

ThresholdSourceRationale
pvalueCutoff = 0.05clusterProfiler defaultfilters on p.adjust by default in enrichResult; standard FDR gate
qvalueCutoff = 0.2clusterProfiler defaultsecondary q-value gate
pAdjustMethod = 'BH'Benjamini-Hochbergvalid FDR control under positive dependence (overlapping sets); Bonferroni over-corrects
minGSSize = 10enrichGO/gseGO defaultdrop tiny sets that overfit
maxGSSize = 500enrichGO/gseGO defaultdrop overly broad sets that always "enrich"
simplify(cutoff = 0.7)GOSemSim semantic similarityGO DAG redundancy cutoff; lower keeps more terms
conversion rate > 0.85practical QC<85% ID conversion flags a wrong ID type/organism
set.seed(123)reproducibilityany fixed seed; the point is to fix the permutation draw

Common Errors

Error / symptomCauseSolution
enrichKEGG returns 0 termsENSEMBL passed (needs kegg-id/ENTREZ), wrong organism code, or KEGG API downconvert with bitr_kegg; check organism; retry (live DB)
--> No gene can be mappedwrong keyType/OrgDb for the input IDsmatch keyType to the actual ID type
gseGO error about namesvector not named or not sorted decreasingbuild a named vector sorted decreasing = TRUE
emapplot/treeplot empty or errorspairwise_termsim() not run first (cnetplot does not need it)run pairwise_termsim() before emapplot/treeplot
simplify fails on ont='ALL'simplify needs one ontologyrun BP/MF/CC separately, then simplify each
different results each runno set.seed, or the live KEGG/WP DB changedset.seed; pin and record the DB version/date
all terms have NA Descriptionreadable/setReadable not appliedset readable = TRUE or call setReadable

References

  • Khatri P, Sirota M, Butte AJ. 2012. Ten years of pathway analysis: current approaches and outstanding challenges. PLoS Comput Biol 8:e1002375.
  • Subramanian A, Tamayo P, Mootha VK, et al. 2005. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. PNAS 102:15545-15550.
  • Goeman JJ, Buhlmann P. 2007. Analyzing gene expression data in terms of gene sets: methodological issues. Bioinformatics 23:980-987.
  • Wu T, Hu E, Xu S, et al. 2021. clusterProfiler 4.0: a universal enrichment tool for interpreting omics data. The Innovation 2:100141.
  • Yu G, He QY. 2016. ReactomePA: an R/Bioconductor package for reactome pathway analysis and visualization. Mol BioSyst 12:477-479.
  • Kanehisa M, Goto S. 2000. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res 28:27-30.
  • Young MD, Wakefield MJ, Smyth GK, Oshlack A. 2010. Gene ontology analysis for RNA-seq: accounting for selection bias. Genome Biol 11:R14.
  • Wijesooriya K, Jadaan SA, Perera KL, Kaur T, Ziemann M. 2022. Urgent need for consistent standards in functional enrichment analysis. PLoS Comput Biol 18:e1009935.
  • pathway-analysis/go-enrichment - GO over-representation, background universe, redundancy reduction, length bias
  • pathway-analysis/gsea - Ranked-list GSEA, named decreasing vector, ranking metric, leading edge, NES
  • pathway-analysis/kegg-pathways - KEGG pathway/module enrichment, live DB, prokaryotes, multi-condition
  • pathway-analysis/reactome-pathways - Reactome curated-pathway ORA and GSEA, ENTREZ IDs, reproducible local DB
  • pathway-analysis/wikipathways - WikiPathways community-pathway enrichment, versioned GMT, broad species
  • pathway-analysis/enrichment-visualization - Dot/bar/cnet/emap/GSEA plots and required pre-steps
  • differential-expression/de-results - Source of the gene list and the ranking statistic

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in workflows/expression-to-pathways of GPTomics/bioSkills.

  • SKILL.md
  • examples/complete_enrichment.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Workflows Expression To Pathways next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Workflows Expression To Pathways compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Workflows Expression To Pathways this skillGPTomics/bioSkills1.2k1 repos~5.5kAutomated safety check: PassMIT
Sandbox Benchvercel/next.js143k—~4.1kAutomated safety check: PassMIT
Statistical Analysisspacering-net/codeg3.8k4 repos~5kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.6k16 repos~4.9kAutomated safety check: PassBSD-3-Clause
Statistical Powerspacering-net/codeg3.8k2 repos~3.6kAutomated safety check: NotesMIT
AI Daily DigestvigorX777/ai-daily-digest1.6k—~1.3kAutomated safety check: PassNone

Similar skills

  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Statistical Analysis

    spacering-net/codeg

    Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting.

    3.8k GitHub starsUsed in 4 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 16 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Power

    spacering-net/codeg

    Sample-size and statistical power calculations for planning studies.

    3.8k GitHub starsUsed in 2 repos~3.6k tokens
    Data & AnalyticsAuto-check: notes
  • AI Daily Digest

    vigorX777/ai-daily-digest

    Fetches RSS feeds from 90 top Hacker News blogs (curated by Karpathy), uses AI to score and filter articles, and generates a daily digest in Markdown with Chinese-translated titles, category…

    1.6k GitHub stars~1.3k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Agent Session Monitor

    higress-group/higress

    Real-time agent conversation monitoring - monitors Higress access logs, aggregates conversations by session, tracks token usage.

    9.5k GitHub stars~3.3k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Workflows Expression To Pathways

What does Bio Workflows Expression To Pathways do?

Orchestrates the full path from differential expression results to redundancy-collapsed functional enrichment: choose ORA vs GSEA, convert gene IDs per method, run…. Bio Workflows Expression To Pathways is an agent skill from GPTomics/bioSkills. Orchestrates the full path from differential expression results to redundancy-collapsed functional enrichment: choose ORA vs GSEA, convert gene IDs per method, run enrichGO/enrichKEGG/enrichPathway/enrichWP or gseGO/gseKEGG (clusterProfiler, ReactomePA, rWikiPathways), and visualize.

When should I use Bio Workflows Expression To Pathways?

Bio Workflows Expression To Pathways fits situations like: A DESeq2/edgeR/limma result must become enriched GO terms; KEGG/Reactome/WikiPathways pathways; A GSEA leading edge; the input is a full ranking for all genes (GSEA.

How do I install Bio Workflows Expression To Pathways in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-workflows-expression-to-pathways -a claude-code`. Or copy the skill folder (workflows/expression-to-pathways in GPTomics/bioSkills) into .claude/skills/bio-workflows-expression-to-pathways in your project. Claude Code loads it when a task matches its description.

How do I install Bio Workflows Expression To Pathways in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-workflows-expression-to-pathways -a codex`. Or copy the skill folder (workflows/expression-to-pathways in GPTomics/bioSkills) into .agents/skills/bio-workflows-expression-to-pathways in your project. Codex loads it when a task matches its description.

Can I use Bio Workflows Expression To Pathways in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-workflows-expression-to-pathways -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-workflows-expression-to-pathways, .gemini/skills/bio-workflows-expression-to-pathways, .github/skills/bio-workflows-expression-to-pathways and .opencode/skills/bio-workflows-expression-to-pathways in your project.

What does Bio Workflows Expression To Pathways need to run?

Going by SKILL.md and its folder, Bio Workflows Expression To Pathways needs R for the scripts in its folder.

Does Bio Workflows Expression To Pathways access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Workflows Expression To Pathways safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Workflows Expression To Pathways use?

Bio Workflows Expression To Pathways is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Workflows Expression To Pathways use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Workflows Expression To Pathways?

Skills that share tags, products or a category with Bio Workflows Expression To Pathways: Sandbox Bench (vercel/next.js, 143k stars), Statistical Analysis (spacering-net/codeg, 3.8k stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.6k stars) and Statistical Power (spacering-net/codeg, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Workflows Expression To Pathways?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.