Agent skill

Bio Pathway Kegg Pathways

by GPTomics in GPTomics/bioSkills

Tests gene lists, ranked vectors, and fold-change vectors against KEGG pathways and modules with clusterProfiler enrichKEGG/enrichMKEGG (ORA), gseKEGG (GSEA), and SPIA/graphite (signed-topology…

MITAuto-check passedBackend & APIs

Install Bio Pathway Kegg Pathways

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-pathway-kegg-pathways -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-pathway-kegg-pathways --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pathway-analysis/kegg-pathways .claude/skills/bio-pathway-kegg-pathways && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-pathway-kegg-pathways
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.4k tokens
SKILL.md length
2,257 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Tests gene lists, ranked vectors, and fold-change vectors against KEGG pathways and modules with clusterProfiler enrichKEGG/enrichMKEGG (ORA), gseKEGG (GSEA), and SPIA/graphite (signed-topology…

  • Works in 2 steps: The query is live, so the result is… → KEGG is the only mainstream database…
  • Finding enriched KEGG pathways
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy (KEGG Across the… and Decision Tree by Scenario, plus 12 more sections
  • Runs R scripts from its folder

What it does

Bio Pathway Kegg Pathways is an agent skill from GPTomics/bioSkills. Tests gene lists, ranked vectors, and fold-change vectors against KEGG pathways and modules with clusterProfiler enrichKEGG/enrichMKEGG (ORA), gseKEGG (GSEA), and SPIA/graphite (signed-topology perturbation) in R. Owns the third pathway-analysis generation because KEGG ships signed directed signaling topology (KGML). Covers why a KEGG result is a timestamped join against a live REST API (irreproducible unless pinned with a gson snapshot, not the stale 2012 KEGG.db), why enrichKEGG keyType is kegg/ncbi-geneid not…

Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `usage-guide.md`).

It sits in Backend & APIs, covering REST APIs. It works with Ensembl and NCBI. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Finding enriched KEGG pathways
  • Scoring signed pathway perturbation
  • Analyzing prokaryotes
  • Non-model organisms via locus tags

Example prompts

  • “Use the bio-pathway-kegg-pathways skill to test gene lists, ranked vectors, and fold-change vectors against KEGG pathways and modules with…”
  • “/bio-pathway-kegg-pathways”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The query is live, so the result is irreproducible unless the release is pinned. enrichKEGG/gseKEGG/SPIA hit the KEGG REST API at call…
  2. KEGG is the only mainstream database shipping signed, directed signaling topology (KGML), which is why this skill owns the third…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • rest.kegg.jp

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Pathway Kegg Pathways loads about 5.4k tokens when it runs. Until then it costs about 249 tokens; SKILL.md has 2,257 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~249
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,257 words, ~5,446 tokens.

Download SKILL.mdSave it as .claude/skills/bio-pathway-kegg-pathways/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-pathway-kegg-pathways
description
Tests gene lists, ranked vectors, and fold-change vectors against KEGG pathways and modules with clusterProfiler enrichKEGG/enrichMKEGG (ORA), gseKEGG (GSEA), and SPIA/graphite (signed-topology perturbation) in R. Owns the third pathway-analysis generation because KEGG ships signed directed signaling topology (KGML). Covers why a KEGG result is a timestamped join against a live REST API (irreproducible unless pinned with a gson snapshot, not the stale 2012 KEGG.db), why enrichKEGG keyType is kegg/ncbi-geneid not OrgDb ENSEMBL/SYMBOL (zero hits), why organism is a KEGG code (hsa, pae) with prokaryotic locus tags, and why SPIA works only on signaling maps. Use when finding enriched KEGG pathways or modules, scoring signed pathway perturbation, analyzing prokaryotes or non-model organisms via locus tags or KO, comparing conditions with compareCluster, or overlaying data with pathview. The hypergeometric universe lives in go-enrichment; the GSEA engine in gsea.
tool_type
r
primary_tool
clusterProfiler

Version Compatibility

Reference examples tested with: clusterProfiler 4.18+, org.Hs.eg.db 3.18+, gson 0.1+ (snapshot pinning), SPIA 2.50+ and graphite 1.56+ (topology section).

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

KEGG is a LIVE DATABASE, not a package. enrichKEGG/enrichMKEGG/gseKEGG query the KEGG REST API (https://rest.kegg.jp/) at call time, so the same code on the same genes returns DIFFERENT pathways months apart as KEGG updates. For any reported result, pin the release with a gson snapshot (below) and record the access date; use_internal_data=TRUE does NOT pin the current KEGG (it loads the deprecated 2012 KEGG.db).

KEGG Pathway and Topology Enrichment

"Which KEGG pathways are perturbed in my data?" -> Join genes to KEGG's curated pathway/module gene sets (ORA or GSEA), or propagate fold-changes through KEGG's signed wiring (SPIA) - and pin the KEGG release, because the result is a timestamped query against a moving curation, not a fact about the biology.

  • R: enrichKEGG(gene, organism, keyType) | gseKEGG(geneList, organism) | spia(de, all, organism)

Scope: KEGG-specific enrichment across all three generations - membership ORA (enrichKEGG/enrichMKEGG), ranked GSEA (gseKEGG), and signed-topology perturbation (SPIA/graphite). KEGG ID mapping (organism codes, keyType, bitr_kegg, prokaryotic locus tags, KO routing), reproducibility/pinning, and pathview map overlay live here. The hypergeometric test and the universe problem -> go-enrichment. The GSEA running-sum engine and ranking-metric choice -> gsea. Reactome/WikiPathways gene sets -> reactome-pathways, wikipathways. Generic dot/cnet/emap plots -> enrichment-visualization. The DE list and fold-changes -> differential-expression/de-results.

The Single Most Important Modern Insight -- A KEGG Result Is a Timestamped Join Against a Moving, Partially-Paywalled Curation, Not a Fact About Biology

Two consequences follow, and both are invisible until someone reruns the analysis.

  1. The query is live, so the result is irreproducible unless the release is pinned. enrichKEGG/gseKEGG/SPIA hit the KEGG REST API at call time; KEGG adds maps, re-annotates genes, and revises edges continuously, so identical code returns a different pathway list next quarter. The fix is a gson snapshot: gson_KEGG('hsa') downloads the current KEGG pathway/module sets into a GSON object, write.gson()/read.gson() persist it, and the generic enricher(gene, gson=k) / GSEA(geneList, gson=k) run frozen and offline against it. Record the access date. use_internal_data=TRUE is NOT this fix - it silently reaches for the deprecated 2012 KEGG.db, which is the wrong, stale snapshot.

  2. KEGG is the only mainstream database shipping signed, directed signaling topology (KGML), which is why this skill owns the third generation of pathway analysis. ORA and GSEA treat a pathway as an unordered bag of exchangeable genes; SPIA asks a question they structurally cannot pose - given where each gene sits in the wiring and the sign of every edge, how perturbed is this pathway? That requires the topology only KEGG (and a few others via graphite) provides. The discipline: choose the generation by the question (membership? rank? signed perturbation?), match keyType/organism to the actual IDs (locus tags for bacteria, KO for non-model), set the universe to the genes that could have been called DE, and pin the release before publishing.

Tool Taxonomy (KEGG Across the Three Generations)

MethodGenerationEngineUses log2FC?Uses topology/direction?Suitable KEGG mapsCitation
enrichKEGG (ORA)1st (over-representation)hypergeometricno (gene list)noallWu 2021 The Innovation 2:100141; Kanehisa & Goto 2000 Nucleic Acids Res 28:27
enrichMKEGG (ORA on modules)1sthypergeometricnonomodules (M-numbers)Wu 2021 The Innovation 2:100141
gseKEGG (GSEA)2nd (functional class scoring)fgsea running sumyes (ranking)noall (as sets)Wu 2021 The Innovation 2:100141; engine -> gsea
SPIA3rd (pathway topology)pNDE (ORA) x pPERT (perturbation) -> pGyes (named log2FC)YES (signed KGML)SIGNALING onlyTarca 2009 Bioinformatics 25:75; Draghici 2007 Genome Res 17:1537
graphite + runSPIA3rdSPIA over harmonized graphsyesYESsignaling (KEGG/Reactome)Sales 2012 BMC Bioinformatics 13:20

The three-generations framing (ORA -> FCS -> pathway topology) is Khatri 2012 PLoS Comput Biol 8:e1002375; this skill is the KEGG instantiation of all three (the category README compares the generations across databases).

Decision Tree by Scenario

ScenarioRecommendedWhy
Pre-selected gene list, "which KEGG pathways"enrichKEGG (ORA), set the universeno ranking available; membership test
All genes carry a DE statistic, no clear cutoffgseKEGG -> gseauses the full ranking; no arbitrary cutoff
Want WHERE in a broad pathway the signal sitsenrichMKEGG (modules)M-numbers are tighter functional units
Have named log2FC + want signed perturbation on a SIGNALING mapSPIA (or graphite + runSPIA)propagates fold-changes through the wiring; uses direction
Metabolic-pathway question (glycolysis, TCA)enrichKEGG / gseKEGGmetabolic maps are compound-mediated; SPIA is undefined there
Human / mouse / model eukaryotebitr -> Entrez, keyType='ncbi-geneid'KEGG gene ID == Entrez for these organisms
Bacterial / prokaryotic datalocus tags, keyType='kegg', NO OrgDb/bitrbacterial KEGG IDs ARE locus tags; no org.*.eg.db exists
Non-model organism with no KEGG genomemap to KO, organism='ko'the universal escape hatch into KEGG pathway space
Result must be reproducible / publishedgson_KEGG snapshot + enricher/GSEA, record datelive unpinned queries drift; use_internal_data pins the WRONG 2012 db
Multiple conditions to compare side by sidecompareCluster(fun='enrichKEGG')one model, faceted dotplot; never compare raw p-values
Overlay per-gene data on the KEGG map imagepathview -> rendera KEGG-specific operation; generic plots -> enrichment-visualization
The DE list / fold-changes themselves-> differential-expression/de-resultsupstream, not enrichment

Prepare the Gene IDs (the Join That Decides Everything)

Goal: Get the query genes and the universe into the exact ID type KEGG expects for the organism, because every KEGG failure is a join failure.

Approach: For model eukaryotes convert SYMBOL/ENSEMBL to Entrez (KEGG's gene ID for hsa/mmu/rno) and pass keyType='ncbi-geneid'. For prokaryotes pass locus tags directly with keyType='kegg' and no OrgDb. Convert the universe the same way. Passing ENSEMBL/SYMBOL to enrichKEGG returns zero hits silently.

r
library(clusterProfiler)
library(org.Hs.eg.db)

de <- read.csv('de_results.csv')   # DE list source -> differential-expression/de-results
sig_symbols <- de$gene[de$padj < 0.05 & abs(de$log2FoldChange) > 1]   # padj is the DESeq2 adjusted-p column
sig_entrez  <- bitr(sig_symbols, fromType='SYMBOL', toType='ENTREZID', OrgDb=org.Hs.eg.db)$ENTREZID

# universe = genes that COULD have been called DE (non-NA test statistic), same ID type
universe <- bitr(de$gene[!is.na(de$pvalue)], fromType='SYMBOL', toType='ENTREZID', OrgDb=org.Hs.eg.db)$ENTREZID

bitr_kegg(geneID, fromType, toType, organism) converts among KEGG's own ID flavors ('kegg', 'ncbi-geneid', 'ncbi-proteinid', 'uniprot') via the REST conv endpoint - use it when starting from UniProt or NCBI protein IDs. Check KEGG coverage of an organism with search_kegg_organism('Pseudomonas aeruginosa', by='scientific_name').

Run KEGG ORA (enrichKEGG / enrichMKEGG)

Goal: Find KEGG pathways (or modules) over-represented among the query genes relative to the measured universe.

Approach: Run enrichKEGG with the correct organism code, keyType, and an explicit universe; enrichKEGG has no readable argument, so translate the geneID column to symbols afterward with setReadable (eukaryotes only).

r
kk <- enrichKEGG(gene=sig_entrez, organism='hsa', keyType='ncbi-geneid',
                 universe=universe, pvalueCutoff=0.05, pAdjustMethod='BH',
                 minGSSize=10, maxGSSize=500, qvalueCutoff=0.2)
kk <- setReadable(kk, OrgDb=org.Hs.eg.db, keyType='ENTREZID')   # eukaryotes only; no OrgDb -> keep raw IDs
head(as.data.frame(kk))   # ID, Description, GeneRatio, BgRatio, pvalue, p.adjust, qvalue, geneID, Count

mkk <- enrichMKEGG(gene=sig_entrez, organism='hsa', keyType='ncbi-geneid', universe=universe)   # KEGG MODULES (M-numbers)

Report p.adjust/qvalue, not raw pvalue. Fold enrichment = GeneRatio / BgRatio. enrichMKEGG tests smaller, sparser sets: higher resolution (which sub-process is hit) but lower power and many genes belong to no module.

Run KEGG GSEA (gseKEGG)

Goal: Find KEGG sets whose genes shift coordinately across the full ranking, with no cutoff.

Approach: Build a named numeric vector sorted DECREASING by the ranking metric, fix the seed (gseKEGG defaults seed=FALSE), then run gseKEGG. The running-sum engine and the ranking-metric choice are owned by gsea; only the KEGG arguments (organism, keyType) are KEGG-specific.

r
geneList <- de$log2FoldChange; names(geneList) <- de$entrez   # names = Entrez IDs
geneList <- sort(geneList[!is.na(geneList)], decreasing=TRUE)
set.seed(123)   # gseKEGG seed=FALSE by default; fix it so permutation p-values are reproducible
kk2 <- gseKEGG(geneList=geneList, organism='hsa', keyType='ncbi-geneid', minGSSize=10, maxGSSize=500, pvalueCutoff=0.05)

Run Signed-Topology Perturbation (SPIA) -- the Third Generation

Goal: Score how perturbed each SIGNALING pathway is given both the over-representation of DE genes and the propagation of their fold-changes through the signed wiring.

Approach: SPIA combines pNDE (the classical over-representation evidence) with pPERT (the probability of the observed total accumulated perturbation tA, computed by propagating log2 fold-changes through KGML activation/inhibition edges) into a single global pG, then FDR-corrects it. It needs a NAMED vector of DE fold-changes plus the universe, and is defined only for signaling maps. graphite is the modern route: it harmonizes node IDs, resolves complexes/families, removes compounds, and can run SPIA over Reactome topology too.

r
library(SPIA)
sig <- de[de$padj < 0.05, ]   # DE genes only
map <- bitr(sig$gene, 'SYMBOL', 'ENTREZID', org.Hs.eg.db)   # bitr drops/many-to-one: MERGE, never assign as names
de_vec <- setNames(sig$log2FoldChange[match(map$SYMBOL, sig$gene)], map$ENTREZID)
de_vec <- de_vec[!duplicated(names(de_vec))]
res <- spia(de=de_vec, all=universe, organism='hsa', nB=2000, plots=FALSE)   # nB=2000 bootstraps for pPERT
# output cols: Name, ID, pSize, NDE, pNDE, tA, pPERT, pG, pGFdr, pGFWER, Status, KEGGLINK
# Status reports inferred Activated / Inhibited from the sign of tA

# graphite route (decouples from KEGG's bundled data; works on Reactome too)
library(graphite)
db <- pathways('hsapiens', 'kegg')
db <- convertIdentifiers(db, 'ENTREZID')
prepareSPIA(db, 'kegg_hsa_spia')              # writes the pathway dataset file
gr <- runSPIA(de=de_vec, all=universe, 'kegg_hsa_spia')

SPIA aborts if more than ~1% of the DE IDs are absent from all, so build the universe from the same ID space. The standalone SPIA package also ships a frozen hsaSPIA data object that is an OLDER snapshot than a live enrichKEGG query - do not mix the two in one comparison.

Pin the KEGG Release for Reproducibility

Goal: Freeze the KEGG data a result depends on so the analysis is reproducible and runs offline.

Approach: Snapshot the current KEGG sets into a GSON object, persist it, and run enrichment against the snapshot with the generic enricher/GSEA (which accept a gson argument); record the access date. Do NOT use use_internal_data=TRUE for this.

r
library(gson)                                      # GSON class + write.gson/read.gson
k <- gson_KEGG('hsa')                              # gson_KEGG is exported by clusterProfiler; downloads current KEGG sets
k@accessed_date <- as.character(Sys.Date())        # the accessed_date slot survives write/read; a base attr() does not
write.gson(k, file.path(tempdir(), 'kegg_hsa.gson'))
k <- read.gson(file.path(tempdir(), 'kegg_hsa.gson'))

kk_pinned  <- enricher(sig_entrez, gson=k, universe=universe)   # frozen ORA, offline, reproducible
gsea_pinned <- GSEA(geneList, gson=k)                            # frozen GSEA against the snapshot
Show full SKILL.md (919 more words)Show less

Compare Multiple Conditions

Goal: See shared and condition-specific KEGG pathways across groups in one faceted figure.

Approach: Pass named gene lists to compareCluster with fun='enrichKEGG'; it fits one model and produces a faceted dotplot. Compare pathway-ID SETS across conditions, never raw p-values (they depend on sample size, DE gene count, and the KEGG release).

r
clusters <- list(up=up_entrez, down=down_entrez)
ck <- compareCluster(geneClusters=clusters, fun='enrichKEGG', organism='hsa', keyType='ncbi-geneid')
ck <- setReadable(ck, OrgDb=org.Hs.eg.db, keyType='ENTREZID')
# dotplot(ck) -> enrichment-visualization for the plot grammar

Overlay Data on the KEGG Map (pathview)

pathview downloads a KEGG pathway's KGML and image, joins per-gene values to the nodes, and writes a colored map PNG/PDF (a KEGG-specific operation owned here; generic dot/cnet/emap plots route to enrichment-visualization). It writes files to the working directory and queries KEGG live.

r
library(pathview)
vals <- setNames(de$log2FoldChange, de$entrez)
pathview(gene.data=vals, pathway.id='hsa04110', species='hsa', gene.idtype='entrez')   # writes hsa04110.pathview.png

Per-Method Failure Modes

ENSEMBL/SYMBOL passed to enrichKEGG

Trigger: feeding OrgDb-style ENSEMBL or SYMBOL IDs to enrichKEGG/gseKEGG. Mechanism: KEGG's keyType is 'kegg'/'ncbi-geneid'/'ncbi-proteinid'/'uniprot', not an OrgDb keytype, so no IDs join. Symptom: zero enriched pathways, no error. Fix: bitr to Entrez and set keyType='ncbi-geneid' (eukaryotes), or pass locus tags with keyType='kegg' (prokaryotes).

Live-query result treated as reproducible

Trigger: reporting an enrichKEGG/gseKEGG/SPIA result without pinning the release. Mechanism: the REST query returns the CURRENT KEGG, which changes over time. Symptom: a rerun months later yields a different pathway list. Fix: snapshot with gson_KEGG, run enricher/GSEA against the gson, and record the access date.

use_internal_data=TRUE believed to pin current KEGG

Trigger: setting use_internal_data=TRUE for reproducibility. Mechanism: it loads the deprecated 2012 KEGG.db, not a current pin (and may simply fail). Symptom: stale or absent pathways unlike the live result. Fix: use a gson snapshot instead; treat KEGG.db as legacy-only.

SPIA on metabolic maps

Trigger: running SPIA/graphite topology on glycolysis or other metabolic maps. Mechanism: metabolic maps are compound-mediated and give no clean signed gene->gene graph. Symptom: meaningless perturbation scores. Fix: restrict SPIA to signaling maps; use enrichKEGG/gseKEGG for metabolism.

Whole-database universe in ORA

Trigger: omitting universe. Mechanism: the default background is all KEGG-annotated genes, biased toward well-studied, metabolically central genes. Symptom: inflated significance for pathways enriched in measured/expressed genes (the tissue-specificity artifact). Fix: set universe to the genes that could have been called DE, in the same ID type.

Locus-tag / strain mismatch in prokaryotes

Trigger: locus tags from a re-annotated genome or a different strain than KEGG's reference. Mechanism: the gene-ID join is exact; drifted locus tags do not match KEGG's pae/eco genome. Symptom: many genes silently dropped, weak or empty enrichment. Fix: confirm the organism code and reference genome with search_kegg_organism; align locus tags to KEGG's annotation, or route through KO.

bitr/OrgDb forced onto bacteria

Trigger: running bitr() or setReadable() on a prokaryote. Mechanism: no org.*.eg.db exists for most bacteria and there is no Entrez==KEGG identity. Symptom: bitr fails or empties the gene list; setReadable errors. Fix: pass locus tags directly with keyType='kegg'; keep raw IDs (no setReadable).

Quantitative Thresholds

ThresholdSourceRationale
pvalueCutoff=0.05enrichKEGG/gseKEGG defaultfilters on p.adjust by default; standard FDR gate
qvalueCutoff=0.2clusterProfiler defaultsecondary q-value gate on enrichResult
pAdjustMethod='BH'clusterProfiler defaultBenjamini-Hochberg FDR; less conservative than Bonferroni for discovery
minGSSize=10enrichKEGG defaultdrop tiny sets that overfit and give unstable p-values
maxGSSize=500enrichKEGG defaultdrop very broad sets that always 'enrich'
nB=2000SPIA defaultbootstrap replicates for the pPERT null; raise for stable small p-values
SPIA aborts if >1% of DE IDs absent from allTarca 2009 Bioinformatics 25:75the perturbation null requires the DE genes live in the universe
set.seed before gseKEGG/SPIAreproducibilitygseKEGG seed=FALSE and SPIA bootstrap are stochastic; fix the seed
ID-conversion loss > ~15%practice heuristicreport the bitr conversion rate; heavy loss makes the result unreliable

Common Errors

Error / symptomCauseSolution
enrichKEGG returns 0 pathwaysENSEMBL/SYMBOL passed, or wrong organism code, or KEGG API unreachablebitr to Entrez + keyType='ncbi-geneid'; verify code with search_kegg_organism; check network
setReadable errorsno OrgDb for the organism (prokaryote)skip setReadable; keep raw KEGG IDs
gson= rejected by enrichKEGGenrichKEGG/gseKEGG have no gson argumentpass the gson to the generic enricher()/GSEA() instead
Different pathways on rerunlive KEGG changed between runspin with a gson snapshot and record the access date
SPIA: "more than 1% of de IDs not in all"DE IDs not a subset of the universebuild de and all from the same ID space
SPIA gives nonsense on glycolysistopology on a metabolic mapuse enrichKEGG/gseKEGG; SPIA is signaling-only
Bacterial list gives 0 hitsEntrez/bitr forced onto a prokaryotepass locus tags with keyType='kegg', no OrgDb

References

  • Kanehisa M, Goto S. 2000. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res 28:27-30.
  • Kanehisa M, Furumichi M, Sato Y, et al. 2023. KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res 51:D587-D592.
  • Wu T, Hu E, Xu S, et al. 2021. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. The Innovation 2:100141.
  • Tarca AL, Draghici S, Khatri P, et al. 2009. A novel signaling pathway impact analysis (SPIA). Bioinformatics 25:75-82.
  • Draghici S, Khatri P, Tarca AL, et al. 2007. A systems biology approach for pathway level analysis. Genome Res 17:1537-1545.
  • Sales G, Calura E, Cavalieri D, Romualdi C. 2012. graphite - a Bioconductor package to convert pathway topology to gene network. BMC Bioinformatics 13:20.
  • Luo W, Brouwer C. 2013. Pathview: an R/Bioconductor package for pathway-based data integration and visualization. Bioinformatics 29:1830-1831.
  • Khatri P, Sirota M, Butte AJ. 2012. Ten years of pathway analysis: current approaches and outstanding challenges. PLoS Comput Biol 8:e1002375.
  • go-enrichment - Hypergeometric ORA and the background-universe problem
  • gsea - GSEA running-sum engine and ranking-metric choice (gseKEGG)
  • reactome-pathways - Reactome curated-pathway enrichment (reproducible local DB)
  • wikipathways - WikiPathways community-pathway enrichment
  • enrichment-visualization - Dot/bar/cnet/emap/ridge plots of enrichment results
  • differential-expression/de-results - Source of the gene list and the fold-changes
  • workflows/expression-to-pathways - End-to-end DE-to-enrichment pipeline

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in pathway-analysis/kegg-pathways of GPTomics/bioSkills.

  • SKILL.md
  • examples/kegg_enrichment.R
  • examples/kegg_spia_topology.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Pathway Kegg Pathways next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Pathway Kegg Pathways compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Pathway Kegg Pathways this skillGPTomics/bioSkills1.2k1 repos~5.4kAutomated safety check: PassMIT
Bio Clinical Databases Clinvar LookupFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~1.4kAutomated safety check: PassNone
Ensembl REST APIwentorai/research-plugins2981 repos~2kAutomated safety check: PassMIT
Ncbi Blast APIwentorai/research-plugins2981 repos~1.6kAutomated safety check: PassMIT
Snpeff Variant Annotationjaechang-hits/SciAgent-Skills3741 repos~5.4kAutomated safety check: PassMIT
Ensembl Databaseaipoch/medical-research-skills1.9k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Bio Clinical Databases Clinvar Lookup

    FreedomIntelligence/OpenClaw-Medical-Skills

    Query ClinVar for variant pathogenicity classifications, review status, and disease associations via REST API or local VCF.

    3.1k GitHub stars~1.4k tokensUpdated 2 mo ago
    Backend & APIsAuto-check passed
  • Ensembl REST API

    wentorai/research-plugins

    Query gene, sequence, and variant data via the Ensembl REST API

    298 GitHub starsUsed in 1 repo~2k tokens
    Backend & APIsAuto-check passed
  • Ncbi Blast API

    wentorai/research-plugins

    Run sequence similarity searches via the NCBI BLAST REST API

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Backend & APIsAuto-check passed
  • Snpeff Variant Annotation

    jaechang-hits/SciAgent-Skills

    Annotate and filter VCF variants with SnpEff and SnpSift. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~5.4k tokens
    Research & ScienceAuto-check passed
  • Ensembl Database

    aipoch/medical-research-skills

    Access Ensembl REST API for vertebrate genomic data; use when you need gene/ID lookups, sequence retrieval, variant effect prediction (VEP), or homology/assembly coordinate mapping.

    1.9k GitHub stars~1.5k tokensUpdated 23 days ago
    Research & ScienceAuto-check passed
  • Mouse Phenome Database

    jaechang-hits/SciAgent-Skills

    Retrieve mouse phenotype data from the Jackson Laboratory Mouse Phenome Database (MPD) via its REST API.

    374 GitHub starsUsed in 1 repo~6.8k tokens
    Knowledge ManagementAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Categories

Questions about Bio Pathway Kegg Pathways

What does Bio Pathway Kegg Pathways do?

Tests gene lists, ranked vectors, and fold-change vectors against KEGG pathways and modules with clusterProfiler enrichKEGG/enrichMKEGG (ORA), gseKEGG (GSEA), and SPIA/graphite (signed-topology…. Bio Pathway Kegg Pathways is an agent skill from GPTomics/bioSkills. Tests gene lists, ranked vectors, and fold-change vectors against KEGG pathways and modules with clusterProfiler enrichKEGG/enrichMKEGG (ORA), gseKEGG (GSEA), and SPIA/graphite (signed-topology perturbation) in R.

When should I use Bio Pathway Kegg Pathways?

Bio Pathway Kegg Pathways fits situations like: finding enriched KEGG pathways; scoring signed pathway perturbation; analyzing prokaryotes; non-model organisms via locus tags.

How do I install Bio Pathway Kegg Pathways in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-pathway-kegg-pathways -a claude-code`. Or copy the skill folder (pathway-analysis/kegg-pathways in GPTomics/bioSkills) into .claude/skills/bio-pathway-kegg-pathways in your project. Claude Code loads it when a task matches its description.

How do I install Bio Pathway Kegg Pathways in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-pathway-kegg-pathways -a codex`. Or copy the skill folder (pathway-analysis/kegg-pathways in GPTomics/bioSkills) into .agents/skills/bio-pathway-kegg-pathways in your project. Codex loads it when a task matches its description.

Can I use Bio Pathway Kegg Pathways in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-pathway-kegg-pathways -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-pathway-kegg-pathways, .gemini/skills/bio-pathway-kegg-pathways, .github/skills/bio-pathway-kegg-pathways and .opencode/skills/bio-pathway-kegg-pathways in your project.

What does Bio Pathway Kegg Pathways need to run?

Going by SKILL.md and its folder, Bio Pathway Kegg Pathways needs R for the scripts in its folder.

Does Bio Pathway Kegg Pathways access the network?

SKILL.md names 1 domain. As links in the text: rest.kegg.jp. This is read from the text; nothing was executed.

Is Bio Pathway Kegg Pathways safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Pathway Kegg Pathways use?

Bio Pathway Kegg Pathways is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Pathway Kegg Pathways use?

About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Pathway Kegg Pathways?

Skills that share tags, products or a category with Bio Pathway Kegg Pathways: Bio Clinical Databases Clinvar Lookup (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Ensembl REST API (wentorai/research-plugins, 298 stars), Ncbi Blast API (wentorai/research-plugins, 298 stars) and Snpeff Variant Annotation (jaechang-hits/SciAgent-Skills, 374 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Pathway Kegg Pathways?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.