Hypothesis Generation
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
Guide to KEGG pathway enrichment for DEG results. An agent skill from jaechang-hits/SciAgent-Skills.
$ npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills kegg-pathway-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/systems-biology-multiomics/kegg-pathway-analysis .claude/skills/kegg-pathway-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "kegg-pathway-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/kegg-pathway-analysis into .claude/skills/kegg-pathway-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kegg-pathway-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/kegg-pathway-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills kegg-pathway-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/systems-biology-multiomics/kegg-pathway-analysis .agents/skills/kegg-pathway-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "kegg-pathway-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/kegg-pathway-analysis into .agents/skills/kegg-pathway-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kegg-pathway-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills kegg-pathway-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/systems-biology-multiomics/kegg-pathway-analysis .cursor/skills/kegg-pathway-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "kegg-pathway-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/kegg-pathway-analysis into .cursor/skills/kegg-pathway-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kegg-pathway-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/systems-biology-multiomics/kegg-pathway-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills kegg-pathway-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/systems-biology-multiomics/kegg-pathway-analysis .gemini/skills/kegg-pathway-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "kegg-pathway-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/kegg-pathway-analysis into .gemini/skills/kegg-pathway-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kegg-pathway-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills kegg-pathway-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/systems-biology-multiomics/kegg-pathway-analysis .github/skills/kegg-pathway-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "kegg-pathway-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/kegg-pathway-analysis into .github/skills/kegg-pathway-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kegg-pathway-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills kegg-pathway-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/systems-biology-multiomics/kegg-pathway-analysis .opencode/skills/kegg-pathway-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "kegg-pathway-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/kegg-pathway-analysis into .opencode/skills/kegg-pathway-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kegg-pathway-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
kegg-pathway-analysisGuide to KEGG pathway enrichment for DEG results. An agent skill from jaechang-hits/SciAgent-Skills.
Kegg Pathway Analysis is an agent skill from jaechang-hits/SciAgent-Skills. Guide to KEGG pathway enrichment for DEG results. Covers ORA vs GSEA, mandatory directionality splitting, KEGG organism codes, API failure handling with offline fallbacks, cross-condition comparisons, and answer-first reporting. Consult when running enrichment with clusterProfiler or gseapy.
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are r, python and yaml).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
rest.kegg.jpAlso links to:
kegg.jpdoi.orggseapy.readthedocs.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Kegg Pathway Analysis loads about 4.8k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 1,663 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 1,663 words, ~4,842 tokens.
.claude/skills/kegg-pathway-analysis/SKILL.md (or your agent's skills folder).KEGG (Kyoto Encyclopedia of Genes and Genomes) pathway enrichment analysis identifies biological pathways that are statistically over-represented among differentially expressed genes. This guide covers the two main enrichment approaches (ORA and GSEA), critical workflow decisions such as splitting genes by directionality, tool selection between R clusterProfiler and Python gseapy, and strategies for handling the notoriously unreliable KEGG REST API. It addresses recurring failure modes that produce incorrect pathway counts or stalled analyses.
The three most common errors in KEGG pathway analysis are: (1) combining up-regulated and down-regulated genes into a single enrichment run, which masks true pathway signals; (2) analysis failures caused by KEGG REST API timeouts with no fallback strategy; and (3) delaying result reporting while attempting cosmetic pathway name lookups that may never complete. This guide provides concrete solutions for each.
Over-Representation Analysis (ORA) and Gene Set Enrichment Analysis (GSEA) are the two primary methods for pathway enrichment, and they differ in both input and statistical approach.
ORA takes a pre-filtered gene list (e.g., genes with padj < 0.05 and |log2FC| > 1.5) and tests whether KEGG pathway members are over-represented in that list relative to a background universe. ORA uses a hypergeometric test (Fisher's exact test). It is straightforward but discards magnitude information and depends heavily on the significance cutoff chosen.
GSEA takes a ranked list of all genes (typically ranked by log2 fold change or a signed significance statistic) without any cutoff. It computes a running enrichment score by walking down the ranked list and identifies pathways whose members cluster toward the top or bottom of the ranking. GSEA captures subtle coordinated changes that ORA may miss.
In practice, ORA via enrichKEGG() (clusterProfiler) or gp.enrichr() (gseapy) is the more common starting point. GSEA via gseKEGG() or gp.prerank() is preferred when you want to avoid arbitrary cutoffs or when effect sizes are small.
When performing ORA, gene directionality -- whether a gene is up-regulated or down-regulated -- is critical. A single pathway can contain genes regulated in opposite directions. If up-regulated and down-regulated genes are combined into one list, their opposing signals cancel out, diluting the enrichment signal and masking genuinely enriched pathways. Running enrichment separately for up-regulated and down-regulated gene sets produces more accurate and interpretable results. This splitting is mandatory for ORA. GSEA inherently handles directionality through the signed ranking, though interpreting leading-edge genes by direction is still important.
KEGG uses three-letter (or four-letter) organism codes to identify species-specific pathway databases. Using the wrong code silently returns empty results. Common codes:
| Organism | Code |
|---|---|
| Human | hsa |
| Mouse | mmu |
| Rat | rno |
| Zebrafish | dre |
| Drosophila | dme |
| C. elegans | cel |
| E. coli K-12 | eco |
| P. aeruginosa PA14 | pau |
| P. aeruginosa PAO1 | pae |
| S. cerevisiae | sce |
| A. thaliana | ath |
Gene ID format also varies by organism: eukaryotic species typically require Entrez gene IDs, while bacterial species use locus tags. Mismatched ID types are a silent failure mode.
The KEGG REST API (rest.kegg.jp) is rate-limited, frequently slow, and prone to timeouts. Both clusterProfiler::enrichKEGG() and direct HTTP requests to KEGG can fail unpredictably. Planning for API failures is not optional -- it is a necessary part of any KEGG-based workflow. Strategies include pre-fetching and caching pathway data, using offline gene set databases bundled with gseapy, and implementing retry logic with timeouts.
Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: organismCode
kind: required
source: upstream
ask: "Which species should pathways be looked up for?"
default: "carried from the dataset"
- id: D2
param: identifierMapping
kind: required
source: data
depends_on: [D1]
ask: "Which identifier type do the genes carry, and how should unmapped ones be handled?"
default: null
- id: D3
param: testingApproach
kind: required
source: upstream
ask: "Test a significant-gene list against pathways (over-representation), or the whole ranked list without a cut (enrichment)?"
default: "over-representation when the input is a gene list"
- id: D4
param: backgroundSet
kind: required
source: upstream
depends_on: [D3]
ask: "Which genes form the background - everything measured in this experiment, or the whole annotated genome?"
default: "the genes tested in this experiment"
skip_if: "ranked-list enrichment, which uses the ranking rather than a background"
- id: D5
param: pathwayScope
kind: optional
source: user
ask: "Should all pathway categories be tested, or only metabolism, signalling, or disease maps?"
default: "all categories"
- id: D6
param: significanceThreshold
kind: required
source: user
ask: "How strong must a pathway's evidence be, after correcting for the number of pathways tested?"
default: "adjusted p-value 0.05"D4 is the decision most often skipped and the one that most changes the answer. Using the whole genome as background when the experiment only measured expressed genes makes every tissue-specific pathway look enriched - the enrichment is against genes that were never measurable, not against the experiment.
Question: What enrichment analysis do you need?
|
+-- Have a pre-filtered DEG list (with cutoffs applied)?
| +-- Yes --> ORA
| | +-- Using R? --> clusterProfiler::enrichKEGG()
| | +-- Using Python? --> gseapy.enrichr()
| | +-- KEGG API failing? --> gseapy with offline gene sets
| +-- No, want cutoff-free analysis --> GSEA
| +-- Using R? --> clusterProfiler::gseKEGG()
| +-- Using Python? --> gseapy.prerank()
|
+-- Need to split by direction?
| +-- ORA --> YES, always split up/down (mandatory)
| +-- GSEA --> No split needed (direction encoded in ranking)
|
+-- KEGG API unreliable?
+-- Try cached/pre-fetched data first
+-- Fall back to gseapy offline databases
+-- Use retry logic with short timeouts| Scenario | Recommended Approach | Rationale |
|---|---|---|
| Standard ORA with R | clusterProfiler::enrichKEGG(), split by direction | Most widely used, integrates with Bioconductor ecosystem |
| Standard ORA with Python | gseapy.enrichr() with KEGG_2021_Human | Offline gene sets avoid API dependency |
| Cutoff-free enrichment | GSEA via gseKEGG() or gp.prerank() | Captures subtle coordinated changes, no arbitrary threshold |
| KEGG API is down | Switch to gseapy offline databases | gseapy bundles KEGG gene sets locally |
| Comparing conditions | Run separate up/down enrichment per condition | Enables direction-aware set operations across conditions |
| Non-model organism | Verify organism code, use KEGGREST to check availability | Wrong code silently returns empty results |
Always split ORA by gene direction. Run enrichKEGG() or gp.enrichr() separately for up-regulated and down-regulated genes. Combining them inflates the gene list, dilutes enrichment signal, and produces incorrect pathway counts. Report the union of significant pathways from both directions.
Specify the background universe explicitly. Set the universe to all tested genes (the full set from your differential expression analysis), not just the significant ones. Omitting the universe defaults to all genes in the KEGG database, which inflates significance for well-studied pathways.
Pre-fetch and cache KEGG data before running enrichment. Download pathway-gene mappings at the start of the analysis and save them locally. This avoids mid-analysis failures when the KEGG API becomes unresponsive and makes the analysis reproducible.
Report the numeric answer before resolving pathway names. Once you have computed the count or list of significant pathway IDs, emit that result immediately. Resolving IDs to human-readable names via additional KEGG API calls is cosmetic and can timeout, losing the primary result.
Apply multiple testing correction consistently. Use adjusted p-values (p.adjust < 0.05, typically BH method) rather than raw p-values. Both clusterProfiler and gseapy apply correction by default, but always verify the cutoff is on the adjusted value.
Verify gene ID format matches the organism. Eukaryotic KEGG pathways expect Entrez gene IDs; bacterial species expect locus tags. A mismatch silently returns zero enriched pathways. Use bitr() in clusterProfiler or equivalent ID conversion if your input uses gene symbols.
Use gseapy as a fallback when clusterProfiler fails. When enrichKEGG() fails due to KEGG API issues, gseapy's enrichr() function with bundled offline gene sets (e.g., KEGG_2021_Human) provides equivalent ORA results without any network dependency.
Combining up-regulated and down-regulated genes into a single enrichment run. Pathways with genes regulated in opposite directions cancel out, producing fewer significant pathways than the true count. Results from combined lists are unreliable. Example of the anti-pattern:
# WRONG: combining up and down genes into one list
all_sig_genes <- rownames(subset(res, padj < 0.05 & abs(log2FoldChange) > 1.5))
ekegg <- enrichKEGG(gene = all_sig_genes, ...) # Will miss pathwaysUsing the wrong KEGG organism code. KEGG silently returns empty results for invalid or mismatched organism codes. This is especially common for bacterial species with multiple strain-specific codes (e.g., pae for PAO1 vs pau for PA14).
Gene ID type mismatch. Providing gene symbols when KEGG expects Entrez IDs (or locus tags for bacteria) silently yields zero enriched pathways with no error message.
clusterProfiler::bitr() or equivalent to convert gene symbols to Entrez IDs before enrichment.Not handling KEGG API timeouts. The KEGG REST API frequently times out, causing enrichKEGG() to fail mid-analysis. Without error handling, the entire analysis is lost.
Delaying the answer to resolve pathway names. Calling keggGet() to convert pathway IDs to human-readable names after computing results can timeout, losing the numeric answer entirely. Example of the anti-pattern:
# BAD: answer delayed by name lookup that may hang
result_ids <- setdiff(pathways_condA, pathways_condB)
names <- keggGet(result_ids) # Can timeout -- answer never emitted
count <- length(result_ids)tryCatch() or try/except block.Omitting the background universe. Not specifying the universe parameter defaults to the entire KEGG gene database for that organism, inflating statistical significance for pathways containing well-annotated housekeeping genes.
universe = rownames(res) (all tested genes from the DE analysis) to enrichKEGG().Using raw p-values instead of adjusted p-values for filtering. Reporting pathways with p < 0.05 without multiple testing correction dramatically increases false positives.
p.adjust < 0.05. Verify that the column you are filtering is the adjusted value, not the raw p-value.Step 1: Prepare gene lists
library(clusterProfiler)
# Filter significant DEGs
sig_genes <- subset(res, padj < 0.05 & abs(log2FoldChange) > 1.5)
# MANDATORY: Split by direction BEFORE running enrichment
up_genes <- rownames(subset(sig_genes, log2FoldChange > 0))
dn_genes <- rownames(subset(sig_genes, log2FoldChange < 0))
cat("Up-regulated genes:", length(up_genes), "\n")
cat("Down-regulated genes:", length(dn_genes), "\n")Step 2: Pre-fetch KEGG data (recommended)
R (KEGGREST package):
library(KEGGREST)
pathway_list <- tryCatch(
keggList("pathway", organism_code),
error = function(e) NULL
)Python (requests with caching):
import requests
import json
import os
cache_file = f"kegg_{organism_code}_pathways.json"
if os.path.exists(cache_file):
with open(cache_file) as f:
pathway_map = json.load(f)
else:
resp = requests.get(
f"https://rest.kegg.jp/list/pathway/{organism_code}", timeout=30
)
if resp.ok:
pathway_map = dict(
line.split("\t") for line in resp.text.strip().split("\n")
)
with open(cache_file, 'w') as f:
json.dump(pathway_map, f)Step 3: Run enrichment separately for each direction
R (clusterProfiler):
ekegg_up <- enrichKEGG(gene = up_genes, organism = organism_code,
universe = rownames(res), pvalueCutoff = 0.05)
ekegg_dn <- enrichKEGG(gene = dn_genes, organism = organism_code,
universe = rownames(res), pvalueCutoff = 0.05)
up_pathways <- subset(as.data.frame(ekegg_up), p.adjust < 0.05)$ID
dn_pathways <- subset(as.data.frame(ekegg_dn), p.adjust < 0.05)$ID
all_sig_pathways <- union(up_pathways, dn_pathways)
cat("Significant pathways (up):", length(up_pathways), "\n")
cat("Significant pathways (down):", length(dn_pathways), "\n")
cat("Total unique significant pathways:", length(all_sig_pathways), "\n")Python (gseapy):
import gseapy as gp
up_genes = sig_genes[sig_genes['log2FoldChange'] > 0].index.tolist()
dn_genes = sig_genes[sig_genes['log2FoldChange'] < 0].index.tolist()
enr_up = gp.enrichr(gene_list=up_genes, gene_sets='KEGG_2021_Human',
organism='human', outdir=None)
enr_dn = gp.enrichr(gene_list=dn_genes, gene_sets='KEGG_2021_Human',
organism='human', outdir=None)Step 4: Report results immediately
result_ids <- setdiff(pathways_condA, pathways_condB)
count <- length(result_ids)
cat("Answer:", count, "pathways unique to condition A\n")
cat("Pathway IDs:", paste(result_ids, collapse = ", "), "\n")
# Optional: resolve names (non-blocking)
names <- tryCatch(keggGet(result_ids), error = function(e) NULL)Step 5: Cross-condition comparison (if applicable)
# Condition A
up_pathways_A <- subset(as.data.frame(ekegg_up_A), p.adjust < 0.05)$ID
dn_pathways_A <- subset(as.data.frame(ekegg_dn_A), p.adjust < 0.05)$ID
# Condition B
up_pathways_B <- subset(as.data.frame(ekegg_up_B), p.adjust < 0.05)$ID
dn_pathways_B <- subset(as.data.frame(ekegg_dn_B), p.adjust < 0.05)$ID
# Find pathways unique to condition A in each direction
unique_up <- setdiff(up_pathways_A, up_pathways_B)
unique_dn <- setdiff(dn_pathways_A, dn_pathways_B)
# Union
unique_to_A <- union(unique_up, unique_dn)
cat("Pathways in A but not B:", length(unique_to_A), "\n")
cat(" From up-regulated:", length(unique_up), "-",
paste(unique_up, collapse = ", "), "\n")
cat(" From down-regulated:", length(unique_dn), "-",
paste(unique_dn, collapse = ", "), "\n")Step 6: Handle API failures with retry logic
fetch_kegg_with_retry <- function(organism_code, max_retries = 3,
timeout_sec = 30) {
for (i in seq_len(max_retries)) {
result <- tryCatch({
R.utils::withTimeout(
keggList("pathway", organism_code),
timeout = timeout_sec
)
}, error = function(e) NULL)
if (!is.null(result)) return(result)
Sys.sleep(2)
}
warning("KEGG API unreachable after retries. Proceeding without pathway names.")
return(NULL)
}gseapy-gene-enrichment -- Python-based gene set enrichment analysis; use as a fallback when clusterProfiler KEGG API calls fail, or as the primary tool for Python-based workflowsdeseq2-differential-expression / pydeseq2-differential-expression -- Upstream differential expression analysis that produces the DEG lists used as input to KEGG pathway enrichment© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/systems-biology-multiomics/kegg-pathway-analysis of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Kegg Pathway Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Kegg Pathway Analysis this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~4.8k | Automated safety check: Pass | CC-BY-4.0 | |
| Hypothesis Generationspacering-net/codeg | 3.9k | 14 repos | ~3.6k | Automated safety check: Notes | MIT | |
| GitHub Deep Researchbytedance/deer-flow | 84k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Nature Paper CardYuan1z0825/nature-skills | 47k | 2 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Content Research Writerweapp-tailwindcss/weapp-tailwindcss | 1.9k | 25 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Last30daysmvanhorn/last30days-skill | 64k | — | ~7.9k | Automated safety check: Notes | MIT |
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
bytedance/deer-flow
Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.
Yuan1z0825/nature-skills
Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.
weapp-tailwindcss/weapp-tailwindcss
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.
mvanhorn/last30days-skill
Research what people actually say about any topic in the last 30 days.
spacering-net/codeg
Structured manuscript/grant review with checklist-based evaluation.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
Guide to KEGG pathway enrichment for DEG results. An agent skill from jaechang-hits/SciAgent-Skills. Kegg Pathway Analysis is an agent skill from jaechang-hits/SciAgent-Skills. Guide to KEGG pathway enrichment for DEG results.
Kegg Pathway Analysis fits situations like: research & Science work in your project.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a claude-code`. Or copy the skill folder (skills/systems-biology-multiomics/kegg-pathway-analysis in jaechang-hits/SciAgent-Skills) into .claude/skills/kegg-pathway-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a codex`. Or copy the skill folder (skills/systems-biology-multiomics/kegg-pathway-analysis in jaechang-hits/SciAgent-Skills) into .agents/skills/kegg-pathway-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill kegg-pathway-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kegg-pathway-analysis, .gemini/skills/kegg-pathway-analysis, .github/skills/kegg-pathway-analysis and .opencode/skills/kegg-pathway-analysis in your project.
SKILL.md names no scripts, command-line tools or credentials: Kegg Pathway Analysis is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 4 domains. In commands or code: rest.kegg.jp; the agent is likely to contact it when it follows the instructions. As links in the text: kegg.jp, doi.org and gseapy.readthedocs.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Kegg Pathway Analysis is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Kegg Pathway Analysis: Hypothesis Generation (spacering-net/codeg, 3.9k stars), GitHub Deep Research (bytedance/deer-flow, 84k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars) and Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.