Agent skill

Sc Enrichment

by TianGzlab in TianGzlab/OmicsClaw

Load when running bulk-style pathway enrichment (ORA / GSEA / GSEA-R / GSVA-R) on a per-group ranked DE / marker list against a gene-set library.

Apache-2.0Auto-check passedData & Analytics

Install Sc Enrichment

skills CLI
$ npx skills add TianGzlab/OmicsClaw --skill sc-enrichment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TianGzlab/OmicsClaw sc-enrichment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TianGzlab/OmicsClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/singlecell/scrna/sc-enrichment .claude/skills/sc-enrichment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sc-enrichment
GitHub stars
161
Token cost
~2.4k tokens
SKILL.md length
999 words
Files
13 (incl. references)
Skills in repo
88
Repo updated
First seen
Licence
Apache-2.0

At a glance

Load when running bulk-style pathway enrichment (ORA / GSEA / GSEA-R / GSVA-R) on a per-group ranked DE / marker list against a gene-set library.

  • Tasks that involve Bioinformatics
  • SKILL.md covers When to use, Steps, Inputs & Outputs and Gotchas, plus 4 more sections
  • Runs Python and R scripts from its folder; calls python

What it does

Sc Enrichment is an agent skill from TianGzlab/OmicsClaw. Load when running bulk-style pathway enrichment (ORA / GSEA / GSEA-R / GSVA-R) on a per-group ranked DE / marker list against a gene-set library. Skip when computing per-cell pathway scores in-place (use sc-pathway-scoring); de-novo gene-program discovery (use sc-gene-programs).

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including reference files (for example `_api.py`, `examples/example_step.py` and `references/methodology.md`).

It sits in Data & Analytics, covering Bioinformatics. It works with Python. The repository describes itself as: Conversational & memory-enabled AI research partner for multi-omics analysis. CLI + Desktop App (installers in Releases). From biological idea to full research paper. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/sc-enrichment”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 90a3bec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and R), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sc Enrichment loads about 2.4k tokens when it runs, and up to ~5.5k if it reads all its reference files. Until then it costs about 73 tokens; SKILL.md has 999 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TianGzlab/OmicsClaw at commit 90a3bec, republished under its Apache-2.0 licence (© TianGzlab). 999 words, ~2,354 tokens.

Download SKILL.mdSave it as .claude/skills/sc-enrichment/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
sc-enrichment
description
Load when running bulk-style pathway enrichment (ORA / GSEA / GSEA-R / GSVA-R) on a per-group ranked DE / marker list against a gene-set library. Skip when computing per-cell pathway scores in-place (use sc-pathway-scoring); de-novo gene-program discovery (use sc-gene-programs).
trigger
sc enrichment, single-cell enrichment, GO enrichment, KEGG enrichment, GSEA, ORA, pathway enrichment
tags
singlecell, scrna, enrichment, gsea, ora, gsva, decoupler, pathway-enrichment

sc-enrichment

When to use

Test named gene sets against per-group marker/DE rankings with ORA or GSEA. GSVA scores mean expression per group in R. For scores per cell use sc-pathway-scoring; for programmes discovered without gene sets use sc-gene-programs.

Steps

Start a step with python skills/_sdk/notebook/run.py new enrichment. Load enrichment = load_skill("sc-enrichment"), then call rank_groups, load_gene_sets (GMT, JSON or Enrichr library), and ora or gsea. Pass the complete tested gene universe as background for ORA; its API default is all genes in the supplied ranking. Write returned tables and Figures with write_output. See examples/example_step.py.

The default engine is Python. Explicit engine="auto" prefers R when its packages are available; engine="r" requires clusterProfiler. The legacy gsea_r method uses GO_BP/KEGG/Reactome annotation databases, not custom GMT. gsva accepts a gene-set mapping, or None for the GO_BP/KEGG R database route.

Inputs & Outputs

Rankings need gene/names and a score or effect, with an optional group column. rank_groups reads log-normalised X, never raw. For the processed PBMC demo, use .raw.to_adata() in new steps; the CLI retains its old scaled-X demo ranking for compatibility. The CLI accepts H5AD or an upstream output directory containing processed.h5ad and marker/DE tables.

ORA/GSEA CLI writes processed.h5ad, report.md, result.json, reproducibility/, and five tables: enrichment_results.csv, enrichment_significant.csv, group_summary.csv, ranking_input.csv, top_terms.csv. Python GSEA adds tables/gsea_running_scores.csv when running curves can be built. figure_data/ mirrors plot tables. Figures depend on available terms. GSVA uses tables/gsva_r_scores.csv and figures/gsva_r_heatmap.png instead. See references/output_contract.md for R-specific files.

Gotchas

  • result.json.summary.resolved_engine records the engine actually used. The default is now python; ask for --engine auto to retain auto-selection.
  • tables/enrichment_results.csv records Python GSEA's local fallback in engine, with the reason in result.json.summary.warnings. This is not interchangeable with gseapy's permutation implementation.
  • tables/enrichment_significant.csv retains the legacy fixed FDR 0.05 filter. --fdr-threshold controls group_summary.csv and the report summary.
  • _api.py:37: rank_groups does not infer a suitable matrix. Using scaled X can yield invalid log-fold changes; use a log-normalised snapshot instead.
  • _api.py:103: empty gene sets fail validation. Failed R annotation mapping raises an error; the R scripts no longer manufacture pathways to fill an empty result.
  • _api.py:205: GSVA's R database route supports GO_BP and KEGG, not Enrichr's Hallmark alias. For Hallmark, explicitly load gene sets and pass the mapping to gsva.

Key CLI

bash
python skills/singlecell/scrna/sc-enrichment/sc_enrichment.py --demo --output /tmp/sc_enrich_demo
python skills/singlecell/scrna/sc-enrichment/sc_enrichment.py --input clustered.h5ad --output results/ora --gene-set-db hallmark --groupby cell_type
python skills/singlecell/scrna/sc-enrichment/sc_enrichment.py --input clustered.h5ad --output results/gsea --method gsea --gene-sets pathways.gmt --gsea-seed 123
python skills/singlecell/scrna/sc-enrichment/sc_enrichment.py --input clustered.h5ad --output results/gsva --method gsva_r --groupby cell_type --gene-sets pathways.gmt

API

<!-- api:begin generated from _api.py; regenerate with run.py api <skill dir> --write -->
rank_groups(adata, *, groupby: str, method: str='wilcoxon') -> pd.DataFrame

Rank each group's genes against the rest using log-normalised X.

method is wilcoxon (default), t-test or logreg. This uses X, not raw; convert a scaled object with adata.raw.to_adata() first when raw holds log-normalised values. Scanpy ranking statistics remain in uns. Return a table with group, gene and available scores/p-values.

load_gene_sets(source, *, species: str='human', universe=None) -> dict[str, list[str]]

Load GMT/JSON or fetch an Enrichr library; optionally match genes to a universe.

source accepts a local path or hallmark/kegg/reactome/go_bp aliases and full Enrichr library names. Remote sources need network access and gseapy; local files do not. species is human (default) or mouse. Return a term-to-gene-list mapping; no output files are written.

demo_gene_sets(*, species: str='human') -> dict[str, list[str]]

Return the six named PBMC demo signatures, human by default or mouse symbols.

marker_gene_sets(markers: pd.DataFrame, *, groups=None, top_n: int | None=100, universe=None) -> dict[str, list[str]]

Convert marker groups to gene sets, sorted by adjusted p-value or score.

markers needs group and names (or gene). groups=None selects all; top_n=None keeps all genes. universe optionally restricts and canonicalises gene symbols. Raise ValueError when no sets remain.

ora(ranking: pd.DataFrame, gene_sets, *, background=None, engine: str='python', padj_cutoff: float=0.05, log2fc_cutoff: float=0.25, max_genes: int=200, source: str='custom', library_mode: str='local', n_top: int=18) -> pd.DataFrame

Test over-representation with a local hypergeometric test or clusterProfiler.

ranking accepts marker/DE columns and optional group (default all). background=None uses every ranked gene; pass all tested genes to make the universe explicit. Positive genes passing adjusted p-value <= 0.05 and log2FC >= 0.25 are kept, up to 200 per group, by default. Score-only rankings keep positive scores. engine is python (default), r or auto (R when available). source and library_mode label result rows; n_top controls R's optional plots. Return a sorted enrichment table; :func:run_info returns warnings and the resolved engine.

Show full SKILL.md (352 more words)Show less
gsea(ranking: pd.DataFrame, gene_sets, *, engine: str='python', ranking_metric: str='auto', min_size: int=5, max_size: int=500, permutation_num: int=100, weight: float=1.0, random_state: int=123, source: str='custom', library_mode: str='local', n_top: int=18, species: str='human', gene_set_db: str | None=None) -> pd.DataFrame

Run preranked GSEA and return sorted terms with NES, p-values and leading edges.

engine is python (default), r (clusterProfiler GMT), auto, or gsea_r (the retained annotation-database R bridge; pass gene_sets=None and select GO_BP, KEGG or Reactome with gene_set_db). Python uses gseapy when available and reports its existing local rank-based fallback. ranking_metric='auto' chooses stat, scores or logfoldchanges. Defaults: gene-set sizes 5..500, 100 permutations, weight 1.0, seed 123. Permutation count and weight configure Python; the retained R bridge uses fgsea's multilevel defaults and exponent 1. source/library_mode label rows; n_top controls R plot selection; species selects human or mouse annotation for gsea_r. Seeds are passed to Python and R. No gene sets are fabricated when input or mapping fails.

gsva(adata, gene_sets, *, groupby: str, species: str='human', gene_set_db: str='GO_BP', method: str='gsva', min_size: int=5, max_size: int=500) -> pd.DataFrame

Score mean log-normalised expression per group with R GSVA, returning a long table.

Supply a term-to-genes mapping, or None for the retained GO_BP/KEGG R annotation bridge selected by gene_set_db and human/mouse species. method is gsva (default), ssgsea or zscore; gene-set size defaults are 5..500. R runs serially in a temporary directory. Missing annotation or gene sets raises an error rather than producing synthetic pathways.

run_info(results: pd.DataFrame, *, keep: bool=True) -> dict

Return a JSON-serialisable diagnostics summary, excluding rankings and R artifacts.

top_terms(results: pd.DataFrame, *, n_top: int=18, per_group: int=3) -> pd.DataFrame

Select up to n_top terms, first reserving per_group rows for each group.

group_summary(enrich_df: pd.DataFrame, *, fdr_threshold: float=0.05) -> pd.DataFrame

Summarise term counts, FDR-significant terms and the top term per group.

top_terms_figure(results: pd.DataFrame, *, n_top: int=18)

Return a horizontal score bar Figure; the caller saves and closes it.

<!-- api:end -->

See also

sc-markers and sc-de provide rankings. marker_gene_sets can convert their tables into a signature library; avoid circular validation against the same cells used to select the markers. CLI flags are in references/parameters.md.

Dependencies

Python packages this skill's script needs. They are not installed for you — check before a long run.

adjustText, anndata, gseapy, matplotlib, networkx, numpy, pandas, scanpy, scipy, seaborn

© TianGzlab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (references) in skills/singlecell/scrna/sc-enrichment of TianGzlab/OmicsClaw.

  • SKILL.md
  • _api.py
  • examples/example_step.py
  • references/methodology.md
  • references/output_contract.md
  • references/parameters.md
  • references/r_visualization.md
  • rscripts/sc_clusterprofiler_enrichment.R
  • sc_enrichment.py
  • tests/test_enrichment_api.py
  • tests/test_r_enrichment_api.py
  • tests/test_sc_enrichment.py
  • tests/test_sc_enrichment_methods.py

Open the folder on GitHubat commit 90a3bec

Compare with similar skills

Sc Enrichment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sc Enrichment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sc Enrichment this skillTianGzlab/OmicsClaw161—~2.4kAutomated safety check: PassApache-2.0
Pyopenmsdavila7/claude-code-templates33k11 repos~1.4kAutomated safety check: PassMIT
Pydeseqaipoch/medical-research-skills1.9k—~1.8kAutomated safety check: PassMIT
Pyopenms Skillaipoch/medical-research-skills1.9k—~852Automated safety check: PassMIT
Bio Metagenomics VisualizationGPTomics/bioSkills1.2k1 repos~3.7kAutomated safety check: PassMIT
Bio Proteomics Differential AbundanceGPTomics/bioSkills1.2k1 repos~5.7kAutomated safety check: PassMIT

Similar skills

  • Pyopenms

    davila7/claude-code-templates

    Python interface to OpenMS for mass spectrometry data analysis.

    33k GitHub starsUsed in 11 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Pydeseq

    aipoch/medical-research-skills

    Differential gene expression analysis for bulk RNA-seq count matrices using a DESeq2-like workflow in Python; use when you need Wald tests, FDR correction, and optional LFC shrinkage for…

    1.9k GitHub stars~1.8k tokensUpdated 24 days ago
    Data & AnalyticsAuto-check passed
  • Pyopenms Skill

    aipoch/medical-research-skills

    Comprehensive tool for computational mass spectrometry using PyOpenMS; use when you need to read/write MS formats (mzML/mzXML/MGF), run signal processing (smoothing/peak picking), detect isotope…

    1.9k GitHub stars~852 tokensUpdated 24 days ago
    Data & AnalyticsAuto-check passed
  • Turns a shotgun profiler table (MetaPhlAn relative abundance, Bracken counts, HUMAnN function tables) into honest figures and defensible community statistics with phyloseq, vegan, microViz, and…

    1.2k GitHub starsUsed in 1 repo~3.7k tokens
    Data & AnalyticsAuto-check passed
  • Tests for differentially abundant proteins between conditions with limma/DEqMS empirical-Bayes moderation, proDA/msqrob2/MSstats missingness modeling, and Python Welch+BH alternatives.

    1.2k GitHub starsUsed in 1 repo~5.7k tokens
    Data & AnalyticsAuto-check passed
  • Scikit Bio

    aipoch/medical-research-skills

    A Python bioinformatics toolkit for sequence, phylogeny, and microbiome/community-ecology analysis; use it when you need to compute diversity/ordination/statistics from biological data and standard…

    1.9k GitHub stars~1.4k tokensUpdated 24 days ago
    Data & AnalyticsAuto-check passed

More from TianGzlab/OmicsClaw

All 88 skills in this repo
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna Batch Correction

    TianGzlab/OmicsClaw

    Load when correcting batch effects in bulk expression using R sva ComBat or the legacy Python parametric approximation.

    161 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna Coexpression

    TianGzlab/OmicsClaw

    Load when discovering bulk gene co-expression modules and hub genes with R WGCNA.

    161 GitHub stars~1.3k tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna De

    TianGzlab/OmicsClaw

    Load when comparing gene expression between two conditions in bulk RNA-seq count data.

    161 GitHub stars~867 tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna Deconvolution

    TianGzlab/OmicsClaw

    Load when estimating cell-type proportions in bulk RNA-seq samples from a single-cell or signature-matrix reference.

    161 GitHub stars~757 tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna Enrichment

    TianGzlab/OmicsClaw

    Load when running pathway / GO term enrichment on a bulk RNA-seq DE result list.

    161 GitHub stars~860 tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Sc Enrichment

What does Sc Enrichment do?

Load when running bulk-style pathway enrichment (ORA / GSEA / GSEA-R / GSVA-R) on a per-group ranked DE / marker list against a gene-set library. Sc Enrichment is an agent skill from TianGzlab/OmicsClaw. Load when running bulk-style pathway enrichment (ORA / GSEA / GSEA-R / GSVA-R) on a per-group ranked DE / marker list against a gene-set library.

When should I use Sc Enrichment?

Sc Enrichment fits situations like: tasks that involve Bioinformatics.

How do I install Sc Enrichment in Claude Code?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-enrichment -a claude-code`. Or copy the skill folder (skills/singlecell/scrna/sc-enrichment in TianGzlab/OmicsClaw) into .claude/skills/sc-enrichment in your project. Claude Code loads it when a task matches its description.

How do I install Sc Enrichment in Codex?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-enrichment -a codex`. Or copy the skill folder (skills/singlecell/scrna/sc-enrichment in TianGzlab/OmicsClaw) into .agents/skills/sc-enrichment in your project. Codex loads it when a task matches its description.

Can I use Sc Enrichment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TianGzlab/OmicsClaw --skill sc-enrichment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sc-enrichment, .gemini/skills/sc-enrichment, .github/skills/sc-enrichment and .opencode/skills/sc-enrichment in your project.

What does Sc Enrichment need to run?

Going by SKILL.md and its folder, Sc Enrichment needs Python and R for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Sc Enrichment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sc Enrichment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sc Enrichment use?

Sc Enrichment is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sc Enrichment use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.

What are the alternatives to Sc Enrichment?

Skills that share tags, products or a category with Sc Enrichment: Pyopenms (davila7/claude-code-templates, 33k stars), Pydeseq (aipoch/medical-research-skills, 1.9k stars), Pyopenms Skill (aipoch/medical-research-skills, 1.9k stars) and Bio Metagenomics Visualization (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sc Enrichment?

TianGzlab (a GitHub organization) maintains it in TianGzlab/OmicsClaw, which has 161 GitHub stars. The repository holds 88 skills in this directory. The repository was last updated on October 7, 2026.

Source: TianGzlab/OmicsClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.