Agent skill

Sc De

by TianGzlab in TianGzlab/OmicsClaw

Load when finding marker genes per cluster or comparing condition expression in single-cell RNA-seq.

Apache-2.0Auto-check passedResearch & Science

Install Sc De

skills CLI
$ npx skills add TianGzlab/OmicsClaw --skill sc-de -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TianGzlab/OmicsClaw sc-de --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TianGzlab/OmicsClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/singlecell/scrna/sc-de .claude/skills/sc-de && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sc-de
GitHub stars
161
Token cost
~2.3k tokens
SKILL.md length
1,017 words
Files
11 (incl. references)
Skills in repo
88
Repo updated
First seen
Licence
Apache-2.0

At a glance

Load when finding marker genes per cluster or comparing condition expression in single-cell RNA-seq.

  • Tasks that involve Bioinformatics
  • SKILL.md covers When to use, Use from a step, API and Methods and parameters, plus 5 more sections
  • Runs Python scripts from its folder; calls python

What it does

Sc De is an agent skill from TianGzlab/OmicsClaw. Load when finding marker genes per cluster or comparing condition expression in single-cell RNA-seq. Skip when the data is bulk (use bulkrna-de); spatial (use spatial-de); cluster-only markers without conditions (use sc-markers).

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `_api.py`, `examples/example_step.py` and `references/methodology.md`).

It sits in Research & Science, covering Bioinformatics. It works with Scanpy. The repository describes itself as: Conversational & memory-enabled AI research partner for multi-omics analysis. CLI + Desktop App (installers in Releases). From biological idea to full research paper. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/sc-de”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 90a3bec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sc De loads about 2.3k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 1,017 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TianGzlab/OmicsClaw at commit 90a3bec, republished under its Apache-2.0 licence (© TianGzlab). 1,017 words, ~2,283 tokens.

Download SKILL.mdSave it as .claude/skills/sc-de/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
sc-de
description
Load when finding marker genes per cluster or comparing condition expression in single-cell RNA-seq. Skip when the data is bulk (use bulkrna-de); spatial (use spatial-de); cluster-only markers without conditions (use sc-markers).
trigger
differential expression, marker genes, de analysis, wilcoxon, pseudo-bulk
tags
singlecell, differential-expression, markers, wilcoxon, deseq2

sc-de

When to use

The user has a preprocessed scRNA-seq AnnData and wants to know either (a) which genes mark each cluster (Wilcoxon / t-test / logreg ranking) or (b) which genes change between conditions in a replicate-aware way (pseudobulk_de, DESeq2 in R). Mixing normalized expression and raw counts across these paths is the most common silent-wrong-answer failure mode, so each path reads the matrix it needs.

Use from a step

python
de = load_skill("sc-de")
adata = read_input("results/04_annotation/intermediate/adata_annotated.h5ad")
table = de.rank_genes(adata, groupby="cell_type", method="wilcoxon")
write_output(table, "tables/de_full.csv")
write_output(de.top_genes(table, n_top=10), "tables/markers_top.csv")
write_output(de.volcano_figure(table, group="B cell"), "figures/volcano_b_cell.png")

A complete step that runs on demo data: examples/example_step.py.

API

<!-- api:begin generated from _api.py; regenerate with run.py api <skill dir> --write -->
rank_genes(adata, *, groupby: str='leiden', method: str='wilcoxon', group1: str | None=None, group2: str | None=None, logreg_solver: str='lbfgs') -> pd.DataFrame

Rank genes per group of cells: every group against the rest, or group1 against group2.

wilcoxon, t-test and logreg run scanpy's rank_genes_groups on X (use_raw=False, with the fraction of expressing cells) and also leave its result in uns['rank_genes_groups']; mast runs the MAST hurdle model in R. Cells are the units here, so p-values overstate the evidence for a difference between conditions; use :func:pseudobulk_de for that.

:param groupby: The obs column defining the groups. Default "leiden"; when it is missing and louvain exists, louvain is used and recorded. :param method: "wilcoxon" (default; scanpy's recommended test), "t-test", "logreg" or "mast". :param group1: Compare only this group ... :param group2: ... against this one. Default: each group against the rest. :param logreg_solver: The scikit-learn solver for logreg. Default "lbfgs". :returns: One row per gene and group. scanpy methods: names, scores, logfoldchanges, pvals, pvals_adj, pct_nz_group, pct_nz_reference, group; mast: gene, group, pvalue, padj and effect columns. :raises ValueError: an unknown method, or the group column is missing. :raises RuntimeError: mast and R or MAST is missing.

pseudobulk_de(adata, *, condition_key: str, group1: str, group2: str, sample_key: str='sample_id', celltype_key: str='cell_type', min_cells: int=10, min_counts: int=1000) -> pd.DataFrame

DESeq2 on pseudobulk counts: one test of group1 against group2 per cell type, run in R.

Counts are summed per sample and cell type from layers['counts'], raw or a count-like X; bins below the thresholds are dropped. Samples, not cells, are the units, so this is the test for condition effects.

:param condition_key: The obs column holding the condition. :param group1: The condition of interest (numerator of the fold change). :param group2: The reference condition. :param sample_key: The obs column naming biological replicates. Default "sample_id". :param celltype_key: The obs column with cell types. Default "cell_type". :param min_cells: Minimum cells per sample and cell type bin. Default 10. :param min_counts: Minimum total counts per bin. Default 1000. :returns: Columns gene, log2FoldChange, pvalue, padj (and DESeq2's others) plus cell_type. :raises ValueError: a missing column, missing groups, or no count-like matrix. :raises RuntimeError: R or DESeq2 is missing, or no bin passes the thresholds.

run_info(adata, *, keep: bool=True) -> dict

What the last :func:rank_genes or :func:pseudobulk_de call recorded under summary.

summary has method, groupby (the column actually used), n_groups, n_genes_tested and expression_source.

:param keep: Leave the record in adata.uns; False removes it. :returns: The record, or an empty dict when neither has run on adata.

top_genes(table: pd.DataFrame, *, n_top: int=10) -> pd.DataFrame

The first n_top genes of each group.

A scanpy table is already ranked within each group. A table with padj (MAST) is sorted by padj then pvalue first.

:param table: What :func:rank_genes returned. :param n_top: Genes per group. Default 10, the CLI's default. :returns: The selected rows, in the table's columns.

volcano_figure(table: pd.DataFrame, *, padj_threshold: float=0.05, log2fc_threshold: float=1.0, group: str | None=None)

A volcano plot of a DE table, significant genes coloured.

:param table: What :func:rank_genes or :func:pseudobulk_de returned. :param padj_threshold: Adjusted p-value cut. Default 0.05. :param log2fc_threshold: Absolute log2 fold-change cut. Default 1.0. :param group: Plot only this group (or cell type). Default: all rows. :returns: A matplotlib Figure.

<!-- api:end -->
Show full SKILL.md (426 more words)Show less

Methods and parameters

QuestionFunctionUnitsNeeds
Which genes mark each cluster or cell typerank_genes(method="wilcoxon") (default)cellsnormalised X
Same, other testsrank_genes(method="t-test" / "logreg")cellsnormalised X
Same, hurdle modelrank_genes(method="mast")cellslog-normalised X; R with MAST
Which genes change between conditionspseudobulk_de(...)samplesraw counts; a replicate column; R with DESeq2

Defaults and their sources: method="wilcoxon" is scanpy's recommended marker test; n_top=10 and the volcano cuts (padj_threshold=0.05, log2fc_threshold=1.0) are the CLI's defaults; min_cells=10 and min_counts=1000 per pseudobulk bin are the CLI's defaults, a common choice for 10x data. A comparison between conditions with biological replicates is a pseudobulk question; ask the user for the replicate column when it is not obvious.

Gotchas

  • Cell-level tests are not replicate-aware. Wilcoxon, t-test, logreg and MAST treat each cell as independent, which inflates false positives for treated-versus-control comparisons. Use pseudobulk_de whenever there are biological replicates and the question is about the condition.
  • pseudobulk_de needs raw counts. It takes layers["counts"], then raw, then X, each only if it looks count-like; check run_info(adata)["summary"]["expression_source"] after the run: layers.counts or adata.raw, not adata.X.
  • pseudobulk_de needs group1 and group2. The cell-level tests compare each group with the rest when they are absent.
  • Small bins drop whole cell types. Sample-by-cell-type bins under min_cells cells are skipped without a note; check the per-sample cell counts before reading "no DEGs" as a biological null.
  • sample_key is the statistical design. It must name biological replicates with at least two per condition; a non-replicate column gives nonsense and is not caught.
  • groupby="leiden" falls back to louvain when only louvain exists; run_info(adata)["summary"]["groupby"] says which column was used.

Inputs and outputs

  • rank_genes reads X and obs[groupby], writes uns['rank_genes_groups'] (scanpy methods) and returns the full table. pseudobulk_de reads the count matrix and obs[condition_key], obs[sample_key], obs[celltype_key], and returns one table for all cell types.
  • top_genes returns a DataFrame; volcano_figure returns a matplotlib Figure.

CLI

sc_de.py runs the same functions outside a project and writes a report, figures, tables and processed.h5ad: python <skill directory>/sc_de.py --help. --demo ranks PBMC3k's louvain clusters.

See also

  • references/parameters.md — every CLI flag and per-method tuning hint
  • references/methodology.md — the DE paths, scope boundary, input expectations, workflow
  • references/output_contract.md — the CLI's output directory layout + visualization contract
  • references/r_visualization.md — five R-enhanced renderers
  • Adjacent skills: sc-clustering (upstream cluster discovery), sc-cell-annotation (upstream cell type labels for celltype_key), sc-markers (lighter cluster-marker-only path), sc-enrichment (downstream pathway enrichment of DEG lists), bulkrna-de / spatial-de (sibling DE skills for the other two data modalities)

Dependencies

Python packages this skill's script needs. They are not installed for you — check before a long run.

adjustText, anndata, matplotlib, numpy, pandas, pydeseq2, scanpy, scipy, seaborn

© TianGzlab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/singlecell/scrna/sc-de of TianGzlab/OmicsClaw.

  • SKILL.md
  • _api.py
  • examples/example_step.py
  • references/methodology.md
  • references/output_contract.md
  • references/parameters.md
  • references/r_visualization.md
  • sc_de.py
  • tests/__init__.py
  • tests/test_mast_exchange.py
  • tests/test_sc_de.py

Open the folder on GitHubat commit 90a3bec

Compare with similar skills

Sc De next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sc De compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sc De this skillTianGzlab/OmicsClaw161—~2.3kAutomated safety check: PassApache-2.0
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
Single Cell Rna AnalysisPKU-YuanGroup/OpenAI4S617—~1.3kAutomated safety check: PassMIT
Anndatadavila7/claude-code-templates32k12 repos~2.5kAutomated safety check: PassMIT
Cellxgene Censusdavila7/claude-code-templates32k11 repos~3.8kAutomated safety check: PassMIT
Bulk RnaseqK-Dense-AI/scientific-agent-skills48k1 repos~4.2kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Single Cell Rna Analysis

    PKU-YuanGroup/OpenAI4S

    Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…

    617 GitHub stars~1.3k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    32k GitHub starsUsed in 12 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Cellxgene Census

    davila7/claude-code-templates

    Query CZ CELLxGENE Census (61M+ cells). An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.8k tokens
    Research & ScienceAuto-check passed
  • Bulk Rnaseq

    K-Dense-AI/scientific-agent-skills

    Prepares bulk RNA-seq FASTQ, Salmon, STAR or featureCounts output for gene-level differential expression.

    48k GitHub starsUsed in 1 repo~4.2k tokens
    Research & ScienceAuto-check passed
  • Pathway Enrichment

    K-Dense-AI/scientific-agent-skills

    Performs pathway and gene-set enrichment analysis on gene lists or ranked gene data and interprets the results.

    48k GitHub starsUsed in 1 repo~4.2k tokens
    Research & ScienceAuto-check passed

More from TianGzlab/OmicsClaw

All 88 skills in this repo
  • Bulkrna De

    TianGzlab/OmicsClaw

    Load when comparing gene expression between two conditions in bulk RNA-seq count data.

    161 GitHub starsUsed in 1 repo~867 tokens
    Auto-check passed
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated yesterday
    Auto-check passed
  • Bulkrna Qc

    TianGzlab/OmicsClaw

    Load when checking a bulk RNA-seq count matrix for library-size outliers, gene detection rates, and sample-sample correlation before DE.

    161 GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Sc Fastq Qc

    TianGzlab/OmicsClaw

    Load when checking raw single-cell FASTQ read quality (Phred / GC / adapter / length) before counting.

    161 GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Sc Filter

    TianGzlab/OmicsClaw

    Load when removing low-quality cells and lowly-detected genes from a single-cell AnnData using QC-derived thresholds or tissue presets.

    161 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Sc Markers

    TianGzlab/OmicsClaw

    Load when ranking cluster-level marker genes from a clustered single-cell AnnData via Scanpy Wilcoxon / t-test / logreg or COSG specificity.

    161 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Works with

Questions about Sc De

What does Sc De do?

Load when finding marker genes per cluster or comparing condition expression in single-cell RNA-seq. Sc De is an agent skill from TianGzlab/OmicsClaw. Load when finding marker genes per cluster or comparing condition expression in single-cell RNA-seq.

When should I use Sc De?

Sc De fits situations like: tasks that involve Bioinformatics.

How do I install Sc De in Claude Code?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-de -a claude-code`. Or copy the skill folder (skills/singlecell/scrna/sc-de in TianGzlab/OmicsClaw) into .claude/skills/sc-de in your project. Claude Code loads it when a task matches its description.

How do I install Sc De in Codex?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-de -a codex`. Or copy the skill folder (skills/singlecell/scrna/sc-de in TianGzlab/OmicsClaw) into .agents/skills/sc-de in your project. Codex loads it when a task matches its description.

Can I use Sc De in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TianGzlab/OmicsClaw --skill sc-de -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sc-de, .gemini/skills/sc-de, .github/skills/sc-de and .opencode/skills/sc-de in your project.

What does Sc De need to run?

Going by SKILL.md and its folder, Sc De needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Sc De access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sc De safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sc De use?

Sc De is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sc De use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Sc De?

Skills that share tags, products or a category with Sc De: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), Single Cell Rna Analysis (PKU-YuanGroup/OpenAI4S, 617 stars), Anndata (davila7/claude-code-templates, 32k stars) and Cellxgene Census (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sc De?

TianGzlab (a GitHub organization) maintains it in TianGzlab/OmicsClaw, which has 161 GitHub stars. The repository holds 88 skills in this directory. The repository was last updated on October 7, 2026.

Source: TianGzlab/OmicsClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.