Scanpy
K-Dense-AI/scientific-agent-skills
Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…
scRNA-seq with Scanpy: QC, normalization, HVG selection, PCA, neighborhood graph, UMAP/t-SNE, Leiden clustering, markers, cell annotation, trajectory inference.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scanpy-scrna-seq --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq .claude/skills/scanpy-scrna-seq && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scanpy-scrna-seq" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq into .claude/skills/scanpy-scrna-seq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scanpy-scrna-seq", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seqType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scanpy-scrna-seq --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq .agents/skills/scanpy-scrna-seq && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scanpy-scrna-seq" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq into .agents/skills/scanpy-scrna-seq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scanpy-scrna-seq", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scanpy-scrna-seq --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq .cursor/skills/scanpy-scrna-seq && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scanpy-scrna-seq" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq into .cursor/skills/scanpy-scrna-seq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scanpy-scrna-seq", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scanpy-scrna-seq --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq .gemini/skills/scanpy-scrna-seq && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scanpy-scrna-seq" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq into .gemini/skills/scanpy-scrna-seq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scanpy-scrna-seq", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills scanpy-scrna-seqInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq .github/skills/scanpy-scrna-seq && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scanpy-scrna-seq" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq into .github/skills/scanpy-scrna-seq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scanpy-scrna-seq", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills scanpy-scrna-seq --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq .opencode/skills/scanpy-scrna-seq && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scanpy-scrna-seq" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq into .opencode/skills/scanpy-scrna-seq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scanpy-scrna-seq", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scanpy-scrna-seqscRNA-seq with Scanpy: QC, normalization, HVG selection, PCA, neighborhood graph, UMAP/t-SNE, Leiden clustering, markers, cell annotation, trajectory inference.
Scanpy Scrna Seq is an agent skill from jaechang-hits/SciAgent-Skills. scRNA-seq with Scanpy: QC, normalization, HVG selection, PCA, neighborhood graph, UMAP/t-SNE, Leiden clustering, markers, cell annotation, trajectory inference. Standard scRNA-seq exploration.
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/api_reference.md`, `references/plotting_guide.md` and `references/standard_workflow.md`).
It sits in Research & Science, covering Bioinformatics. It works with Scanpy, UMAP and Python. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
doi.orgscanpy.readthedocs.ioscanpy-tutorials.readthedocs.iotraining.galaxyproject.orgscverse.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scanpy Scrna Seq loads about 4.7k tokens when it runs, and up to ~9.6k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 1,023 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 1,023 words, ~4,719 tokens.
.claude/skills/scanpy-scrna-seq/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Scanpy is a scalable Python toolkit for analyzing single-cell RNA-seq data built on the AnnData format. This skill covers the end-to-end standard workflow: quality control, normalization, highly variable gene selection, dimensionality reduction, clustering, marker gene identification, and cell type annotation. It produces annotated datasets and publication-quality visualizations.
sc.pl.*scanpy>=1.10, leidenalg, igraph, anndatafiltered_feature_bc_matrix/), .h5ad, .h5, .csvpip install "scanpy[leiden]" anndataSettle these with the user before writing any analysis code.
decisions:
- id: D1
param: pct_counts_mt
kind: required
source: data
ask: "Dying cells leak cytoplasmic RNA and read as high-mitochondrial - where should the cut sit for this tissue?"
default: "20% (PBMC ~5%, solid tumour ~20%; propose from the observed distribution)"
- id: D2
param: batch_key
kind: required
source: data
ask: "Which variable separates samples that were processed apart (donor, run, 10X lane), so its effect is not read as biology?"
default: null
skip_if: "all cells come from one sample and one processing run"
- id: D3
param: resolution
kind: optional
source: user
ask: "How finely should cells be split - broad lineages, or subtypes within them?"
default: "0.8 (0.1-0.3 broad lineages, 1.0-2.0 fine subtypes)"
- id: D4
param: n_top_genes
kind: optional
source: user
ask: "How many variable genes should drive the embedding? More captures subtle states but adds noise."
default: 2000
- id: D5
param: n_pcs
kind: optional_conditional
source: data
depends_on: [D4]
ask: "How many principal components carry real structure, by the variance-ratio elbow?"
default: 40
- id: D6
param: min_genes, min_cells
kind: optional_conditional
source: data
ask: "Where should empty droplets and undetected genes be cut off?"
default: "cells with <200 genes, genes in <3 cells"
- id: D6b
param: annotationStrategy
kind: required
source: user
ask: "After clustering, should clusters be named from canonical markers by hand, or handed to a reference-based annotator?"
default: "manual marker-based within this skill; see `single-cell-annotation-guide` to choose"
- id: D6c
param: markerPanel
kind: required
source: literature
depends_on: [D6b]
ask: "Which canonical markers define the populations expected in this tissue and state?"
default: null
skip_if: "annotation delegated to a reference-based tool"
- id: D7
param: method (rank_genes_groups)
kind: optional
source: user
ask: "Which test should rank marker genes per cluster?"
default: "wilcoxon"
- id: D8
param: target_sum, flavor
kind: optional
source: user
ask: "Should the standard library-size normalization and Seurat v3 highly-variable-gene selection be used?"
default: "target_sum 1e4, flavor seurat_v3"
- id: D9
param: random_state
kind: never_ask
source: data
reason: "Fixes the draw for reproducibility; does not change what the data supports"
default: 0D5 hangs on D4 because the elbow is read off a PCA computed over the selected
variable genes — a different gene count moves where it sits. D2 shapes the
pipeline rather than a single call: naming a batch adds a correction stage
(harmony-batch-correction or ComBat) before the neighbor graph.
Configure scanpy settings and read the count matrix into an AnnData object. Ensure gene names are unique.
import scanpy as sc
import numpy as np
import pandas as pd
# Configure settings
sc.settings.verbosity = 3 # errors=0, warnings=1, info=2, hints=3
sc.settings.set_figure_params(dpi=80, facecolor="white")
# Load data — pick the reader matching your format
adata = sc.read_10x_mtx(
"path/to/filtered_feature_bc_matrix/",
var_names="gene_symbols",
cache=True,
)
# Alternatives:
# adata = sc.read_h5ad("data.h5ad")
# adata = sc.read_10x_h5("data.h5")
adata.var_names_make_unique()
print(f"Loaded: {adata.n_obs} cells x {adata.n_vars} genes")
print(f"Sparsity: {1 - adata.X.nnz / (adata.n_obs * adata.n_vars):.1%}")Annotate mitochondrial/ribosomal genes, compute QC metrics, visualize distributions, and filter low-quality cells.
# Annotate gene groups
adata.var["mt"] = adata.var_names.str.startswith("MT-") # human; use "mt-" for mouse
adata.var["ribo"] = adata.var_names.str.startswith(("RPS", "RPL"))
# Calculate QC metrics
sc.pp.calculate_qc_metrics(
adata, qc_vars=["mt", "ribo"], percent_top=None, log1p=False, inplace=True
)
# Visualize QC distributions
sc.pl.violin(
adata,
["n_genes_by_counts", "total_counts", "pct_counts_mt"],
jitter=0.4,
multi_panel=True,
)
sc.pl.scatter(adata, x="total_counts", y="pct_counts_mt")
sc.pl.scatter(adata, x="total_counts", y="n_genes_by_counts")
# Filter — adjust thresholds based on the QC plots above
sc.pp.filter_cells(adata, min_genes=200)
sc.pp.filter_genes(adata, min_cells=3)
adata = adata[adata.obs.n_genes_by_counts < 5000, :] # remove potential doublets
adata = adata[adata.obs.pct_counts_mt < 20, :] # remove dying cells
print(f"After QC: {adata.n_obs} cells x {adata.n_vars} genes")Normalize library sizes, log-transform, and identify highly variable genes (HVGs). Store raw counts for later differential expression.
# Preserve raw counts in a layer
adata.layers["counts"] = adata.X.copy()
# Normalize to 10,000 counts per cell, then log-transform
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
# Save normalized data as .raw for visualization
adata.raw = adata
# Identify highly variable genes
sc.pp.highly_variable_genes(adata, n_top_genes=2000, flavor="seurat_v3", layer="counts")
sc.pl.highly_variable_genes(adata)
# Subset to HVGs
adata = adata[:, adata.var.highly_variable]
print(f"HVG subset: {adata.n_obs} cells x {adata.n_vars} genes")Regress out confounders and scale gene expression values. This prepares data for PCA.
# Regress out unwanted sources of variation
sc.pp.regress_out(adata, ["total_counts", "pct_counts_mt"])
# Scale to unit variance, clip extreme values
sc.pp.scale(adata, max_value=10)
print(f"Scaled matrix: mean={adata.X.mean():.4f}, std={adata.X.std():.4f}")Compute PCA, inspect the variance ratio elbow plot, then build a neighborhood graph and embed with UMAP.
# PCA
sc.tl.pca(adata, svd_solver="arpack", n_comps=50)
sc.pl.pca_variance_ratio(adata, log=True, n_pcs=50)
# Determine n_pcs from elbow plot (typically 30-50)
n_pcs = 40
# Compute k-nearest neighbor graph
sc.pp.neighbors(adata, n_neighbors=15, n_pcs=n_pcs)
# UMAP embedding
sc.tl.umap(adata)
sc.pl.umap(adata, color=["n_genes_by_counts", "total_counts", "pct_counts_mt"])
print(f"UMAP shape: {adata.obsm['X_umap'].shape}")Run Leiden clustering at multiple resolutions and select the best granularity by visual inspection.
# Test multiple resolutions
for res in [0.3, 0.5, 0.8, 1.0, 1.2]:
sc.tl.leiden(adata, resolution=res, key_added=f"leiden_res{res}", flavor="igraph", n_iterations=2)
# Compare resolutions side-by-side
sc.pl.umap(
adata,
color=["leiden_res0.3", "leiden_res0.5", "leiden_res0.8", "leiden_res1.0"],
ncols=2,
legend_loc="on data",
)
# Pick final resolution based on biological plausibility
adata.obs["leiden"] = adata.obs["leiden_res0.5"]
n_clusters = adata.obs["leiden"].nunique()
print(f"Selected resolution=0.5: {n_clusters} clusters")Find differentially expressed genes per cluster using Wilcoxon rank-sum test on raw counts.
# Rank genes per cluster
sc.tl.rank_genes_groups(adata, groupby="leiden", method="wilcoxon", use_raw=True)
# Visualization
sc.pl.rank_genes_groups(adata, n_genes=20, sharey=False)
sc.pl.rank_genes_groups_dotplot(adata, n_genes=5)
sc.pl.rank_genes_groups_heatmap(adata, n_genes=10, use_raw=True, show_gene_labels=True)
# Extract as DataFrame
markers_df = sc.get.rank_genes_groups_df(adata, group=None)
top_markers = markers_df[markers_df["pvals_adj"] < 0.05].groupby("group").head(10)
print(f"Significant markers: {len(top_markers)} genes across {top_markers['group'].nunique()} clusters")
print(top_markers[["group", "names", "logfoldchanges", "pvals_adj"]].head(15))Assign cell types based on canonical marker genes, visualize, and save all results.
# Define canonical markers (example: human PBMC)
marker_genes = {
"CD4 T cells": ["CD3D", "CD3E", "IL7R"],
"CD8 T cells": ["CD8A", "CD8B", "GZMK"],
"B cells": ["MS4A1", "CD79A", "CD79B"],
"NK cells": ["NKG7", "GNLY", "KLRD1"],
"CD14+ Monocytes": ["CD14", "LYZ", "S100A9"],
"FCGR3A+ Monocytes":["FCGR3A", "MS4A7"],
"Dendritic cells": ["FCER1A", "CST3"],
"Platelets": ["PPBP", "PF4"],
}
# Dot plot — rows = clusters, columns = markers
sc.pl.dotplot(adata, var_names=marker_genes, groupby="leiden", use_raw=True)
# Map clusters → cell types (adjust based on your dot plot)
cluster_to_celltype = {
"0": "CD4 T cells",
"1": "CD14+ Monocytes",
"2": "B cells",
"3": "CD8 T cells",
"4": "NK cells",
"5": "FCGR3A+ Monocytes",
"6": "Dendritic cells",
"7": "Platelets",
}
adata.obs["cell_type"] = adata.obs["leiden"].map(cluster_to_celltype).fillna("Unknown")
sc.pl.umap(adata, color="cell_type", legend_loc="on data", frameon=False)
# Save results
adata.write("results/annotated_data.h5ad", compression="gzip")
adata.obs.to_csv("results/cell_metadata.csv")
markers_df.to_csv("results/marker_genes.csv", index=False)
print(f"Saved: {adata.n_obs} cells with {adata.obs['cell_type'].nunique()} cell types")| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
min_genes (filter_cells) | 200 | 100-500 | Minimum genes detected per cell; lower keeps more cells |
min_cells (filter_genes) | 3 | 3-10 | Minimum cells expressing a gene; higher is stricter |
pct_counts_mt threshold | 20 | 5-30 | Max mitochondrial %; tissue-dependent (5% for PBMCs, 20% for tumors) |
n_top_genes (HVG) | 2000 | 1000-5000 | Number of highly variable genes; more captures subtle variation |
flavor (HVG) | "seurat_v3" | "seurat", "seurat_v3", "cell_ranger" | HVG detection algorithm |
n_pcs (neighbors) | 40 | 20-50 | Number of PCs for neighbor graph; check elbow plot |
n_neighbors | 15 | 5-50 | k for k-NN graph; lower emphasizes local structure |
resolution (leiden) | 0.5 | 0.1-2.0 | Clustering granularity; higher → more clusters |
method (rank_genes) | "wilcoxon" | "wilcoxon", "t-test", "logreg" | DE test; Wilcoxon recommended for publication |
target_sum (normalize) | 1e4 | 1e4-1e5 | Target library size per cell |
When to use: multiple samples/batches with visible batch effects on the UMAP.
# Requires 'batch' column in adata.obs
sc.pp.combat(adata, key="batch")
# Re-run PCA, neighbors, UMAP, clustering after correction
sc.tl.pca(adata, svd_solver="arpack")
sc.pp.neighbors(adata, n_pcs=40)
sc.tl.umap(adata)
sc.pl.umap(adata, color=["batch", "leiden"])When to use: cells follow a developmental continuum rather than discrete clusters.
# PAGA graph abstraction
sc.tl.paga(adata, groups="leiden")
sc.pl.paga(adata, color="leiden", threshold=0.03)
# Reinitialize UMAP with PAGA layout
sc.tl.umap(adata, init_pos="paga")
# Diffusion pseudotime — set root cell
adata.uns["iroot"] = np.flatnonzero(adata.obs["leiden"] == "0")[0]
sc.tl.diffmap(adata)
sc.tl.dpt(adata)
sc.pl.umap(adata, color=["dpt_pseudotime", "leiden"])When to use: generating figures for manuscripts or presentations.
sc.settings.set_figure_params(dpi=300, frameon=False, figsize=(5, 5))
# UMAP with custom palette
sc.pl.umap(
adata, color="cell_type",
palette="Set2",
legend_loc="on data",
legend_fontsize=10,
legend_fontoutline=2,
title="",
save="_celltype_publication.pdf",
)
# Stacked violin plot of top markers
flat_markers = [g for genes in marker_genes.values() for g in genes[:2]]
sc.pl.stacked_violin(
adata, var_names=flat_markers, groupby="cell_type",
use_raw=True, swap_axes=True,
save="_markers_violin.pdf",
)When to use: comparing treated vs control within a specific cell type.
# Subset to cell type of interest
adata_t = adata[adata.obs["cell_type"] == "CD4 T cells"].copy()
# DE between conditions
sc.tl.rank_genes_groups(adata_t, groupby="condition", groups=["treated"], reference="control", method="wilcoxon", use_raw=True)
sc.pl.rank_genes_groups_volcano(adata_t, groups=["treated"])
de_results = sc.get.rank_genes_groups_df(adata_t, group="treated")
sig_genes = de_results[de_results["pvals_adj"] < 0.05]
print(f"Significant DE genes: {len(sig_genes)}")results/annotated_data.h5ad — Full AnnData with embeddings (PCA, UMAP), cluster labels, cell type annotations, and raw counts layerresults/cell_metadata.csv — Per-cell table: barcode, cluster, cell type, QC metrics (n_genes, total_counts, pct_mt)results/marker_genes.csv — DE genes per cluster: gene name, log fold change, adjusted p-value| Problem | Cause | Solution |
|---|---|---|
ModuleNotFoundError: leidenalg | Leiden package not installed | pip install leidenalg igraph |
MemoryError during PCA | Dense matrix exceeds RAM | Use sc.pp.pca(adata, chunked=True, chunk_size=20000) |
| UMAP shows single blob | Over-filtering or too few HVGs | Relax QC thresholds; increase n_top_genes to 3000-5000 |
| Too many/few clusters | Resolution mismatch | Sweep resolution from 0.1 to 2.0 and compare UMAPs |
KeyError: 'MT-' not found | Wrong species prefix | Use mt- (lowercase) for mouse, MT- for human |
| Batch effects dominate UMAP | Uncorrected technical variation | Apply ComBat recipe or use Harmony/scVI for integration |
adata.raw is None error | Forgot to save raw before subsetting | Set adata.raw = adata after normalization, before HVG subset |
| Marker genes non-specific | Resolution too low merging cell types | Increase resolution or use logistic_regression method |
Slow regress_out | Large cell count | Skip regress_out; use sc.pp.scale() alone (often sufficient) |
ValueError in rank_genes_groups | Cluster with too few cells | Remove clusters with <10 cells before DE: adata = adata[adata.obs.groupby('leiden').filter(lambda x: len(x)>=10).index] |
This skill includes reference files for deeper lookup. Read these on demand when the main SKILL.md needs more detail.
Quick-lookup table of all scanpy functions organized by module (sc.pp., sc.tl., sc.pl.*), with full signatures, common parameters, and AnnData structure reference. Use when: you need the exact function name or parameter for a specific operation.
Comprehensive visualization guide: QC plots, UMAP styling, marker gene plots (dot plot, heatmap, stacked violin), trajectory plots, multi-panel figures, publication customization, and color palette recommendations. Use when: creating figures for publications or presentations.
Detailed step-by-step reference for each pipeline stage with decision points, parameter rationale, and alternative approaches (MAD-based filtering, scran normalization, automated annotation). Use when: performing a full analysis from scratch and needing deeper guidance than the Workflow section above.
© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Scanpy Scrna Seq next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scanpy Scrna Seq this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~4.7k | Automated safety check: Pass | CC-BY-4.0 | |
| ScanpyK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~5.1k | Automated safety check: Pass | BSD-3-Clause | |
| Bio Single Cell ClusteringFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | 1 repos | ~2k | Automated safety check: Pass | None | |
| Bio Single Cell ClusteringGPTomics/bioSkills | 1.2k | 1 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Anndatadavila7/claude-code-templates | 33k | 11 repos | ~2.5k | Automated safety check: Pass | MIT | |
| AnndataK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.9k | Automated safety check: Notes | BSD-3-Clause |
K-Dense-AI/scientific-agent-skills
Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…
FreedomIntelligence/OpenClaw-Medical-Skills
Dimensionality reduction and clustering for single-cell RNA-seq using Seurat (R) and Scanpy (Python).
GPTomics/bioSkills
Dimensionality reduction and graph-based clustering for single-cell RNA-seq with Scanpy (Python) and Seurat (R).
davila7/claude-code-templates
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…
K-Dense-AI/scientific-agent-skills
Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.
FreedomIntelligence/OpenClaw-Medical-Skills
Read, write, and create single-cell data objects using Seurat (R) and Scanpy (Python).
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
scRNA-seq with Scanpy: QC, normalization, HVG selection, PCA, neighborhood graph, UMAP/t-SNE, Leiden clustering, markers, cell annotation, trajectory inference. Scanpy Scrna Seq is an agent skill from jaechang-hits/SciAgent-Skills. scRNA-seq with Scanpy: QC, normalization, HVG selection, PCA, neighborhood graph, UMAP/t-SNE, Leiden clustering, markers, cell annotation, trajectory inference.
Scanpy Scrna Seq fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq in jaechang-hits/SciAgent-Skills) into .claude/skills/scanpy-scrna-seq in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/scanpy-scrna-seq in jaechang-hits/SciAgent-Skills) into .agents/skills/scanpy-scrna-seq in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill scanpy-scrna-seq -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scanpy-scrna-seq, .gemini/skills/scanpy-scrna-seq, .github/skills/scanpy-scrna-seq and .opencode/skills/scanpy-scrna-seq in your project.
Going by SKILL.md and its folder, Scanpy Scrna Seq needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 5 domains. As links in the text: doi.org, scanpy.readthedocs.io, scanpy-tutorials.readthedocs.io, training.galaxyproject.org and scverse.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Scanpy Scrna Seq is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Scanpy Scrna Seq: Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars), Bio Single Cell Clustering (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Bio Single Cell Clustering (GPTomics/bioSkills, 1.2k stars) and Anndata (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.