Scanpy Single-Cell Analysis
davila7/claude-code-templates
Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.
Automated scRNA-seq cell type annotation via pre-trained logistic regression.
$ npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills celltypist-cell-annotation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation .claude/skills/celltypist-cell-annotation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "celltypist-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation into .claude/skills/celltypist-cell-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "celltypist-cell-annotation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills celltypist-cell-annotation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation .agents/skills/celltypist-cell-annotation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "celltypist-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation into .agents/skills/celltypist-cell-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "celltypist-cell-annotation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills celltypist-cell-annotation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation .cursor/skills/celltypist-cell-annotation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "celltypist-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation into .cursor/skills/celltypist-cell-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "celltypist-cell-annotation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills celltypist-cell-annotation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation .gemini/skills/celltypist-cell-annotation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "celltypist-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation into .gemini/skills/celltypist-cell-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "celltypist-cell-annotation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills celltypist-cell-annotationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation .github/skills/celltypist-cell-annotation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "celltypist-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation into .github/skills/celltypist-cell-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "celltypist-cell-annotation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills celltypist-cell-annotation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation .opencode/skills/celltypist-cell-annotation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "celltypist-cell-annotation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation into .opencode/skills/celltypist-cell-annotation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "celltypist-cell-annotation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
celltypist-cell-annotationAutomated scRNA-seq cell type annotation via pre-trained logistic regression.
Celltypist Cell Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Automated scRNA-seq cell type annotation via pre-trained logistic regression. 45+ models: immune, gut, lung, brain, fetal, cancer microenvironments. Input normalized AnnData; outputs per-cell labels, majority-vote cluster labels, confidence scores. Use for fast, reference-backed annotation without manual marker inspection.
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with AnnData. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
celltypist.readthedocs.iogithub.comdoi.orgcelltypist.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Celltypist Cell Annotation loads about 5.4k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 1,326 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its MIT licence (© jaechang-hits). 1,326 words, ~5,355 tokens.
.claude/skills/celltypist-cell-annotation/SKILL.md (or your agent's skills folder).CellTypist is an automated cell type classifier for single-cell RNA-seq data built on logistic regression models trained on curated reference atlases. Given a normalized AnnData object, it predicts cell type labels at the single-cell level and optionally applies majority voting within user-defined clusters to produce consensus, biologically coherent annotations. The tool ships with 45+ ready-to-use models spanning pan-immune, organ-specific, and developmental contexts, and supports training custom models from labeled data.
sc.pl.*celltypist>=1.6, scanpy>=1.9, anndataadata.X (10,000 UMIs per cell target sum). Raw counts must be normalized before calling CellTypistpip install celltypist "scanpy[leiden]" anndataSettle these with the user before writing any analysis code.
decisions:
- id: D1
param: annotationStrategy
kind: required
source: user
ask: "Name cell types from canonical markers by hand over the clusters, transfer them from a pre-trained model, or do both and compare the two?"
default: "transfer from a model, then confirm against markers"
- id: D2
param: tissueContext
kind: required
source: user
ask: "Which tissue is this, and in what state - healthy adult, fetal or developmental, diseased, or perturbed?"
default: null
- id: D3
param: model
kind: required
source: literature
depends_on: [D2]
ask: "Which pre-trained model matches that tissue AND that state?"
default: null
- id: D4
param: markerValidation
kind: required
source: user
depends_on: [D1]
ask: "Should the assigned labels be checked against canonical marker expression before they are accepted?"
default: "checked - a mismatched model still labels every cell confidently"
- id: D5
param: validationMarkerPanel
kind: required
source: literature
depends_on: [D2, D4]
ask: "Which canonical markers should the assigned types be confirmed against?"
default: null
skip_if: "marker validation declined"
- id: D6
param: labelGranularity
kind: required
source: user
depends_on: [D2]
ask: "Name broad lineages, or subtypes and activation states within them?"
default: null
- id: D7
param: majorityVoting
kind: required
source: user
ask: "Assign labels cell by cell, or smooth them to a consensus within each cluster?"
default: "per-cell, with voting recommended once clusters are trusted"
- id: D8
param: clusteringKey
kind: derived
source: upstream
depends_on: [D7]
ask: "Which clustering should the consensus be taken over?"
default: "the clustering computed upstream"
skip_if: "majority voting disabled"
- id: D9
param: assignmentThreshold
kind: required
source: user
ask: "Should every cell receive its best-matching label, or should uncertain cells be left unassigned?"
default: "best match regardless of confidence"
- id: D10
param: minClusterProportion
kind: optional
source: user
depends_on: [D7]
ask: "How much of a cluster must agree before the consensus label is applied?"
default: "no minimum"
skip_if: "majority voting disabled"D1 is asked even though the user has arrived at an automated annotator,
because the alternative is not visible from here: manual marker naming and
model transfer fail in opposite directions, and "both, compared" is the right
answer more often than either alone. If the answer is manual only, this skill
is not the tool — see single-cell-annotation-guide.
D2 precedes D3 because a model is specific to a tissue and a state. An adult immune model applied to fetal or tumour tissue returns a confident label for every cell; the classifier has no way to say "these cells are not in my reference". D4 is what makes that failure visible, and it is separate from D9 — a confident assignment and a correct one are not the same thing, and only markers tell them apart.
Minimal pipeline — annotate a preprocessed AnnData with the pan-immune model:
import celltypist
import scanpy as sc
# Load a preprocessed AnnData (normalized + log1p, Leiden clusters already in adata.obs)
adata = sc.read_h5ad("preprocessed_pbmc.h5ad")
# Run annotation with majority voting across Leiden clusters
predictions = celltypist.annotate(
adata,
model="Immune_All_Low.pkl",
majority_voting=True,
)
adata = predictions.to_adata()
print(adata.obs[["predicted_labels", "majority_voting", "conf_score"]].head(10))
# predicted_labels majority_voting conf_score
# CD4+ T cells CD4+ T cells 0.92
# ...Install CellTypist and download pre-trained models. Models are cached locally after the first download.
pip install celltypist "scanpy[leiden]" anndataimport celltypist
from celltypist import models
# Download all available models (only needed once; ~2 GB total)
models.download_models(force_update=False)
# List available models with metadata
models_df = models.models_description()
print(models_df[["model", "description", "n_celltypes", "n_cells"]].to_string())
# Output (excerpt):
# model description n_celltypes n_cells
# Immune_All_Low.pkl Pan-immune low-hierarchy (98 cell types) 98 324,320
# Immune_All_High.pkl Pan-immune high-hierarchy (30 cell types) 30 324,320
# Human_Lung_Atlas.pkl Lung cell types from Human Lung Atlas 61 584,944CellTypist requires normalized, log1p-transformed counts in adata.X. Run normalization before annotation. Raw counts must be stored separately.
import scanpy as sc
# Load raw count matrix
adata = sc.read_h5ad("raw_counts.h5ad")
# Alternatively from 10X:
# adata = sc.read_10x_mtx("filtered_feature_bc_matrix/")
# adata.var_names_make_unique()
# Store raw counts before normalization
adata.layers["counts"] = adata.X.copy()
# Normalize to 10,000 UMIs per cell and log1p-transform
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
print(f"Prepared: {adata.n_obs} cells x {adata.n_vars} genes")
print(f"adata.X mean: {adata.X.mean():.3f} (expected ~0.5–2.0 after log1p normalization)")Choose the model that best matches your tissue type and desired annotation resolution.
from celltypist import models
# Show full model table with filtering
models_df = models.models_description()
# Filter to human immune models
immune_models = models_df[models_df["description"].str.contains("immune|Immune", case=False)]
print(immune_models[["model", "description", "n_celltypes"]].to_string())
# Load a specific model to inspect its cell type labels
model = models.Model.load("Immune_All_Low.pkl")
print(f"Model cell types ({len(model.cell_types)}):")
print(model.cell_types[:20]) # first 20 labelsAvailable models (key selection guide):
| Model | Cell Types | Best For |
|---|---|---|
Immune_All_Low.pkl | 98 | Pan-immune with fine subtypes (e.g., MAIT, Tfh, cDC1) |
Immune_All_High.pkl | 30 | Pan-immune major lineages (T, B, NK, monocyte, DC) |
Human_Lung_Atlas.pkl | 61 | Lung: alveolar, stromal, immune, endothelial |
Pan_Fetal_Human.pkl | 139 | Fetal human multi-organ development |
Developing_Human_Brain.pkl | 51 | Brain development: progenitors, neurons, glia |
Human_Colorectal_Cancer.pkl | 62 | Colorectal cancer cells + tumor microenvironment |
Run celltypist.annotate() with majority_voting=True for cluster-level consensus labels alongside per-cell predictions.
import celltypist
import scanpy as sc
# Ensure Leiden clusters exist for majority voting
# If not already computed:
sc.pp.highly_variable_genes(adata, n_top_genes=2000)
sc.pp.pca(adata)
sc.pp.neighbors(adata, n_pcs=30)
sc.tl.leiden(adata, resolution=0.5, key_added="leiden")
# Run CellTypist annotation
predictions = celltypist.annotate(
adata,
model="Immune_All_Low.pkl",
majority_voting=True, # cluster-level consensus
over_clustering="leiden", # clustering key for majority voting
p_thres=0.5, # cells below threshold → "Unassigned"
mode="best match", # assign the single highest-probability label
)
# Inspect prediction object
print(type(predictions)) # celltypist.classifier.AnnotationResult
print(predictions.predicted_labels.head())
print(predictions.probability_matrix.shape) # (n_cells, n_cell_types)Transfer predictions back to the AnnData object and review confidence scores.
# Merge predictions into adata.obs
adata = predictions.to_adata()
# Key result columns:
# adata.obs["predicted_labels"] — per-cell best-match label
# adata.obs["majority_voting"] — cluster-level consensus label
# adata.obs["conf_score"] — probability of the predicted label (0–1)
print(adata.obs[["predicted_labels", "majority_voting", "conf_score"]].head(10))
print(f"\nCell type distribution (majority voting):")
print(adata.obs["majority_voting"].value_counts().head(15))
# Flag low-confidence cells
low_conf = adata.obs["conf_score"] < 0.5
print(f"\nLow-confidence cells (conf_score < 0.5): {low_conf.sum()} ({low_conf.mean():.1%})")
adata.obs["high_conf"] = ~low_confPlot predictions on UMAP, validate with canonical marker genes, and confirm annotation quality.
import scanpy as sc
import matplotlib.pyplot as plt
# Compute UMAP if not already done
if "X_umap" not in adata.obsm:
sc.tl.umap(adata)
# UMAP colored by annotation results
fig, axes = plt.subplots(1, 3, figsize=(21, 6))
sc.pl.umap(adata, color="majority_voting", legend_loc="on data",
legend_fontsize=7, title="Majority Voting", ax=axes[0], show=False)
sc.pl.umap(adata, color="predicted_labels", legend_loc="right margin",
legend_fontsize=7, title="Per-Cell Prediction", ax=axes[1], show=False)
sc.pl.umap(adata, color="conf_score", cmap="RdYlGn",
title="Confidence Score", ax=axes[2], show=False)
plt.tight_layout()
plt.savefig("celltypist_annotation.png", dpi=150, bbox_inches="tight")
plt.show()
print("Saved celltypist_annotation.png")
# Validate with canonical immune markers
marker_genes = {
"CD4+ T": ["CD3D", "CD4", "IL7R"],
"CD8+ T": ["CD3D", "CD8A", "GZMK"],
"B cells": ["MS4A1", "CD79A"],
"NK cells": ["GNLY", "NKG7"],
"CD14 Mono": ["CD14", "LYZ"],
}
sc.pl.dotplot(adata, var_names=marker_genes, groupby="majority_voting",
use_raw=False, standard_scale="var",
save="_celltypist_markers.png")| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
model | — | Any .pkl filename or path | Selects the reference atlas for annotation; must match tissue/species |
majority_voting | False | True, False | When True, smooths per-cell labels to cluster consensus; requires a clustering key in over_clustering |
over_clustering | None | Any adata.obs key, "leiden", "louvain" | Clustering column used for majority voting; auto-detected if common keys present |
p_thres | 0.5 | 0.0–1.0 | Minimum probability to assign a label; cells below threshold are labeled "Unassigned" |
mode | "best match" | "best match", "prob match" | "best match": top label regardless of threshold; "prob match": applies p_thres |
min_prop | 0.0 | 0.0–1.0 | For majority voting: minimum fraction of cluster cells with the consensus label; rare labels may be suppressed |
Each CellTypist model is a one-vs-rest logistic regression classifier trained on a curated cell atlas. Key properties:
Majority voting applies a two-stage correction after per-cell prediction:
majority_voting labelmin_prop is setMajority voting is recommended when individual cells have noisy expression but the cluster is biologically coherent. Disable it when cells within a cluster are biologically heterogeneous (e.g., transitional states).
CellTypist automatically intersects the model's training genes with the input AnnData's gene names. Genes present in the model but absent from the query are zero-filled. Annotations degrade if fewer than ~60% of model genes are present — check with model.cell_types and adata.var_names.
When to use: your tissue or species is not covered by an existing model, and you have a labeled reference dataset.
import celltypist
import scanpy as sc
# Load labeled reference AnnData (must be normalized + log1p)
ref = sc.read_h5ad("labeled_reference.h5ad")
# ref.obs["cell_type"] must contain string cell type labels
# Train custom model
new_model = celltypist.train(
ref,
labels="cell_type", # obs column with training labels
n_jobs=4, # parallel workers
max_iter=200, # logistic regression iterations
use_SGD=False, # use full L-BFGS-B solver (recommended for <100k cells)
top_genes=500, # number of most informative genes per class
)
# Save for reuse
new_model.write("custom_tissue_model.pkl")
print(f"Trained model: {len(new_model.cell_types)} cell types")
# Apply to query
predictions = celltypist.annotate(query_adata, model="custom_tissue_model.pkl",
majority_voting=True)When to use: uncertain which model best matches your dataset; run multiple models and compare agreement.
import celltypist
import pandas as pd
model_names = ["Immune_All_High.pkl", "Immune_All_Low.pkl", "Human_Lung_Atlas.pkl"]
results = {}
for model_name in model_names:
preds = celltypist.annotate(adata, model=model_name, majority_voting=True)
adata_tmp = preds.to_adata()
key = model_name.replace(".pkl", "")
results[key] = adata_tmp.obs["majority_voting"].values
comparison = pd.DataFrame(results, index=adata.obs_names)
print("Agreement between Immune_All_High and Immune_All_Low:")
agreement = (comparison["Immune_All_High"] == comparison["Immune_All_Low"]).mean()
print(f" {agreement:.1%} of cells agree")
print(comparison.head(10))When to use: saving annotated data with all prediction metadata for downstream differential expression or trajectory analysis.
import scanpy as sc
import pandas as pd
# Save full annotated AnnData
adata.write_h5ad("annotated_celltypist.h5ad", compression="gzip")
print(f"Saved annotated_celltypist.h5ad ({adata.n_obs} cells)")
# Export cell type table
cell_table = adata.obs[[
"predicted_labels", "majority_voting", "conf_score", "leiden"
]].copy()
cell_table.to_csv("celltypist_annotations.csv")
# Cell type proportions per sample
if "sample" in adata.obs.columns:
props = (adata.obs.groupby(["sample", "majority_voting"])
.size().unstack(fill_value=0))
props_norm = props.div(props.sum(axis=1), axis=0)
props_norm.to_csv("celltypist_proportions.csv")
print(f"Cell type proportions saved (shape: {props_norm.shape})")| Output | Description |
|---|---|
adata.obs["predicted_labels"] | Per-cell best-match label from logistic regression |
adata.obs["majority_voting"] | Cluster-consensus label (when majority_voting=True) |
adata.obs["conf_score"] | Probability of the predicted label (0–1); >0.5 = confident |
adata.obsm["X_umap"] | UMAP embedding (if computed in preprocessing step) |
celltypist_annotation.png | UMAP panels: majority voting label, per-cell label, confidence scores |
celltypist_annotations.csv | Per-cell annotation table with predicted labels and confidence |
| Problem | Cause | Solution |
|---|---|---|
ValueError: adata.X does not appear to be log1p normalized | Raw counts passed directly | Run sc.pp.normalize_total(adata, target_sum=1e4) then sc.pp.log1p(adata) before calling celltypist.annotate() |
Many cells labeled "Unassigned" | p_thres too high or model species mismatch | Lower p_thres to 0.3; verify model matches species and tissue; check conf_score distribution |
KeyError for over_clustering key | Clustering column name not found in adata.obs | Run sc.tl.leiden(adata, key_added="leiden") first, or set over_clustering="leiden" explicitly |
| Implausible labels (e.g., immune labels on neurons) | Wrong model selected for tissue | Choose a tissue-specific model (e.g., Developing_Human_Brain.pkl for brain data); list options with models.models_description() |
MemoryError on large datasets (>500k cells) | Full probability matrix held in RAM | Subsample to 200k cells for annotation, then transfer labels via KNN; or use mode="best match" to skip storing full probability matrix |
Low overall conf_score (<0.4 median) | Dataset is poorly represented by the reference model | Train a custom model from a matched reference or use popv-cell-annotation for ensemble voting |
Model not found error on download | Network issue or wrong model name | Run models.download_models(force_update=True); verify name with models.models_description()["model"].tolist() |
© jaechang-hits, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Celltypist Cell Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Celltypist Cell Annotation this skilljaechang-hits/SciAgent-Skills | 374 | 2 repos | ~5.4k | Automated safety check: Pass | MIT | |
| Scanpy Single-Cell Analysisdavila7/claude-code-templates | 33k | 15 repos | ~2.8k | Automated safety check: Pass | MIT | |
| ScgptJimLiu/science-skills | 228 | 4 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| PyDESeq2 Differential Expressiondavila7/claude-code-templates | 33k | 11 repos | ~4k | Automated safety check: Pass | MIT | |
| Anndatadavila7/claude-code-templates | 33k | 11 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Single-Cell Initial AnalysisLigphiDonk/Oh-my--paper | 739 | 1 repos | ~1.4k | Automated safety check: Pass | MIT |
davila7/claude-code-templates
Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.
JimLiu/science-skills
Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.
davila7/claude-code-templates
Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.
davila7/claude-code-templates
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…
LigphiDonk/Oh-my--paper
Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.
harrisongzhang/TheVirtualBiotech
Single-cell RNA-seq data preparation and quality control pipeline.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Automated scRNA-seq cell type annotation via pre-trained logistic regression. Celltypist Cell Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Automated scRNA-seq cell type annotation via pre-trained logistic regression.
Celltypist Cell Annotation fits situations like: reference-backed annotation without manual marker inspection; tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation in jaechang-hits/SciAgent-Skills) into .claude/skills/celltypist-cell-annotation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/celltypist-cell-annotation in jaechang-hits/SciAgent-Skills) into .agents/skills/celltypist-cell-annotation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill celltypist-cell-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/celltypist-cell-annotation, .gemini/skills/celltypist-cell-annotation, .github/skills/celltypist-cell-annotation and .opencode/skills/celltypist-cell-annotation in your project.
Going by SKILL.md and its folder, Celltypist Cell Annotation needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 4 domains. As links in the text: celltypist.readthedocs.io, github.com, doi.org and celltypist.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Celltypist Cell Annotation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.4k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Celltypist Cell Annotation: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 33k stars), Scgpt (JimLiu/science-skills, 228 stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 33k stars) and Anndata (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.