Scanpy
K-Dense-AI/scientific-agent-skills
Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…
Harmony batch correction for scRNA-seq and other omics. An agent skill from jaechang-hits/SciAgent-Skills.
$ npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills harmony-batch-correction --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/harmony-batch-correction .claude/skills/harmony-batch-correction && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "harmony-batch-correction" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/harmony-batch-correction into .claude/skills/harmony-batch-correction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harmony-batch-correction", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/harmony-batch-correctionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills harmony-batch-correction --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/harmony-batch-correction .agents/skills/harmony-batch-correction && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "harmony-batch-correction" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/harmony-batch-correction into .agents/skills/harmony-batch-correction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harmony-batch-correction", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills harmony-batch-correction --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/harmony-batch-correction .cursor/skills/harmony-batch-correction && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "harmony-batch-correction" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/harmony-batch-correction into .cursor/skills/harmony-batch-correction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harmony-batch-correction", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/single-cell/harmony-batch-correction--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills harmony-batch-correction --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/harmony-batch-correction .gemini/skills/harmony-batch-correction && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "harmony-batch-correction" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/harmony-batch-correction into .gemini/skills/harmony-batch-correction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harmony-batch-correction", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills harmony-batch-correctionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/harmony-batch-correction .github/skills/harmony-batch-correction && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "harmony-batch-correction" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/harmony-batch-correction into .github/skills/harmony-batch-correction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harmony-batch-correction", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills harmony-batch-correction --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/harmony-batch-correction .opencode/skills/harmony-batch-correction && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "harmony-batch-correction" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/single-cell/harmony-batch-correction into .opencode/skills/harmony-batch-correction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harmony-batch-correction", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
harmony-batch-correctionHarmony batch correction for scRNA-seq and other omics. An agent skill from jaechang-hits/SciAgent-Skills.
Harmony Batch Correction is an agent skill from jaechang-hits/SciAgent-Skills. Harmony batch correction for scRNA-seq and other omics. Removes batch effects from PCA embeddings while preserving biology. Run after PCA, before UMAP. Scales to millions of cells. Python (harmonypy, scanpy) and R (Seurat).
Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with Scanpy, UMAP and Python. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comdoi.orgportals.broadinstitute.orgscanpy.readthedocs.iosc-best-practices.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Harmony Batch Correction loads about 5.6k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 1,275 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its MIT licence (© jaechang-hits). 1,275 words, ~5,611 tokens.
.claude/skills/harmony-batch-correction/SKILL.md (or your agent's skills folder).Harmony is a fast, scalable algorithm for batch integration in single-cell data. It takes a PCA embedding (cells × PCs) as input and returns a corrected embedding from which batch effects have been regressed out via iterative soft-clustering and per-cluster linear regression. The corrected embedding is then used to compute neighbors, UMAP, and downstream clustering — the raw count matrix is never modified. Harmony works for single-cell RNA-seq, ATAC-seq, and other omics modalities where a PCA-like embedding is available.
sc.pl.*harmonypy, scanpy>=1.10, leidenalg, igraph, anndata, pandas, matplotlibharmony, Seurat>=4.0 (for R workflow in Step 7)adata.obs, and PCA embedding already computed (adata.obsm["X_pca"])pip install harmonypy "scanpy[leiden]" anndata pandas matplotlib# R installation (for Step 7 / Seurat integration)
install.packages("harmony")
# If using Seurat:
install.packages("Seurat")Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: batchKey
kind: required
source: data
ask: "Which variable separates samples that were processed apart - donor, run, chemistry, or a combination?"
default: null
- id: D2
param: inputEmbedding
kind: derived
source: upstream
ask: "Which reduced space should be corrected?"
default: "the PCA embedding computed upstream"
- id: D3
param: diversityPenalty
kind: required
source: user
depends_on: [D1]
ask: "How hard should batches be pushed together? Too hard erases real differences between samples."
default: "2.0; raise toward 4-6 when batch sizes are very uneven"
- id: D4
param: numClusters
kind: optional_conditional
source: data
ask: "Does the dataset contain rare populations that need more soft clusters to survive correction?"
default: "automatic, min(100, cells/30)"
- id: D5
param: maxIterations
kind: optional_conditional
source: data
ask: "Did the correction converge, or does it need more iterations?"
default: 10
- id: D6
param: smallClusterProtection
kind: optional_conditional
source: user
ask: "Are small populations being over-split during correction?"
default: "off"
- id: D7
param: kernelBandwidth
kind: optional
source: user
ask: "Should Harmony use its standard soft-clustering bandwidth, or be tuned for unusually heterogeneous batches?"
default: 0.1D3 is where batch correction goes wrong in the direction nobody checks. Under-correction is visible - batches stay separated in the embedding. Over-correction is not: samples mix beautifully because a real treatment effect was removed along with the batch effect. Ask what the batch variable is confounded with before raising it.
Minimal pipeline — load preprocessed data, run Harmony via scanpy, produce UMAP:
import scanpy as sc
# Load preprocessed AnnData (counts normalized, HVGs selected, PCA run)
adata = sc.read_h5ad("preprocessed.h5ad") # must have adata.obs["batch"]
# Run Harmony batch correction (corrects adata.obsm["X_pca"] → "X_pca_harmony")
sc.external.pp.harmony_integrate(adata, key="batch")
# Build neighborhood graph on corrected embedding, then UMAP
sc.pp.neighbors(adata, use_rep="X_pca_harmony")
sc.tl.umap(adata)
sc.tl.leiden(adata, resolution=0.5)
# Visualize
sc.pl.umap(adata, color=["batch", "leiden"], ncols=2)
print(f"Clusters: {adata.obs['leiden'].nunique()} | Cells: {adata.n_obs}")Filter, normalize, select highly variable genes (HVGs), and run PCA. Harmony is applied to the PCA embedding produced here.
import scanpy as sc
import numpy as np
sc.settings.set_figure_params(dpi=80, facecolor="white")
# Load multi-batch data — adata.obs must contain a batch column
adata = sc.read_h5ad("multi_batch_counts.h5ad")
# Alternatively, concatenate multiple AnnData objects:
# adata = sc.concat([adata1, adata2, adata3], label="batch",
# keys=["batch1", "batch2", "batch3"])
print(f"Loaded: {adata.n_obs} cells × {adata.n_vars} genes")
print(f"Batches: {adata.obs['batch'].value_counts().to_dict()}")
# QC: annotate mitochondrial genes and filter low-quality cells
adata.var["mt"] = adata.var_names.str.startswith("MT-")
sc.pp.calculate_qc_metrics(adata, qc_vars=["mt"], inplace=True)
sc.pp.filter_cells(adata, min_genes=200)
sc.pp.filter_cells(adata, max_genes=6000)
sc.pp.filter_genes(adata, min_cells=3)
adata = adata[adata.obs["pct_counts_mt"] < 20].copy()
print(f"After QC: {adata.n_obs} cells × {adata.n_vars} genes")
# Normalize and log-transform (store raw counts first)
adata.layers["counts"] = adata.X.copy()
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
# Select HVGs and compute PCA
sc.pp.highly_variable_genes(adata, n_top_genes=3000, batch_key="batch")
sc.pp.pca(adata, n_comps=50, use_highly_variable=True)
print(f"PCA shape: {adata.obsm['X_pca'].shape}") # (n_cells, 50)The scanpy wrapper calls harmonypy under the hood and stores the corrected embedding in adata.obsm["X_pca_harmony"].
import scanpy as sc
# Run Harmony — corrects adata.obsm["X_pca"] in-place and stores result as "X_pca_harmony"
sc.external.pp.harmony_integrate(
adata,
key="batch", # obs column with batch labels
basis="X_pca", # source embedding (default)
adjusted_basis="X_pca_harmony", # destination key (default)
max_iter_harmony=20, # maximum Harmony iterations (default 10; increase for complex batches)
theta=2.0, # diversity penalty per batch variable (default 2)
sigma=0.1, # width of soft-clustering Gaussian kernel
random_state=42,
)
print(f"Corrected embedding: {adata.obsm['X_pca_harmony'].shape}")
# Expected: (n_cells, 50) — same shape as X_pcaUse the lower-level harmonypy API when you need finer control, access to the HarmonyObject, or integration outside of scanpy.
import harmonypy
import pandas as pd
import numpy as np
# Extract PCA embedding and metadata
pca_embeddings = adata.obsm["X_pca"] # numpy array (n_cells, n_PCs)
meta_data = adata.obs[["batch", "donor"]].copy() # DataFrame with batch variables
# Run Harmony
ho = harmonypy.run_harmony(
pca_embeddings,
meta_data,
vars_use=["batch"], # list of columns to correct for
theta=2.0, # diversity penalty strength (per variable)
sigma=0.1, # soft-clustering bandwidth
nclust=None, # number of clusters; None → auto (min(100, n_cells/30))
tau=0, # protection against over-clustering small clusters
block_size=0.05, # fraction of cells per mini-batch
max_iter_harmony=20,
max_iter_kmeans=20,
epsilon_cluster=1e-5,
epsilon_harmony=1e-4,
random_state=42,
verbose=True,
)
corrected = ho.Z_corr.T # (n_cells, n_PCs) corrected embedding
adata.obsm["X_pca_harmony"] = corrected
print(f"Harmony converged. Corrected embedding shape: {corrected.shape}")Always use use_rep="X_pca_harmony" so that downstream graph construction is based on the corrected embedding, not the original PCA.
import scanpy as sc
# Build k-nearest neighbor graph using corrected embedding
sc.pp.neighbors(
adata,
n_neighbors=15, # number of neighbors (15-30 typical for scRNA-seq)
n_pcs=None, # use all PCs in the embedding
use_rep="X_pca_harmony", # IMPORTANT: use corrected embedding
random_state=42,
)
# Compute UMAP for visualization
sc.tl.umap(adata, min_dist=0.3, spread=1.0, random_state=42)
print(f"UMAP computed: {adata.obsm['X_umap'].shape}") # (n_cells, 2)Clustering is performed on the neighbor graph built from the Harmony-corrected embedding in Step 4.
import scanpy as sc
# Run Leiden community detection at multiple resolutions
for resolution in [0.3, 0.5, 0.8]:
sc.tl.leiden(adata, resolution=resolution, key_added=f"leiden_{resolution}")
# Select final resolution
adata.obs["leiden"] = adata.obs["leiden_0.5"]
print(f"Clusters at resolution 0.5: {adata.obs['leiden'].nunique()}")
print(adata.obs["leiden"].value_counts().head())
# Find marker genes per cluster
sc.tl.rank_genes_groups(adata, groupby="leiden", method="wilcoxon", n_genes=50)
sc.pl.rank_genes_groups_dotplot(adata, n_genes=5, groupby="leiden")Plot UMAP colored by batch and by cell type / cluster to visually confirm batch effects are removed while biological structure is preserved.
import scanpy as sc
import matplotlib.pyplot as plt
fig, axes = plt.subplots(1, 3, figsize=(18, 5))
# Before correction: color by batch using uncorrected PCA-based UMAP
sc.pp.neighbors(adata, use_rep="X_pca", key_added="uncorrected_neighbors", random_state=42)
sc.tl.umap(adata, neighbors_key="uncorrected_neighbors", random_state=42)
adata.obsm["X_umap_uncorrected"] = adata.obsm["X_umap"].copy()
# Restore Harmony UMAP
sc.pp.neighbors(adata, use_rep="X_pca_harmony", random_state=42)
sc.tl.umap(adata, random_state=42)
# Panel 1: Corrected UMAP colored by batch
sc.pl.umap(adata, color="batch", title="Harmony: colored by batch",
ax=axes[0], show=False)
# Panel 2: Corrected UMAP colored by leiden cluster
sc.pl.umap(adata, color="leiden", title="Harmony: Leiden clusters",
ax=axes[1], show=False)
# Panel 3: Uncorrected UMAP colored by batch (for comparison)
sc.pl.embedding(adata, basis="X_umap_uncorrected", color="batch",
title="Uncorrected PCA: batch separation",
ax=axes[2], show=False)
plt.tight_layout()
plt.savefig("harmony_batch_evaluation.png", dpi=150, bbox_inches="tight")
plt.show()
print("Batch evaluation figure saved to harmony_batch_evaluation.png")For users working in R, Harmony integrates directly with Seurat via RunHarmony().
library(Seurat)
library(harmony)
# Load Seurat object with batch metadata in metadata column "batch"
seurat_obj <- readRDS("multi_batch_seurat.rds")
# Ensure PCA is computed
seurat_obj <- NormalizeData(seurat_obj)
seurat_obj <- FindVariableFeatures(seurat_obj, nfeatures = 3000)
seurat_obj <- ScaleData(seurat_obj)
seurat_obj <- RunPCA(seurat_obj, npcs = 50)
# Run Harmony — corrects the "pca" reduction and adds "harmony" reduction
seurat_obj <- RunHarmony(
seurat_obj,
group.by.vars = "batch", # metadata column(s) to correct for
dims.use = 1:30, # number of PCs to use
theta = 2, # diversity penalty
sigma = 0.1, # soft-clustering bandwidth
max.iter.harmony = 20,
plot_convergence = TRUE, # plot objective function convergence
verbose = TRUE
)
# Build neighbor graph and UMAP using Harmony embedding
seurat_obj <- FindNeighbors(seurat_obj, reduction = "harmony", dims = 1:30)
seurat_obj <- FindClusters(seurat_obj, resolution = 0.5)
seurat_obj <- RunUMAP(seurat_obj, reduction = "harmony", dims = 1:30)
# Visualize
DimPlot(seurat_obj, reduction = "umap", group.by = "batch")
DimPlot(seurat_obj, reduction = "umap", group.by = "seurat_clusters")
cat(sprintf("Clusters: %d | Cells: %d\n",
length(unique(seurat_obj$seurat_clusters)),
ncol(seurat_obj)))| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
theta | 2.0 | 0–10 | Diversity penalty per batch variable. Higher values force more mixing; set to 0 to disable penalty. Increase to 4–6 for very uneven batch sizes. |
sigma | 0.1 | 0.01–0.5 | Bandwidth of the soft-clustering Gaussian kernel. Smaller values → sharper cluster assignments; larger values → smoother. Rarely needs changing. |
nclust | None (auto) | 5–200 | Number of soft clusters used for correction. Auto-set to min(100, n_cells/30). Increase for datasets with many rare cell types. |
max_iter_harmony | 10 | 5–50 | Maximum Harmony iterations. Increase to 20–30 for datasets with strong or complex batch effects where convergence is slow. |
tau | 0 | 0–10 | Overclustering protection. Positive values penalize very small clusters; use tau=5 when tiny populations are over-splitting. |
n_neighbors (scanpy) | 15 | 5–50 | Number of nearest neighbors for the post-Harmony graph. Increase for larger datasets (>100k cells) to stabilize clusters. |
block_size | 0.05 | 0.01–0.1 | Fraction of cells per mini-batch for the K-means step. Smaller values are faster but less stable; default is usually appropriate. |
Harmony correction proceeds in three repeated phases until convergence:
K clusters based on Euclidean distance in PC space, weighted by the batch-diversity penalty theta.epsilon_harmony.The corrected embedding lives in the same PC space as the input PCA but with batch-specific directions subtracted. The raw count matrix (adata.X) and all gene-level information are untouched.
Harmony assumes the input PCA was computed with batch_key awareness. In scanpy, always pass batch_key="batch" to sc.pp.highly_variable_genes() before PCA so that HVGs are selected independently per batch and then intersected, preventing batch-specific HVGs from dominating the embedding.
Harmony corrects the embedding, not the count matrix. Differential expression analysis should still use the raw or normalized counts (stored in adata.layers["counts"] or adata.raw), not the corrected PCs.
Correct for multiple confounding variables simultaneously by passing a list to vars_use. Each variable gets its own diversity penalty theta.
import harmonypy
import pandas as pd
# Prepare metadata with multiple batch variables
meta_data = adata.obs[["batch", "donor", "platform"]].copy()
# Run Harmony correcting for all three variables
ho = harmonypy.run_harmony(
adata.obsm["X_pca"],
meta_data,
vars_use=["batch", "donor", "platform"], # correct for all three
theta=[2.0, 1.0, 2.0], # per-variable diversity penalty
# theta can also be a scalar (same for all variables)
max_iter_harmony=30,
random_state=42,
verbose=True,
)
adata.obsm["X_pca_harmony"] = ho.Z_corr.T
print(f"Multi-variable correction complete. Shape: {adata.obsm['X_pca_harmony'].shape}")
# Verify batch mixing improved for all variables
import scanpy as sc
sc.pp.neighbors(adata, use_rep="X_pca_harmony")
sc.tl.umap(adata, random_state=42)
sc.pl.umap(adata, color=["batch", "donor", "platform"], ncols=3)If Harmony over-corrects (merges biologically distinct populations), canonical marker genes will lose cell-type specificity on the UMAP. Check this before downstream analysis.
import scanpy as sc
import matplotlib.pyplot as plt
# Define known canonical marker genes for your tissue
marker_genes = {
"T cells": ["CD3D", "CD3E", "TRAC"],
"B cells": ["MS4A1", "CD79A", "CD19"],
"NK cells": ["GNLY", "NKG7", "KLRD1"],
"Monocytes": ["LYZ", "CD14", "CST3"],
"Dendritic cells": ["FCER1A", "CLEC10A"],
}
# Plot canonical markers on the Harmony-corrected UMAP
all_markers = [g for genes in marker_genes.values() for g in genes
if g in adata.var_names]
sc.pl.umap(
adata,
color=all_markers[:6], # plot first 6 markers
ncols=3,
vmax="p99", # clip color scale at 99th percentile
frameon=False,
save="_marker_genes.png",
)
# If markers look diffuse or don't separate cell types cleanly,
# reduce theta (e.g., from 2.0 to 1.0) or reduce max_iter_harmony
print("If marker gene expression is diffuse across clusters, reduce theta.")
print("If batch effects remain visible, increase theta or max_iter_harmony.")Save the corrected object with all embeddings for downstream analysis without re-running Harmony.
import scanpy as sc
# Save with all embeddings
adata.write_h5ad("harmony_corrected.h5ad", compression="gzip")
print(f"Saved to harmony_corrected.h5ad ({adata.n_obs} cells)")
# Reload and verify
adata_loaded = sc.read_h5ad("harmony_corrected.h5ad")
print(f"Loaded. Keys in obsm: {list(adata_loaded.obsm.keys())}")
# Expected: ['X_pca', 'X_pca_harmony', 'X_umap']| Output | Type | Description |
|---|---|---|
adata.obsm["X_pca_harmony"] | np.ndarray (cells × PCs) | Harmony-corrected PCA embedding; input for neighbors/UMAP/clustering |
adata.obsm["X_umap"] | np.ndarray (cells × 2) | 2D UMAP coordinates from corrected embedding |
adata.obs["leiden"] | pd.Categorical | Leiden cluster labels |
adata.uns["neighbors"] | dict | KNN graph metadata (connectivities, distances) |
adata.uns["rank_genes_groups"] | dict | Per-cluster marker gene statistics (scores, p-values, log fold changes) |
harmony_batch_evaluation.png | PNG figure | Side-by-side UMAP panels: batch labels, cluster labels, uncorrected comparison |
| Problem | Cause | Solution |
|---|---|---|
| Harmony does not converge (max iterations reached) | Strong batch effects or too few iterations | Increase max_iter_harmony to 30–50; check that PCA was computed with batch_key HVG selection |
| Batches still separate on UMAP after correction | theta too low or too few PCs | Increase theta to 3–4; use more PCs (40–50); ensure use_rep="X_pca_harmony" is set in sc.pp.neighbors() |
| Over-correction: biologically distinct populations merge | theta too high or too many iterations | Reduce theta to 1.0; reduce max_iter_harmony; verify marker gene expression is cell-type-specific |
KeyError: 'X_pca_harmony' in sc.pp.neighbors() | Step 2 (harmony_integrate) not run yet | Run sc.external.pp.harmony_integrate(adata, key="batch") before calling sc.pp.neighbors() |
AttributeError or import error for sc.external.pp.harmony_integrate | harmonypy not installed | Run pip install harmonypy; verify import harmonypy succeeds |
ValueError: vars_use not in meta_data columns | Batch column name mismatch | Check adata.obs.columns and confirm the column name passed to key= exists |
| Memory error on large datasets (>1M cells) | Full PCA matrix loaded into memory | Use harmonypy directly with chunked PCA; consider subsampling to 200k cells for parameter tuning first |
| Clusters appear identical before and after correction | PCA was not computed with HVGs selected per batch | Re-run sc.pp.highly_variable_genes(adata, batch_key="batch") and sc.pp.pca() before Harmony |
© jaechang-hits, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/single-cell/harmony-batch-correction of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Harmony Batch Correction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Harmony Batch Correction this skilljaechang-hits/SciAgent-Skills | 370 | 2 repos | ~5.6k | Automated safety check: Pass | MIT | |
| ScanpyK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~5.1k | Automated safety check: Pass | BSD-3-Clause | |
| Bio Single Cell ClusteringGPTomics/bioSkills | 1.2k | 1 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Bio Single Cell Clusteringmajiayu000/claude-skill-registry | 666 | 2 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Anndatadavila7/claude-code-templates | 32k | 12 repos | ~2.5k | Automated safety check: Pass | MIT | |
| AnndataK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.9k | Automated safety check: Notes | BSD-3-Clause |
K-Dense-AI/scientific-agent-skills
Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…
GPTomics/bioSkills
Dimensionality reduction and graph-based clustering for single-cell RNA-seq with Scanpy (Python) and Seurat (R).
majiayu000/claude-skill-registry
Dimensionality reduction and clustering for single-cell RNA-seq using Seurat (R) and Scanpy (Python).
davila7/claude-code-templates
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…
K-Dense-AI/scientific-agent-skills
Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.
FreedomIntelligence/OpenClaw-Medical-Skills
Read, write, and create single-cell data objects using Seurat (R) and Scanpy (Python).
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
Harmony batch correction for scRNA-seq and other omics. An agent skill from jaechang-hits/SciAgent-Skills. Harmony Batch Correction is an agent skill from jaechang-hits/SciAgent-Skills. Harmony batch correction for scRNA-seq and other omics.
Harmony Batch Correction fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/harmony-batch-correction in jaechang-hits/SciAgent-Skills) into .claude/skills/harmony-batch-correction in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/harmony-batch-correction in jaechang-hits/SciAgent-Skills) into .agents/skills/harmony-batch-correction in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill harmony-batch-correction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harmony-batch-correction, .gemini/skills/harmony-batch-correction, .github/skills/harmony-batch-correction and .opencode/skills/harmony-batch-correction in your project.
Going by SKILL.md and its folder, Harmony Batch Correction needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 5 domains. As links in the text: github.com, doi.org, portals.broadinstitute.org, scanpy.readthedocs.io and sc-best-practices.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Harmony Batch Correction is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Harmony Batch Correction: Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars), Bio Single Cell Clustering (GPTomics/bioSkills, 1.2k stars), Bio Single Cell Clustering (majiayu000/claude-skill-registry, 666 stars) and Anndata (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 163 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.