Agent skill

Popv Cell Annotation

by jaechang-hits in jaechang-hits/SciAgent-Skills

Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via…

BSD-3-ClauseAuto-check passedResearch & Science

Install Popv Cell Annotation

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill popv-cell-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills popv-cell-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/single-cell/popv-cell-annotation .claude/skills/popv-cell-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
popv-cell-annotation
GitHub stars
374
Used in
2 other repos
Token cost
~6.9k tokens
SKILL.md length
1,595 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via…

  • Works in 5 steps: Check gene overlap before running: popV… → Use raw counts as input: pass raw… → Match reference granularity to query… → …
  • Single-method annotation is insufficient
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 10 more sections
  • Calls pip

What it does

Popv Cell Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via majority voting. Outputs per-method labels, consensus, agreement score. Use when single-method annotation is insufficient or you need ensemble uncertainty for novel states.

Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics, Machine learning and Creative writing and fiction. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.

When your agent uses it

  • Single-method annotation is insufficient
  • You need ensemble uncertainty for novel states

Example prompts

  • “/popv-cell-annotation”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Check gene overlap before running: popV performs best with >70% gene overlap between reference and query. If overlap is <50%, annotation…
  2. Use raw counts as input: pass raw (un-normalized) counts in adata.X to Process_Query. popV internally applies its own normalization…
  3. Match reference granularity to query biology: if your query contains subtypes not in the reference, no method will correctly assign them…
  4. Exclude slow methods when speed matters: scanvi_popv and onclass are the slowest. For a quick first-pass, run only knn_harmony, knn_bbknn…
  5. Save trained models for repeated queries: Process_Query stores scVI/SCANVI models in save_path_trained_models. Reuse these when annotating…

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org
    • github.com
    • popv.readthedocs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Popv Cell Annotation loads about 6.9k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 1,595 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,595 words, ~6,928 tokens.

Download SKILL.mdSave it as .claude/skills/popv-cell-annotation/SKILL.md (or your agent's skills folder).
name
popv-cell-annotation
description
Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via majority voting. Outputs per-method labels, consensus, agreement score. Use when single-method annotation is insufficient or you need ensemble uncertainty for novel states.
license
BSD-3-Clause

popV Multi-Method Cell Type Transfer

Overview

popV (Population Voting for single-cell annotation) annotates a query scRNA-seq dataset by running 10+ independent classification algorithms against a labeled reference atlas and aggregating results via majority voting. Each method produces its own label; the final popv_prediction is the consensus across all methods, and the popv_agreement score quantifies how many methods agree. This ensemble strategy is robust to individual method failures on unusual datasets and provides a principled uncertainty estimate: low agreement highlights novel cell states or annotation gaps.

When to Use

  • Annotating a query dataset by transferring labels from a well-curated reference atlas when you want a consensus rather than a single model's judgment
  • Identifying novel or ambiguous cell states as cells where methods disagree (low popv_agreement score)
  • Benchmarking annotation reliability by comparing per-method labels to detect systematic disagreements
  • Annotating large atlas datasets (>100k cells) where batch effects between reference and query are substantial
  • Producing annotation for downstream analyses that require high-confidence labels (clinical data, regulatory submissions)
  • Use omics-plotting SKILL for confidence bar charts and method-agreement heatmaps from exported tables (UMAPs stay in scanpy sc.pl.*)
  • Use CellTypist (celltypist-cell-annotation) instead when speed matters and a pre-trained model matches your tissue; popV is slower because it trains multiple models on your reference
  • Use scANVI (scvi-tools-single-cell) instead when you need a single probabilistic deep generative model with formal uncertainty quantification and do not require the ensemble

Prerequisites

  • Python packages: popv>=0.6, scanpy>=1.9, anndata, scvi-tools>=1.0, harmonypy, bbknn, celltypist
  • Data requirements: Two AnnData objects — a labeled reference (adata_ref) with cell type labels in obs, and an unlabeled query (adata_query). Both must be from the same species and have overlapping gene sets. Raw counts in adata.X (popV applies its own normalization internally)
  • Environment: Python 3.9+; GPU recommended for scVI/SCANVI methods (falls back to CPU); 32 GB RAM recommended for >200k reference cells
bash
pip install popv scvi-tools harmonypy bbknn celltypist

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: annotationStrategy
    kind: required
    source: user
    ask: "Name cell types from canonical markers by hand over the clusters, transfer them from an annotated reference by ensemble consensus, or do both and compare the two?"
    default: "ensemble transfer, then confirm against markers"

  - id: D2
    param: tissueContext
    kind: required
    source: user
    ask: "Which tissue is this, and in what state - healthy adult, fetal or developmental, diseased, or perturbed?"
    default: null

  - id: D3
    param: referenceAtlas
    kind: required
    source: literature
    depends_on: [D2]
    ask: "Which annotated reference matches that tissue AND that state?"
    default: null

  - id: D4
    param: referenceLabelColumn
    kind: required
    source: data
    depends_on: [D3]
    ask: "Which column of the reference holds the cell type labels, and at what granularity?"
    default: null

  - id: D5
    param: markerValidation
    kind: required
    source: user
    depends_on: [D1]
    ask: "Should the consensus labels be checked against canonical marker expression before they are accepted?"
    default: "checked - method agreement measures consistency, not correctness"

  - id: D6
    param: validationMarkerPanel
    kind: required
    source: literature
    depends_on: [D2, D5]
    ask: "Which canonical markers should the consensus types be confirmed against?"
    default: null
    skip_if: "marker validation declined"

  - id: D7
    param: consensusThreshold
    kind: required
    source: user
    ask: "How many of the annotation methods must agree before a label is trusted?"
    default: "80% agreement"

  - id: D8
    param: methodSubset
    kind: required
    source: user
    ask: "Run the full method panel, or drop the slow model-based ones?"
    default: "all methods"

  - id: D9
    param: variableGeneCount
    kind: optional
    source: user
    ask: "How many variable genes should the embedding and nearest-neighbour methods use?"
    default: 4000

  - id: D10
    param: trainingEpochs
    kind: optional_conditional
    source: data
    depends_on: [D8]
    ask: "Do the model-based methods need longer training for a large or complex query?"
    default: "50 unsupervised, 20 semi-supervised"
    skip_if: "model-based methods excluded"

  - id: D11
    param: gpuUse
    kind: never_ask
    source: data
    reason: "Falls back to CPU automatically; affects runtime only"
    default: true

D1 is asked even though the user has arrived at an ensemble annotator, because the alternative is not visible from here: manual marker naming and reference transfer fail in opposite directions. If the answer is manual only, this skill is not the tool — see single-cell-annotation-guide.

D5 matters more here than it looks. The agreement score in D7 measures whether the methods concur, not whether they are right — several methods sharing one unsuitable reference agree with each other confidently. Markers are the only outside check. D4 sets the granularity of every downstream comparison: a reference labelled at lineage level cannot produce subtype calls.

Quick Start

Minimal pipeline from labeled reference and unlabeled query to annotated result:

python
import popv
import scanpy as sc

# Load reference (labeled) and query (unlabeled) AnnData objects
adata_ref = sc.read_h5ad("reference_atlas.h5ad")  # adata_ref.obs["cell_type"] exists
adata_query = sc.read_h5ad("query_dataset.h5ad")

# Prepare combined object with popV preprocessing
adata = popv.preprocessing.Process_Query(
    adata_ref,
    adata_query,
    ref_labels_key="cell_type",
    ref_batch_key="batch",
    query_batch_key="batch",
    unknown_celltype_label="unknown",
    save_path_trained_models="./popv_models/",
    n_epochs_unsupervised=50,
)

# Run all annotation methods
popv.annotation.annotate_data(adata)

# Inspect consensus results for query cells
query_mask = adata.obs["_dataset"] == "query"
print(adata[query_mask].obs[["popv_prediction", "popv_agreement"]].head(10))

Core API

Module 1: Reference and Query Data Setup

Both AnnData objects must share a gene space and have required metadata columns. popV will subset to the intersection of genes automatically.

python
import anndata as ad
import scanpy as sc
import numpy as np

# Reference: must have cell type labels and (optionally) batch metadata
adata_ref = sc.read_h5ad("reference_atlas.h5ad")
print(f"Reference: {adata_ref.n_obs} cells x {adata_ref.n_vars} genes")
print(f"Cell types: {adata_ref.obs['cell_type'].nunique()} unique labels")
print(f"Reference cell type counts:\n{adata_ref.obs['cell_type'].value_counts().head(10)}")

# Query: no labels required; batch metadata optional
adata_query = sc.read_h5ad("query_dataset.h5ad")
print(f"\nQuery: {adata_query.n_obs} cells x {adata_query.n_vars} genes")

# Check gene overlap (popV will handle subsetting but >70% overlap is recommended)
shared_genes = adata_ref.var_names.intersection(adata_query.var_names)
pct_shared = len(shared_genes) / adata_ref.n_vars
print(f"\nShared genes: {len(shared_genes)} ({pct_shared:.1%} of reference genes)")
if pct_shared < 0.5:
    print("WARNING: <50% gene overlap — annotation quality may be reduced")
python
# Verify required fields before popV setup
assert "cell_type" in adata_ref.obs.columns, "Reference needs cell type labels"

# Add batch column if absent (popV requires it even for single-batch data)
if "batch" not in adata_ref.obs.columns:
    adata_ref.obs["batch"] = "ref_batch"
if "batch" not in adata_query.obs.columns:
    adata_query.obs["batch"] = "query_batch"

print("Reference obs columns:", adata_ref.obs.columns.tolist())
print("Query obs columns:    ", adata_query.obs.columns.tolist())
Module 2: POPV Object Creation (Process_Query)

Process_Query combines reference and query, normalizes counts, selects HVGs, and prepares the joint embedding needed by all annotation methods.

python
import popv

# Create processed combined AnnData
adata = popv.preprocessing.Process_Query(
    adata_ref,
    adata_query,
    ref_labels_key="cell_type",      # obs column with reference labels
    ref_batch_key="batch",           # obs column with reference batch info
    query_batch_key="batch",         # obs column with query batch info
    unknown_celltype_label="unknown",# label to use for query cells before annotation
    save_path_trained_models="./popv_models/",  # directory for scVI/SCANVI model checkpoints
    n_epochs_unsupervised=50,        # scVI training epochs (increase to 100–200 for large datasets)
    n_epochs_semisupervised=20,      # scANVI fine-tuning epochs
    use_gpu=True,                    # GPU for scVI/SCANVI (falls back to CPU if unavailable)
    hvg=4000,                        # number of highly variable genes to use
)

print(f"Combined object: {adata.n_obs} cells x {adata.n_vars} genes")
print(f"Dataset labels: {adata.obs['_dataset'].value_counts().to_dict()}")
# Expected: {'ref': N_ref, 'query': N_query}
Module 3: Running the Method Ensemble

annotate_data runs all selected methods sequentially and adds per-method label columns plus the consensus to adata.obs.

python
import popv

# Run annotation with default set of methods
popv.annotation.annotate_data(
    adata,
    methods=[
        "knn_harmony",    # KNN on Harmony-corrected embedding
        "knn_bbknn",      # KNN on BBKNN cross-batch graph
        "knn_scvi",       # KNN on scVI latent space
        "scanvi_popv",    # Semi-supervised scANVI label transfer
        "celltypist_popv",# CellTypist logistic regression
        "rf",             # Random Forest on HVG expression
        "xgboost",        # XGBoost classifier
        "svm",            # Support Vector Machine
        "onclass",        # ONCLASS (ontology-guided)
    ],
)

# Inspect per-method result columns (all end in "_popv")
query_mask = adata.obs["_dataset"] == "query"
popv_cols = adata.obs.filter(like="_popv").columns.tolist()
print(f"Per-method columns: {popv_cols}")
print(adata[query_mask].obs[popv_cols + ["popv_prediction", "popv_agreement"]].head(10))
Module 4: Consensus Results and Agreement Scoring

popv_prediction is the majority-vote consensus; popv_agreement is the fraction of methods that agreed on the winning label.

python
import pandas as pd

query_mask = adata.obs["_dataset"] == "query"
query_obs = adata[query_mask].obs.copy()

# Consensus label distribution
print("Consensus cell type distribution:")
print(query_obs["popv_prediction"].value_counts().head(15))

# Agreement score statistics
print(f"\npopv_agreement statistics:")
print(query_obs["popv_agreement"].describe())
# agreement = 1.0 → all methods agree; agreement = 0.2 → only 2/10 methods agree

# Cells with high confidence (>80% method agreement)
high_conf = query_obs["popv_agreement"] >= 0.8
print(f"\nHigh-confidence cells (agreement >= 0.8): {high_conf.sum()} ({high_conf.mean():.1%})")

# Cells with low confidence — candidate novel states or annotation gaps
low_conf = query_obs["popv_agreement"] < 0.5
print(f"Low-confidence cells  (agreement <  0.5): {low_conf.sum()} ({low_conf.mean():.1%})")
Module 5: Visualization

popV provides built-in UMAP and heatmap visualization of per-method agreement and consensus labels.

python
import popv
import scanpy as sc
import matplotlib.pyplot as plt

# Compute UMAP on the joint reference+query embedding (if not already present)
if "X_umap" not in adata.obsm:
    sc.tl.umap(adata)

# popV built-in visualization: UMAP panel showing consensus + agreement
popv.visualization.predict_celltypes_umap(
    adata,
    save="popv_annotation_umap.png",
)
print("Saved popv_annotation_umap.png")

# Custom UMAP panels
fig, axes = plt.subplots(1, 3, figsize=(21, 6))
sc.pl.umap(adata, color="popv_prediction", ax=axes[0],
           title="popV Consensus", legend_loc="on data",
           legend_fontsize=6, show=False)
sc.pl.umap(adata, color="popv_agreement", ax=axes[1],
           cmap="RdYlGn", vmin=0, vmax=1,
           title="Method Agreement Score", show=False)
sc.pl.umap(adata, color="_dataset", ax=axes[2],
           title="Reference vs Query", show=False)
plt.tight_layout()
plt.savefig("popv_custom_umap.png", dpi=150, bbox_inches="tight")
print("Saved popv_custom_umap.png")

Key Concepts

Method Ensemble and Majority Voting

popV runs each method independently; the final prediction is determined by plurality vote across all methods. The popv_agreement score equals the fraction of methods that voted for the winning label (e.g., 0.7 = 7/10 methods agreed). This design has several properties:

  • Robustness: if one method fails or produces outlier labels, the consensus is unaffected if the remaining methods agree
  • Uncertainty signal: low agreement does not mean the annotation is wrong — it often flags biologically interesting cells (transitional states, rare populations) that differ from all reference cell types
  • Method independence: KNN-based methods depend on the embedding quality; tree-based methods (RF, XGBoost) work directly on expression; SVM works in feature space; CellTypist uses a separate logistic regression. Together they span multiple algorithmic families
Method Comparison
MethodBatch CorrectionSpeedBest For
knn_harmonyHarmonyFastModerate batch effects, large datasets
knn_bbknnBBKNNFastDiverse multi-tissue references
knn_scanoramaScanoramaFastMultiple heterogeneous batches
knn_scviscVI VAEMediumComplex batch effects, probabilistic embedding
scanvi_popvscVI+labelsSlowSemi-supervised; most accurate when reference is clean
celltypist_popvNone (logistic)FastImmune cells; works well without batch correction
rfNoneMediumBalanced class distributions; interpretable feature importance
xgboostNoneMediumHigh-confidence predictions on well-separated cell types
svmNoneMediumHigh-dimensional gene expression; linear boundaries
onclassNoneMediumOntology-aware; handles unseen cell types via CL ontology
ONCLASS and Ontology-Aware Annotation

ONCLASS uses the Cell Ontology (CL) to represent cell types as nodes in a knowledge graph and predict unseen cell types by propagating similarity through the ontology. Unlike other methods, ONCLASS can predict a cell type that was not present in the training reference if it is ontologically adjacent to known types. Enable it by including "onclass" in the methods list.

Reference Quality Requirements

popV annotation quality scales directly with reference quality:

  • Minimum cell count per type: 50–100 cells per label; rare types with <20 cells may be missed by KNN methods
  • Balanced representation: highly imbalanced references (one type is 80% of cells) cause tree methods to be biased toward the majority class
  • Label granularity: coarse labels (10 types) annotate reliably; fine-grained labels (100+ types) require a larger, matched reference

Common Workflows

Workflow 1: Standard Reference-Query Annotation

Goal: Annotate an unlabeled query dataset using a curated reference atlas end-to-end.

python
import popv
import scanpy as sc
import pandas as pd

# 1. Load data
adata_ref = sc.read_h5ad("reference_atlas.h5ad")   # has obs["cell_type"] and obs["batch"]
adata_query = sc.read_h5ad("query_dataset.h5ad")   # no cell type labels
if "batch" not in adata_query.obs.columns:
    adata_query.obs["batch"] = "query"

# 2. Preprocess: build joint normalized object
adata = popv.preprocessing.Process_Query(
    adata_ref,
    adata_query,
    ref_labels_key="cell_type",
    ref_batch_key="batch",
    query_batch_key="batch",
    unknown_celltype_label="unknown",
    save_path_trained_models="./popv_models/",
    n_epochs_unsupervised=100,
    n_epochs_semisupervised=30,
    use_gpu=True,
    hvg=4000,
)
print(f"Prepared: {adata.n_obs} total cells")

# 3. Run ensemble annotation
popv.annotation.annotate_data(adata)

# 4. Extract query results
query_mask = adata.obs["_dataset"] == "query"
query_annotations = adata[query_mask].obs[[
    "popv_prediction", "popv_agreement",
    "knn_harmony_popv", "scanvi_popv", "rf_popv", "xgboost_popv"
]].copy()

# 5. Transfer back to original query object
adata_query.obs = adata_query.obs.join(
    query_annotations, how="left"
)
print(f"Annotated {query_mask.sum()} query cells")
print(query_annotations["popv_prediction"].value_counts().head(10))

# 6. Save annotated query
adata_query.write_h5ad("annotated_query.h5ad", compression="gzip")
query_annotations.to_csv("popv_annotations.csv")
print("Saved annotated_query.h5ad and popv_annotations.csv")
Workflow 2: Confidence Filtering and Novel Cell State Detection

Goal: Separate high-confidence annotations from ambiguous cells; flag candidate novel or transitional states for manual review.

python
import popv
import scanpy as sc
import pandas as pd
import matplotlib.pyplot as plt

# Assume adata has been annotated (as in Workflow 1)
query_mask = adata.obs["_dataset"] == "query"
query_obs = adata[query_mask].obs.copy()

# Tier cells by agreement score
bins = [0.0, 0.5, 0.8, 1.01]
labels = ["low (<0.5)", "medium (0.5–0.8)", "high (≥0.8)"]
query_obs["confidence_tier"] = pd.cut(
    query_obs["popv_agreement"], bins=bins, labels=labels, right=False
)
print("Cells per confidence tier:")
print(query_obs["confidence_tier"].value_counts())

# High-confidence subset: use popv_prediction directly
high_conf_mask = query_obs["popv_agreement"] >= 0.8
print(f"\nHigh-confidence annotations ({high_conf_mask.mean():.1%} of query cells):")
print(query_obs[high_conf_mask]["popv_prediction"].value_counts().head(10))

# Low-confidence subset: inspect per-method disagreement
low_conf = query_obs[query_obs["popv_agreement"] < 0.5]
popv_method_cols = [c for c in query_obs.columns if c.endswith("_popv") and
                    c not in ("popv_prediction", "popv_agreement")]
print(f"\nLow-confidence cells sample (showing per-method labels):")
print(low_conf[popv_method_cols + ["popv_prediction"]].head(10).to_string())

# Save agreement + confidence-tier summaries; render the tier bar chart with the omics-plotting
# SKILL (skills/data-visualization/omics-plotting/SKILL.md) "Box / Violin / Bar" recipe -> figures/popv_confidence_distribution.png
query_obs[["popv_agreement", "confidence_tier"]].to_csv("popv_confidence.csv")
print(query_obs["confidence_tier"].value_counts().to_string())
Show full SKILL.md (662 more words)Show less

Key Parameters

ParameterModuleDefaultRange / OptionsEffect
ref_labels_keyProcess_Query—Any obs columnColumn in adata_ref.obs containing training cell type labels
n_epochs_unsupervisedProcess_Query5020–500scVI training epochs; increase for better embedding on large/complex datasets
n_epochs_semisupervisedProcess_Query2010–100scANVI fine-tuning epochs on top of scVI
hvgProcess_Query40002000–8000Highly variable genes used for embedding and KNN methods
use_gpuProcess_QueryTrueTrue, FalseGPU acceleration for scVI/SCANVI; falls back to CPU automatically if no GPU
methodsannotate_dataallList of method namesSubset of methods to run; excluding slow methods (scanvi, onclass) speeds up pipeline
unknown_celltype_labelProcess_Query"unknown"Any stringLabel assigned to query cells before annotation; used to separate reference labels from query
popv_agreement(output)—0.0–1.0Fraction of methods agreeing on consensus label; >=0.8 recommended for high confidence

Best Practices

  1. Check gene overlap before running: popV performs best with >70% gene overlap between reference and query. If overlap is <50%, annotation quality degrades significantly — consider using a different reference or imputing missing genes.

    python
    shared = adata_ref.var_names.intersection(adata_query.var_names)
    print(f"Gene overlap: {len(shared) / adata_ref.n_vars:.1%}")
  2. Use raw counts as input: pass raw (un-normalized) counts in adata.X to Process_Query. popV internally applies its own normalization. Pre-normalized data can distort the scVI/SCANVI latent space.

  3. Match reference granularity to query biology: if your query contains subtypes not in the reference, no method will correctly assign them — they will appear as low-agreement cells. Either add them to the reference or accept that the consensus will assign the nearest parent type.

  4. Exclude slow methods when speed matters: scanvi_popv and onclass are the slowest. For a quick first-pass, run only knn_harmony, knn_bbknn, rf, xgboost, and celltypist_popv.

    python
    popv.annotation.annotate_data(adata, methods=["knn_harmony", "knn_bbknn", "rf", "xgboost", "celltypist_popv"])
  5. Save trained models for repeated queries: Process_Query stores scVI/SCANVI models in save_path_trained_models. Reuse these when annotating additional query batches against the same reference to avoid retraining.

Common Recipes

Recipe: Subset to High-Confidence Annotations Only

When to use: downstream analyses (DE, trajectory) require clean labels; exclude ambiguous cells.

python
import scanpy as sc

# Annotate as in Workflow 1 first
query_mask = adata.obs["_dataset"] == "query"
adata_query_annotated = adata[query_mask].copy()

# Keep only high-confidence cells
high_conf = adata_query_annotated[adata_query_annotated.obs["popv_agreement"] >= 0.8].copy()
print(f"High-confidence cells: {high_conf.n_obs} / {adata_query_annotated.n_obs} "
      f"({high_conf.n_obs/adata_query_annotated.n_obs:.1%})")
print(high_conf.obs["popv_prediction"].value_counts())

# Recompute UMAP on high-confidence subset for visualization
sc.pp.neighbors(high_conf, use_rep="X_scVI")  # use scVI embedding stored by popV
sc.tl.umap(high_conf)
sc.pl.umap(high_conf, color="popv_prediction", save="_high_conf_celltypes.png")
Recipe: Per-Method Label Comparison Heatmap

When to use: understanding where methods disagree to identify systematic biases or novel populations.

python
import pandas as pd

query_mask = adata.obs["_dataset"] == "query"
query_obs = adata[query_mask].obs.copy()

# Collect per-method columns
method_cols = [c for c in query_obs.columns
               if c.endswith("_popv") and c not in ("popv_prediction", "popv_agreement")]

# Cross-tabulate two key methods
ct = pd.crosstab(
    query_obs["knn_harmony_popv"],
    query_obs["scanvi_popv"],
    margins=False,
)
# Normalize rows
ct_norm = ct.div(ct.sum(axis=1), axis=0)
ct_norm.to_csv("popv_method_agreement.csv")
print(f"Method agreement matrix: {ct_norm.shape} -> popv_method_agreement.csv")
# Render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Expression heatmap" recipe (Blues, 0..1)
# -> figures/popv_method_agreement_heatmap.png
Recipe: Fast Annotation Without Deep Learning Methods

When to use: quick annotation without GPU or when scVI/SCANVI training is prohibitively slow (>500k cells).

python
import popv

# Process without training deep generative models (scVI not needed for KNN-Harmony)
adata = popv.preprocessing.Process_Query(
    adata_ref,
    adata_query,
    ref_labels_key="cell_type",
    ref_batch_key="batch",
    query_batch_key="batch",
    unknown_celltype_label="unknown",
    save_path_trained_models="./popv_models/",
    n_epochs_unsupervised=0,   # skip scVI training
    n_epochs_semisupervised=0, # skip scANVI training
    use_gpu=False,
    hvg=3000,
)

# Run only fast non-DL methods
popv.annotation.annotate_data(
    adata,
    methods=["knn_harmony", "knn_bbknn", "knn_scanorama", "rf", "xgboost", "svm", "celltypist_popv"],
)

query_mask = adata.obs["_dataset"] == "query"
print(adata[query_mask].obs[["popv_prediction", "popv_agreement"]].describe())

Troubleshooting

ProblemCauseSolution
KeyError: ref_labels_key not in adata_ref.obsReference lacks a cell type columnVerify the column name: print(adata_ref.obs.columns.tolist()); update ref_labels_key accordingly
Gene space mismatch errorReference and query have very few shared genesCheck adata_ref.var_names.intersection(adata_query.var_names); if <50% overlap, use a different reference or match gene panels
CUDA out-of-memory for scVI/SCANVIGPU VRAM insufficient for batch sizeSet use_gpu=False or reduce n_epochs_unsupervised; scVI falls back to CPU automatically on most systems
onclass_popv failures on small datasetsONCLASS requires sufficient label coverageRemove "onclass" from the methods list when reference has <10 cell types or <500 cells per type
Very slow annotation (>2 hours)scVI/SCANVI training on large referenceSubsample reference to 50k cells per type; exclude "scanvi_popv" and "onclass" from methods
All cells receive same consensus labelReference highly imbalanced toward one typeBalance reference by subsampling the dominant type or upsampling rare types before running popV
popv_agreement is 0 for many cellsMany methods returning different labelsInspect per-method columns; consider whether reference covers the query biology; add methods or retrain with a better reference
  • celltypist-cell-annotation — single-model annotation with pre-trained logistic regression; faster but lacks ensemble uncertainty
  • scanpy-scrna-seq — preprocessing pipeline (QC, normalization, clustering) that produces AnnData inputs for popV
  • scvi-tools-single-cell — scANVI for probabilistic label transfer with a single deep generative model; use when you prefer a formal variational framework over ensemble voting
  • harmony-batch-correction — Harmony embedding used by knn_harmony method internally; understand it to tune popV's KNN-based methods

References

© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/single-cell/popv-cell-annotation of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Popv Cell Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Popv Cell Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Popv Cell Annotation this skilljaechang-hits/SciAgent-Skills3742 repos~6.9kAutomated safety check: PassBSD-3-Clause
Gtars Genomic Interval Toolkitdavila7/claude-code-templates33k11 repos~1.9kAutomated safety check: PassMIT
Bio Spatial Transcriptomics Spatial PreprocessingFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2kAutomated safety check: PassNone
Bio Clip Seq M6a ClipGPTomics/bioSkills1.2k2 repos~5.7kAutomated safety check: PassMIT
External Model Validationaipoch/medical-research-skills1.9k—~3.2kAutomated safety check: PassMIT
Bio Imaging Mass Cytometry Data PreprocessingGPTomics/bioSkills1.2k1 repos~4.2kAutomated safety check: PassMIT

Similar skills

  • Gtars Genomic Interval Toolkit

    davila7/claude-code-templates

    Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.

    33k GitHub starsUsed in 11 repos~1.9k tokens
    Research & ScienceAuto-check passed
  • Bio Spatial Transcriptomics Spatial Preprocessing

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Bio Clip Seq M6a Clip

    GPTomics/bioSkills

    Map N6-methyladenosine (m6A) RNA modifications at single-nucleotide resolution using miCLIP (Linder 2015), miCLIP2 + m6Aboost machine learning (Kortel 2021), GLORI (Liu 2023, antibody-free chemical…

    1.2k GitHub starsUsed in 2 repos~5.7k tokens
    Research & ScienceAuto-check passed
  • External Model Validation

    aipoch/medical-research-skills

    A skill your agent uses when validating an existing prognostic risk signature on an external bulk expression cohort with survival outcomes, producing risk scores, Kaplan-Meier curves, risk…

    1.9k GitHub stars~3.2k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Load and preprocess imaging mass cytometry (IMC) and MIBI data from raw MCD/TXT through hot-pixel removal, spillover compensation, and variance-stabilizing transformation, covering readimc/steinbock…

    1.2k GitHub starsUsed in 1 repo~4.2k tokens
    Research & ScienceAuto-check passed
  • Infers directed, time-delayed gene regulatory edges from BULK time-series expression using Granger causality (statsmodels VAR F-test), dynGENIE3 (tree ensembles regressing ODE-derived derivatives…

    1.2k GitHub starsUsed in 1 repo~5k tokens
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Popv Cell Annotation

What does Popv Cell Annotation do?

Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via…. Popv Cell Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via majority voting.

When should I use Popv Cell Annotation?

Popv Cell Annotation fits situations like: single-method annotation is insufficient; you need ensemble uncertainty for novel states.

How do I install Popv Cell Annotation in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill popv-cell-annotation -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/popv-cell-annotation in jaechang-hits/SciAgent-Skills) into .claude/skills/popv-cell-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Popv Cell Annotation in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill popv-cell-annotation -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/single-cell/popv-cell-annotation in jaechang-hits/SciAgent-Skills) into .agents/skills/popv-cell-annotation in your project. Codex loads it when a task matches its description.

Can I use Popv Cell Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill popv-cell-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/popv-cell-annotation, .gemini/skills/popv-cell-annotation, .github/skills/popv-cell-annotation and .opencode/skills/popv-cell-annotation in your project.

What does Popv Cell Annotation need to run?

Going by SKILL.md and its folder, Popv Cell Annotation needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Popv Cell Annotation access the network?

SKILL.md names 3 domains. As links in the text: doi.org, github.com and popv.readthedocs.io. This is read from the text; nothing was executed.

Is Popv Cell Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Popv Cell Annotation use?

Popv Cell Annotation is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Popv Cell Annotation use?

About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Popv Cell Annotation?

Skills that share tags, products or a category with Popv Cell Annotation: Gtars Genomic Interval Toolkit (davila7/claude-code-templates, 33k stars), Bio Spatial Transcriptomics Spatial Preprocessing (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Bio Clip Seq M6a Clip (GPTomics/bioSkills, 1.2k stars) and External Model Validation (aipoch/medical-research-skills, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Popv Cell Annotation?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.