Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Decision framework for manual marker-based, automated (CellTypist), and reference-based (popV) cell type annotation in scRNA-seq.
$ npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install BioTender-max/awesome-bio-agent-skills single-cell-annotation-guide --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/BioTender-max/awesome-bio-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sciagent/single-cell-annotation-guide .claude/skills/single-cell-annotation-guide && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "single-cell-annotation-guide" agent skill from https://github.com/BioTender-max/awesome-bio-agent-skills/tree/main/skills/sciagent/single-cell-annotation-guide into .claude/skills/single-cell-annotation-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-annotation-guide", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/BioTender-max/awesome-bio-agent-skills/tree/main/skills/sciagent/single-cell-annotation-guideType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install BioTender-max/awesome-bio-agent-skills single-cell-annotation-guide --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BioTender-max/awesome-bio-agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/sciagent/single-cell-annotation-guide .agents/skills/single-cell-annotation-guide && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "single-cell-annotation-guide" agent skill from https://github.com/BioTender-max/awesome-bio-agent-skills/tree/main/skills/sciagent/single-cell-annotation-guide into .agents/skills/single-cell-annotation-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-annotation-guide", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install BioTender-max/awesome-bio-agent-skills single-cell-annotation-guide --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BioTender-max/awesome-bio-agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/sciagent/single-cell-annotation-guide .cursor/skills/single-cell-annotation-guide && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "single-cell-annotation-guide" agent skill from https://github.com/BioTender-max/awesome-bio-agent-skills/tree/main/skills/sciagent/single-cell-annotation-guide into .cursor/skills/single-cell-annotation-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-annotation-guide", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/BioTender-max/awesome-bio-agent-skills.git --path skills/sciagent/single-cell-annotation-guide--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install BioTender-max/awesome-bio-agent-skills single-cell-annotation-guide --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BioTender-max/awesome-bio-agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/sciagent/single-cell-annotation-guide .gemini/skills/single-cell-annotation-guide && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "single-cell-annotation-guide" agent skill from https://github.com/BioTender-max/awesome-bio-agent-skills/tree/main/skills/sciagent/single-cell-annotation-guide into .gemini/skills/single-cell-annotation-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-annotation-guide", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install BioTender-max/awesome-bio-agent-skills single-cell-annotation-guideInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/BioTender-max/awesome-bio-agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/sciagent/single-cell-annotation-guide .github/skills/single-cell-annotation-guide && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "single-cell-annotation-guide" agent skill from https://github.com/BioTender-max/awesome-bio-agent-skills/tree/main/skills/sciagent/single-cell-annotation-guide into .github/skills/single-cell-annotation-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-annotation-guide", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install BioTender-max/awesome-bio-agent-skills single-cell-annotation-guide --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BioTender-max/awesome-bio-agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/sciagent/single-cell-annotation-guide .opencode/skills/single-cell-annotation-guide && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "single-cell-annotation-guide" agent skill from https://github.com/BioTender-max/awesome-bio-agent-skills/tree/main/skills/sciagent/single-cell-annotation-guide into .opencode/skills/single-cell-annotation-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "single-cell-annotation-guide", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
single-cell-annotation-guideDecision framework for manual marker-based, automated (CellTypist), and reference-based (popV) cell type annotation in scRNA-seq.
Single Cell Annotation Guide is an agent skill from BioTender-max/awesome-bio-agent-skills. Decision framework for manual marker-based, automated (CellTypist), and reference-based (popV) cell type annotation in scRNA-seq. Three-tier strategy: Tier 1 manual markers, Tier 2 CellTypist, Tier 3 popV ensemble transfer. Use when planning or troubleshooting annotation.
Its SKILL.md is about 7.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: A curated collection of AI agent skills for biomedical research, covering genomics, proteomics, single-cell analysis, clinical AI, and protein design. The licence is CC-BY-4.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 8cbdd18. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
doi.orgcelltypist.orggithub.comhumancellatlas.orgscanpy.readthedocs.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Single Cell Annotation Guide loads about 7.7k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 3,893 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from BioTender-max/awesome-bio-agent-skills at commit 8cbdd18, republished under its CC-BY-4.0 licence (© BioTender-max). 3,893 words, ~7,724 tokens.
.claude/skills/single-cell-annotation-guide/SKILL.md (or your agent's skills folder).Cell type annotation is the process of assigning biological identities to computationally defined clusters in single-cell RNA-seq data. It is one of the most consequential analytical decisions in a scRNA-seq project: annotation errors propagate into downstream analyses of differential expression, trajectory inference, and cell-cell communication. This guide presents a three-tier decision strategy — manual marker-based annotation first, automated reference-free classification second, and ensemble reference-based label transfer third — and explains when each approach is most appropriate.
The guide synthesizes community know-how on CellTypist (Dominguez Conde et al., Science 2022), popV (Luecken et al., Nature Methods 2024), and classical marker-based approaches, following standards established by the Human Cell Atlas project.
The three tiers represent a progression from effort-intensive but transparent (manual) to efficient and scalable (automated). They are not mutually exclusive: best practice is to run automated annotation first to generate hypotheses and then validate with manual marker inspection. For high-stakes biological claims — rare cell types, novel disease states, clinical applications — all three tiers should be used in parallel and discordant results resolved explicitly before publication.
A cell type marker is a gene whose expression distinguishes one cell population from all others in a dataset. Canonical markers are those validated across many studies and tissues — for example, CD3E for T cells, CD19 and MS4A1 (CD20) for B cells, CD14 for classical monocytes, and EPCAM for epithelial cells. Effective markers fulfill three criteria: they are highly expressed in the target cell type (high sensitivity), they are absent or very low in all other cell types (high specificity), and their identity is confirmed by at least two independent markers.
Clusters produced by algorithms such as Leiden or Louvain represent groups of transcriptionally similar cells; they do not inherently correspond to biological cell types. A single true cell type may appear as multiple clusters if the resolution parameter is too high (overclustering), and biologically distinct cell types may be merged if resolution is too low. Annotation quality depends on both the quality of the clustering and the quality of the evidence used to assign identities.
Reference atlases are large, curated scRNA-seq datasets with expert-validated cell type labels that serve as a "ground truth" for annotation of new query datasets. Prominent examples include the Human Cell Atlas (HCA), the Human Lung Cell Atlas (HLCA, 2.4M cells), Tabula Sapiens (500K cells across 28 tissues), and the Human BM Atlas for bone marrow. Label transfer is the computational process of projecting query cells onto the reference space and assigning labels based on proximity.
Label transfer quality depends critically on biological match between query and reference. Adult tissue query data annotated with a fetal reference will produce systematic errors for cell types that differ between developmental stages. Similarly, a blood-derived reference will poorly annotate tumor-infiltrating immune cells, which have altered transcriptional programs compared to their circulating counterparts.
Automated annotation tools produce confidence scores (probability or agreement rate) that quantify the certainty of each cell's label. These scores should never be ignored. Cells with low confidence scores may represent novel cell types absent from the reference, doublets (two cells captured together), or transitional states along a differentiation continuum.
Validation refers to verifying that automated labels are consistent with independent biological evidence — typically canonical marker gene expression. Even when an automated method reports high confidence, a few minutes spent visualizing marker genes in a UMAP or dotplot is essential to catch systematic misannotations, especially for rare or disease-specific cell populations absent from training data.
Doublets are droplets containing two or more cells, which appear as a single cell with an anomalously high gene count and mixed transcriptional identity. Annotating doublets before removing them produces spurious "cell types" that combine markers from two real populations (e.g., an apparent NK-B cell hybrid). Tools such as Scrublet and DoubletFinder should be applied before annotation; doublets should be marked and excluded from the annotation workflow.
Batch effects — systematic technical differences between samples processed at different times, with different reagents, or on different platforms — can mimic cell type differences. If one batch contains more stressed cells (higher stress gene expression), they may cluster separately and be misannotated as a distinct cell type. Batch correction with Harmony, scVI, or BBKNN should be applied before annotation when multiple batches are present.
Cell ontologies are formal vocabularies that define cell type names, synonyms, and parent-child relationships. The Cell Ontology (CL), maintained by the OBO Foundry, is the standard used by the Human Cell Atlas and most major atlases. Using ontology-compliant cell type names (e.g., "CD8-positive, alpha-beta T cell" rather than "CD8 T cell") enables cross-dataset comparison and interoperability. CellTypist and popV return labels that map to the Cell Ontology when tissue-matched models are used.
Annotation hierarchies reflect that cell identities exist at multiple levels of granularity. At the coarsest level, cells are classified as broad lineages (immune, epithelial, stromal). At intermediate granularity, they are classified as cell types (T cell, B cell, macrophage). At the finest granularity, they are classified as cell states or subtypes (exhausted CD8 T cell, tissue-resident memory T cell, M1 macrophage). Choosing the right granularity for an annotation depends on the biological question and the depth of sequencing: underpowered studies should not attempt fine-grained annotation.
Marker gene evidence is not all equally reliable. Evidence categories, from most to least reliable, are:
sc.tl.rank_genes_groups. These are hypothesis-generating and require biological interpretation.rank_genes_groups in your dataset. This is the strongest evidence — it means the marker's biology is reflected in your actual data.When building a marker panel for annotation, prioritize cross-validated markers (Category 3) and use canonical published markers (Category 1) as anchors.
When planning a cell type annotation strategy, choose the tier based on dataset characteristics and scientific context:
Is the tissue and cell type composition well-characterized?
├── YES, standard tissue with clear markers (blood, lung, gut, brain)
│ ├── Dataset < 5,000 cells OR small team experiment
│ │ └── Tier 1: Manual marker-based annotation
│ │ (dotplot + violin + UMAP visualization)
│ └── Dataset ≥ 5,000 cells OR high-throughput study
│ ├── Standard tissue with pre-trained atlas available
│ │ └── Tier 2: CellTypist automated annotation
│ │ (select tissue-matched model, use majority_voting)
│ └── Need cross-atlas integration or rare cell types
│ └── Tier 3: popV ensemble label transfer
│ (requires reference AnnData + tissue label)
└── NO, novel tissue, disease state, or poorly characterized system
├── Reference atlas exists for related tissue
│ └── Tier 3: popV ensemble label transfer
│ (multiple methods + consensus = more robust)
└── No good reference exists
└── Tier 1: Manual annotation (cautious, exploratory)
+ Tier 2: CellTypist as validation
(report as "putative" cell types, note limitations)| Scenario | Recommended Approach | Rationale |
|---|---|---|
| Small dataset (<5,000 cells), well-characterized tissue | Tier 1: Manual markers | Automated tools provide marginal advantage; manual review is faster and more transparent |
| Large dataset (>50,000 cells), blood or immune tissue | Tier 2: CellTypist with Immune_All_High model | CellTypist has the most complete immune training data; scales without added complexity |
| Large dataset, lung or airway tissue | Tier 2: CellTypist with Human_Lung_Atlas model | HLCA-trained model provides granular lung-specific annotation |
| Novel disease condition or patient-derived samples | Tier 3: popV with tissue-matched reference | Ensemble consensus is more robust to distribution shift; produces a method-agreement score |
| Integration with published atlas (e.g., HLCA, HBA) | Tier 3: popV | Designed for atlas-query integration; maintains label ontology compatibility |
| Rare cell types (<1% of dataset) | Tier 3: popV + manual inspection | Rare populations benefit from multi-method ensemble; manual inspection catches misannotation |
| Fetal or developmental tissue | Tier 1 + Tier 3 with fetal reference | Fetal cell types are distinct from adult; use a matched fetal reference (Tabula Sapiens fetal) |
| Mixed or uncertain tissue composition | Tier 2 (CellTypist) then Tier 1 validation | CellTypist provides a quick hypothesis; validate each predicted type with markers |
| Benchmark or methods comparison study | All three tiers in parallel | Triangulating annotation from multiple strategies increases credibility |
Remove doublets before annotation: Scrublet or DoubletFinder should be applied as part of QC, before any clustering or annotation step. Doublets score high on gene count and UMI metrics. Annotating them produces hybrid "cell types" that confuse downstream analyses and inflate the apparent cell type diversity. Apply doublet detection, mark predicted doublets in adata.obs, and exclude them before proceeding to annotation.
Visualize canonical markers after any automated annotation: Automated tools, including CellTypist and popV, can be wrong. After obtaining automated labels, always inspect canonical marker expression using a dotplot (sc.pl.dotplot), violin plot (sc.pl.violin), or feature plot on UMAP (sc.pl.umap). If a population labeled "B cell" does not express MS4A1 and CD19, the label is suspect. This step takes 15–30 minutes and catches the most common annotation errors.
Use multiple independent methods and compare consensus: No single annotation method is universally superior. Comparing at least two methods (e.g., CellTypist and manual markers, or CellTypist and popV) builds confidence in labels where methods agree and flags cells requiring deeper inspection where they disagree. popV's built-in method-agreement score (number of methods that agree divided by total methods) directly quantifies this; a score below 0.5 suggests the cell may be a novel or ambiguous type.
Select reference atlases that match tissue, species, and developmental stage: The most common cause of poor label transfer is biological mismatch between query and reference. An adult lung query dataset should be annotated with an adult lung reference (HLCA), not a pan-tissue reference trained on different proportions of cell types. If the reference uses different developmental stage, health status, or sequencing technology, label transfer accuracy drops significantly. When no perfectly matched reference exists, use popV's ensemble, which is more tolerant of imperfect matches than single-method transfer.
Check annotation confidence scores and flag low-confidence cells: Both CellTypist (confidence score, 0–1) and popV (popv_score, fraction of methods agreeing) provide per-cell confidence metrics. Cells with CellTypist confidence below 0.5 or popV score below 0.5 should be flagged for manual review rather than automatically assigned a label. These cells often represent rare populations, transitional states, or doublets that passed the initial QC filter. Publishing results without confidence thresholds obscures annotation uncertainty.
Correct for batch effects before reference-based annotation: When query data contains multiple batches or samples, integrate them (Harmony, scVI, or BBKNN) before label transfer. Uncorrected batch effects create spurious cluster structure that cross-sample annotation methods cannot distinguish from true biology. Run batch correction → compute joint UMAP → apply annotation, not annotation → batch correction.
Use tissue-specific resolution when constructing annotation hierarchy: Cell type granularity should match the biological question. For a study of broad immune composition, major lineage labels (T cell, B cell, myeloid, NK) may suffice. For studies of T cell function, sub-type resolution (CD4 naive, CD4 Treg, CD8 effector, CD8 exhausted) is needed. Use Leiden clustering at multiple resolutions and annotate at the granularity appropriate for the question — not the maximum possible granularity.
Iterate between clustering resolution and annotation: Overclustering (too many clusters) creates redundant splits within a single cell type; underclustering merges distinct cell types. After initial annotation, if two clusters receive the same label from multiple methods, consider merging them. If a labeled cluster is heterogeneous on marker inspection, increase resolution and re-cluster. This iteration is normal and should be documented in the methods section.
Annotating with a single marker gene per cell type: Assigning "T cell" based solely on CD3E expression will misannotate cells that briefly co-express CD3E during differentiation, doublets of CD3E+ T cells with other cell types, or low-quality cells with ambient RNA contamination.
Skipping doublet removal before annotation: Doublets containing, for example, a T cell and a B cell will express both CD3E and CD19. They form their own cluster and may be annotated as an unusual "co-expressing" cell type that does not exist biologically.
scrub.scrub_doublets()) or DoubletFinder immediately after QC filtering, before any normalization or clustering. Set the expected doublet rate to the multiplet rate reported by the 10x Genomics cell ranger summary (typically 1–4% per thousand cells loaded). Filter cells with predicted_doublet == True from adata.obs before proceeding.Using an atlas from a mismatched tissue or developmental stage: Annotating adult colon cells with a fetal intestine reference, or annotating diseased tissue with a healthy tissue reference, produces systematic label transfer errors that can affect all clusters in the dataset.
celltypist.models.models_description()) or the popV reference documentation to confirm match. If no matched reference exists, use Tier 1 manual annotation as the primary approach and flag the uncertainty.Ignoring batch effects that mimic cell type differences: A dataset with two batches, one processed fresh and one processed after a freeze-thaw cycle, may show stress-responsive gene up-regulation in the freeze-thaw batch. These cells cluster separately and can be mistaken for a "stressed" or "activated" cell subtype.
Overclustering before annotation: Running Leiden at resolution 3.0 on a dataset of 20,000 cells may produce 80 clusters, many of which represent the same cell type at different transcriptional states or are simply noise-driven splits. Annotating all 80 clusters is laborious and produces an illusion of biological diversity.
Not checking for cell-cycle state as a confounding factor: Rapidly dividing cells (in S or G2/M phase) can cluster separately from resting cells of the same type because proliferation-related genes (MKI67, TOP2A, PCNA) drive cluster identity rather than lineage markers.
sc.tl.score_genes_cell_cycle with established gene lists. Visualize cell cycle scores on UMAP. If a cluster is primarily defined by proliferation genes, label it as "[Cell Type] proliferating" rather than treating it as a distinct lineage. Optionally, regress out cell cycle scores if you want clusters to reflect lineage rather than cell state.Failing to validate low-frequency automated predictions with marker evidence: Some automated annotation tools produce plausible-sounding labels for every cell, even for rare populations with only 50–200 cells. Without marker validation, a CellTypist label of "plasmablast" for a small cluster cannot be trusted.
Quality control and preprocessing:
sc.pp.normalize_total, sc.pp.log1p).sc.pp.highly_variable_genes).Doublet detection:
adata.obs["doublet_score"] and boolean flags in adata.obs["predicted_doublet"].Clustering:
sc.pp.neighbors, k=15–30).sc.tl.umap).Tier 1 — Manual marker-based annotation:
Tier 2 — CellTypist automated annotation (if applicable):
celltypist.models.models_description().celltypist.annotate(adata, model=model, majority_voting=True).adata.obs["celltypist_cell_type"] and adata.obs["celltypist_conf_score"].Tier 3 — popV ensemble label transfer (if applicable):
cell_type containing validated labels.adata.X.popv.process_query(query_adata, ref_adata, ...) and popv.annotate_data(query_adata).adata.obs["popv_prediction"] and agreement scores in adata.obs["popv_score"].Cross-validation and refinement:
Final annotation and export:
adata.obs["cell_type"].adata.obs["annotation_method"] (manual/celltypist/popv/manual_refined)..h5ad for downstream analysis.Canonical marker reference for common tissue types: For blood and immune tissue, use CD3E/CD3D (T cells), CD4/FOXP3 (CD4 T/Tregs), CD8A (CD8 T), NKG7/GNLY (NK), CD19/MS4A1 (B cells), MZB1/JCHAIN (plasma), CD14/S100A8 (classical monocytes), FCGR3A (non-classical monocytes), FCER1A/CST3 (conventional dendritic cells), CLEC4C/IL3RA (plasmacytoid dendritic cells). For lung, add EPCAM/KRT18 (epithelial), PECAM1/VWF (endothelial), COL1A1/DCN (fibroblast), and SFTPC/SFTPB (alveolar type II epithelial).
CellTypist model selection: Match the model to the tissue. For immune datasets use Immune_All_High_Resolution or Immune_All_Low_Resolution. For lung use Human_Lung_Atlas. For gut use Human_Colon_Cell_Atlas or Cells_of_the_Human_Intestine. For fetal use Fetal_Human_Pancreas or the appropriate fetal model. Run celltypist.models.models_description() for the full list with tissue metadata.
popV reference selection: The Human Lung Cell Atlas (HLCA), Tabula Sapiens, and the Human BM Atlas are the most commonly used references. Download them as AnnData from CELLxGENE Census or Zenodo. Subset the reference to the tissue matching the query before running popV to reduce computation time and improve transfer specificity.
Reporting annotation in publications: Methods sections should specify: (a) software and version for each annotation tool used, (b) which markers were used for manual validation, (c) which CellTypist model was used, (d) which reference atlas was used for label transfer, (e) how discordant labels between methods were resolved, and (f) the confidence threshold applied to filter uncertain annotations. Failing to report these details makes annotation results non-reproducible.
Marker gene reference table for common cell types:
| Cell Type | Positive Markers | Negative Markers |
|---|---|---|
| CD4 T cell | CD3E, CD4, IL7R | CD8A, CD19, CD14 |
| CD8 T cell | CD3E, CD8A, GZMK | CD4, CD19, CD14 |
| NK cell | NKG7, GNLY, KLRB1 | CD3E, CD19, CD14 |
| B cell | CD19, MS4A1, CD79A | CD3E, CD14, NKG7 |
| Plasma cell | MZB1, JCHAIN, IGHG1 | CD19, CD3E, CD14 |
| Classical monocyte | CD14, S100A8, LYZ | CD3E, CD19, FCGR3A |
| Non-classical monocyte | FCGR3A, CX3CR1, LST1 | CD14, CD3E, CD19 |
| Conventional DC | FCER1A, CST3, CD1C | CD3E, CD19, CD14 |
| Plasmacytoid DC | CLEC4C, IL3RA, JCHAIN | CD3E, CD19, CD14 |
| Alveolar macrophage | MARCO, FABP4, PPARG | CD14, CD3E, CD19 |
| AT2 epithelial (lung) | SFTPC, SFTPB, LAMP3 | CD45/PTPRC, VWF |
| Endothelial cell | PECAM1, VWF, CDH5 | EPCAM, CD45/PTPRC |
| Fibroblast | COL1A1, DCN, PDGFRA | EPCAM, CD45/PTPRC |
| Epithelial cell | EPCAM, KRT18, CDH1 | CD45/PTPRC, VWF |
Tools for annotation QC and inter-rater reliability: When multiple annotators independently assign labels to the same dataset, measure agreement using Cohen's kappa or Fleiss' kappa. Agreement above 0.8 indicates high consistency; agreement below 0.6 suggests the markers or clustering resolution used are insufficient to distinguish the cell types in question. The scArches and scPhere frameworks provide additional tools for mapping query cells onto reference manifolds and quantifying mapping quality.
Ambient RNA decontamination before annotation: In droplet-based scRNA-seq, ambient RNA from lysed cells contaminates all droplets. This means cells that do not express a marker gene may appear to express it at low levels due to contamination. Tools such as SoupX or CellBender estimate and remove ambient RNA contamination. Running decontamination before annotation reduces false-positive marker signals that can cause misannotation, particularly for rare cell types whose canonical markers appear as low-level "contamination" in neighboring clusters.
Cell type composition as a sanity check: After annotation, the expected proportion of each cell type should be biologically plausible. For a peripheral blood mononuclear cell (PBMC) dataset, expect 50–70% T cells, 10–15% B cells, 10–20% monocytes, and 5–10% NK cells. If annotation yields 90% monocytes in a PBMC sample, the annotation is almost certainly wrong. Compare proportions against published reference ranges for the tissue and experimental condition.
scanpy-scrna-seq — Full scRNA-seq analysis pipeline in Scanpy covering QC, normalization, clustering, and basic marker-based annotation; use as the computational foundation before applying this guide's annotation strategycelltypist-cell-annotation — Tier 2 automated annotation tool using pre-trained logistic regression models for 100+ cell types across blood, lung, gut, brain, and other tissuespopv-cell-annotation — Tier 3 ensemble label transfer tool using 10+ methods with majority voting; use for robust annotation of novel or heterogeneous datasetsscvi-tools-single-cell — Deep generative models for batch correction (scVI) and semi-supervised annotation (scANVI); use scANVI as an alternative to popV for joint integration and annotationharmony-batch-correction — Batch correction for PCA embeddings; apply before annotation when integrating samples from multiple donors or processing batchesanndata-data-structure — Core data container for all scRNA-seq analysis; required for understanding how annotations are stored and propagated in AnnData objectscellxgene-census — Source of large curated reference atlases (Human Cell Atlas, Tabula Sapiens) for use as label transfer references in popV workflowsmuon-multiomics-singlecell — Multi-modal single-cell analysis; annotation of multi-omics data (RNA+ATAC, CITE-seq) requires considering the additional modalities as supplementary evidencemofaplus-multi-omics — Multi-omics factor analysis; latent factors from MOFA+ can inform cell state annotation when transcriptional signatures are not sufficientcellchat-cell-communication — Downstream analysis of annotated single-cell data; ligand-receptor interaction inference depends entirely on having accurate cell type labels from annotation© BioTender-max, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/sciagent/single-cell-annotation-guide of BioTender-max/awesome-bio-agent-skills.
Open the folder on GitHubat commit 8cbdd18
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in BioTender-max/awesome-bio-agent-skills, which our catalogue first saw on October 7, 2026.
Single Cell Annotation Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Single Cell Annotation Guide this skillBioTender-max/awesome-bio-agent-skills | 199 | 1 repos | ~7.7k | Automated safety check: Pass | CC-BY-4.0 | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
BioTender-max/awesome-bio-agent-skills
Critically review, score, compare, and rank one or more AI scientist outputs for biology, bioinformatics, computational life science, or adjacent research tasks.
BioTender-max/awesome-bio-agent-skills
Web toolkit powered by Exa, tuned for scientific and technical content.
BioTender-max/awesome-bio-agent-skills
Queries JGI Lakehouse (Dremio) for genomics metadata from GOLD, IMG, Mycocosm, Phytozome.
BioTender-max/awesome-bio-agent-skills
Operator toolkit for nf-core/pacsomatic matched tumor-normal workflows from BAM inputs.
BioTender-max/awesome-bio-agent-skills
Assess paper and journal impact using OpenAlex citation counts, optional Altmetric data, and curated journal impact-factor references.
BioTender-max/awesome-bio-agent-skills
Search arXiv preprints through the official arXiv API and turn arXiv IDs into local Markdown summaries.
Categories
Decision framework for manual marker-based, automated (CellTypist), and reference-based (popV) cell type annotation in scRNA-seq. Single Cell Annotation Guide is an agent skill from BioTender-max/awesome-bio-agent-skills. Decision framework for manual marker-based, automated (CellTypist), and reference-based (popV) cell type annotation in scRNA-seq.
Single Cell Annotation Guide fits situations like: troubleshooting annotation; tasks that involve Bioinformatics.
Run `npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a claude-code`. Or copy the skill folder (skills/sciagent/single-cell-annotation-guide in BioTender-max/awesome-bio-agent-skills) into .claude/skills/single-cell-annotation-guide in your project. Claude Code loads it when a task matches its description.
Run `npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a codex`. Or copy the skill folder (skills/sciagent/single-cell-annotation-guide in BioTender-max/awesome-bio-agent-skills) into .agents/skills/single-cell-annotation-guide in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BioTender-max/awesome-bio-agent-skills --skill single-cell-annotation-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/single-cell-annotation-guide, .gemini/skills/single-cell-annotation-guide, .github/skills/single-cell-annotation-guide and .opencode/skills/single-cell-annotation-guide in your project.
SKILL.md names no scripts, command-line tools or credentials: Single Cell Annotation Guide is instructions for the agent only.
SKILL.md names 5 domains. As links in the text: doi.org, celltypist.org, github.com, humancellatlas.org and scanpy.readthedocs.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Single Cell Annotation Guide is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.7k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Single Cell Annotation Guide: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
BioTender-max (a GitHub user) maintains it in BioTender-max/awesome-bio-agent-skills, which has 199 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on July 1, 2026.
Source: BioTender-max/awesome-bio-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.