Agent skill

Bio Machine Learning Atlas Mapping

by GPTomics in GPTomics/bioSkills

Maps query single-cell data onto reference atlases and transfers cell-type labels using scArches surgery (scVI/scANVI), Symphony, Azimuth, CellTypist, scPoli, popV, and foundation models, with…

MITAuto-check passedData & Analytics

Install Bio Machine Learning Atlas Mapping

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-atlas-mapping -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-machine-learning-atlas-mapping --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/machine-learning/atlas-mapping .claude/skills/bio-machine-learning-atlas-mapping && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-machine-learning-atlas-mapping
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.2k tokens
SKILL.md length
2,144 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Maps query single-cell data onto reference atlases and transfers cell-type labels using scArches surgery (scVI/scANVI), Symphony, Azimuth, CellTypist, scPoli, popV, and foundation models, with…

  • Annotating new single-cell datasets against a pre-trained reference
  • SKILL.md covers Version Compatibility, The Single Most Important…, Methods Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Python scripts from its folder; calls pip
  • Deciding which mapping method fits

What it does

Bio Machine Learning Atlas Mapping is an agent skill from GPTomics/bioSkills. Maps query single-cell data onto reference atlases and transfers cell-type labels using scArches surgery (scVI/scANVI), Symphony, Azimuth, CellTypist, scPoli, popV, and foundation models, with explicit out-of-distribution and label-transfer uncertainty. Use when annotating new single-cell datasets against a pre-trained reference, deciding which mapping method fits, or judging whether transferred labels are trustworthy. For de novo clustering and manual annotation see single-cell/cell-annotation; for batch…

Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/ood_gating_demo.py`, `examples/scarches_annotation.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Bioinformatics and Machine learning. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Annotating new single-cell datasets against a pre-trained reference
  • Deciding which mapping method fits
  • Judging whether transferred labels are trustworthy

Example prompts

  • “/bio-machine-learning-atlas-mapping”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Machine Learning Atlas Mapping loads about 5.2k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 2,144 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~153
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,144 words, ~5,154 tokens.

Download SKILL.mdSave it as .claude/skills/bio-machine-learning-atlas-mapping/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-machine-learning-atlas-mapping
description
Maps query single-cell data onto reference atlases and transfers cell-type labels using scArches surgery (scVI/scANVI), Symphony, Azimuth, CellTypist, scPoli, popV, and foundation models, with explicit out-of-distribution and label-transfer uncertainty. Use when annotating new single-cell datasets against a pre-trained reference, deciding which mapping method fits, or judging whether transferred labels are trustworthy. For de novo clustering and manual annotation see single-cell/cell-annotation; for batch integration without a reference see single-cell/batch-integration.
tool_type
python
primary_tool
scvi-tools

Version Compatibility

Reference examples tested with: anndata 0.10+, scanpy 1.10+, scvi-tools 1.1+, scikit-learn 1.3+, celltypist 1.6+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

scvi-tools 1.x has had API churn around minified/registered models and the semi-supervised loader. Confirm scvi.__version__ and help(scvi.model.SCANVI.load_query_data) before relying on argument names. If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Reference Mapping and Label Transfer for Single-Cell Data

"Annotate my scRNA-seq query against a reference atlas" -> Project query cells into a fixed reference latent space and transfer labels, then gate every label on an out-of-distribution signal.

  • scArches surgery: scvi.model.SCANVI.prepare_query_anndata() -> load_query_data() -> train(plan_kwargs={'weight_decay': 0.0}) -> predict(soft=True)
  • Linear/fixed reference: Symphony (R) or Azimuth/Seurat anchor transfer
  • Pure classifier (no embedding): CellTypist logistic regression

The Single Most Important Modern Insight -- Mapping Is a Projection, Not an Annotation Oracle

Reference mapping indicates where in a fixed reference manifold a query cell lands; it does not establish whether that location is biologically meaningful for the query. The projection always succeeds geometrically: every query cell is assigned its nearest reference label whether or not that label is true. A softmax over reference classes is normalized to sum to 1 and therefore cannot express "I have never seen this cell" -- a hepatocyte handed to an immune reference gets confidently called a T cell. Every failure mode below is a corollary of confusing projection with annotation.

The operational consequence: a transferred label is untrustworthy until paired with an out-of-distribution / label-transfer-uncertainty signal that measures whether the cell belongs to the reference at all, which is a different quantity from the prediction probability that measures which reference label. Conflating these two is the field's most common error.

Methods Taxonomy

MethodModel classNeeds reference model?Novel-cell-type signalBest whenFails when
scVI + scArches surgeryConditional VAE; surgery freezes reference weights, fits query-batch nodesYes (saved scVI model)None intrinsic; add kNN uncertainty / OOD distanceUnseen batch; want a de novo joint embedding then cluster query yourselfExpected to label (it only embeds); reference lacks query biology
scANVI + scArches surgerySemi-supervised VAE (scVI latent + label classifier head)Yes (scANVI, often from_scvi_model)Classifier softmax (overconfident); kNN-on-latent uncertaintyReference well-labeled, query is the same tissue/biologySemi-supervised leakage carves latent; novel states confidently mislabeled
SymphonyLinear Harmony soft-cluster mixture; query projected into fixed referenceYes (compressed reference object)Per-cell Mahalanobis distance to soft-cluster centroidsSeconds-scale, deterministic, CPU-only, reproducible/clinicalStrong nonlinear batch the reference never saw
Azimuth / Seurat anchor transferCCA/PCA anchors; supervised PCA projectionYes (precomputed ref)prediction.score.max, mapping.scoreMultimodal refs (CITE-seq/WNN), curated tissue atlases, R shopFiltered-anchor pathology when query is very divergent
scPoliConditional VAE + learnable sample embeddings + cell-type prototypesYes (built on scArches)Prototype distance + uncertaintyWant sample-level (patient) embeddings too, many small batchesFew samples (condition embedding underdetermined)
popVEnsemble of methods + ontology-aware votingMixed (wraps several)Cross-method disagreement = uncertaintyHigh-stakes atlas annotation; distrust any single methodCompute-heavy; consensus can be confidently wrong if all share reference bias
CellTypistLogistic regression (pre-trained models)No embedding (ships models)Low max-probability = ambiguous; no true OODFast immune annotation, no integration neededTreated as a mapper (no shared embedding, no batch handling)
treeArches / scHPLscArches + hierarchical classifier with rejectionYesExplicit rejection -> "unseen" nodeExpect novel subtypes, want hierarchy-aware fallbackMis-specified hierarchy propagates error down branches
Foundation models (scGPT, Geneformer)Transformer pretrained on 10s of millions of cellsCheckpoint; fine-tune needs labelsNone intrinsic; OOD poorly characterizedFine-tuned on the target task; cross-modality/species; data-scarceZero-shot: underperform scVI/Harmony/HVG-PCA (Kedzierska 2025)

Cross-cutting: linear methods (Symphony, Azimuth-sPCA) give a fixed, reproducible reference embedding (the query never perturbs the reference); VAE surgery fine-tunes and can drift. Reproducibility-critical or clinical pipelines lean linear; maximal batch-effect flexibility leans VAE.

Decision Tree by Scenario

ScenarioRecommended approachWhy
Same tissue as a published scVI/scANVI atlas; want one embedding + labelsscArches surgery onto the scANVI model; then gate labels on kNN uncertaintyBuilt for this; reuses the learned manifold; only query fine-tuned
Seconds-scale, deterministic, CPU-only, must re-run identically (clinical)Symphony (or Azimuth if curated)Fixed reference embedding, no fine-tuning drift, built-in Mahalanobis OOD
Query likely contains cell types/states NOT in the reference (disease, new niche)treeArches/scHPL rejection, or scArches + explicit OOD distanceDefault kNN/softmax confidently mislabels novel cells
High-stakes annotation; distrust any single methodpopV ensembleCross-method disagreement is a more honest uncertainty than any softmax
Want patient/sample-level structure tooscPoliOnly method learning sample (condition) embeddings jointly with cell prototypes
Just need fast immune labels, no integrationCellTypist (Immune_All_Low, majority_voting=True)Calibrated classifier, no embedding needed; QC the query first
Multimodal reference (CITE-seq/ATAC)Azimuth/Seurat WNN or totalVI+scArchesAnchor framework natively weights modalities
Considering scGPT/GeneformerOnly if fine-tuning with labels, or cross-modality/species, or data too scarceZero-shot foundation embeddings are not a justified default for same-tissue transfer
De novo clustering, no reference, or manual marker annotation-> single-cell/markers-annotation, single-cell/clusteringOut of scope here

scArches Surgery: Embedding (scVI)

Goal: Project query cells into a pre-trained reference latent space without retraining on combined data.

Approach: Align query genes to the reference exactly, load into the frozen reference model, and fine-tune only query-specific parameters with zero weight decay so the shared manifold does not drift.

python
import scvi
import scanpy as sc

ref_model = scvi.model.SCVI.load('reference_model/')          # saved with save_anndata=True (or minified)

adata_query = sc.read_h5ad('query.h5ad')
# Align genes to the reference EXACTLY: zero-pad missing, reorder. Mandatory and silent if skipped.
scvi.model.SCVI.prepare_query_anndata(adata_query, 'reference_model/')

query_model = scvi.model.SCVI.load_query_data(adata_query, 'reference_model/')
# weight_decay=0.0 + frozen reference weights make surgery a query-only fine-tune;
# non-zero decay drifts the shared latent and breaks cross-query comparability.
query_model.train(max_epochs=200, plan_kwargs={'weight_decay': 0.0}, check_val_every_n_epoch=10)
adata_query.obsm['X_scVI'] = query_model.get_latent_representation()

scANVI Label Transfer

Goal: Transfer reference cell-type labels to an unlabeled query.

Approach: Build a semi-supervised scANVI head on the reference, map the query by surgery, then read hard labels and per-class probabilities -- treating the probability as "which label," not "does it belong."

python
# Reference side (once): scANVI from a trained scVI model. unlabeled_category is REQUIRED.
ref_scanvi = scvi.model.SCANVI.from_scvi_model(ref_vae, unlabeled_category='Unknown', labels_key='cell_type')
ref_scanvi.train(max_epochs=20, n_samples_per_label=100)
ref_scanvi.save('ref_scanvi/', save_anndata=True)

# Query side (surgery):
scvi.model.SCANVI.prepare_query_anndata(adata_query, 'ref_scanvi/')
query_scanvi = scvi.model.SCANVI.load_query_data(adata_query, 'ref_scanvi/')
query_scanvi.train(max_epochs=100, plan_kwargs={'weight_decay': 0.0})

adata_query.obs['predicted_label'] = query_scanvi.predict()          # hard labels
adata_query.obsm['X_scANVI'] = query_scanvi.get_latent_representation()
proba = query_scanvi.predict(soft=True)                             # per-class probabilities (which label)

Out-of-Distribution Gating (the step that makes labels trustworthy)

Goal: Decide whether each query cell belongs to the reference, separately from which label it would get.

Approach: Compute a distance/entropy signal on the shared latent. The canonical scArches/HLCA approach is a weighted-kNN label-transfer uncertainty (neighbor disagreement in the reference latent), thresholded at 0.2 to set cells to "Unknown." A portable kNN-entropy version is shown; the softmax proba is NOT this signal.

python
import numpy as np
from sklearn.neighbors import KNeighborsClassifier

ref_latent = ref_scanvi.get_latent_representation()                  # reference cells in latent
knn = KNeighborsClassifier(n_neighbors=15, weights='distance').fit(ref_latent, adata_ref.obs['cell_type'])
query_latent = adata_query.obsm['X_scANVI']

neighbor_proba = knn.predict_proba(query_latent)                    # weighted neighbor label distribution
# Uncertainty = 1 - max neighbor agreement. HLCA sets cells above 0.2 to 'Unknown'.
uncertainty = 1.0 - neighbor_proba.max(axis=1)
adata_query.obs['transfer_uncertainty'] = uncertainty
adata_query.obs.loc[uncertainty > 0.2, 'predicted_label'] = 'Unknown'   # gate, do not trust ungated labels
print(f'Flagged Unknown: {(uncertainty > 0.2).mean():.1%}')

Per-Method Failure Modes

scANVI / kNN -- forcing the query onto reference labels
  • Trigger: Query contains a population absent from the reference (novel type, disease state, perturbed program).
  • Mechanism: The classifier/kNN assigns every query cell to its nearest reference label; there is no "none of the above" unless added.
  • Symptom: A coherent novel cluster split across 2-3 reference labels, each with high probability; the UMAP looks "well integrated."
  • Fix: Always gate on transfer uncertainty / OOD distance (above). Inspect query-only clusters for marker genes independent of transferred labels.
Softmax overconfidence vs label-transfer uncertainty conflated
  • Trigger: Reporting "confidence" as the predict(soft=True) max.
  • Mechanism: The softmax measures which reference label conditional on belonging; it is normalized away from distance and cannot say "far from everything." A cell can be 0.99 "T cell" and be a hepatocyte.
  • Symptom: Pipeline filters on softmax >= 0.5 and still passes OOD cells.
  • Fix: Threshold the weighted-kNN uncertainty (HLCA 0.2) or a Mahalanobis/ensemble OOD signal for the "Unknown" decision; use the softmax only to disambiguate among in-distribution labels.
scANVI -- semi-supervised label leakage / latent carving
  • Trigger: scANVI reference where labels strongly drive latent geometry; trajectory or novel-state query.
  • Mechanism: The classifier head back-propagates label structure into the latent, carving it to separate reference types; query cells are pulled toward that structure even when their biology lies between/outside it.
  • Symptom: A query continuum (differentiation trajectory) collapses onto discrete reference clusters; intermediate states vanish.
  • Fix: For trajectory/novel-state queries prefer unsupervised scVI surgery, annotate the query independently, and cross-check against the scVI latent.
Reference composition / missing-biology bias
  • Trigger: Reference from healthy/limited donors; query from disease, different age, ancestry, or tissue region.
  • Mechanism: The reference manifold is the entire hypothesis space; off-manifold cells are projected onto the nearest on-manifold point and rare reference populations act as attractors.
  • Symptom: Disease-specific states labeled as the closest healthy type; ancestry/age effects read as "batch."
  • Fix: Audit reference composition before mapping; prefer references covering the query's expected biology; treat mapping as hypothesis generation and validate query findings de novo; consider extending the reference (treeArches).
Show full SKILL.md (830 more words)Show less
Query QC artifacts laundered into confident labels
  • Trigger: Query not QC'd to the reference's standard (empty droplets, ambient RNA, doublets, high-MT).
  • Mechanism: A doublet sits between two reference types and maps to a spurious "intermediate"; ambient RNA shifts profiles toward the dominant type.
  • Symptom: Artifactual "transitional" populations; doublet clusters labeled as rare real types.
  • Fix: Run full query QC before mapping (single-cell/doublet-detection, ambient correction, MT/count filters matched to the reference). Mapping does not clean data.
Feature-space / gene-set mismatch
  • Trigger: Query missing reference HVGs; different gene annotation/version; prepare_query_anndata skipped.
  • Mechanism: The encoder expects the exact reference gene order; missing genes are zero-padded and reordering silently corrupts the input.
  • Symptom: Garbage latent, everything maps to one blob, or a silent accuracy cliff (no error raised).
  • Fix: Always prepare_query_anndata(query, reference_model); verify the shared-gene fraction; too few shared HVGs is a hard stop.
Good integration metrics, wrong labels
  • Trigger: Judging mapping by scIB integration scores alone.
  • Mechanism: Integration metrics reward query/reference mixing; mixing OOD cells into the wrong neighborhood raises the batch-removal score, and bio-conservation uses reference labels (circular for the query).
  • Symptom: Top scIB total score with biologically wrong annotation.
  • Fix: Integration metrics validate the embedding, not labels. Evaluate transfer on held-out labeled query cells (per-type F1, especially rare types), OOD detection on spiked-in unseen types, and marker-gene sanity checks.
Zero-shot foundation-model embedding as a mapper
  • Trigger: Using scGPT/Geneformer zero-shot embeddings for clustering/transfer expecting SOTA.
  • Mechanism: The masked-gene pretraining objective does not guarantee a label- or batch-aware embedding; zero-shot embeddings are not batch-corrected.
  • Symptom: Worse AvgBio/integration than scVI or even HVG-PCA + Harmony.
  • Fix: Fine-tune with task labels, or use an established mapper; reserve foundation models for cross-modality/species/data-scarce cases (Kedzierska 2025).

Reconciliation: When Methods Disagree

PatternLikely causeAction
scANVI label confident but kNN uncertainty highOOD cell forced onto nearest labelTrust the uncertainty; set Unknown and inspect markers
Symphony Mahalanobis flags OOD but scANVI does notscANVI latent carved to absorb the cellPrefer the distance-based flag; novel biology likely
popV members disagreeGenuine ambiguity or granularity mismatchRoute to manual review; report the disagreement, do not force a leaf
High scIB score, poor per-type F1 on held-out labelsEmbedding mixes well but labels wrongBelieve the F1; integration score is not a label metric

Quantitative Thresholds

ThresholdSourceRationale
Transfer uncertainty > 0.2 -> "Unknown"Sikkema 2023 (HLCA)Weighted-kNN neighbor-disagreement cutoff bounding false labels; recalibrate per reference
Surgery weight_decay=0.0, ~100-200 epochsscvi-tools scArches tutorialFrozen reference weights + no decay keep the shared latent fixed
scIB total = 0.6bio + 0.4batchLuecken 2022Benchmark weighting; scores the embedding, NOT query labels
CellTypist input = log1p of CP10kCellTypist docsWrong normalization silently degrades accuracy

Common Errors

Error / symptomCauseSolution
Everything maps to one blobprepare_query_anndata skipped; gene mismatchRun it before load_query_data; check shared-gene fraction
OOD cells pass a 0.5 softmax filterThresholding the wrong quantityGate on weighted-kNN uncertainty / OOD distance, not softmax
Reference cells move between runsNon-zero weight_decay or over-training in surgerySet weight_decay=0.0, keep freeze_* defaults, modest epochs
load_query_data errors on labelsscANVI needs unlabeled_categoryPass it to from_scvi_model; query labels filled with that category
CellTypist labels look randomRaw or wrongly normalized countsFeed log1p CP10k input

References

  • Lopez R, Regier J, Cole MB, Jordan MI, Yosef N. 2018. Deep generative modeling for single-cell transcriptomics. Nat Methods 15:1053-1058.
  • Xu C, Lopez R, Mehlman E, Regier J, Jordan MI, Yosef N. 2021. Probabilistic harmonization and annotation of single-cell transcriptomics data with deep generative models. Mol Syst Biol 17:e9620.
  • Lotfollahi M, Naghipourfar M, Luecken MD, et al. 2022. Mapping single-cell data to reference atlases by transfer learning. Nat Biotechnol 40:121-130.
  • Kang JB, Nathan A, Weinand K, et al. 2021. Efficient and precise single-cell reference atlas mapping with Symphony. Nat Commun 12:5890.
  • Hao Y, Hao S, Andersen-Nissen E, et al. 2021. Integrated analysis of multimodal single-cell data. Cell 184:3573-3587.
  • De Donno C, Hediyeh-Zadeh S, Moinfar AA, et al. 2023. Population-level integration of single-cell datasets enables multi-scale analysis across samples. Nat Methods 20:1683-1692.
  • Ergen C, Xing G, Xin C, et al. 2024. Consensus prediction of cell type labels in single-cell data with popV. Nat Genet 56:2731-2738.
  • Dominguez Conde C, Xu C, Jarvis LB, et al. 2022. Cross-tissue immune cell analysis reveals tissue-specific features in humans. Science 376:eabl5197.
  • Michielsen L, Lotfollahi M, Strobl D, et al. 2023. Single-cell reference mapping to construct and extend cell-type hierarchies. NAR Genom Bioinform 5:lqad070.
  • Luecken MD, Buttner M, Chaichoompu K, et al. 2022. Benchmarking atlas-level data integration in single-cell genomics. Nat Methods 19:41-50.
  • Sikkema L, Ramirez-Suastegui C, Strobl DC, et al. 2023. An integrated cell atlas of the lung in health and disease. Nat Med 29:1563-1577.
  • Kedzierska KZ, Crawford L, Amini AP, Lu AX. 2025. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biol 26:101.
  • single-cell/preprocessing - QC, normalization, and HVG selection the query needs before mapping
  • single-cell/markers-annotation - Manual marker-based cluster annotation when there is no reference
  • single-cell/batch-integration - Integrating datasets without a labeled reference
  • single-cell/doublet-detection - Removing doublets that map to spurious intermediates
  • differential-expression/de-results - Pseudobulk validation of mapping-derived populations

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in machine-learning/atlas-mapping of GPTomics/bioSkills.

  • SKILL.md
  • examples/ood_gating_demo.py
  • examples/scarches_annotation.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Machine Learning Atlas Mapping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Machine Learning Atlas Mapping compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Machine Learning Atlas Mapping this skillGPTomics/bioSkills1.2k1 repos~5.2kAutomated safety check: PassMIT
Comorbidity Common Immune Biomarker Research Planneraipoch/medical-research-skills2k—~4.5kAutomated safety check: PassMIT
Process Related Diagnostic Biomarker Nomogram Research Planneraipoch/medical-research-skills2k—~4.7kAutomated safety check: PassMIT
Bioconductor OrfhunterbioMate-AI/biomate-bioconductor-kb804—~1.9kAutomated safety check: PassCustom licence
tangermeme Genomic Model Analysisjmschrei/tangermeme316—~1.6kAutomated safety check: PassMIT
Gtars Genomic Interval Toolkitdavila7/claude-code-templates32k11 repos~1.9kAutomated safety check: PassMIT

Similar skills

  • Generates complete comorbidity-oriented shared-biomarker bioinformatics research designs from a user-provided disease pair and validation direction.

    2k GitHub stars~4.5k tokensUpdated 22 days ago
    Data & AnalyticsAuto-check passed
  • Generates complete process-related diagnostic biomarker bioinformatics research designs from a user-provided disease context, gene-family or pathway theme, and validation direction.

    2k GitHub stars~4.7k tokensUpdated 22 days ago
    Data & AnalyticsAuto-check passed
  • Bioconductor Orfhunter

    bioMate-AI/biomate-bioconductor-kb

    The ORFhunteR package is a R and C++ library for an automatic determination and annotation of open reading frames (ORF) in a large set of RNA molecules.

    804 GitHub stars~1.9k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.

    316 GitHub stars~1.6k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Gtars Genomic Interval Toolkit

    davila7/claude-code-templates

    Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.

    32k GitHub starsUsed in 11 repos~1.9k tokens
    Research & ScienceAuto-check passed
  • Pathml

    davila7/claude-code-templates

    Computational pathology toolkit for analyzing whole-slide images (WSI) and multiparametric imaging data.

    32k GitHub starsUsed in 11 repos~1.9k tokens
    Documents & OfficeAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Machine Learning Atlas Mapping

What does Bio Machine Learning Atlas Mapping do?

Maps query single-cell data onto reference atlases and transfers cell-type labels using scArches surgery (scVI/scANVI), Symphony, Azimuth, CellTypist, scPoli, popV, and foundation models, with…. Bio Machine Learning Atlas Mapping is an agent skill from GPTomics/bioSkills. Maps query single-cell data onto reference atlases and transfers cell-type labels using scArches surgery (scVI/scANVI), Symphony, Azimuth, CellTypist, scPoli, popV, and foundation models, with explicit out-of-distribution and label-transfer uncertainty.

When should I use Bio Machine Learning Atlas Mapping?

Bio Machine Learning Atlas Mapping fits situations like: annotating new single-cell datasets against a pre-trained reference; deciding which mapping method fits; judging whether transferred labels are trustworthy.

How do I install Bio Machine Learning Atlas Mapping in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-atlas-mapping -a claude-code`. Or copy the skill folder (machine-learning/atlas-mapping in GPTomics/bioSkills) into .claude/skills/bio-machine-learning-atlas-mapping in your project. Claude Code loads it when a task matches its description.

How do I install Bio Machine Learning Atlas Mapping in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-atlas-mapping -a codex`. Or copy the skill folder (machine-learning/atlas-mapping in GPTomics/bioSkills) into .agents/skills/bio-machine-learning-atlas-mapping in your project. Codex loads it when a task matches its description.

Can I use Bio Machine Learning Atlas Mapping in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-atlas-mapping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-machine-learning-atlas-mapping, .gemini/skills/bio-machine-learning-atlas-mapping, .github/skills/bio-machine-learning-atlas-mapping and .opencode/skills/bio-machine-learning-atlas-mapping in your project.

What does Bio Machine Learning Atlas Mapping need to run?

Going by SKILL.md and its folder, Bio Machine Learning Atlas Mapping needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Machine Learning Atlas Mapping access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Machine Learning Atlas Mapping safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Machine Learning Atlas Mapping use?

Bio Machine Learning Atlas Mapping is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Machine Learning Atlas Mapping use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Machine Learning Atlas Mapping?

Skills that share tags, products or a category with Bio Machine Learning Atlas Mapping: Comorbidity Common Immune Biomarker Research Planner (aipoch/medical-research-skills, 2k stars), Process Related Diagnostic Biomarker Nomogram Research Planner (aipoch/medical-research-skills, 2k stars), Bioconductor Orfhunter (bioMate-AI/biomate-bioconductor-kb, 804 stars) and tangermeme Genomic Model Analysis (jmschrei/tangermeme, 316 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Machine Learning Atlas Mapping?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.