Agent skill

Cellxgene Census

by aipoch in aipoch/medical-research-skills

Programmatically query the CZ CELLxGENE Census (61M+ cells) when you need cross-tissue, disease, or cell-type expression data for population-scale queries and reference atlas comparisons.

MITAuto-check passedResearch & Science

Install Cellxgene Census

skills CLI
$ npx skills add aipoch/medical-research-skills --skill cellxgene-census -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills cellxgene-census --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Evidence Insight/cellxgene-census' .claude/skills/cellxgene-census && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cellxgene-census
GitHub stars
2k
Token cost
~1.6k tokens
SKILL.md length
378 words
Files
4 (incl. references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Programmatically query the CZ CELLxGENE Census (61M+ cells) when you need cross-tissue, disease, or cell-type expression data for population-scale queries and reference atlas comparisons.

  • Works in 4 steps: opens a pinned Census version, → explores metadata, → loads a small AnnData slice, and → …
  • Research & Science work in your project
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 1 more section
  • Calls uv

What it does

Cellxgene Census is an agent skill from aipoch/medical-research-skills. Programmatically query the CZ CELLxGENE Census (61M+ cells) when you need cross-tissue, disease, or cell-type expression data for population-scale queries and reference atlas comparisons.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `cellxgene-census_audit_result_v1.json`, `references/census_schema.md` and `references/common_patterns.md`).

It sits in Research & Science. It works with AnnData. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Research & Science work in your project

Example prompts

  • “/cellxgene-census”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. opens a pinned Census version,
  2. explores metadata,
  3. loads a small AnnData slice, and
  4. runs an out-of-core query to compute a simple statistic.

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cellxgene Census loads about 1.6k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 51 tokens; SKILL.md has 378 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 378 words, ~1,625 tokens.

Download SKILL.mdSave it as .claude/skills/cellxgene-census/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
cellxgene-census
description
Programmatically query the CZ CELLxGENE Census (61M+ cells) when you need cross-tissue, disease, or cell-type expression data for population-scale queries and reference atlas comparisons.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

  • Cross-tissue or cross-disease expression comparisons (e.g., macrophages across lung/liver/brain; COVID-19 vs control).
  • Reference atlas lookups to contextualize findings from your own single-cell dataset (marker validation, expected expression patterns).
  • Population-scale metadata exploration (what tissues/cell types/datasets exist; cell counts by cohort attributes).
  • Large-scale expression statistics where results exceed RAM and require out-of-core iteration.
  • Model training on curated atlas data (e.g., cell-type classifiers) using the experimental PyTorch integration.

Key Features

  • Programmatic access to versioned CZ CELLxGENE Census data (human and mouse).
  • Query cell (obs) metadata and gene (var) metadata with expressive filter syntax.
  • Retrieve expression as AnnData for small/medium queries via get_anndata().
  • Perform out-of-core expression access via SOMA axis_query() and chunked iteration.
  • Optional experimental ML utilities (PyTorch dataloaders/datasets).
  • Works well with scanpy workflows after loading AnnData.

Dependencies

  • cellxgene-census (latest)
  • tiledbsoma (latest; required for axis_query() workflows)
  • pyarrow (latest; used for chunked table batches)
  • anndata (latest; for get_anndata() results)
  • scanpy (latest; optional, for downstream analysis)
  • torch (latest; optional, for experimental ML integration)

Install:

bash
uv pip install cellxgene-census

Optional (experimental ML helpers):

bash
uv pip install cellxgene-census[experimental]

Example Usage

The following script is a complete, runnable example that:

  1. opens a pinned Census version,
  2. explores metadata,
  3. loads a small AnnData slice, and
  4. runs an out-of-core query to compute a simple statistic.
python
import numpy as np
import cellxgene_census
import tiledbsoma as soma

def main():
    # Pin a version for reproducibility (replace with a valid release if needed)
    census_version = "2023-07-25"

    with cellxgene_census.open_soma(census_version=census_version) as census:
        # 1) Explore summary info
        summary = census["census_info"]["summary"].read().concat().to_pandas()
        total_cells = int(summary["total_cell_count"].iloc[0])
        print(f"Census version: {census_version}")
        print(f"Total cells: {total_cells:,}")

        # 2) Explore obs metadata (always filter primary data unless you want duplicates)
        obs = cellxgene_census.get_obs(
            census,
            "homo_sapiens",
            value_filter="tissue_general == 'brain' and is_primary_data == True",
            column_names=["cell_type", "tissue_general", "disease", "donor_id"],
        )
        print(f"Brain (primary) cells returned (metadata only): {len(obs):,}")
        print("Top cell types:")
        print(obs["cell_type"].value_counts().head(10))

        # 3) Small/medium query -> AnnData in memory
        adata = cellxgene_census.get_anndata(
            census=census,
            organism="Homo sapiens",
            obs_value_filter=(
                "cell_type == 'T cell' and disease == 'COVID-19' and is_primary_data == True"
            ),
            var_value_filter="feature_name in ['CD4', 'CD8A', 'FOXP3']",
            obs_column_names=["cell_type", "tissue_general", "disease", "donor_id", "sex"],
        )
        print(adata)
        print("AnnData X shape:", adata.X.shape)

        # 4) Large-scale pattern -> out-of-core iteration with axis_query()
        # Example: compute mean of non-zero expression values for a few genes in brain.
        query = census["census_data"]["homo_sapiens"].axis_query(
            measurement_name="RNA",
            obs_query=soma.AxisQuery(
                value_filter="tissue_general == 'brain' and is_primary_data == True"
            ),
            var_query=soma.AxisQuery(
                value_filter="feature_name in ['FOXP2', 'TBR1', 'SATB2']"
            ),
        )

        n = 0
        s = 0.0
        for batch in query.X("raw").tables():
            # batch is a pyarrow.Table with at least: soma_data, soma_dim_0, soma_dim_1
            values = batch["soma_data"].to_numpy(zero_copy_only=False)
            n += values.size
            s += float(values.sum())

        mean_expr = s / n if n else np.nan
        print(f"Out-of-core mean expression (over returned entries): {mean_expr:.6g}")

if __name__ == "__main__":
    main()
Show full SKILL.md (174 more words)Show less

Implementation Details

  • Opening the Census

    • Use a context manager to ensure resources are released:
      • with cellxgene_census.open_soma(...) as census: ...
    • For reproducibility, set census_version="YYYY-MM-DD"; otherwise the latest stable release is used.
  • Data model (high level)

    • Census data is stored in SOMA collections.
    • census["census_info"] provides summary tables (e.g., datasets, counts).
    • census["census_data"][organism] provides the experiment for an organism (e.g., homo_sapiens).
  • Filtering semantics

    • obs_value_filter filters cells (obs); var_value_filter filters genes (var).
    • Combine predicates with and / or; use in [...] for multi-value membership.
    • Best practice: include is_primary_data == True to avoid double-counting cells that appear in multiple source datasets.
  • Choosing an access pattern

    • Use get_anndata() when the result is expected to fit in memory (commonly < ~100k cells, depending on gene count and sparsity).
    • Use axis_query() + query.X("raw").tables() for out-of-core iteration and incremental statistics.
  • Expression layers / matrices

    • Examples commonly use X("raw") to access raw expression.
    • Chunk iteration yields Arrow tables with:
      • soma_data: expression values
      • soma_dim_0: obs (cell) coordinates
      • soma_dim_1: var (gene) coordinates
  • Optional ML integration

    • The cellxgene_census.experimental.ml utilities provide PyTorch-friendly datasets/dataloaders for training workflows, typically driven by the same obs/var filtering concepts used elsewhere.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in scientific-skills/Evidence Insight/cellxgene-census of aipoch/medical-research-skills.

  • SKILL.md
  • cellxgene-census_audit_result_v1.json
  • references/census_schema.md
  • references/common_patterns.md

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Cellxgene Census next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cellxgene Census compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cellxgene Census this skillaipoch/medical-research-skills2k—~1.6kAutomated safety check: PassMIT
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k15 repos~2.8kAutomated safety check: PassMIT
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
PyDESeq2 Differential Expressiondavila7/claude-code-templates32k11 repos~4kAutomated safety check: PassMIT
Anndatadavila7/claude-code-templates32k11 repos~2.5kAutomated safety check: PassMIT
Single-Cell Initial AnalysisLigphiDonk/Oh-my--paper7381 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 15 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    32k GitHub starsUsed in 11 repos~4k tokens
    Research & ScienceAuto-check passed
  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    32k GitHub starsUsed in 11 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    738 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check passed
  • Single Cell Data Prep Qc

    harrisongzhang/TheVirtualBiotech

    Single-cell RNA-seq data preparation and quality control pipeline.

    121 GitHub stars~2.9k tokensUpdated 23 days ago
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 23 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 23 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 23 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 23 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 23 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 23 days ago
    Auto-check passed

Works with

Questions about Cellxgene Census

What does Cellxgene Census do?

Programmatically query the CZ CELLxGENE Census (61M+ cells) when you need cross-tissue, disease, or cell-type expression data for population-scale queries and reference atlas comparisons. Cellxgene Census is an agent skill from aipoch/medical-research-skills. Programmatically query the CZ CELLxGENE Census (61M+ cells) when you need cross-tissue, disease, or cell-type expression data for population-scale queries and reference atlas comparisons.

When should I use Cellxgene Census?

Cellxgene Census fits situations like: research & Science work in your project.

How do I install Cellxgene Census in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill cellxgene-census -a claude-code`. Or copy the skill folder (scientific-skills/Evidence Insight/cellxgene-census in aipoch/medical-research-skills) into .claude/skills/cellxgene-census in your project. Claude Code loads it when a task matches its description.

How do I install Cellxgene Census in Codex?

Run `npx skills add aipoch/medical-research-skills --skill cellxgene-census -a codex`. Or copy the skill folder (scientific-skills/Evidence Insight/cellxgene-census in aipoch/medical-research-skills) into .agents/skills/cellxgene-census in your project. Codex loads it when a task matches its description.

Can I use Cellxgene Census in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill cellxgene-census -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cellxgene-census, .gemini/skills/cellxgene-census, .github/skills/cellxgene-census and .opencode/skills/cellxgene-census in your project.

What does Cellxgene Census need to run?

Going by SKILL.md and its folder, Cellxgene Census needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Cellxgene Census access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Cellxgene Census safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cellxgene Census use?

Cellxgene Census is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cellxgene Census use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.1k tokens, read only when the agent opens those files.

What are the alternatives to Cellxgene Census?

Skills that share tags, products or a category with Cellxgene Census: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), Scgpt (JimLiu/science-skills, 227 stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 32k stars) and Anndata (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cellxgene Census?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,978 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.