Agent skill

Scgpt

by JimLiu in JimLiu/science-skills

Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

Apache-2.0Auto-check passedResearch & Science

Install Scgpt

skills CLI
$ npx skills add JimLiu/science-skills --skill scgpt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JimLiu/science-skills scgpt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JimLiu/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scgpt .claude/skills/scgpt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scgpt
GitHub stars
227
Used in
4 other repos
Token cost
~1.3k tokens
SKILL.md length
337 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
Apache-2.0

At a glance

Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

  • Producing cell embeddings from an AnnData for clustering/integration
  • SKILL.md covers Prerequisites, How to run, Output format and Remote compute, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Fine-tuned cell-type annotation

What it does

Scgpt is an agent skill from JimLiu/science-skills. Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation/GRN tasks. For probabilistic single-cell models (scVI etc.), use the scvi-tools library.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics and Embeddings. It works with AnnData and scvi-tools. The licence is Apache-2.0.

When your agent uses it

  • Producing cell embeddings from an AnnData for clustering/integration
  • Fine-tuned cell-type annotation
  • Gene-level representation for perturbation/GRN tasks

Example prompts

  • “/scgpt”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit fb309c3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scgpt loads about 1.3k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 337 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JimLiu/science-skills at commit fb309c3, republished under its Apache-2.0 licence (© JimLiu). 337 words, ~1,330 tokens.

Download SKILL.mdSave it as .claude/skills/scgpt/SKILL.md (or your agent's skills folder).
name
scgpt
description
Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation/GRN tasks. For probabilistic single-cell models (scVI etc.), use the scvi-tools library.
license
Apache-2.0
category
biomodels
requirements
gpu
metadata.display-name
scGPT

scGPT — Single-Cell Foundation Model

Prerequisites

RequirementMinimumRecommended
Python3.10+3.11
CUDA12.1+12.4+
GPU VRAM16 GB24 GB+

How to run

Loading the vocabulary and checkpoint

scGPT checkpoints are raw directories (args.json, best_model.pt, vocab.json) — not Hugging Face hub repos. Point at the directory, not an HF repo id.

python
from scgpt.tokenizer.gene_tokenizer import GeneVocab
gv = GeneVocab.from_file("/path/to/scgpt-human/vocab.json")
print(len(gv))   # 60697 for the released human checkpoint
Embedding an AnnData
python
import anndata as ad
from scgpt.tasks import embed_data

adata = ad.read_h5ad("dataset.h5ad")        # var must contain a gene-name column
emb = embed_data(
    adata,
    model_dir="/path/to/scgpt-human",
    gene_col="feature_name",
    use_fast_transformer=False,             # see Gotchas
)
# emb is an AnnData with .obsm["X_scGPT"]

Output format

embed_data returns an AnnData whose .obsm["X_scGPT"] is the per-cell embedding (n_cells × emb_dim, 512 by default). Downstream: feed to scanpy.pp.neighbors / scanpy.tl.umap.

Remote compute

Needs ≥24 GB VRAM and the released human checkpoint (~200 MB: args.json, best_model.pt, vocab.json). Read compute_details({provider, mode:'read'}) for an environment with scgpt and a pre-cached checkpoint directory, then:

python
c = host.compute.create(provider)
job = c.submit_job(
    intent="scGPT embed 50k cells — 1×GPU, ~5 min",
    inputs=[
        {"src": "dataset.h5ad", "dst_filename": "dataset.h5ad"},
        {"src": "embed.py", "dst_filename": "embed.py"},
    ],
    command="python3 embed.py",
    environment=...,   # env name from compute_details
    outputs=["embedded.h5ad"],
    timeout_seconds=1800,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

Then call the wait_for_notification brain-tool. When the compute_done notification arrives, act on its payload:

python
save_artifacts(payload["featured_files"])   # paths under hpc/<job_id>/

For the full result dict (output_files, remote_workdir, …), re-enter the kernel and bind the compute handle separately — .close() lives on the handle, not on the job object:

python
h = host.compute.create(provider)
res = h.attach_job(job_id).result()
h.close()

See the remote-compute-ssh / remote-compute-modal skill for the orchestration details.

In embed.py, pass model_dir= the checkpoint path from compute_details. If flash-attn is unavailable in that environment, set use_fast_transformer=False.

Gotchas

  • use_fast_transformer default is True but resolves to a FlashAttention path that may not import in every env. Pass use_fast_transformer=False unless you've confirmed flash_attn loads cleanly.
  • The package historically depended on torchtext.vocab.Vocab; in environments without torchtext a pure-Python shim provides Vocab — functionally identical for GeneVocab, but if you hit AttributeError: 'Vocab' object has no attribute …, you're on a stale shim.
  • Gene names must match the vocab; unmatched genes are dropped. Set gene_col to the column in adata.var that holds symbols.

Troubleshooting

SymptomFix
flash_attn is not installed warning at importHarmless; pass use_fast_transformer=False
'Vocab' object has no attribute 'vocab'Env has an old torchtext shim — update the env
Nearly all genes droppedWrong gene_col; check adata.var.columns
"scgpt not in manifest" / env-detection misses scGPTThe baked env manifest lists the distribution as scGPT (and flash_attn), pip's canonical casing — normalize manifest keys before lookup: name.lower().replace('-', '_')

Next: cluster/annotate the embedding with the scanpy library (sc.pp.neighbors → sc.tl.leiden / sc.tl.umap), or compare to an scvi-tools latent space on the same data.

© JimLiu, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scgpt of JimLiu/science-skills.

Open the folder on GitHubat commit fb309c3

Used in 4 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in JimLiu/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Scgpt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scgpt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scgpt this skillJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
Sc MetacellTianGzlab/OmicsClaw161—~1.3kAutomated safety check: PassApache-2.0
ScanpyK-Dense-AI/scientific-agent-skills48k1 repos~5.1kAutomated safety check: PassBSD-3-Clause
Cellxgene CensusK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: NotesMIT
AnndataK-Dense-AI/scientific-agent-skills48k1 repos~3.9kAutomated safety check: NotesBSD-3-Clause
Anndata Data Structurejaechang-hits/SciAgent-Skills3712 repos~5.8kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Sc Metacell

    TianGzlab/OmicsClaw

    Load when aggregating single cells into metacells (sample-aware coarse-grained pseudo-cells) on a normalised scRNA AnnData via SEACells or KMeans on a low-D embedding.

    161 GitHub stars~1.3k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Scanpy

    K-Dense-AI/scientific-agent-skills

    Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…

    48k GitHub starsUsed in 1 repo~5.1k tokens
    Research & ScienceAuto-check passed
  • Cellxgene Census

    K-Dense-AI/scientific-agent-skills

    Queries the CZ CELLxGENE Census programmatically for versioned public single-cell and spatial transcriptomics data.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Research & ScienceAuto-check: notes
  • Anndata

    K-Dense-AI/scientific-agent-skills

    Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.

    48k GitHub starsUsed in 1 repo~3.9k tokens
    Research & ScienceAuto-check: notes
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 2 repos~5.8k tokens
    Research & ScienceAuto-check passed
  • Scanpy

    aipoch/medical-research-skills

    Standard single-cell RNA-seq analysis pipeline. An agent skill from aipoch/medical-research-skills.

    2k GitHub stars~3.9k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed

More from JimLiu/science-skills

All 27 skills in this repo
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    Auto-check passed
  • Compute Env Setup

    JimLiu/science-skills

    Set up a compute environment on a remote provider so Claude Science jobs can run there.

    227 GitHub starsUsed in 2 repos~4.4k tokens
    Auto-check passed
  • Borzoi

    JimLiu/science-skills

    Predict genome-wide functional tracks (RNA-seq, CAGE, DNase, ChIP) from DNA sequence with Borzoi.

    227 GitHub starsUsed in 4 repos~973 tokens
    Auto-check passed
  • Evo2

    JimLiu/science-skills

    Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check passed
  • Fair Esm2

    JimLiu/science-skills

    Embed proteins with Meta AI's ESM-2 (fair-esm package). An agent skill from JimLiu/science-skills.

    227 GitHub starsUsed in 4 repos~1.2k tokens
    Auto-check passed
  • Openfold3

    JimLiu/science-skills

    Structure prediction using OpenFold3, an open-weights PyTorch reproduction of AlphaFold3 from the AlQuraishi Lab.

    227 GitHub starsUsed in 4 repos~1.8k tokens
    Auto-check passed

Questions about Scgpt

What does Scgpt do?

Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Scgpt is an agent skill from JimLiu/science-skills. Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

When should I use Scgpt?

Scgpt fits situations like: producing cell embeddings from an AnnData for clustering/integration; fine-tuned cell-type annotation; gene-level representation for perturbation/GRN tasks.

How do I install Scgpt in Claude Code?

Run `npx skills add JimLiu/science-skills --skill scgpt -a claude-code`. Or copy the skill folder (skills/scgpt in JimLiu/science-skills) into .claude/skills/scgpt in your project. Claude Code loads it when a task matches its description.

How do I install Scgpt in Codex?

Run `npx skills add JimLiu/science-skills --skill scgpt -a codex`. Or copy the skill folder (skills/scgpt in JimLiu/science-skills) into .agents/skills/scgpt in your project. Codex loads it when a task matches its description.

Can I use Scgpt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JimLiu/science-skills --skill scgpt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scgpt, .gemini/skills/scgpt, .github/skills/scgpt and .opencode/skills/scgpt in your project.

What does Scgpt need to run?

SKILL.md names no scripts, command-line tools or credentials: Scgpt is instructions for the agent only. Our summary lists: Python 3.

Does Scgpt access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scgpt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scgpt use?

Scgpt is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scgpt use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scgpt?

Skills that share tags, products or a category with Scgpt: Sc Metacell (TianGzlab/OmicsClaw, 161 stars), Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars), Cellxgene Census (K-Dense-AI/scientific-agent-skills, 48k stars) and Anndata (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scgpt?

JimLiu (a GitHub user) maintains it in JimLiu/science-skills, which has 227 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on July 1, 2026.

Source: JimLiu/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.