Agent skill

Scarf Single Cell

by NygenAnalytics in NygenAnalytics/scarf

Analyze single-cell data with core Scarf, the out-of-core Zarr DataStore library with immutable artifacts and pipeline runs.

BSD-3-ClauseAuto-check passedResearch & Science

Install Scarf Single Cell

skills CLI
$ npx skills add NygenAnalytics/scarf --skill scarf-single-cell -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NygenAnalytics/scarf scarf-single-cell --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NygenAnalytics/scarf.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scarf-single-cell .claude/skills/scarf-single-cell && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scarf-single-cell
GitHub stars
126
Token cost
~5.6k tokens
SKILL.md length
2,482 words
Files
12 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Analyze single-cell data with core Scarf, the out-of-core Zarr DataStore library with immutable artifacts and pipeline runs.

  • Works in 12 steps: Inspect read-only first.… → Never filter by editing I. Keep the… → Keep every ref you will reuse in a dict… → …
  • A task involves a Scarf .zarr store
  • SKILL.md covers Setup, Mental model, Golden rules and The analysis loop, plus 6 more sections
  • Runs Python scripts from its folder; calls uv, pip and python

What it does

Scarf Single Cell is an agent skill from NygenAnalytics/scarf. Analyze single-cell data with core Scarf, the out-of-core Zarr DataStore library with immutable artifacts and pipeline runs. Covers opening, converting and mounting stores (including Cytebase datasets), QC with removal audits, HVG/PCA/neighbour graphs, Leiden/Paris clustering, UMAP, markers and cautious annotation, batch correction and donor-level comparisons, headless plotting, provenance and export. Use when a task involves a Scarf .zarr store, a Cytebase dataset, scarf.DataStore, ds.pipeline, or converting…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `references/clustering-and-embedding.md`, `references/data-access.md` and `references/features-and-graphs.md`). Compatibility notes: Requires Python 3.12+ and scarf 1.0.0rc20 or newer (pip install "scarf[extra]=1.0.0rc20"; add the cytebase extra and network access for Cytebase datasets, and…

It sits in Research & Science, covering Bioinformatics. It works with Zarr and UMAP. The repository describes itself as: Memory-efficient single-cell analysis in Python. Stream RNA, ATAC, CITE-seq and multi-omics from local or remote Zarr stores, from laptop to atlas scale, with reusable… The licence is BSD-3-Clause.

When your agent uses it

  • A task involves a Scarf .zarr store
  • A Cytebase dataset
  • Scarf.DataStore
  • Converting H5AD/10x/MTX/Seurat data for Scarf

Example prompts

  • “/scarf-single-cell”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.12+ and scarf 1.0.0rc20 or newer (pip install "scarf[extra]>=1.0.0rc20"; add the cytebase extra and network access for Cytebase datasets, and the tsne extra for t-SNE).

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Inspect read-only first. scarf.DataStore(path, zarr_mode="r") writes nothing. A writable
  2. Never filter by editing I. Keep the cell_selection ref a filter returns and pass it on, or
  3. Keep every ref you will reuse in a dict and record them. Never guess artifact IDs or read
  4. Check the matrix, then audit QC removals before accepting a filter. Published matrices can
  5. Labels are immutable and runs cannot be resumed. Use a new label for every variant. A rerun
  6. Read pipeline outputs instead of recomputing them. run["markers"], run["cell_cycle"] and
  7. Align arrays through the cell selection. Use run.cells.fetch(...) (run cells only) or
  8. Budget count passes on mounts. QC metrics, the HVG summary, normalization per selection,
  9. One DataStore per process; it is not thread-safe. Run parallel analyses in separate
  10. Payload names differ by kind. Leiden labels, embeddings, doublet and strength scores live
  11. get_markers defaults hide genes. min_score=0.25, min_frac_exp=0.2; pass -1 for both
  12. Labels are hypotheses. Name a cluster only with two or more specific positive markers plus

What it can do on your machine

Read from SKILL.md and the folder at commit 9c3e99c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • scarf.readthedocs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.12+ and scarf 1.0.0rc20 or newer (pip install "scarf[extra]>=1.0.0rc20"; add the cytebase extra and network access for Cytebase datasets, and the tsne extra for t-SNE).

    From compatibility in the SKILL.md frontmatter.

Context cost

Scarf Single Cell loads about 5.6k tokens when it runs, and up to ~49k if it reads all its reference files. Until then it costs about 157 tokens; SKILL.md has 2,482 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~157
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~49k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NygenAnalytics/scarf at commit 9c3e99c, republished under its BSD-3-Clause licence (© NygenAnalytics). 2,482 words, ~5,600 tokens.

Download SKILL.mdSave it as .claude/skills/scarf-single-cell/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
scarf-single-cell
description
Analyze single-cell data with core Scarf, the out-of-core Zarr DataStore library with immutable artifacts and pipeline runs. Covers opening, converting and mounting stores (including Cytebase datasets), QC with removal audits, HVG/PCA/neighbour graphs, Leiden/Paris clustering, UMAP, markers and cautious annotation, batch correction and donor-level comparisons, headless plotting, provenance and export. Use when a task involves a Scarf .zarr store, a Cytebase dataset, scarf.DataStore, ds.pipeline, or converting H5AD/10x/MTX/Seurat data for Scarf. Does not cover scarf.agent (the automated agent package).
compatibility
Requires Python 3.12+ and scarf 1.0.0rc20 or newer (pip install "scarf[extra]>=1.0.0rc20"; add the cytebase extra and network access for Cytebase datasets, and the tsne extra for t-SNE).
license
BSD-3-Clause
metadata.version
0.4

Scarf single-cell analysis

Scarf streams Zarr-backed count matrices in bounded blocks and records every computation as an immutable, content-addressed artifact. This skill encodes how to drive core Scarf for single-cell RNA analysis: which calls to make, in what order, what evidence to check, and what not to do. It does not use or describe scarf.agent.

Read this file first. Then open only the reference modules you need (index below). Paths in this skill are relative to the skill directory. The maintained copy lives in the Scarf repository at https://github.com/NygenAnalytics/scarf/tree/master/skills/scarf-single-cell. Every recipe in the modules was run against a real store. Numbers quoted there come from the 10x 5K PBMC documentation dataset and are illustrations, not thresholds.

Setup

  • Install into the environment you run Python from: pip install "scarf[extra]>=1.0.0rc20", plus scarf[cytebase] for Cytebase and scarf[tsne] for t-SNE. The explicit pre-release floor matters: a bare scarf[extra] resolves to the old 0.32 series, whose API this skill does not describe.
  • Set resources per process before importing Scarf. Defaults claim all detected RAM and every CPU: SCARF_MEM_BUDGET=8G SCARF_WORKERS=8 (memory specs need a unit; a bare 8 is rejected).
  • Headless: MPLBACKEND=Agg, call plots with show=False, then result.save(path) and result.close().
  • Quiet logs for scripts: scarf.configure_output(level="WARNING", progress=False).
  • scripts/inspect_store.py STORE.zarr prints a read-only first look: cell and feature counts, runs, artifact kinds, whether counts look raw, and a profile of every metadata column with design-like, annotation-like and author-derived columns flagged. Annotation values stay hidden unless you pass --show-annotation-values.
  • Long steps (a pipeline run on tens of thousands of cells over remote counts can take many minutes): write the step as a script that logs to a file and ends with a DONE or FAILED line, run it in the background, and wait with a bounded loop that also stops on a Traceback or a dead process. If you launch through a wrapper such as uv run, put timeout inside it (uv run timeout N python step.py; the reverse can leave stale running records), and track the process by $! rather than pgrep -f. See references/performance-and-export.md. Never poll without a time limit.

Mental model

  • A DataStore is one Zarr directory: a shared cell table ds.cells, one group per assay (ds.RNA, feature table ds.RNA.feats, counts as cell-major counts plus gene-major countsT), artifacts, and pipeline/runs. QC columns are assay-prefixed: RNA_nCounts, RNA_nFeatures, RNA_percentMito, RNA_percentRibo.
  • I is the live boolean cell key. Analytical filters never edit it; they return a frozen cell selection artifact.
  • Every result is an ArtifactRef (scope, kind, artifact_id, assay). Producers take refs and return refs: select_hvgs -> run_normalization -> run_pca -> build_ann_index -> query_neighbors -> build_connectivity_map -> run_leiden_clustering / run_umap -> run_marker_search. An identical call returns the existing artifact without recomputing.
  • ds.pipeline.run(label=...) runs that whole RNA recipe (filtering, cell cycle, HVG, normalization, PCA, graph, UMAP, Leiden at 0.5/0.75/1.0/1.25 with a silhouette pick, Paris, doublet scores, markers) and returns a durable PipelineRun: a mapping of output names to refs plus frozen run.cells / run.features views and run.report().
  • Artifact payload rows follow the artifact's cell selection, not ds.cells.N.
  • A mount is a local writable store whose counts stay in a remote source (matrixSource). Every pass over counts is a network read; artifacts are written locally.

Golden rules

  1. Inspect read-only first. scarf.DataStore(path, zarr_mode="r") writes nothing. A writable open (the default) prepares new stores and permanently removes from I the cells with at most min_features_per_cell (default 10) default-assay features, unless that would remove at least half of the active cells. Reopen existing stores and mounts with scarf.DataStore(path, min_features_per_cell=-1), which keeps every cell.
  2. Never filter by editing I. Keep the cell_selection ref a filter returns and pass it on, or insert a boolean column and use cell_key=. ds.cells.reset_key("I") undoes an accidental open-time filter.
  3. Keep every ref you will reuse in a dict and record them. Never guess artifact IDs or read private Zarr paths. Use ds.inspect_artifact(ref), ds.load_artifact(ref) and ds.list_artifacts(...). Cell selections need scope="datastore" when listing.
  4. Check the matrix, then audit QC removals before accepting a filter. Published matrices can be corrected (for example SCTransform) or already filtered, which makes QC bounds meaningless; scripts/inspect_store.py reports this. Pooled MAD filters routinely remove low-complexity populations (platelets, erythrocytes, neutrophils) and high-RNA ones (plasma, cycling cells). Audit retention per sample and per marker-defined group without author labels, and look at the markers of removed cells (references/quality-control.md).
  5. Labels are immutable and runs cannot be resumed. Use a new label for every variant. A rerun reuses every complete artifact, so it is cheap. Choose snapshot_columns (design columns you want inside run.cells and exports) on the first run, because changing them recomputes everything downstream.
  6. Read pipeline outputs instead of recomputing them. run["markers"], run["cell_cycle"] and run["doublets"] already exist. A direct run_marker_search(run["clusters"], ...) creates a new artifact and rereads all counts.
  7. Align arrays through the cell selection. Use run.cells.fetch(...) (run cells only) or fetch_all(...) (full length, with -1, NaN or "" outside the run). For explicit artifacts use ds.inspect_artifact(ref).input_ref("cell_selection"). Join tables on cell ids, never on row order.
  8. Budget count passes on mounts. QC metrics, the HVG summary, normalization per selection, markers, gene plots and AUCell each read counts. For many passes, make a local copy once: python -m scarf.tools.repack_zarr MOUNT.zarr LOCAL.zarr --mem-budget 4G.
  9. One DataStore per process; it is not thread-safe. Run parallel analyses in separate processes on separate stores, each with its own SCARF_MEM_BUDGET/SCARF_WORKERS.
  10. Payload names differ by kind. Leiden labels, embeddings, doublet and strength scores live under values; Paris (cluster_cut) under labels; PCA under data. get_markers returns string group_id while Leiden labels are integers.
  11. get_markers defaults hide genes. min_score=0.25, min_frac_exp=0.2; pass -1 for both to see negatives and full panels.
  12. Labels are hypotheses. Name a cluster only with two or more specific positive markers plus low lineage-negative markers and a plausible QC and doublet profile. Keep unresolved for mixed, doublet-like or marker-poor clusters.
  13. Cells are not replicates, and Scarf does not check your design. Condition claims need donor-level aggregation and at least two independent donors per group. A donor sampled twice is repeated measures. A donor with samples in two arms must not count in both: run_statistical_testing(sample_by="sample_id") silently does that, while sample_by="donor_id" refuses such a design. run_harmony runs silently on a column confounded with the biology you want to compare, so audit donors x batch x condition first. When inserting design columns, decide on missing values: ds.cells.insert records None, NaN and pd.NA as missing, and so the rows outside key when values covers only the active cells, unless fill_value gives them a value. Harmony raises on a missing batch value.
  14. Hold out annotation columns when they will judge the result. Do not use author labels (cell_type, author_cell_type, cell.type.*, singler, predicted.*, ...) to choose QC, parameters or labels. Use them only in a final, clearly separated comparison.
  15. Pass an explicit HVG blacklist. On stores with current gene symbols the default blacklist misses replication-dependent histones (^HIST matches only pre-2020 names), and its ^CCN family removes the CCN1-6 matricellular genes and non-cell-cycle cyclins. It also leaves Ig V/J genes in, which can split plasma cells by light chain instead of biology. Use the species string from references/gene-blacklists.md (params={"hvg": {"blacklist": ...}} in the pipeline). Add the clonotype add-on only on marker evidence. A blacklist= string replaces the default entirely, and changing it creates new HVG, PCA and graph artifacts.

The analysis loop

Work in this order. After each step, write the decision and its evidence to an analysis log.

StepDoEvidence to recordModule
1. Inspectscripts/inspect_store.py, ds.summary(), existing runs; check raw versus corrected countscells, assays, mounted or local, count type, metadata rolesdata-access.md, quality-control.md
2. Designidentify donor, sample, capture, batch and condition columns; derive missing unitsunit of inference, confounding, held-out columnsintegration-and-comparisons.md
3. QCinspect distributions; compare policies; audit removalschosen policy, cells kept per group, what was removedquality-control.md
4. Baselineds.pipeline.run(label=..., filtering=<chosen>, snapshot_columns=<design>)run id, report, kept cellspipeline-runs-and-artifacts.md
5. Structureresolutions, seed stability, nesting, marker support, QC and doublets per clusterchosen partition and whyclustering-and-embedding.md
6. Branch if neededHVG count and blacklist, PCs, k, subclustering one lineage, Harmony (only if the design allows)which alternative changed whatfeatures-and-graphs.md, gene-blacklists.md, integration-and-comparisons.md
7. Annotatemarker tables, canonical panels, gene-set scoreslabel, markers, negatives, unresolved clustersmarkers-and-annotation.md
8. Comparecomposition per donor, pseudobulk; only with replicateswhat can and cannot be claimedintegration-and-comparisons.md, performance-and-export.md (pseudobulk export)
9. Reportfigures, tables, handoff JSON, lineage, exportfiles written, open questionsplotting.md, performance-and-export.md

Quick start

A complete first pass on an RNA store. Every step after the run reuses the run's artifacts.

python
from pathlib import Path

import matplotlib

matplotlib.use("Agg")  # headless; select the backend before importing scarf
import numpy as np
import pandas as pd
import scarf

scarf.configure_output(level="WARNING", progress=False)
store = "analysis.zarr"
out = Path("analysis_out")
(out / "figures").mkdir(parents=True, exist_ok=True)

ds = scarf.DataStore(store, min_features_per_cell=-1)

# QC: compare candidate filters and audit them before choosing (quality-control.md).
cells = ds.snapshot_cell_selection("I")
mad5 = ds.auto_filter_cells(cell_selection=cells, n_mads=5)
kept = np.asarray(ds.load_artifact(mad5)["values"][:], dtype=bool)
print(f"MAD 5 keeps {kept.sum()} of {ds.cells.N} cells")

# Baseline run with the chosen policy; freeze design columns you will need later.
run = ds.pipeline.run(
    label="baseline_mad5",
    filtering={"method": "mad", "n_mads": 5},
    snapshot_columns=[c for c in ("donor_id", "sample_id") if c in ds.cells.columns],
)
print(run.status, list(run.keys()))
(out / "run_report.md").write_text(run.report(format="markdown"))

# Figures and tables.
umap = ds.plots.embedding(run=run, color_by="clusters", show=False)
umap.save(out / "figures" / "umap_clusters.png", dpi=150)
umap.close()
markers = ds.get_markers(marker=run["markers"])
markers.to_csv(out / "markers_top.csv", index=False)
print(markers.groupby("group_id", sort=False).head(5)
      .groupby("group_id", sort=False)["feature_name"].agg(", ".join))
sizes = pd.Series(run.cells.fetch("clusters")).value_counts().sort_index()
print(sizes.to_dict())

Record the analysis

Keep these in the output directory so another agent or a person can audit and continue:

  • analysis_log.md: the question, the units, and one entry per step (decision, alternatives considered, evidence with numbers, figure paths).
  • handoff.json: run label and ID, and the refs you relied on (ref.to_dict(); restore with scarf.ArtifactRef.from_dict). Add unsupported claims and open questions. See pipeline-runs-and-artifacts.md.
  • run_report.md (run.report(format="markdown")) and lineage.md (ds.lineage({...}).to_markdown()).
  • Figures saved as PNG, and tables as CSV (markers, cluster sizes, composition, label map).
Show full SKILL.md (1,001 more words)Show less

Evaluating against held-out labels

When a dataset ships author annotations, run the whole loop without them. Then add one final section that crosstabs your labels against theirs, reports ARI and per-type F1 after harmonizing vocabularies, and explains disagreements with marker evidence. Report QC retention per author type there too, as an audit of the filter. Never go back and tune parameters to raise agreement. Record any changes you make after seeing the comparison as such.

Reference modules

ModuleRead when
references/data-access.mdopening, inspecting, converting (H5AD, 10x, MTX, Seurat), Cytebase search, open and mount
references/quality-control.mdQC metrics, MAD/manual/per-sample filters, removal audits, doublet scores
references/features-and-graphs.mdHVGs, normalization, PCA dims, ANN, neighbours, graph diagnostics, branching, subclustering
references/gene-blacklists.mdwhat the default HVG blacklist removes, corrected blacklists for human and mouse, Ig/TCR and haemoglobin add-ons
references/clustering-and-embedding.mdLeiden and Paris, choosing a partition, UMAP and t-SNE, labels as arrays
references/markers-and-annotation.mdmarker tables, canonical marker panels, labelling, AUCell/WAGGR, cell cycle, label comparison
references/pipeline-runs-and-artifacts.mdds.pipeline.run options, run reports, artifacts, lineage, failures, handoff record
references/plotting.mdheadless figures, run mode versus ref mode, every ds.plots method
references/integration-and-comparisons.mdstudy design and units, donors in two arms, composition tests, Harmony and its safety, merging, mapping, condition comparisons
references/performance-and-export.mdbudgets, long-running steps, streaming, mount repacking, AnnData/H5AD/MTX/CSV and pseudobulk export, subset stores

Troubleshooting

SymptomLikely causeFix
Assay 'RNA' is not prepared on a read-only openstore just written by a converter or SubsetZarropen once writable, then read-only
Fewer active cells than expected after openingwritable open applied min_features_per_cellds.cells.reset_key("I"); reopen with -1
ValueError on ds.pipeline.run(label=...)label already completednew label; reuse makes it cheap
PermissionError from a producerstore opened with zarr_mode="r" and no matching artifact; feature selections, WAGGR, AUCell, select_prevalent_peaks, and cell-cycle scoring refuse even with onereopen writable
KeyError: 'values' on a Paris refParis stores labelsds.load_artifact(ref)["labels"]
Array length differs from ds.cells.Npayload follows the cell selectionalign with run.cells.fetch_all or the selection mask
Plot raises with run= and a gene or live columnrun mode reads only frozen run fields and run outputs, and genes need an explicit normalizationgene: add normalization=scarf.plotting.NormalizationSpec(); live column: freeze it with snapshot_columns= or use layout=run["umap"]
TypeError from distribution(grouping="col")grouping needs a ref or CellFieldgrouping=scarf.plotting.CellField("col")
Very slow steps on a mounteach count pass is a network readfewer passes; repack locally (rule 8)
MemoryError (CountLayoutMemoryError after 1.0.0rc19) from a converterdefault count layout does not fit mem_budgetlarger mem_budget; else the policy= the message names (references/data-access.md)
ValueError plotting after reopening the storea PipelineRun is bound to the store object that opened itreopen the run from the new ds
list_artifacts(kind="cell_selection") is emptycell selections are datastore-scopedadd scope="datastore"
KeyError for a grouping column of a dot or matrix plot tablesummary tables name grouping columns by role: group, subgroup, sampleres.tables["aggregate"]["group"]; res.provenance.extras["group_by"] names the source columns
A wait for a long step never endsunbounded polling, or the process diedbounded wait that checks the process and the log tail (Setup)
QC bounds look odd or nothing is filteredcounts are corrected or already filteredcheck the matrix first; prefer flag-only or gentle filters
Pipeline fails in stage 'doublets': every selected cell has the same cluster labelparams["leiden"]["selected"] saved a one-cluster partition, and heterotypic doublets need two clustersselect a finer resolution or drop selected; or params={"doublets": {"heterotypic_fraction": 0}}; or doublets=False
Writer raises FileExistsError: Destination ... is not emptythe path holds data, such as an earlier importoverwrite=True (SubsetZarr: overwrite_existing_file=True) to replace an output that no DataStore has opened, or a new path
Writer with overwrite=True raises FileExistsError: Destination ... holds the prepared assays ... or ... not part of a Scarf storethe path holds a store that a DataStore has opened, or other fileswrite to a new path, or delete the old store yourself first
Writer raises ValueError: Destination ... lies inside the Zarr store at ...the path is inside another store, such as s.zarr/RNAwrite each store to its own directory outside other stores
assay_types names assays that are not in the store or ... is not a preseta key is not in ds.assay_names, or a value is not a preset (case-sensitive)keys from ds.assay_names; values such as RNA, ADT, HTO, GeneActivity, Assay
assay_types declares assay ... but the store declares it as ... on zarr_mode="r"a read-only open cannot record a different typeopen once writable with that assay_types, or omit assay_types
The assayTypes attribute of the store is ..., which is not a mappinganother tool wrote a malformed assayTypesopen writable with assay_types naming a preset for every assay
Merge raises Sources declare different types for assay ...sources record different assayTypes for one assayreopen the wrong source writable with the right assay_types, then merge
Merge raises Cell column '<assay>_I' is reserved for the membership of assay ...a plain <assay>_I column, from an earlier import of an exported file or inserted by handimport the H5AD again with this release, or drop the plain column in the source with cells.drop("<assay>_I")
cells.insert/update_key/reset_key/drop raises Cell column '<assay>_I' is reserved ... or ... records which cells assay ... measuredonly imports, merges, and derived assays write membershipstore values under another name; select measured cells with select_measured_cells("<assay>")
UnmeasuredCellsError: ... reads assay '<assay>' over N of M selected cells that it did not measurea merged store whose assay measured only some cells; their zero counts are no measurementnarrow first: ds.select_measured_cells("<assay>", cell_selection=...), which make_bulk, run_statistical_testing, and QC filters take as cell_selection=; labels: snapshot_cluster_labels(labels, cell_selection=...) of that; graphs: rebuild over it; pipeline: a cell_key column of the cells of I that the assay measured, ds.cells.insert("RNA_measured", ds.cells.fetch_all("I") & ds.cells.fetch_all("RNA_I"))
ValueError: None of the s_genes match the assay feature names from ds.pipeline.runfeature names are not gene symbols (Ensembl IDs, synthetic names); matching ignores case, so mouse symbols workcell_cycle=False, or pass lists in the store's naming via params={"cell_cycle": {"s_genes": [...], "g2m_genes": [...]}}

Documentation map

The docs hold the full explanations, online at https://scarf.readthedocs.io/en/latest/: quickstart.html, tutorials/<step>.html (one page per step, for example quality_control, graph_construction, clustering, annotation, batch_correction, pseudobulk_and_differential_expression, cytebase), concepts/ (provenance, memory_and_execution, benchmarks), analysis_with_agents.html (scientific decision loop, task routing, handoff) and reference/api/<module>.html (exact signatures).

© NygenAnalytics, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references) in skills/scarf-single-cell of NygenAnalytics/scarf.

  • SKILL.md
  • references/clustering-and-embedding.md
  • references/data-access.md
  • references/features-and-graphs.md
  • references/gene-blacklists.md
  • references/integration-and-comparisons.md
  • references/markers-and-annotation.md
  • references/performance-and-export.md
  • references/pipeline-runs-and-artifacts.md
  • references/plotting.md
  • references/quality-control.md
  • scripts/inspect_store.py

Open the folder on GitHubat commit 9c3e99c

Compare with similar skills

Scarf Single Cell next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scarf Single Cell compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scarf Single Cell this skillNygenAnalytics/scarf126—~5.6kAutomated safety check: PassBSD-3-Clause
Spatial AteraQING1105/ezST101—~576Automated safety check: PassMIT
Anndata Data Structurejaechang-hits/SciAgent-Skills3702 repos~5.8kAutomated safety check: PassBSD-3-Clause
ScanpyK-Dense-AI/scientific-agent-skills48k1 repos~5.1kAutomated safety check: PassBSD-3-Clause
AnndataK-Dense-AI/scientific-agent-skills48k1 repos~3.9kAutomated safety check: NotesBSD-3-Clause
Spatial S2 Normalize ClusterQING1105/ezST101—~428Automated safety check: PassMIT

Similar skills

  • Spatial Atera

    QING1105/ezST

    Atera platform branch of the spatial transcriptomics workflow — load and validate Atera cell-level output (AnnData + Zarr segmentation) for downstream analysis.

    101 GitHub stars~576 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 2 repos~5.8k tokens
    Research & ScienceAuto-check passed
  • Scanpy

    K-Dense-AI/scientific-agent-skills

    Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…

    48k GitHub starsUsed in 1 repo~5.1k tokens
    Research & ScienceAuto-check passed
  • Anndata

    K-Dense-AI/scientific-agent-skills

    Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.

    48k GitHub starsUsed in 1 repo~3.9k tokens
    Research & ScienceAuto-check: notes
  • Stage 2 of the spatial transcriptomics workflow — normalize 10x Visium data and cluster spatial spots.

    101 GitHub stars~428 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Umap Tsne Analysis

    aipoch/medical-research-skills

    A skill your agent uses when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE…

    2k GitHub stars~2.7k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed

Works with

Questions about Scarf Single Cell

What does Scarf Single Cell do?

Analyze single-cell data with core Scarf, the out-of-core Zarr DataStore library with immutable artifacts and pipeline runs. Scarf Single Cell is an agent skill from NygenAnalytics/scarf. Analyze single-cell data with core Scarf, the out-of-core Zarr DataStore library with immutable artifacts and pipeline runs.

When should I use Scarf Single Cell?

Scarf Single Cell fits situations like: A task involves a Scarf .zarr store; A Cytebase dataset; scarf.DataStore; converting H5AD/10x/MTX/Seurat data for Scarf.

How do I install Scarf Single Cell in Claude Code?

Run `npx skills add NygenAnalytics/scarf --skill scarf-single-cell -a claude-code`. Or copy the skill folder (skills/scarf-single-cell in NygenAnalytics/scarf) into .claude/skills/scarf-single-cell in your project. Claude Code loads it when a task matches its description.

How do I install Scarf Single Cell in Codex?

Run `npx skills add NygenAnalytics/scarf --skill scarf-single-cell -a codex`. Or copy the skill folder (skills/scarf-single-cell in NygenAnalytics/scarf) into .agents/skills/scarf-single-cell in your project. Codex loads it when a task matches its description.

Can I use Scarf Single Cell in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NygenAnalytics/scarf --skill scarf-single-cell -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scarf-single-cell, .gemini/skills/scarf-single-cell, .github/skills/scarf-single-cell and .opencode/skills/scarf-single-cell in your project.

What does Scarf Single Cell need to run?

Going by SKILL.md and its folder, Scarf Single Cell needs Python for the scripts in its folder and the command-line tools its instructions call (uv, pip and python). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.12+ and scarf 1.0.0rc20 or newer (pip install "scarf[extra]>=1.0.0rc20"; add the cytebase extra and network access for Cytebase datasets, and the tsne extra for t-SNE)..

Does Scarf Single Cell access the network?

SKILL.md names 1 domain. As links in the text: scarf.readthedocs.io. This is read from the text; nothing was executed.

Is Scarf Single Cell safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Scarf Single Cell use?

Scarf Single Cell is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scarf Single Cell use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 43k tokens, read only when the agent opens those files.

What are the alternatives to Scarf Single Cell?

Skills that share tags, products or a category with Scarf Single Cell: Spatial Atera (QING1105/ezST, 101 stars), Anndata Data Structure (jaechang-hits/SciAgent-Skills, 370 stars), Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars) and Anndata (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scarf Single Cell?

NygenAnalytics (a GitHub organization) maintains it in NygenAnalytics/scarf, which has 126 GitHub stars. The repository was last updated on October 7, 2026.

Source: NygenAnalytics/scarf on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.