Agent skill

Scarf Single Cell

by sickn33 in sickn33/agentic-awesome-skills

Analyze single-cell RNA-seq at million-cell scale with Scarf: out-of-core Zarr stores on disk or object storage, provenance-tracked artifacts, audited QC, clustering, markers and donor comparisons.

BSD-3-ClauseAuto-check passedResearch & Science

Install Scarf Single Cell

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill scarf-single-cell -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills scarf-single-cell --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scarf-single-cell .claude/skills/scarf-single-cell && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scarf-single-cell
GitHub stars
47k
Token cost
~5.7k tokens
SKILL.md length
2,584 words
Files
12 (incl. references)
Skills in repo
1,493
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

Analyze single-cell RNA-seq at million-cell scale with Scarf: out-of-core Zarr stores on disk or object storage, provenance-tracked artifacts, audited QC, clustering, markers and donor comparisons.

  • Works in 12 steps: Inspect read-only first.… → Never filter by editing I. Keep the… → Keep every ref you will reuse in a dict… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers When to Use This Skill, Setup, Mental model and Golden rules, plus 10 more sections
  • Calls pip, uv and python

What it does

Scarf Single Cell is an agent skill from sickn33/agentic-awesome-skills. Analyze single-cell RNA-seq at million-cell scale with Scarf: out-of-core Zarr stores on disk or object storage, provenance-tracked artifacts, audited QC, clustering, markers and donor comparisons.

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `references/LICENSE.md`, `references/clustering-and-embedding.md` and `references/data-access.md`). Compatibility notes: Requires Python 3.12+ and scarf 1.0.0rc17 or newer (pip install "scarf[extra]=1.0.0rc17"; add the cytebase extra and network access for Cytebase datasets).

It sits in Research & Science, covering Bioinformatics. It works with Zarr. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is BSD-3-Clause.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/scarf-single-cell”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.12+ and scarf 1.0.0rc17 or newer (pip install "scarf[extra]>=1.0.0rc17"; add the cytebase extra and network access for Cytebase datasets).

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Inspect read-only first. scarf.DataStore(path, zarr_mode="r") writes nothing. A writable
  2. Never filter by editing I. Keep the cell_selection ref a filter returns and pass it on, or
  3. Keep every ref you will reuse in a dict and record them. Never guess artifact IDs or read
  4. Check the matrix, then audit QC removals before accepting a filter. Published matrices can
  5. Labels are immutable and runs cannot be resumed. Use a new label for every variant. A rerun
  6. Read pipeline outputs instead of recomputing them. run["markers"], run["cell_cycle"] and
  7. Align arrays through the cell selection. Use run.cells.fetch(...) (run cells only) or
  8. Budget count passes on mounts. QC metrics, the HVG summary, normalization per selection,
  9. One DataStore per process; it is not thread-safe. Run parallel analyses in separate
  10. Payload names differ by kind. Leiden labels, embeddings, doublet and strength scores live
  11. get_markers defaults hide genes. min_score=0.25, min_frac_exp=0.2; pass -1 for both
  12. Labels are hypotheses. Name a cluster only with two or more specific positive markers plus

What it can do on your machine

Read from SKILL.md and the folder at commit 680176d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • uv
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • scarf.readthedocs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.12+ and scarf 1.0.0rc17 or newer (pip install "scarf[extra]>=1.0.0rc17"; add the cytebase extra and network access for Cytebase datasets).

    From compatibility in the SKILL.md frontmatter.

Context cost

Scarf Single Cell loads about 5.7k tokens when it runs, and up to ~44k if it reads all its reference files. Until then it costs about 54 tokens; SKILL.md has 2,584 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~44k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 680176d, republished under its BSD-3-Clause licence (© sickn33). 2,584 words, ~5,695 tokens.

Download SKILL.mdSave it as .claude/skills/scarf-single-cell/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
scarf-single-cell
description
Analyze single-cell RNA-seq at million-cell scale with Scarf: out-of-core Zarr stores on disk or object storage, provenance-tracked artifacts, audited QC, clustering, markers and donor comparisons.
compatibility
Requires Python 3.12+ and scarf 1.0.0rc17 or newer (pip install "scarf[extra]>=1.0.0rc17"; add the cytebase extra and network access for Cytebase datasets).
category
science
risk
critical
source
https://github.com/NygenAnalytics/scarf/tree/master/skills/scarf-single-cell
source_repo
NygenAnalytics/scarf
source_type
official
date_added
2026-10-05
author
NygenAnalytics
license
BSD-3-Clause
license_source
https://github.com/NygenAnalytics/scarf/blob/master/LICENSE
metadata.version
0.4

Scarf single-cell analysis

Scarf streams Zarr-backed count matrices in bounded blocks and records every computation as an immutable, content-addressed artifact. This skill encodes how to drive core Scarf for single-cell RNA analysis: which calls to make, in what order, what evidence to check, and what not to do. It does not use or describe scarf.agent.

Read this file first. Then open only the reference modules you need (index below). Paths in this skill are relative to the skill directory. The maintained copy lives in the Scarf repository at https://github.com/NygenAnalytics/scarf/tree/master/skills/scarf-single-cell. Every recipe in the modules was run against a real store. Numbers quoted there come from the 10x 5K PBMC documentation dataset and are illustrations, not thresholds.

When to Use This Skill

  • Use when a task involves a Scarf .zarr store, scarf.DataStore, ds.pipeline or a PipelineRun.
  • Use when converting H5AD, 10x, MTX or Seurat data into a Scarf store, or opening and mounting a Cytebase dataset.
  • Use for single-cell RNA-seq analysis with core Scarf: QC with removal audits, HVG, PCA and neighbour graphs, Leiden and Paris clustering, UMAP, markers and cautious annotation, batch correction, donor-level comparisons, headless plotting, provenance and export.
  • Use when counts are too large for memory or live on object storage, so they must be streamed under an explicit memory budget.
  • Do not use it for scarf.agent (the automated agent package), which this skill does not cover.

Setup

  • Install into the environment you run Python from: pip install "scarf[extra]>=1.0.0rc17", plus scarf[cytebase] for Cytebase. The explicit pre-release floor matters: a bare scarf[extra] resolves to the old 0.32 series, whose API this skill does not describe.
  • Set resources per process before importing Scarf. Defaults claim all detected RAM and every CPU: SCARF_MEM_BUDGET=8G SCARF_WORKERS=8 (memory specs need a unit; a bare 8 is rejected).
  • Headless: MPLBACKEND=Agg, call plots with show=False, then result.save(path) and result.close().
  • Quiet logs for scripts: scarf.configure_output(level="WARNING", progress=False).
  • First look at a store: open it read-only and follow "Open and inspect a store" (references/data-access.md) and "Check the matrix first" (references/quality-control.md) for cell and feature counts, runs, artifact kinds and whether counts look raw. Then review every metadata column yourself: design columns (donor, sample, batch, condition), author annotations to hold out (rule 14) and author-derived columns. The maintained copy of this skill in the Scarf repository also ships scripts/inspect_store.py, a read-only command that runs these checks and flags those columns by name (annotation values stay hidden unless you pass --show-annotation-values). This catalogue copy is Markdown only and does not include the script.
  • Long steps (a pipeline run on tens of thousands of cells over remote counts can take many minutes): write the step as a script that logs to a file and ends with a DONE or FAILED line, run it in the background, and wait with a bounded loop that also stops on a Traceback or a dead process. If you launch through a wrapper such as uv run, put timeout inside it (uv run timeout N python step.py; the reverse can leave stale running records), and track the process by $! rather than pgrep -f. See references/performance-and-export.md. Never poll without a time limit.

Mental model

  • A DataStore is one Zarr directory: a shared cell table ds.cells, one group per assay (ds.RNA, feature table ds.RNA.feats, counts as cell-major counts plus gene-major countsT), artifacts, and pipeline/runs. QC columns are assay-prefixed: RNA_nCounts, RNA_nFeatures, RNA_percentMito, RNA_percentRibo.
  • I is the live boolean cell key. Analytical filters never edit it; they return a frozen cell selection artifact.
  • Every result is an ArtifactRef (scope, kind, artifact_id, assay). Producers take refs and return refs: select_hvgs -> run_normalization -> run_pca -> build_ann_index -> query_neighbors -> build_connectivity_map -> run_leiden_clustering / run_umap -> run_marker_search. An identical call returns the existing artifact without recomputing.
  • ds.pipeline.run(label=...) runs that whole RNA recipe (filtering, cell cycle, HVG, normalization, PCA, graph, UMAP, Leiden at 0.5/0.75/1.0/1.25 with a silhouette pick, Paris, doublet scores, markers) and returns a durable PipelineRun: a mapping of output names to refs plus frozen run.cells / run.features views and run.report().
  • Artifact payload rows follow the artifact's cell selection, not ds.cells.N.
  • A mount is a local writable store whose counts stay in a remote source (matrixSource). Every pass over counts is a network read; artifacts are written locally.

Golden rules

  1. Inspect read-only first. scarf.DataStore(path, zarr_mode="r") writes nothing. A writable open (the default) prepares new stores and applies min_features_per_cell (default 10) to I permanently. Reopen existing stores and mounts with scarf.DataStore(path, min_features_per_cell=-1).
  2. Never filter by editing I. Keep the cell_selection ref a filter returns and pass it on, or insert a boolean column and use cell_key=. ds.cells.reset_key("I") undoes an accidental open-time filter.
  3. Keep every ref you will reuse in a dict and record them. Never guess artifact IDs or read private Zarr paths. Use ds.inspect_artifact(ref), ds.load_artifact(ref) and ds.list_artifacts(...). Cell selections need scope="datastore" when listing.
  4. Check the matrix, then audit QC removals before accepting a filter. Published matrices can be corrected (for example SCTransform) or already filtered, which makes QC bounds meaningless; the read-only checks in references/quality-control.md ("Check the matrix first") report this. Pooled MAD filters routinely remove low-complexity populations (platelets, erythrocytes, neutrophils) and high-RNA ones (plasma, cycling cells). Audit retention per sample and per marker-defined group without author labels, and look at the markers of removed cells (references/quality-control.md).
  5. Labels are immutable and runs cannot be resumed. Use a new label for every variant. A rerun reuses every complete artifact, so it is cheap. Choose snapshot_columns (design columns you want inside run.cells and exports) on the first run, because changing them recomputes everything downstream.
  6. Read pipeline outputs instead of recomputing them. run["markers"], run["cell_cycle"] and run["doublets"] already exist. A direct run_marker_search(run["clusters"], ...) creates a new artifact and rereads all counts.
  7. Align arrays through the cell selection. Use run.cells.fetch(...) (run cells only) or fetch_all(...) (full length, with -1, NaN or "" outside the run). For explicit artifacts use ds.inspect_artifact(ref).input_ref("cell_selection"). Join tables on cell ids, never on row order.
  8. Budget count passes on mounts. QC metrics, the HVG summary, normalization per selection, markers, gene plots and AUCell each read counts. For many passes, make a local copy once: python -m scarf.tools.repack_zarr MOUNT.zarr LOCAL.zarr --mem-budget 4G.
  9. One DataStore per process; it is not thread-safe. Run parallel analyses in separate processes on separate stores, each with its own SCARF_MEM_BUDGET/SCARF_WORKERS.
  10. Payload names differ by kind. Leiden labels, embeddings, doublet and strength scores live under values; Paris (cluster_cut) under labels; PCA under data. get_markers returns string group_id while Leiden labels are integers.
  11. get_markers defaults hide genes. min_score=0.25, min_frac_exp=0.2; pass -1 for both to see negatives and full panels.
  12. Labels are hypotheses. Name a cluster only with two or more specific positive markers plus low lineage-negative markers and a plausible QC and doublet profile. Keep unresolved for mixed, doublet-like or marker-poor clusters.
  13. Cells are not replicates, and Scarf does not check your design. Condition claims need donor-level aggregation and at least two independent donors per group. A donor sampled twice is repeated measures. A donor with samples in two arms must not count in both: run_statistical_testing(sample_by="sample_id") silently does that, while sample_by="donor_id" refuses such a design. run_harmony runs silently on a column confounded with the biology you want to compare, so audit donors x batch x condition first. When inserting design columns, replace missing values explicitly: ds.cells.insert stores None as "" and pd.NA as "<NA>".
  14. Hold out annotation columns when they will judge the result. Do not use author labels (cell_type, author_cell_type, cell.type.*, singler, predicted.*, ...) to choose QC, parameters or labels. Use them only in a final, clearly separated comparison.
  15. Pass an explicit HVG blacklist. On stores with current gene symbols the default blacklist misses replication-dependent histones (^HIST matches only pre-2020 names), and its ^CCN family removes the CCN1-6 matricellular genes and non-cell-cycle cyclins. It also leaves Ig V/J genes in, which can split plasma cells by light chain instead of biology. Use the species string from references/gene-blacklists.md (params={"hvg": {"blacklist": ...}} in the pipeline). Add the clonotype add-on only on marker evidence. A blacklist= string replaces the default entirely, and changing it creates new HVG, PCA and graph artifacts.

The analysis loop

Work in this order. After each step, write the decision and its evidence to an analysis log.

StepDoEvidence to recordModule
1. Inspectread-only open, ds.summary(), metadata column profile, existing runs; check raw versus corrected countscells, assays, mounted or local, count type, metadata rolesdata-access.md, quality-control.md
2. Designidentify donor, sample, capture, batch and condition columns; derive missing unitsunit of inference, confounding, held-out columnsintegration-and-comparisons.md
3. QCinspect distributions; compare policies; audit removalschosen policy, cells kept per group, what was removedquality-control.md
4. Baselineds.pipeline.run(label=..., filtering=<chosen>, snapshot_columns=<design>)run id, report, kept cellspipeline-runs-and-artifacts.md
5. Structureresolutions, seed stability, nesting, marker support, QC and doublets per clusterchosen partition and whyclustering-and-embedding.md
6. Branch if neededHVG count and blacklist, PCs, k, subclustering one lineage, Harmony (only if the design allows)which alternative changed whatfeatures-and-graphs.md, gene-blacklists.md, integration-and-comparisons.md
7. Annotatemarker tables, canonical panels, gene-set scoreslabel, markers, negatives, unresolved clustersmarkers-and-annotation.md
8. Comparecomposition per donor, pseudobulk; only with replicateswhat can and cannot be claimedintegration-and-comparisons.md, performance-and-export.md (pseudobulk export)
9. Reportfigures, tables, handoff JSON, lineage, exportfiles written, open questionsplotting.md, performance-and-export.md

Quick start

A complete first pass on an RNA store. Every step after the run reuses the run's artifacts.

python
from pathlib import Path

import matplotlib

matplotlib.use("Agg")  # headless; select the backend before importing scarf
import numpy as np
import pandas as pd
import scarf

scarf.configure_output(level="WARNING", progress=False)
store = "analysis.zarr"
out = Path("analysis_out")
(out / "figures").mkdir(parents=True, exist_ok=True)

ds = scarf.DataStore(store, min_features_per_cell=-1)

# QC: compare candidate filters and audit them before choosing (quality-control.md).
cells = ds.snapshot_cell_selection("I")
mad5 = ds.auto_filter_cells(cell_selection=cells, n_mads=5)
kept = np.asarray(ds.load_artifact(mad5)["values"][:], dtype=bool)
print(f"MAD 5 keeps {kept.sum()} of {ds.cells.N} cells")

# Baseline run with the chosen policy; freeze design columns you will need later.
run = ds.pipeline.run(
    label="baseline_mad5",
    filtering={"method": "mad", "n_mads": 5},
    snapshot_columns=[c for c in ("donor_id", "sample_id") if c in ds.cells.columns],
)
print(run.status, list(run.keys()))
(out / "run_report.md").write_text(run.report(format="markdown"))

# Figures and tables.
umap = ds.plots.embedding(run=run, color_by="clusters", show=False)
umap.save(out / "figures" / "umap_clusters.png", dpi=150)
umap.close()
markers = ds.get_markers(marker=run["markers"])
markers.to_csv(out / "markers_top.csv", index=False)
print(markers.groupby("group_id", sort=False).head(5)
      .groupby("group_id", sort=False)["feature_name"].agg(", ".join))
sizes = pd.Series(run.cells.fetch("clusters")).value_counts().sort_index()
print(sizes.to_dict())
Show full SKILL.md (1,046 more words)Show less

Record the analysis

Keep these in the output directory so another agent or a person can audit and continue:

  • analysis_log.md: the question, the units, and one entry per step (decision, alternatives considered, evidence with numbers, figure paths).
  • handoff.json: run label and ID, and the refs you relied on (ref.to_dict(); restore with scarf.ArtifactRef.from_dict). Add unsupported claims and open questions. See pipeline-runs-and-artifacts.md.
  • run_report.md (run.report(format="markdown")) and lineage.md (ds.lineage({...}).to_markdown()).
  • Figures saved as PNG, and tables as CSV (markers, cluster sizes, composition, label map).

Evaluating against held-out labels

When a dataset ships author annotations, run the whole loop without them. Then add one final section that crosstabs your labels against theirs, reports ARI and per-type F1 after harmonizing vocabularies, and explains disagreements with marker evidence. Report QC retention per author type there too, as an audit of the filter. Never go back and tune parameters to raise agreement. Record any changes you make after seeing the comparison as such.

Reference modules

ModuleRead when
references/data-access.mdopening, inspecting, converting (H5AD, 10x, MTX, Seurat), Cytebase search, open and mount
references/quality-control.mdQC metrics, MAD/manual/per-sample filters, removal audits, doublet scores
references/features-and-graphs.mdHVGs, normalization, PCA dims, ANN, neighbours, graph diagnostics, branching, subclustering
references/gene-blacklists.mdwhat the default HVG blacklist removes, corrected blacklists for human and mouse, Ig/TCR and haemoglobin add-ons
references/clustering-and-embedding.mdLeiden and Paris, choosing a partition, UMAP and t-SNE, labels as arrays
references/markers-and-annotation.mdmarker tables, canonical marker panels, labelling, AUCell/WAGGR, cell cycle, label comparison
references/pipeline-runs-and-artifacts.mdds.pipeline.run options, run reports, artifacts, lineage, failures, handoff record
references/plotting.mdheadless figures, run mode versus ref mode, every ds.plots method
references/integration-and-comparisons.mdstudy design and units, donors in two arms, composition tests, Harmony and its safety, merging, mapping, condition comparisons
references/performance-and-export.mdbudgets, long-running steps, streaming, mount repacking, AnnData/H5AD/MTX/CSV and pseudobulk export, subset stores

Troubleshooting

SymptomLikely causeFix
Assay 'RNA' is not prepared on a read-only openstore just written by a converter or SubsetZarropen once writable, then read-only
Fewer active cells than expected after openingwritable open applied min_features_per_cellds.cells.reset_key("I"); reopen with -1
ValueError on ds.pipeline.run(label=...)label already completednew label; reuse makes it cheap
PermissionError from a producerstore opened with zarr_mode="r" and no matching artifactreopen writable
KeyError: 'values' on a Paris refParis stores labelsds.load_artifact(ref)["labels"]
Array length differs from ds.cells.Npayload follows the cell selectionalign with run.cells.fetch_all or the selection mask
Plot raises with run= and a gene or live columnrun mode accepts one frozen field onlylayout=run["umap"], color_by=[...]
TypeError from distribution(grouping="col")grouping needs a ref or CellFieldgrouping=scarf.plotting.CellField("col")
Very slow steps on a mounteach count pass is a network readfewer passes; repack locally (rule 8)
MemoryError (CountLayoutMemoryError after 1.0.0rc19) from a converterdefault count layout does not fit mem_budgetlarger mem_budget; else the policy= the message names (references/data-access.md)
ValueError plotting after reopening the storea PipelineRun is bound to the store object that opened itreopen the run from the new ds
list_artifacts(kind="cell_selection") is emptycell selections are datastore-scopedadd scope="datastore"
KeyError: 'groups' in a dot plot tablewith group_by= the column is named after the grouping columnread res.tables["aggregate"].columns first
A wait for a long step never endsunbounded polling, or the process diedbounded wait that checks the process and the log tail (Setup)
QC bounds look odd or nothing is filteredcounts are corrected or already filteredcheck the matrix first; prefer flag-only or gentle filters
ValueError: None of the s_genes match the assay feature names from ds.pipeline.runfeature names are not gene symbols (Ensembl IDs, synthetic names); matching ignores case, so mouse symbols workcell_cycle=False, or pass lists in the store's naming via params={"cell_cycle": {"s_genes": [...], "g2m_genes": [...]}}

Limitations

  • Covers core Scarf only. It does not use or describe scarf.agent.
  • Needs Python 3.12+ and scarf 1.0.0rc17 or newer, which is a pre-release. A bare scarf[extra] install resolves to the 0.32 series, whose API differs from what this skill describes.
  • Scarf does not check the study design. Donor and replicate structure, confounding and held-out annotation columns are the analyst's responsibility (rules 13 and 14).
  • Cluster labels from this workflow are hypotheses backed by marker evidence, not a validated cell-type classifier.
  • Numbers in the reference modules come from the 10x 5K PBMC documentation dataset. They are illustrations, not thresholds.
  • Runtime and memory depend on hardware, storage and data. Measured end-to-end runs are on the benchmarks page (concepts/benchmarks.html); they are not hardware guarantees.
  • A DataStore is not thread-safe: use one per process.
  • Cytebase datasets need the cytebase extra and network access, and a mounted store reads counts over the network on every pass.
  • This catalogue copy leaves out scripts/inspect_store.py from the maintained copy; use the read-only recipes in the reference modules instead.

Security & Safety Notes

  • pip install "scarf[extra]>=1.0.0rc17" installs a pre-release from PyPI into the active Python environment. Ask the user before installing, and prefer a virtual environment.
  • Writable opens and every producer write into the store: artifacts, pipeline runs and, unless you pass min_features_per_cell=-1, the open-time filter on I (rule 1). Inspect with zarr_mode="r" first, and copy a store before the first writable open when the original must stay unchanged.
  • Converters, repack_zarr, subset stores and exports write new files that can be as large as the counts. Check free disk space and confirm output paths before running them.
  • Object-storage paths, mounts and Cytebase read over the network. Cytebase can use an existing Hugging Face login for private buckets. Never print, log or copy access tokens into analysis logs, handoff files or reports.
  • By default Scarf claims all detected RAM and every CPU. On shared machines set SCARF_MEM_BUDGET and SCARF_WORKERS before importing it.
  • Run long steps in the background with a bounded wait that also stops on a Traceback or a dead process (Setup). Never poll without a time limit.

Documentation map

The docs hold the full explanations, online at https://scarf.readthedocs.io/en/latest/: quickstart.html, tutorials/<step>.html (one page per step, for example quality_control, graph_construction, clustering, annotation, batch_correction, pseudobulk_and_differential_expression, cytebase), concepts/ (provenance, memory_and_execution, benchmarks), analysis_with_agents.html (scientific decision loop, task routing, handoff) and reference/api/<module>.html (exact signatures).

Source and License

Copyright (c) 2026, Nygen Analytics AB. Licensed under the BSD 3-Clause License; the full text is in references/LICENSE.md. The maintained copy lives in NygenAnalytics/scarf at https://github.com/NygenAnalytics/scarf/tree/master/skills/scarf-single-cell. This catalogue copy adds catalogue frontmatter and the "When to Use This Skill", "Limitations", "Security & Safety Notes" and "Source and License" sections, and leaves out scripts/inspect_store.py. The analysis guidance is otherwise unchanged.

© sickn33, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (references) in skills/scarf-single-cell of sickn33/agentic-awesome-skills.

  • SKILL.md
  • references/LICENSE.md
  • references/clustering-and-embedding.md
  • references/data-access.md
  • references/features-and-graphs.md
  • references/gene-blacklists.md
  • references/integration-and-comparisons.md
  • references/markers-and-annotation.md
  • references/performance-and-export.md
  • references/pipeline-runs-and-artifacts.md
  • references/plotting.md
  • references/quality-control.md

Open the folder on GitHubat commit 680176d

Compare with similar skills

Scarf Single Cell next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scarf Single Cell compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scarf Single Cell this skillsickn33/agentic-awesome-skills47k—~5.7kAutomated safety check: PassBSD-3-Clause
Scarf Single CellNygenAnalytics/scarf126—~5.6kAutomated safety check: PassBSD-3-Clause
Spatial AteraQING1105/ezST101—~576Automated safety check: PassMIT
AnndataK-Dense-AI/scientific-agent-skills48k1 repos~3.9kAutomated safety check: NotesBSD-3-Clause
Anndata Data Structurejaechang-hits/SciAgent-Skills3712 repos~5.8kAutomated safety check: PassBSD-3-Clause
Bio Single Cell Data IoGPTomics/bioSkills1.2k1 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Scarf Single Cell

    NygenAnalytics/scarf

    Analyze single-cell data with core Scarf, the out-of-core Zarr DataStore library with immutable artifacts and pipeline runs.

    126 GitHub stars~5.6k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Spatial Atera

    QING1105/ezST

    Atera platform branch of the spatial transcriptomics workflow — load and validate Atera cell-level output (AnnData + Zarr segmentation) for downstream analysis.

    101 GitHub stars~576 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Anndata

    K-Dense-AI/scientific-agent-skills

    Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.

    48k GitHub starsUsed in 1 repo~3.9k tokens
    Research & ScienceAuto-check: notes
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 2 repos~5.8k tokens
    Research & ScienceAuto-check passed
  • Bio Single Cell Data Io

    GPTomics/bioSkills

    Read, write, create, and convert single-cell objects across AnnData (Python), Seurat (R), and SingleCellExperiment (R).

    1.2k GitHub starsUsed in 1 repo~3.3k tokens
    Research & ScienceAuto-check passed
  • Lamindb Data Management

    jaechang-hits/SciAgent-Skills

    Open-source FAIR biology data framework. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 2 repos~4k tokens
    Research & ScienceAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,493 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Questions about Scarf Single Cell

What does Scarf Single Cell do?

Analyze single-cell RNA-seq at million-cell scale with Scarf: out-of-core Zarr stores on disk or object storage, provenance-tracked artifacts, audited QC, clustering, markers and donor comparisons. Scarf Single Cell is an agent skill from sickn33/agentic-awesome-skills. Analyze single-cell RNA-seq at million-cell scale with Scarf: out-of-core Zarr stores on disk or object storage, provenance-tracked artifacts, audited QC, clustering, markers and donor comparisons.

When should I use Scarf Single Cell?

Scarf Single Cell fits situations like: tasks that involve Bioinformatics.

How do I install Scarf Single Cell in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill scarf-single-cell -a claude-code`. Or copy the skill folder (skills/scarf-single-cell in sickn33/agentic-awesome-skills) into .claude/skills/scarf-single-cell in your project. Claude Code loads it when a task matches its description.

How do I install Scarf Single Cell in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill scarf-single-cell -a codex`. Or copy the skill folder (skills/scarf-single-cell in sickn33/agentic-awesome-skills) into .agents/skills/scarf-single-cell in your project. Codex loads it when a task matches its description.

Can I use Scarf Single Cell in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill scarf-single-cell -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scarf-single-cell, .gemini/skills/scarf-single-cell, .github/skills/scarf-single-cell and .opencode/skills/scarf-single-cell in your project.

What does Scarf Single Cell need to run?

Going by SKILL.md and its folder, Scarf Single Cell needs the command-line tools its instructions call (pip, uv and python). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.12+ and scarf 1.0.0rc17 or newer (pip install "scarf[extra]>=1.0.0rc17"; add the cytebase extra and network access for Cytebase datasets)..

Does Scarf Single Cell access the network?

SKILL.md names 2 domains. As links in the text: github.com and scarf.readthedocs.io. This is read from the text; nothing was executed.

Is Scarf Single Cell safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scarf Single Cell use?

Scarf Single Cell is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scarf Single Cell use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 38k tokens, read only when the agent opens those files.

What are the alternatives to Scarf Single Cell?

Skills that share tags, products or a category with Scarf Single Cell: Scarf Single Cell (NygenAnalytics/scarf, 126 stars), Spatial Atera (QING1105/ezST, 101 stars), Anndata (K-Dense-AI/scientific-agent-skills, 48k stars) and Anndata Data Structure (jaechang-hits/SciAgent-Skills, 371 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scarf Single Cell?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,379 GitHub stars. The repository holds 1,493 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.