Agent skill

Sc Qc

by TianGzlab in TianGzlab/OmicsClaw

Load when computing per-cell QC metrics (ngenes, total counts, mt%, ribo%) on a single-cell AnnData before filtering.

Apache-2.0Auto-check passedResearch & Science

Install Sc Qc

skills CLI
$ npx skills add TianGzlab/OmicsClaw --skill sc-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TianGzlab/OmicsClaw sc-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TianGzlab/OmicsClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/singlecell/scrna/sc-qc .claude/skills/sc-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sc-qc
GitHub stars
161
Token cost
~2k tokens
SKILL.md length
892 words
Files
9 (incl. references)
Skills in repo
88
Repo updated
First seen
Licence
Apache-2.0

At a glance

Load when computing per-cell QC metrics (ngenes, total counts, mt%, ribo%) on a single-cell AnnData before filtering.

  • Tasks that involve Bioinformatics
  • SKILL.md covers When to use, Use from a step, API and Methods and parameters, plus 5 more sections
  • Runs Python scripts from its folder; calls python

What it does

Sc Qc is an agent skill from TianGzlab/OmicsClaw. Load when computing per-cell QC metrics (ngenes, total counts, mt%, ribo%) on a single-cell AnnData before filtering. Skip when reads are still raw FASTQ (use sc-fastq-qc); you want to filter cells now (use sc-filter).

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `_api.py`, `examples/example_step.py` and `references/methodology.md`).

It sits in Research & Science, covering Bioinformatics. It works with AnnData. The repository describes itself as: Conversational & memory-enabled AI research partner for multi-omics analysis. CLI + Desktop App (installers in Releases). From biological idea to full research paper. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/sc-qc”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 90a3bec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sc Qc loads about 2k tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 56 tokens; SKILL.md has 892 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TianGzlab/OmicsClaw at commit 90a3bec, republished under its Apache-2.0 licence (© TianGzlab). 892 words, ~1,981 tokens.

Download SKILL.mdSave it as .claude/skills/sc-qc/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
sc-qc
description
Load when computing per-cell QC metrics (n_genes, total counts, mt%, ribo%) on a single-cell AnnData before filtering. Skip when reads are still raw FASTQ (use sc-fastq-qc); you want to filter cells now (use sc-filter).
trigger
scRNA QC, single-cell QC, quality control, mitochondrial percentage, ribosomal percentage, QC violin, QC scatter, n genes per cell
tags
singlecell, scrna, qc, mitochondrial, ribosomal

sc-qc

When to use

The user has a single-cell AnnData (post-counting / post-standardisation) and wants to review cell quality — counts, detected genes, mitochondrial percentage, ribosomal percentage — before any filtering. This skill reports, it does not remove cells. Use sc-filter to actually drop cells based on these metrics.

Use from a step

python
qc = load_skill("sc-qc")
adata = qc.calculate_qc(read_input("data/pbmc.h5ad"), species="human")
write_output(qc.qc_summary(adata), "tables/qc_metrics_summary.csv")
write_output(qc.qc_figure(adata), "figures/qc_histograms.png")
write_output(adata, "intermediate/adata_qc.h5ad")

A complete step that runs on demo data: examples/example_step.py.

API

<!-- api:begin generated from _api.py; regenerate with run.py api <skill dir> --write -->
calculate_qc(adata, *, species: str='human', calculate_ribo: bool=True)

Bring the input into the OmicsClaw scRNA contract and add per-cell QC metrics to obs.

The count-like matrix becomes X and layers['counts'], raw keeps a counts snapshot, and gene names are made unique. Then scanpy's calculate_qc_metrics adds n_genes_by_counts, total_counts, pct_counts_mt and pct_counts_ribo (when the gene prefixes match), plus log10_total_counts and log10_n_genes_by_counts; var gets mt and ribo flags. How the input was read is recorded for :func:run_info.

:param species: "human" (MT-, RPS/RPL prefixes) or "mouse" (mt-, Rps/Rpl). Default "human", the CLI default. Set it to the organism of the data: with the wrong one no mitochondrial gene matches and pct_counts_mt is not computed at all. :param calculate_ribo: Also compute pct_counts_ribo. Default True, as the CLI does. :returns: A new AnnData (the standardized copy), with the metrics in obs. :raises ValueError: no count-like matrix can be found in X, layers or raw.

run_info(adata, *, keep: bool=True) -> dict

What :func:calculate_qc recorded about the input it prepared.

Keys: species, calculate_ribo, expression_source (the matrix used as counts), gene_name_source, warnings, input_contract, matrix_contract and qc_obs_columns.

:param keep: Leave the record in adata.uns; False removes it. :returns: The record, or an empty dict when calculate_qc has not run on adata.

qc_summary(adata) -> pd.DataFrame

One row per QC metric with min, max, mean, median, std, q25 and q75.

:returns: Columns metric, min, max, mean, median, std, q25, q75, for the metrics :func:calculate_qc added.

qc_metrics_table(adata) -> pd.DataFrame

The per-cell QC metrics, one row per cell.

:returns: Column cell_id followed by the QC metrics present in obs.

highest_expressed_genes(adata, *, n_top: int=20) -> pd.DataFrame

The genes with the highest mean expression in X.

A few genes dominating the counts (mitochondrial, ribosomal, MALAT1) point to stressed cells or ambient RNA.

:param n_top: How many genes to return. Default 20, the CLI's table length. :returns: Columns gene and mean_expression, highest first.

barcode_rank_table(adata) -> pd.DataFrame

Library sizes ranked from largest to smallest, for a barcode-rank (knee) plot.

:returns: Columns rank, total_counts, log10_rank and log10_total_counts. :raises KeyError: total_counts is not in obs; run :func:calculate_qc first.

qc_correlation_table(adata, *, metrics: list[str] | None=None) -> pd.DataFrame

Pearson correlations between QC metrics across cells.

:param metrics: The obs columns to correlate. Default: the QC metrics present. :returns: A square table with a metric column naming each row.

qc_figure(adata, *, metrics: list[str] | None=None)

Histograms of the QC metrics, one panel per metric, with the median marked.

:param metrics: The obs columns to plot. Default: the QC metrics present. :returns: A matplotlib Figure.

<!-- api:end -->
Show full SKILL.md (436 more words)Show less

Methods and parameters

There is one method: scanpy's calculate_qc_metrics on the count matrix, after the input is brought into the OmicsClaw scRNA contract (counts in X and layers['counts'], a counts snapshot in raw, unique gene names).

ParameterDefaultWhere it comes fromWhen to change it
species"human"the CLI defaultAlways set it to the organism of the data. It picks the gene prefixes: MT- and RPS/RPL for human, mt- and Rps/Rpl for mouse. Ask the user when the organism is not in the data description.
calculate_riboTruethe CLI's fixed behaviourLeave it on; set False only when ribosomal content is irrelevant to the question.
n_top of highest_expressed_genes20the CLI's table lengthRaise it to look for more contaminating genes.

QC thresholds are not chosen here. Read qc_summary and the figure, then choose thresholds for sc-filter or sc-preprocessing with the user; the usual starting points per tissue are in references/methodology.md.

Gotchas

  • No filtering happens here. calculate_qc keeps every cell and gene. Filtering is sc-filter (or the thresholds of sc-preprocessing).
  • A wrong species loses the mitochondrial metric. With no gene matching the prefix, pct_counts_mt is not computed at all, and qc_summary has no row for it. Check that the row is there before reporting mitochondrial content.
  • run_info(adata)["expression_source"] says which matrix was used as counts: layers.counts, adata.raw or adata.X. QC fractions only mean something on a count-like source; report it when the input came from outside OmicsClaw.
  • calculate_qc returns a new object. Keep the return value (adata = qc.calculate_qc(adata)); the input AnnData is left as it was.
  • Non-count input raises ValueError. A log-normalised X with no counts layer or count-like raw cannot be QC'd; ask for the raw counts.

Inputs and outputs

  • calculate_qc reads the count-like matrix from layers['counts'], raw or X, and gene names from var (a gene-symbol column when there is one). It writes obs: n_genes_by_counts, total_counts, pct_counts_mt, pct_counts_ribo, log10_total_counts, log10_n_genes_by_counts; var: mt, ribo; layers['counts'], raw, and the contract entries in uns.
  • The table functions read those obs columns and return DataFrames; qc_figure returns a matplotlib Figure.

CLI

sc_qc.py runs the same functions outside a project and writes a report, figures, tables and processed.h5ad: python <skill directory>/sc_qc.py --help. --demo runs it on PBMC3K.

See also

  • references/parameters.md — every CLI flag and tuning hint
  • references/methodology.md — mt/ribo gene-pattern detection, scanpy QC parameters, tissue thresholds
  • references/output_contract.md — the CLI's table column schemas and figure roles
  • Adjacent skills: sc-standardize-input (upstream — required if input is external), sc-filter (next step — actually removes cells), sc-doublet-detection (parallel — finds doublets)

Dependencies

Python packages this skill's script needs. They are not installed for you — check before a long run.

anndata, matplotlib, numpy, pandas, scanpy, scipy, seaborn

© TianGzlab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/singlecell/scrna/sc-qc of TianGzlab/OmicsClaw.

  • SKILL.md
  • _api.py
  • examples/example_step.py
  • references/methodology.md
  • references/output_contract.md
  • references/parameters.md
  • references/r_visualization.md
  • sc_qc.py
  • tests/test_sc_qc.py

Open the folder on GitHubat commit 90a3bec

Compare with similar skills

Sc Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sc Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sc Qc this skillTianGzlab/OmicsClaw161—~2kAutomated safety check: PassApache-2.0
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k15 repos~2.8kAutomated safety check: PassMIT
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
PyDESeq2 Differential Expressiondavila7/claude-code-templates32k11 repos~4kAutomated safety check: PassMIT
Anndatadavila7/claude-code-templates32k11 repos~2.5kAutomated safety check: PassMIT
Single-Cell Initial AnalysisLigphiDonk/Oh-my--paper7381 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 15 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    32k GitHub starsUsed in 11 repos~4k tokens
    Research & ScienceAuto-check passed
  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    32k GitHub starsUsed in 11 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    738 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check passed
  • Single Cell Data Prep Qc

    harrisongzhang/TheVirtualBiotech

    Single-cell RNA-seq data preparation and quality control pipeline.

    121 GitHub stars~2.9k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed

More from TianGzlab/OmicsClaw

All 88 skills in this repo
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated 2 days ago
    Auto-check passed
  • Bulkrna Batch Correction

    TianGzlab/OmicsClaw

    Load when correcting batch effects in bulk expression using R sva ComBat or the legacy Python parametric approximation.

    161 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Bulkrna Coexpression

    TianGzlab/OmicsClaw

    Load when discovering bulk gene co-expression modules and hub genes with R WGCNA.

    161 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Bulkrna De

    TianGzlab/OmicsClaw

    Load when comparing gene expression between two conditions in bulk RNA-seq count data.

    161 GitHub stars~867 tokensUpdated 2 days ago
    Auto-check passed
  • Bulkrna Deconvolution

    TianGzlab/OmicsClaw

    Load when estimating cell-type proportions in bulk RNA-seq samples from a single-cell or signature-matrix reference.

    161 GitHub stars~757 tokensUpdated 2 days ago
    Auto-check passed
  • Bulkrna Enrichment

    TianGzlab/OmicsClaw

    Load when running pathway / GO term enrichment on a bulk RNA-seq DE result list.

    161 GitHub stars~860 tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Sc Qc

What does Sc Qc do?

Load when computing per-cell QC metrics (ngenes, total counts, mt%, ribo%) on a single-cell AnnData before filtering. Sc Qc is an agent skill from TianGzlab/OmicsClaw. Load when computing per-cell QC metrics (ngenes, total counts, mt%, ribo%) on a single-cell AnnData before filtering.

When should I use Sc Qc?

Sc Qc fits situations like: tasks that involve Bioinformatics.

How do I install Sc Qc in Claude Code?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-qc -a claude-code`. Or copy the skill folder (skills/singlecell/scrna/sc-qc in TianGzlab/OmicsClaw) into .claude/skills/sc-qc in your project. Claude Code loads it when a task matches its description.

How do I install Sc Qc in Codex?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-qc -a codex`. Or copy the skill folder (skills/singlecell/scrna/sc-qc in TianGzlab/OmicsClaw) into .agents/skills/sc-qc in your project. Codex loads it when a task matches its description.

Can I use Sc Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TianGzlab/OmicsClaw --skill sc-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sc-qc, .gemini/skills/sc-qc, .github/skills/sc-qc and .opencode/skills/sc-qc in your project.

What does Sc Qc need to run?

Going by SKILL.md and its folder, Sc Qc needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Sc Qc access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sc Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sc Qc use?

Sc Qc is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sc Qc use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Sc Qc?

Skills that share tags, products or a category with Sc Qc: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), Scgpt (JimLiu/science-skills, 227 stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 32k stars) and Anndata (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sc Qc?

TianGzlab (a GitHub organization) maintains it in TianGzlab/OmicsClaw, which has 161 GitHub stars. The repository holds 88 skills in this directory. The repository was last updated on October 7, 2026.

Source: TianGzlab/OmicsClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.