Agent skill

Sc Filter

by TianGzlab in TianGzlab/OmicsClaw

Load when removing low-quality cells and lowly-detected genes from a single-cell AnnData using QC-derived thresholds or tissue presets.

Apache-2.0Auto-check passedResearch & Science

Install Sc Filter

skills CLI
$ npx skills add TianGzlab/OmicsClaw --skill sc-filter -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TianGzlab/OmicsClaw sc-filter --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TianGzlab/OmicsClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/singlecell/scrna/sc-filter .claude/skills/sc-filter && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sc-filter
GitHub stars
161
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
870 words
Files
10 (incl. references)
Skills in repo
88
Repo updated
First seen
Licence
Apache-2.0

At a glance

Load when removing low-quality cells and lowly-detected genes from a single-cell AnnData using QC-derived thresholds or tissue presets.

  • Tasks that involve Bioinformatics
  • SKILL.md covers When to use, Use from a step, API and Methods and parameters, plus 5 more sections
  • Runs Python scripts from its folder; calls python

What it does

Sc Filter is an agent skill from TianGzlab/OmicsClaw. Load when removing low-quality cells and lowly-detected genes from a single-cell AnnData using QC-derived thresholds or tissue presets. Skip when the full normalize→HVG→PCA→cluster pipeline (use sc-preprocessing); reads are still raw FASTQ (use sc-fastq-qc).

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `_api.py`, `examples/example_step.py` and `references/methodology.md`).

It sits in Research & Science, covering Bioinformatics. It works with AnnData. The repository describes itself as: Conversational & memory-enabled AI research partner for multi-omics analysis. CLI + Desktop App (installers in Releases). From biological idea to full research paper. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/sc-filter”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 90a3bec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sc Filter loads about 2.1k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 870 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TianGzlab/OmicsClaw at commit 90a3bec, republished under its Apache-2.0 licence (© TianGzlab). 870 words, ~2,127 tokens.

Download SKILL.mdSave it as .claude/skills/sc-filter/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
sc-filter
description
Load when removing low-quality cells and lowly-detected genes from a single-cell AnnData using QC-derived thresholds or tissue presets. Skip when the full normalize→HVG→PCA→cluster pipeline (use sc-preprocessing); reads are still raw FASTQ (use sc-fastq-qc).
trigger
filter cells, cell filtering, gene filtering, remove low quality, qc filtering, tissue-specific thresholds
tags
singlecell, scrna, filter, qc, mitochondrial

sc-filter

When to use

The user has reviewed sc-qc output and now wants to actually drop low-quality cells and lowly-detected genes — by per-cell thresholds (--min-genes, --max-genes, --max-mt-percent, --min-counts, --max-counts, --min-cells) or tissue-specific presets (--tissue brain / pbmc / etc.). This skill removes cells; it does not normalise, cluster, or annotate.

Use from a step

python
filtering = load_skill("sc-filter")
before = read_input("results/01_qc/intermediate/adata_qc.h5ad")
after = filtering.filter_cells(before, tissue="pbmc")
write_output(filtering.filter_summary(after), "tables/filter_summary.csv")
write_output(filtering.filter_stats_table(after), "tables/filter_stats.csv")
write_output(after, "intermediate/adata_filtered.h5ad")

Run doublet detection before filtering when doublets should be removed. In the next preprocessing step, call preprocess(after, apply_filters=False) to keep these cells and genes. A complete demo step is in examples/example_step.py.

API

<!-- api:begin generated from _api.py; regenerate with run.py api <skill dir> --write -->
filter_cells(adata, *, min_genes: int=200, max_genes: int | None=None, min_counts: int | None=None, max_counts: int | None=None, max_mt_percent: float | None=20.0, min_cells: int=3, tissue: str | None=None, remove_doublets: bool=True, doublet_score_threshold: float=0.25)

Return a filtered copy of an AnnData, with counts and matrix contracts.

Existing QC columns are reused. Otherwise counts are selected from the input's counts layer, raw snapshot or X and missing QC metrics are computed. A normalized X with existing QC metrics is preserved. Doublet removal reads existing obs columns; it does not run doublet detection.

:param min_genes: Minimum detected genes per cell. Default 200. :param max_genes: Maximum detected genes per cell; None leaves it uncapped. :param min_counts: Minimum counts per cell; None disables this threshold. :param max_counts: Maximum counts per cell; None disables this threshold. :param max_mt_percent: Maximum mitochondrial percentage. Default 20.0; None disables the threshold. :param min_cells: Minimum retained cells expressing a gene. Default 3. :param tissue: Preset from tissue_presets(); overrides min_genes, max_genes and max_mt_percent. None uses the supplied thresholds. :param remove_doublets: Drop cells marked by predicted_doublet, or by doublet_score when the boolean column is absent. Default True. :param doublet_score_threshold: Score cutoff for the score-only case. Default 0.25. :returns: A new AnnData with retained cells and genes; run_info records effective thresholds, the count source and the filtering summary. :raises ValueError: Input requiring QC metrics has no count-like matrix.

run_info(adata, *, keep: bool=True) -> dict

Read the filtering record from the returned AnnData.

:param keep: False removes the record from uns; default True keeps it. :returns: summary, effective_params, input_contract and matrix_contract; an empty dict when filter_cells has not run. outliers_flagged counts existing outlier flags, which do not themselves remove cells.

filter_summary(adata) -> pd.DataFrame

Return retention counts and effective thresholds as metric/value rows.

:param adata: The result of filter_cells. :returns: A DataFrame matching the CLI's tables/filter_summary.csv.

filter_stats_table(adata) -> pd.DataFrame

Return metric/value rows for threshold removals and existing outlier flags.

Threshold counts can overlap. outliers_flagged is a count of flags, not removals. doublets_removed counts doublets remaining after QC thresholds.

retention_table(adata) -> pd.DataFrame

Return Cells and Genes rows with before/after counts from filter_cells.

filter_state_table(before, after) -> pd.DataFrame

Return QC metrics and Retained/Removed labels for every input cell.

Missing QC columns are calculated on a copy using filter_cells' input preparation. The index contains the original cell names.

tissue_presets() -> dict

Return copies of the shared QC presets, including their descriptions.

Each preset has min_genes, max_genes and max_mt (a percentage). These replace the corresponding filter_cells thresholds when tissue is set.

filter_figure(before, after)

Return a matplotlib Figure comparing cell and gene counts before and after filtering.

<!-- api:end -->
Show full SKILL.md (377 more words)Show less

Methods and parameters

filter_cells applies thresholds to cells, removes existing doublet calls, then drops genes expressed in too few retained cells.

  • min_genes=200 and min_cells=3 follow the existing OmicsClaw CLI defaults.
  • max_mt_percent=20.0 is the existing permissive ceiling. Choose the threshold from the QC distribution and record the reason in the step.
  • max_genes, min_counts and max_counts default to None, disabling those limits.
  • tissue=None uses the supplied thresholds. tissue_presets() returns the OmicsClaw heuristics: PBMC uses 200 to 2,500 genes and at most 5% MT counts.
  • remove_doublets=True uses predicted_doublet first; if only doublet_score exists, the existing default cutoff is 0.25. Without either column it does no doublet filtering.

Gotchas

  • run_info(after)["effective_params"] records the thresholds after a tissue preset overrides min_genes, max_genes and max_mt_percent. To set these independently, leave tissue=None; editing documentation does not change a preset.
  • run_info(after)["summary"]["input_preparation"] identifies the count source. QC columns already in obs are reused, including on normalized inputs; check how they were computed before applying count and MT thresholds.
  • filter_stats_table(after) counts overlapping threshold failures. Its rows need not sum to the number removed. outliers_flagged counts pre-existing outlier flags; flags alone do not remove cells.
  • _api.py:53 (filter_cells) needs pct_counts_mt when MT filtering is enabled. If no mitochondrial features matched and this column is absent, inspect the gene names or explicitly disable that threshold with max_mt_percent=None.
  • tissue_presets() has no lung preset. The legacy CLI accepts --tissue lung but the shared helper uses the default preset and logs a warning.

Inputs and outputs

The functions accept AnnData. filter_cells reads counts and QC columns, returns a new object with retained cells and genes, preserves a normalized X when existing QC metrics permit it, and stores the run record in uns. The summary functions return DataFrames; filter_figure returns a matplotlib Figure. Use write_output for files.

The CLI writes processed.h5ad, report.md, result.json, three tables (filter_stats.csv, filter_summary.csv, retention_summary.csv), and its plot data under figure_data/. The figure and optional R-output inventory is in references/output_contract.md.

Key CLI

bash
# Demo
python skills/singlecell/scrna/sc-filter/sc_filter.py --demo --output /tmp/sc_filter_demo

# Threshold-based (typical PBMC defaults)
python skills/singlecell/scrna/sc-filter/sc_filter.py \
  --input qc_output.h5ad --output results/ \
  --min-genes 200 --max-mt-percent 20 --min-cells 3

# Tissue preset (overrides matching CLI flags)
python skills/singlecell/scrna/sc-filter/sc_filter.py \
  --input qc_output.h5ad --output results/ --tissue pbmc

See also

  • references/parameters.md — every CLI flag and tuning hint
  • references/methodology.md — tissue preset definitions, threshold semantics
  • references/output_contract.md — processed.h5ad + table schemas
  • Adjacent skills: sc-qc supplies metrics; sc-doublet-detection supplies doublet calls; sc-preprocessing normalises the filtered AnnData with apply_filters=False.

Dependencies

Python packages this skill's script needs. They are not installed for you — check before a long run.

anndata, matplotlib, numpy, pandas, scanpy, scipy, seaborn

© TianGzlab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in skills/singlecell/scrna/sc-filter of TianGzlab/OmicsClaw.

  • SKILL.md
  • _api.py
  • examples/example_step.py
  • references/methodology.md
  • references/output_contract.md
  • references/parameters.md
  • references/r_visualization.md
  • sc_filter.py
  • tests/test_filter_api.py
  • tests/test_sc_filter.py

Open the folder on GitHubat commit 90a3bec

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in TianGzlab/OmicsClaw, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sc Filter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sc Filter compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sc Filter this skillTianGzlab/OmicsClaw1611 repos~2.1kAutomated safety check: PassApache-2.0
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
PyDESeq2 Differential Expressiondavila7/claude-code-templates32k12 repos~4kAutomated safety check: PassMIT
Anndatadavila7/claude-code-templates32k12 repos~2.5kAutomated safety check: PassMIT
Single-Cell Initial AnalysisLigphiDonk/Oh-my--paper7381 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    32k GitHub starsUsed in 12 repos~4k tokens
    Research & ScienceAuto-check passed
  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    32k GitHub starsUsed in 12 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    738 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check passed
  • Single Cell Data Prep Qc

    harrisongzhang/TheVirtualBiotech

    Single-cell RNA-seq data preparation and quality control pipeline.

    120 GitHub stars~2.9k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed

More from TianGzlab/OmicsClaw

All 88 skills in this repo
  • Bulkrna De

    TianGzlab/OmicsClaw

    Load when comparing gene expression between two conditions in bulk RNA-seq count data.

    161 GitHub starsUsed in 1 repo~867 tokens
    Auto-check passed
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated yesterday
    Auto-check passed
  • Bulkrna Qc

    TianGzlab/OmicsClaw

    Load when checking a bulk RNA-seq count matrix for library-size outliers, gene detection rates, and sample-sample correlation before DE.

    161 GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Sc Fastq Qc

    TianGzlab/OmicsClaw

    Load when checking raw single-cell FASTQ read quality (Phred / GC / adapter / length) before counting.

    161 GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Sc Markers

    TianGzlab/OmicsClaw

    Load when ranking cluster-level marker genes from a clustered single-cell AnnData via Scanpy Wilcoxon / t-test / logreg or COSG specificity.

    161 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Sc Qc

    TianGzlab/OmicsClaw

    Load when computing per-cell QC metrics (ngenes, total counts, mt%, ribo%) on a single-cell AnnData before filtering.

    161 GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed

Works with

Questions about Sc Filter

What does Sc Filter do?

Load when removing low-quality cells and lowly-detected genes from a single-cell AnnData using QC-derived thresholds or tissue presets. Sc Filter is an agent skill from TianGzlab/OmicsClaw. Load when removing low-quality cells and lowly-detected genes from a single-cell AnnData using QC-derived thresholds or tissue presets.

When should I use Sc Filter?

Sc Filter fits situations like: tasks that involve Bioinformatics.

How do I install Sc Filter in Claude Code?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-filter -a claude-code`. Or copy the skill folder (skills/singlecell/scrna/sc-filter in TianGzlab/OmicsClaw) into .claude/skills/sc-filter in your project. Claude Code loads it when a task matches its description.

How do I install Sc Filter in Codex?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-filter -a codex`. Or copy the skill folder (skills/singlecell/scrna/sc-filter in TianGzlab/OmicsClaw) into .agents/skills/sc-filter in your project. Codex loads it when a task matches its description.

Can I use Sc Filter in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TianGzlab/OmicsClaw --skill sc-filter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sc-filter, .gemini/skills/sc-filter, .github/skills/sc-filter and .opencode/skills/sc-filter in your project.

What does Sc Filter need to run?

Going by SKILL.md and its folder, Sc Filter needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Sc Filter access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sc Filter safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sc Filter use?

Sc Filter is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sc Filter use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Sc Filter?

Skills that share tags, products or a category with Sc Filter: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), Scgpt (JimLiu/science-skills, 227 stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 32k stars) and Anndata (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sc Filter?

TianGzlab (a GitHub organization) maintains it in TianGzlab/OmicsClaw, which has 161 GitHub stars. The repository holds 88 skills in this directory. The repository was last updated on October 7, 2026.

Source: TianGzlab/OmicsClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.