Agent skill

Sc Ambient Removal

by TianGzlab in TianGzlab/OmicsClaw

Load when removing ambient RNA contamination from droplet-based scRNA-seq using a simple subtraction path, CellBender, or SoupX.

Apache-2.0Auto-check passedResearch & Science

Install Sc Ambient Removal

skills CLI
$ npx skills add TianGzlab/OmicsClaw --skill sc-ambient-removal -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TianGzlab/OmicsClaw sc-ambient-removal --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TianGzlab/OmicsClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/singlecell/scrna/sc-ambient-removal .claude/skills/sc-ambient-removal && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sc-ambient-removal
GitHub stars
161
Token cost
~2.2k tokens
SKILL.md length
900 words
Files
9 (incl. references)
Skills in repo
88
Repo updated
First seen
Licence
Apache-2.0

At a glance

Load when removing ambient RNA contamination from droplet-based scRNA-seq using a simple subtraction path, CellBender, or SoupX.

  • Works in 6 steps: Load filtered AnnData; optionally load… → Validate --contamination is in [0, 1)… → Run the chosen --method against… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers When to use, Use from a step, API and Methods and parameters, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

Sc Ambient Removal is an agent skill from TianGzlab/OmicsClaw. Load when removing ambient RNA contamination from droplet-based scRNA-seq using a simple subtraction path, CellBender, or SoupX. Skip when the contamination is multiplet barcodes (use sc-doublet-detection); before counts exist (use sc-count).

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `_api.py`, `examples/example_step.py` and `references/methodology.md`).

It sits in Research & Science, covering Bioinformatics. It works with AnnData. The repository describes itself as: Conversational & memory-enabled AI research partner for multi-omics analysis. CLI + Desktop App (installers in Releases). From biological idea to full research paper. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/sc-ambient-removal”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Load filtered AnnData; optionally load raw matrix (SoupX requires both).
  2. Validate --contamination is in [0, 1) and --expected-cells is positive when set.
  3. Run the chosen --method against METHOD_REGISTRY.
  4. If the requested backend is unavailable, fall back deterministically to simple.
  5. Stash the pre-correction matrix in layers["counts"] and overwrite adata.X with the corrected counts; record the run params in…
  6. Render diagnostic figures + emit report.md + result.json.

What it can do on your machine

Read from SKILL.md and the folder at commit 90a3bec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sc Ambient Removal loads about 2.2k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 900 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TianGzlab/OmicsClaw at commit 90a3bec, republished under its Apache-2.0 licence (© TianGzlab). 900 words, ~2,238 tokens.

Download SKILL.mdSave it as .claude/skills/sc-ambient-removal/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
sc-ambient-removal
description
Load when removing ambient RNA contamination from droplet-based scRNA-seq using a simple subtraction path, CellBender, or SoupX. Skip when the contamination is multiplet barcodes (use sc-doublet-detection); before counts exist (use sc-count).
trigger
ambient RNA, ambient removal, cellbender, contamination, background RNA
tags
singlecell, scrna, ambient, cellbender, soupx, contamination

sc-ambient-removal

When to use

The user has filtered (or raw + filtered) droplet-based scRNA-seq counts and suspects ambient RNA from cell-free droplets is inflating per-cell expression — typical for 10X data with high droplet density. Three backends share the CLI: simple (a deterministic ambient-profile subtraction, default), cellbender (Python, requires GPU for sensible runtime), and soupx (Rscript; needs raw + filtered matrices). Doublets are a different problem — use sc-doublet-detection for multiplet barcodes.

Use from a step

python
ambient = load_skill("sc-ambient-removal")
adata = ambient.remove_ambient(read_input("counts.h5ad"), contamination=0.05)
write_output(ambient.correction_summary(adata), "tables/correction_summary.csv")
write_output(ambient.correction_figure(adata), "figures/counts_comparison.png")
write_output(adata, "intermediate/adata_corrected.h5ad")

With raw droplets, use remove_ambient_soupx(filtered, raw=raw_droplets). CellBender remains CLI-only because it uses an external process and files. The simple-method example is in examples/example_step.py.

API

<!-- api:begin generated from _api.py; regenerate with run.py api <skill dir> --write -->
remove_ambient(adata, *, contamination: float=0.05)

Subtract a mean ambient profile from count-like expression in place.

Select layers['counts'], aligned raw or X; retain that matrix in layers['counts'] and replace X with nonnegative corrected values. The mean cell profile is only an approximation when empty droplets are absent.

:param adata: AnnData with a count-like expression matrix. :param contamination: Fraction to subtract before clipping at zero, in [0, 1). Default 0.05, matching the CLI. :returns: The same AnnData; run_info reports the matrix source and reduction. :raises ValueError: No count-like matrix exists or contamination is invalid.

remove_ambient_soupx(adata, *, raw)

Return SoupX-corrected cells, using raw droplet counts to estimate ambient RNA.

Both objects select counts from layers['counts'], aligned raw or X and exchange temporary 10x matrices with R. This backend does not expose a seed through its wrapper; results vary between runs. Errors propagate.

:param adata: AnnData containing the filtered cells. :param raw: AnnData containing raw droplets, with matching feature names. :returns: A new AnnData aligned to SoupX's retained cells and genes, with original counts in layers['counts'] and corrected X. :raises RuntimeError: R, Seurat or SoupX is unavailable, or the method fails. :raises ValueError: Counts are missing or the result cannot be aligned.

run_info(adata, *, keep: bool=True) -> dict

Read JSON correction diagnostics from adata.uns.

:param adata: AnnData returned by a correction function. :param keep: False removes the diagnostics after reading; default True. :returns: Matrix source, method and before/after count summaries.

correction_summary(adata) -> pd.DataFrame

Return the recorded count reduction as a one-row table.

:param adata: AnnData returned by a correction function. :returns: Mean counts, reduction_pct, contamination_estimate and method.

counts_comparison_table(adata) -> pd.DataFrame

Compare per-cell original counts in layers['counts'] with corrected X.

:param adata: AnnData returned by a correction function. :returns: cell_id, counts_before and counts_after in observation order.

ambient_profile_table(adata) -> pd.DataFrame

Return the mean original-count profile used by simple subtraction.

:param adata: Corrected AnnData with original counts in layers['counts']. :returns: gene and fraction columns. This is not SoupX's estimated profile.

correction_figure(adata)

Plot original versus corrected counts per cell and return the Figure.

:param adata: AnnData returned by a correction function. :returns: A matplotlib Figure for write_output.

<!-- api:end -->

Methods and parameters

remove_ambient selects counts from layers["counts"], aligned raw, then X, saves them in layers["counts"] and replaces X in place. The contamination=0.05 default retains the CLI's fixed subtraction fraction; it is not an estimated biological contamination rate. The profile is the mean across input cells, an approximation when no empty droplets exist. This path is deterministic and needs no optional backend.

remove_ambient_soupx requires R, Seurat and SoupX. It exchanges compressed 10x matrices in a temporary directory and returns a new, aligned AnnData. Its wrapper exposes no seed, so results may vary. Unlike the compatibility CLI, the API propagates backend errors; it does not fall back to simple subtraction. run_info records the method, count source and reduction.

Show full SKILL.md (346 more words)Show less

Inputs & Outputs

Inputs

  • Input kinds: file, directory
  • Modalities: scrna
  • File types: .h5ad, .h5, .loom, .csv, .tsv

Outputs

  • figures/barcode_rank.png
  • figures/count_distribution.png
  • figures/counts_comparison.png
  • figure_data/ and its manifest, including correction summary and gene expression
  • figures/r_enhanced/ only for successful --r-enhanced renders
  • cellbender_output/ only when CellBender runs; files depend on its backend output
  • processed.h5ad
  • report.md
  • result.json
  • Processed AnnData (saves_h5ad) — adds layers: counts; uns: ambient_correction, soupx, cellbender

Flow

  1. Load filtered AnnData; optionally load raw matrix (SoupX requires both).
  2. Validate --contamination is in [0, 1) and --expected-cells is positive when set.
  3. Run the chosen --method against METHOD_REGISTRY.
  4. If the requested backend is unavailable, fall back deterministically to simple.
  5. Stash the pre-correction matrix in layers["counts"] and overwrite adata.X with the corrected counts; record the run params in uns["ambient_correction"|"soupx"|"cellbender"].
  6. Render diagnostic figures + emit report.md + result.json.

Gotchas

  • The CLI can fall back to simple when a backend or its inputs are missing. Check result.json["summary"]["executed_method"] and fallback_reason. The SoupX API instead raises on failure.
  • --contamination is bounded to [0, 1) (left-inclusive). sc_ambient.py checks 0 <= float(args.contamination) < 1 and raises ValueError("--contamination must be between 0 and 1 (for example 0.05).") otherwise. 0 is allowed (degenerate no-op); 1 and 5.0 (the common typo for 0.05) both fail loudly.
  • --expected-cells must be a positive integer. sc_ambient.py raises ValueError. Zero or negative values fail loudly here rather than producing a degenerate run.
  • SoupX without both --raw-matrix-dir and --filtered-matrix-dir silently falls back to simple. sc_ambient.py logs "SoupX requires --raw-matrix-dir and --filtered-matrix-dir. Falling back to simple subtraction." and continues with the simple path. result.json records the fallback in summary["fallback_reason"]; CellBender uses just the filtered matrix and the simple path uses neither.

Key CLI

bash
# Demo (simple subtraction)
python skills/singlecell/scrna/sc-ambient-removal/sc_ambient.py --demo --output /tmp/sc_ambient_demo

# CellBender on a 10X-filtered AnnData
python skills/singlecell/scrna/sc-ambient-removal/sc_ambient.py \
  --input filtered.h5ad --output results/ \
  --method cellbender --raw-h5 raw_feature_bc_matrix.h5 --expected-cells 8000

# SoupX with explicit raw + filtered matrices
python skills/singlecell/scrna/sc-ambient-removal/sc_ambient.py \
  --input filtered.h5ad --output results/ \
  --method soupx \
  --raw-matrix-dir cellranger_out/raw_feature_bc_matrix \
  --filtered-matrix-dir cellranger_out/filtered_feature_bc_matrix

See also

  • references/parameters.md — every CLI flag and per-method tuning hint
  • references/methodology.md — when each backend wins, ambient profile derivation, R/Python tradeoffs
  • references/output_contract.md — layers["counts"] (pre-correction) and .X (corrected) semantics, per-method uns diagnostics
  • Adjacent skills: sc-doublet-detection (parallel — multiplet barcodes, complementary contamination class), sc-filter (upstream — cell QC), sc-preprocessing (downstream — normalise/HVG/PCA on the cleaned counts)

Dependencies

Python packages this skill's script needs. They are not installed for you — check before a long run.

anndata, cellbender, matplotlib, numpy, pandas, scanpy, scipy, torch

© TianGzlab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/singlecell/scrna/sc-ambient-removal of TianGzlab/OmicsClaw.

  • SKILL.md
  • _api.py
  • examples/example_step.py
  • references/methodology.md
  • references/output_contract.md
  • references/parameters.md
  • references/r_visualization.md
  • sc_ambient.py
  • tests/test_ambient_api.py

Open the folder on GitHubat commit 90a3bec

Compare with similar skills

Sc Ambient Removal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sc Ambient Removal compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sc Ambient Removal this skillTianGzlab/OmicsClaw161—~2.2kAutomated safety check: PassApache-2.0
Scanpy Single-Cell Analysisdavila7/claude-code-templates33k15 repos~2.8kAutomated safety check: PassMIT
PyDESeq2 Differential Expressiondavila7/claude-code-templates33k11 repos~4kAutomated safety check: PassMIT
Single-Cell Initial AnalysisLigphiDonk/Oh-my--paper7391 repos~1.4kAutomated safety check: PassMIT
Anndataaipoch/medical-research-skills1.9k—~1.7kAutomated safety check: PassMIT
Multiomics StatisticsVectorSpaceLab/AREX-Skill331—~1kAutomated safety check: PassGPL-3.0

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    33k GitHub starsUsed in 15 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    33k GitHub starsUsed in 11 repos~4k tokens
    Research & ScienceAuto-check passed
  • Single-Cell Initial Analysis

    LigphiDonk/Oh-my--paper

    Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

    739 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check passed
  • Anndata

    aipoch/medical-research-skills

    Data structure for annotated matrices in single-cell analysis; use when reading/writing .h5ad (or zarr) and exchanging data with the scverse ecosystem.

    1.9k GitHub stars~1.7k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Multiomics Statistics

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for OmicVerse bulk RNA-seq, enrichment/signature scoring, metabolomics, proteomics, microbiome, and statistical table workflows.

    331 GitHub stars~1k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    228 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed

More from TianGzlab/OmicsClaw

All 88 skills in this repo
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna Batch Correction

    TianGzlab/OmicsClaw

    Load when correcting batch effects in bulk expression using R sva ComBat or the legacy Python parametric approximation.

    161 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna Coexpression

    TianGzlab/OmicsClaw

    Load when discovering bulk gene co-expression modules and hub genes with R WGCNA.

    161 GitHub stars~1.3k tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna De

    TianGzlab/OmicsClaw

    Load when comparing gene expression between two conditions in bulk RNA-seq count data.

    161 GitHub stars~867 tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna Deconvolution

    TianGzlab/OmicsClaw

    Load when estimating cell-type proportions in bulk RNA-seq samples from a single-cell or signature-matrix reference.

    161 GitHub stars~757 tokensUpdated 3 days ago
    Auto-check passed
  • Bulkrna Enrichment

    TianGzlab/OmicsClaw

    Load when running pathway / GO term enrichment on a bulk RNA-seq DE result list.

    161 GitHub stars~860 tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Sc Ambient Removal

What does Sc Ambient Removal do?

Load when removing ambient RNA contamination from droplet-based scRNA-seq using a simple subtraction path, CellBender, or SoupX. Sc Ambient Removal is an agent skill from TianGzlab/OmicsClaw. Load when removing ambient RNA contamination from droplet-based scRNA-seq using a simple subtraction path, CellBender, or SoupX.

When should I use Sc Ambient Removal?

Sc Ambient Removal fits situations like: tasks that involve Bioinformatics.

How do I install Sc Ambient Removal in Claude Code?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-ambient-removal -a claude-code`. Or copy the skill folder (skills/singlecell/scrna/sc-ambient-removal in TianGzlab/OmicsClaw) into .claude/skills/sc-ambient-removal in your project. Claude Code loads it when a task matches its description.

How do I install Sc Ambient Removal in Codex?

Run `npx skills add TianGzlab/OmicsClaw --skill sc-ambient-removal -a codex`. Or copy the skill folder (skills/singlecell/scrna/sc-ambient-removal in TianGzlab/OmicsClaw) into .agents/skills/sc-ambient-removal in your project. Codex loads it when a task matches its description.

Can I use Sc Ambient Removal in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TianGzlab/OmicsClaw --skill sc-ambient-removal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sc-ambient-removal, .gemini/skills/sc-ambient-removal, .github/skills/sc-ambient-removal and .opencode/skills/sc-ambient-removal in your project.

What does Sc Ambient Removal need to run?

Going by SKILL.md and its folder, Sc Ambient Removal needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Sc Ambient Removal access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sc Ambient Removal safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sc Ambient Removal use?

Sc Ambient Removal is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sc Ambient Removal use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Sc Ambient Removal?

Skills that share tags, products or a category with Sc Ambient Removal: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 33k stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 33k stars), Single-Cell Initial Analysis (LigphiDonk/Oh-my--paper, 739 stars) and Anndata (aipoch/medical-research-skills, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sc Ambient Removal?

TianGzlab (a GitHub organization) maintains it in TianGzlab/OmicsClaw, which has 161 GitHub stars. The repository holds 88 skills in this directory. The repository was last updated on October 7, 2026.

Source: TianGzlab/OmicsClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.