Agent skill

Bio Rna Quantification Count Matrix Qc

by GPTomics in GPTomics/bioSkills

Quality control and exploration of RNA-seq count matrices before differential expression.

MITAuto-check passedResearch & Science

Install Bio Rna Quantification Count Matrix Qc

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-rna-quantification-count-matrix-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-rna-quantification-count-matrix-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/rna-quantification/count-matrix-qc .claude/skills/bio-rna-quantification-count-matrix-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-rna-quantification-count-matrix-qc
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.6k tokens
SKILL.md length
1,082 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Quality control and exploration of RNA-seq count matrices before differential expression.

  • Works in 4 steps: A sample clusters away from its group on… → A size factor far from 1, or a library… → Near-zero correlation of a sample to its… → …
  • Checking library sizes and composition
  • SKILL.md covers Version Compatibility, Load and Inspect, Filtering: the principled cut and Normalization and transformation, plus 8 more sections
  • Runs Python and R scripts from its folder; calls pip

What it does

Bio Rna Quantification Count Matrix Qc is an agent skill from GPTomics/bioSkills. Quality control and exploration of RNA-seq count matrices before differential expression. Use when checking library sizes and composition, choosing VST vs rlog for visualization, running PCA and sample correlation, detecting outliers with Cook's distance, deciding how to handle known vs unknown batch effects, screening for sample swaps, or judging whether a sample or design is too compromised to test.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/qc_analysis.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Checking library sizes and composition
  • Choosing VST vs rlog for visualization
  • Running PCA and sample correlation
  • Detecting outliers with Cooks distance

Example prompts

  • “/bio-rna-quantification-count-matrix-qc”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. A sample clusters away from its group on the VST-PCA (and concentrates Cook's-flagged genes).
  2. A size factor far from 1, or a library an order of magnitude off the cohort.
  3. Near-zero correlation of a sample to its replicates.
  4. Condition (near-)perfectly confounded with batch, lane, or run.

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Rna Quantification Count Matrix Qc loads about 2.6k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 1,082 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,082 words, ~2,639 tokens.

Download SKILL.mdSave it as .claude/skills/bio-rna-quantification-count-matrix-qc/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-rna-quantification-count-matrix-qc
description
Quality control and exploration of RNA-seq count matrices before differential expression. Use when checking library sizes and composition, choosing VST vs rlog for visualization, running PCA and sample correlation, detecting outliers with Cook's distance, deciding how to handle known vs unknown batch effects, screening for sample swaps, or judging whether a sample or design is too compromised to test.
tool_type
mixed
primary_tool
DESeq2

Version Compatibility

Reference examples tested with: DESeq2 1.42+, edgeR 4.0+, ggplot2 3.5+, pheatmap 1.0+, matplotlib 3.8+, numpy 1.26+, pandas 2.2+, scikit-learn 1.4+, scipy 1.12+, seaborn 0.13+

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Count Matrix QC

"Check my count matrix for outliers and batch effects" -> Assess depth, composition, sample relationships, and outliers on appropriately transformed data, then decide what (if anything) to remove or model before differential expression.

  • R: DESeq2::vst() -> plotPCA(), sample-distance heatmap, Cook's distance
  • Python: sklearn.decomposition.PCA, seaborn.clustermap (with the low-count caveat below)

Two principles govern this whole skill. First, DE testing runs on raw counts with a size-factor offset; the transformed matrices here are for QC and visualization only, never fed back into the count model. Second, raw counts confound depth, composition, and biology, so QC must look at the right scale: a variance-stabilized matrix for clustering/PCA, and the size factors and Cook's distances from the count model for normalization and outliers.

Load and Inspect

Goal: Get counts into a model object and read off depth and detection per sample.

Approach: Build a DESeqDataSet (from tximport or a matrix), then summarize library size and genes detected.

r
library(DESeq2)
counts <- read.csv('count_matrix.csv', row.names = 1)
coldata <- data.frame(condition = factor(c('ctrl', 'ctrl', 'treat', 'treat')),
                      row.names = colnames(counts))
dds <- DESeqDataSetFromMatrix(countData = counts, colData = coldata, design = ~ condition)

colSums(counts(dds))        # library size per sample
colSums(counts(dds) > 0)    # genes detected per sample
python
import pandas as pd, numpy as np
counts = pd.read_csv('count_matrix.csv', index_col=0)
metadata = pd.read_csv('sample_info.csv', index_col=0)
print(counts.sum()); print((counts > 0).sum())

Filtering: the principled cut

Goal: Drop genes with too little signal to test, in a depth- and design-aware way.

Approach: Prefer edgeR filterByExpr (keeps genes with enough counts in at least the smallest group's worth of samples) over an arbitrary CPM > 1 rule.

r
library(edgeR)
keep <- filterByExpr(counts(dds), group = dds$condition)
dds <- dds[keep, ]
python
min_counts, min_samples = 10, 3   # 10 reads in >=3 samples; ~smallest group size
counts_filt = counts[(counts >= min_counts).sum(axis=1) >= min_samples]

In DESeq2, pre-filtering is mainly for speed and to drop all-zero rows; the inferential filter is independent filtering done automatically inside results() (it picks a mean-count threshold maximizing discoveries at the chosen alpha). Keep pre-filtering light. For edgeR/limma-voom, filterByExpr is the filter.

Normalization and transformation

Composition bias is the reason depth scaling is not enough: if a few genes dominate a library, every other gene looks depressed at unchanged absolute output. DESeq2 median-of-ratios and edgeR TMM each estimate one size factor per sample assuming most genes are not DE, then apply it as an offset on the raw counts. CPM and TPM do NOT correct composition (they rescale by a within-sample total) -- the same reason TPM is invalid for cross-sample comparison upstream -- so they are for visualization, not DE normalization. For matrices with many structural zeros (single-cell, metagenomics), use the poscounts size-factor estimator.

For QC visualization the matrix must be homoskedastic. log2(CPM + 1) is not: at low counts the log amplifies sampling noise, so PCA on it is driven by noisy near-zero genes. Use a variance-stabilizing transform instead.

TransformSpeedUse when
vst()FastDefault, especially medium-to-large n (>30)
rlog()SlowSmall n (roughly < 30) and heterogeneous designs; but can over-shrink when size factors span a very wide range (then prefer vst)
r
vsd <- vst(dds, blind = TRUE)    # blind=TRUE for unsupervised QC; FALSE only after DESeq() for plotting
mat <- assay(vsd)

PCA and sample relationships

Goal: See whether replicates cluster and whether PC1 is biology or a technical artifact.

Approach: PCA on the VST matrix (top variable genes), then read PC1 against depth and batch.

r
plotPCA(vsd, intgroup = 'condition')                     # uses top 500 most-variable genes
sampleDists <- dist(t(assay(vsd)))
pheatmap::pheatmap(as.matrix(sampleDists))
python
from sklearn.decomposition import PCA
# log-CPM PCA is a quick look only: low-count heteroskedasticity can drive the PCs.
# For publication QC, compute VST in R and bring the matrix into Python.
cpm = counts_filt * 1e6 / counts_filt.sum()
log_cpm = np.log2(cpm + 1)
pcs = PCA(n_components=2).fit_transform(log_cpm.T)

If PC1 correlates with library size or detected-gene count rather than condition, it is a depth artifact (color the PCA by log10 library size to confirm). A common pattern is PC1 = batch, PC2 = condition, which is a design problem, not a normalization fix.

Outlier detection with Cook's distance

Goal: Distinguish a single bad count in one gene from a globally bad sample.

Approach: Read per-gene-per-sample Cook's distances from the fitted model; treat single-gene outliers and whole-sample outliers differently.

r
dds <- DESeq(dds)
cooks <- assays(dds)[['cooks']]          # per gene x sample; NOT results(dds)$cooksd
boxplot(log10(cooks), las = 2, main = "Cook's distance")
# results() flags a gene whose max Cook's exceeds qf(0.99, p, m-p) by setting its p-value to NA.
# With >= 7 replicates per group (minReplicatesForReplace) DESeq2 replaces the outlier count instead.

A single-gene-in-one-sample outlier is exactly what Cook's filtering and replaceOutliers are for; let DESeq2 handle it. A whole-sample outlier (many flagged genes in one sample, that sample far on the VST-PCA, low correlation to its replicates, an anomalous size factor) is not rescuable by replaceOutliers. Investigate, and remove only with a documented technical cause, since post-hoc cherry-picking inflates false positives.

Show full SKILL.md (425 more words)Show less

Batch effects

Known batch goes in the design; the engine estimates and removes it on raw counts while propagating uncertainty:

r
design(dds) <- ~ batch + condition       # condition last = contrast of interest

Do NOT run removeBatchEffect() or ComBat and feed the adjusted matrix into DESeq2/edgeR; those engines model batch internally, and pre-adjusting double-corrects and breaks the count model. limma::removeBatchEffect(assay(vsd), batch = vsd$batch) is for visualization only. For unknown/unmeasured structure, estimate surrogate variables (sva/svaseq) or factors of unwanted variation (RUVSeq: RUVg control genes, RUVs replicate samples, RUVr residuals) and add them to the design.

The fatal case: if batch is correlated with condition, regressing it out removes biology too; a perfect confound (all treated in batch 1, all control in batch 2) is statistically unfixable. Cross-tabulate batch against condition before fitting.

Library-level QC and sample swaps

r
sf <- sizeFactors(estimateSizeFactors(dds))   # a size factor far from 1 (< ~0.3 or > ~3) is a red flag

Do not deduplicate standard RNA-seq: high duplication is expected from highly expressed genes, and position-based dedup discards real signal (deduplicate only with UMIs). Screen for sample swaps cheaply with sex-linked genes (XIST high in XX; RPS4Y1/UTY/DDX3Y high in XY) against recorded sex, and confirm identity with genotype concordance tools (VerifyBamID, somalier) when available.

Red flags that should halt a DE analysis

  1. A sample clusters away from its group on the VST-PCA (and concentrates Cook's-flagged genes).
  2. A size factor far from 1, or a library an order of magnitude off the cohort.
  3. Near-zero correlation of a sample to its replicates.
  4. Condition (near-)perfectly confounded with batch, lane, or run.

Common Errors

SymptomCauseFix
results(dds)$cooksd is NULLCook's distance is not a results columnRead assays(dds)[['cooks']]
PCA driven by a few noisy genesPCA run on log2(CPM+1) or raw countsUse VST/rlog; restrict to top-variable genes
Batch effect persists after correctionremoveBatchEffect output fed to DESeq2Put batch in the design instead; keep correction for plots only
Every gene significant, or noneSample swap / confounded batch / wrong normalizationCheck metadata, batch x condition table, and size factors first
One transform behaves oddly with wide size factorsrlog over-shrinksSwitch to vst
  • rna-quantification/featurecounts-counting - Generate the count matrix
  • rna-quantification/tximport-workflow - Import transcript counts with the length offset
  • differential-expression/deseq2-basics - DE testing after QC
  • differential-expression/de-visualization - Downstream result visualization
  • read-qc/rnaseq-qc - Upstream read-level QC (rRNA, degradation, contamination)

References

  • Love MI, Huber W, Anders S. 2014. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15(12):550. doi:10.1186/s13059-014-0550-8
  • Robinson MD, Oshlack A. 2010. A scaling normalization method for differential expression analysis of RNA-seq data. Genome Biol 11(3):R25. doi:10.1186/gb-2010-11-3-r25
  • Risso D, Ngai J, Speed TP, Dudoit S. 2014. Normalization of RNA-seq data using factor analysis of control genes or samples. Nat Biotechnol 32(9):896-902. doi:10.1038/nbt.2931

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in rna-quantification/count-matrix-qc of GPTomics/bioSkills.

  • SKILL.md
  • examples/qc_analysis.py
  • examples/qc_report.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Rna Quantification Count Matrix Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Rna Quantification Count Matrix Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Rna Quantification Count Matrix Qc this skillGPTomics/bioSkills1.2k1 repos~2.6kAutomated safety check: PassMIT
Scanpy Single-Cell Analysisdavila7/claude-code-templates33k15 repos~2.8kAutomated safety check: PassMIT
Bulkrna Cosinor RhythmTianGzlab/OmicsClaw161—~840Automated safety check: PassApache-2.0
deepTools NGS Toolkitdavila7/claude-code-templates33k12 repos~4.5kAutomated safety check: PassMIT
LaminDB Biological Data Managementdavila7/claude-code-templates33k12 repos~3.6kAutomated safety check: PassMIT
PyDESeq2 Differential Expressiondavila7/claude-code-templates33k11 repos~4kAutomated safety check: PassMIT

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    33k GitHub starsUsed in 15 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • deepTools NGS Toolkit

    davila7/claude-code-templates

    Guides use of deepTools on sequencing data: BAM to bigWig conversion, QC, sample correlation, and heatmaps or profiles around TSS and peaks for ChIP-seq, RNA-seq and ATAC-seq.

    33k GitHub starsUsed in 12 repos~4.5k tokens
    Research & ScienceAuto-check passed
  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    33k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    33k GitHub starsUsed in 11 repos~4k tokens
    Research & ScienceAuto-check passed
  • Gtars Genomic Interval Toolkit

    davila7/claude-code-templates

    Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.

    33k GitHub starsUsed in 11 repos~1.9k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Rna Quantification Count Matrix Qc

What does Bio Rna Quantification Count Matrix Qc do?

Quality control and exploration of RNA-seq count matrices before differential expression. Bio Rna Quantification Count Matrix Qc is an agent skill from GPTomics/bioSkills. Quality control and exploration of RNA-seq count matrices before differential expression.

When should I use Bio Rna Quantification Count Matrix Qc?

Bio Rna Quantification Count Matrix Qc fits situations like: checking library sizes and composition; choosing VST vs rlog for visualization; running PCA and sample correlation; detecting outliers with Cooks distance.

How do I install Bio Rna Quantification Count Matrix Qc in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-count-matrix-qc -a claude-code`. Or copy the skill folder (rna-quantification/count-matrix-qc in GPTomics/bioSkills) into .claude/skills/bio-rna-quantification-count-matrix-qc in your project. Claude Code loads it when a task matches its description.

How do I install Bio Rna Quantification Count Matrix Qc in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-count-matrix-qc -a codex`. Or copy the skill folder (rna-quantification/count-matrix-qc in GPTomics/bioSkills) into .agents/skills/bio-rna-quantification-count-matrix-qc in your project. Codex loads it when a task matches its description.

Can I use Bio Rna Quantification Count Matrix Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-count-matrix-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-rna-quantification-count-matrix-qc, .gemini/skills/bio-rna-quantification-count-matrix-qc, .github/skills/bio-rna-quantification-count-matrix-qc and .opencode/skills/bio-rna-quantification-count-matrix-qc in your project.

What does Bio Rna Quantification Count Matrix Qc need to run?

Going by SKILL.md and its folder, Bio Rna Quantification Count Matrix Qc needs Python and R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Rna Quantification Count Matrix Qc access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Rna Quantification Count Matrix Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Rna Quantification Count Matrix Qc use?

Bio Rna Quantification Count Matrix Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Rna Quantification Count Matrix Qc use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Rna Quantification Count Matrix Qc?

Skills that share tags, products or a category with Bio Rna Quantification Count Matrix Qc: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 33k stars), Bulkrna Cosinor Rhythm (TianGzlab/OmicsClaw, 161 stars), deepTools NGS Toolkit (davila7/claude-code-templates, 33k stars) and LaminDB Biological Data Management (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Rna Quantification Count Matrix Qc?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.