Agent skill

Pydeseq

by aipoch in aipoch/medical-research-skills

Differential gene expression analysis for bulk RNA-seq count matrices using a DESeq2-like workflow in Python; use when you need Wald tests, FDR correction, and optional LFC shrinkage for…

MITAuto-check passedData & Analytics

Install Pydeseq

skills CLI
$ npx skills add aipoch/medical-research-skills --skill pydeseq -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills pydeseq --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Data Analysis/pydeseq2' .claude/skills/pydeseq && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pydeseq
GitHub stars
1.9k
Token cost
~1.8k tokens
SKILL.md length
434 words
Files
5 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Differential gene expression analysis for bulk RNA-seq count matrices using a DESeq2-like workflow in Python; use when you need Wald tests, FDR correction, and optional LFC shrinkage for…

  • Works in 5 steps: Case vs control bulk RNA-seq from a raw… → Multi-factor designs to adjust for batch… → DESeq2 migration when converting an R… → …
  • You need Wald tests
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 1 more section
  • Runs Python scripts from its folder; calls uv

What it does

Pydeseq is an agent skill from aipoch/medical-research-skills. Differential gene expression analysis for bulk RNA-seq count matrices using a DESeq2-like workflow in Python; use when you need Wald tests, FDR correction, and optional LFC shrinkage for condition/batch/covariate designs.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `pydeseq2_audit_result_v1.json`, `references/api_reference.md` and `references/workflow_guide.md`).

It sits in Data & Analytics, covering Bioinformatics. It works with Python, AnnData and pandas. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • You need Wald tests
  • Optional LFC shrinkage for condition/batch/covariate designs

Example prompts

  • “/pydeseq”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Case vs control bulk RNA-seq from a raw integer count matrix (e.g., treated vs control).
  2. Multi-factor designs to adjust for batch effects or covariates (e.g., ~ batch + condition, ~ age + condition).
  3. DESeq2 migration when converting an R DESeq2 workflow into a Python pipeline.
  4. Pipeline integration where results must stay in Python objects (pandas/AnnData) for downstream QC, plots, or reporting.
  5. Requests mentioning “DESeq2”, “differential expression”, “Wald test”, “FDR/padj”, “volcano plot”, “MA plot”, or “PyDESeq2”.

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pydeseq loads about 1.8k tokens when it runs, and up to ~7.2k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 434 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 434 words, ~1,841 tokens.

Download SKILL.mdSave it as .claude/skills/pydeseq/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
pydeseq
description
Differential gene expression analysis for bulk RNA-seq count matrices using a DESeq2-like workflow in Python; use when you need Wald tests, FDR correction, and optional LFC shrinkage for condition/batch/covariate designs.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

Use this skill when you need to run DESeq2-style differential expression in Python, especially in these scenarios:

  1. Case vs control bulk RNA-seq from a raw integer count matrix (e.g., treated vs control).
  2. Multi-factor designs to adjust for batch effects or covariates (e.g., ~ batch + condition, ~ age + condition).
  3. DESeq2 migration when converting an R DESeq2 workflow into a Python pipeline.
  4. Pipeline integration where results must stay in Python objects (pandas/AnnData) for downstream QC, plots, or reporting.
  5. Requests mentioning “DESeq2”, “differential expression”, “Wald test”, “FDR/padj”, “volcano plot”, “MA plot”, or “PyDESeq2”.

Key Features

  • End-to-end DESeq2-like workflow: normalization (size factors), dispersion estimation/shrinkage, LFC fitting, outlier handling.
  • Wald tests for differential expression with Benjamini–Hochberg FDR (padj).
  • Design formulas in Wilkinson/R-style notation (single-factor and multi-factor).
  • Contrast-based comparisons: [variable, test_group, reference_group].
  • Optional Cook’s distance outlier filtering and refitting.
  • Optional LFC shrinkage (apeGLM-style) for visualization/ranking.
  • Works naturally with pandas and can interoperate with AnnData.

Dependencies

Minimum environment (as documented in the source material):

  • Python 3.10–3.11
  • pydeseq2 (install via pip/uv)
  • pandas >= 1.4.3
  • numpy >= 1.23.0
  • scipy >= 1.11.0
  • scikit-learn >= 1.1.1
  • anndata >= 0.8.0 (optional, for AnnData I/O)

Optional plotting:

  • matplotlib (recommended)
  • seaborn (optional)

Installation:

bash
uv pip install pydeseq2

Example Usage

The following script is a complete, runnable example for a standard treated-vs-control analysis.

python
import pandas as pd
import numpy as np

from pydeseq2.dds import DeseqDataSet
from pydeseq2.ds import DeseqStats

# -----------------------------
# 1) Load inputs
# -----------------------------
# counts.csv is commonly stored as genes x samples; transpose to samples x genes.
counts_df = pd.read_csv("counts.csv", index_col=0).T
metadata = pd.read_csv("metadata.csv", index_col=0)

# Ensure sample alignment
common = counts_df.index.intersection(metadata.index)
counts_df = counts_df.loc[common]
metadata = metadata.loc[common]

# -----------------------------
# 2) Basic filtering
# -----------------------------
# Remove genes with very low total counts
min_total_counts = 10
genes_to_keep = counts_df.columns[counts_df.sum(axis=0) >= min_total_counts]
counts_df = counts_df[genes_to_keep]

# Drop samples with missing condition
metadata = metadata.dropna(subset=["condition"])
counts_df = counts_df.loc[metadata.index]

# -----------------------------
# 3) Fit DESeq2 model
# -----------------------------
dds = DeseqDataSet(
    counts=counts_df,
    metadata=metadata,
    design="~ condition",
    refit_cooks=True,
    n_cpus=1,
)
dds.deseq2()

# -----------------------------
# 4) Wald test with contrast
# -----------------------------
ds = DeseqStats(
    dds,
    contrast=["condition", "treated", "control"],
    alpha=0.05,
    cooks_filter=True,
    independent_filter=True,
)
ds.summary()

# -----------------------------
# 5) Results + optional shrinkage
# -----------------------------
res = ds.results_df.copy()
sig = res[res["padj"] < 0.05].sort_values("padj")
print(f"Significant genes (padj < 0.05): {len(sig)}")

# Optional: shrink LFC for visualization/ranking (p-values do not change)
ds.lfc_shrink()
res_shrunk = ds.results_df.copy()

# Export
res.to_csv("deseq2_results.csv")
res_shrunk.to_csv("deseq2_results_shrunk_lfc.csv")
sig.to_csv("significant_genes.csv")

# -----------------------------
# 6) Minimal volcano plot (optional)
# -----------------------------
try:
    import matplotlib.pyplot as plt

    plot_df = res.copy()
    plot_df["neglog10_padj"] = -np.log10(plot_df["padj"].clip(lower=1e-300))
    is_sig = plot_df["padj"] < 0.05

    plt.figure(figsize=(9, 5))
    plt.scatter(
        plot_df.loc[~is_sig, "log2FoldChange"],
        plot_df.loc[~is_sig, "neglog10_padj"],
        s=10,
        alpha=0.3,
        c="gray",
        label="Not significant",
    )
    plt.scatter(
        plot_df.loc[is_sig, "log2FoldChange"],
        plot_df.loc[is_sig, "neglog10_padj"],
        s=10,
        alpha=0.6,
        c="red",
        label="padj < 0.05",
    )
    plt.axhline(-np.log10(0.05), linestyle="--", color="blue", alpha=0.5)
    plt.xlabel("Log2 Fold Change")
    plt.ylabel("-Log10(adjusted p-value)")
    plt.title("Volcano Plot")
    plt.legend()
    plt.tight_layout()
    plt.savefig("volcano_plot.png", dpi=300)
except ImportError:
    pass

Implementation Details

Inputs and orientation
  • Counts matrix must be samples × genes with non-negative integer counts.
  • Many files are stored as genes × samples; transpose with .T after loading.
Design formula (Wilkinson/R-style)
  • Use strings like:
    • ~ condition (single factor)
    • ~ batch + condition (batch-adjusted)
    • ~ age + condition (continuous covariate)
    • ~ group + condition + group:condition (interaction)
  • Put adjustment variables first (e.g., ~ batch + condition) so the primary effect is interpreted cleanly.
Show full SKILL.md (166 more words)Show less
What dds.deseq2() does (high level)

The fitting pipeline typically includes:

  1. Size factor estimation (library-size normalization)
  2. Gene-wise dispersion estimation
  3. Dispersion trend fitting and prior estimation
  4. MAP dispersion shrinkage
  5. Log2 fold change fitting under the specified design
  6. Cook’s distance outlier detection
  7. Optional refitting after outlier handling (refit_cooks=True)
Statistical testing and multiple testing correction
  • DeseqStats(...).summary() runs Wald tests for the requested coefficient/contrast.
  • Output columns commonly include:
    • baseMean: mean normalized expression
    • log2FoldChange, lfcSE, stat
    • pvalue: raw p-value
    • padj: Benjamini–Hochberg FDR adjusted p-value
  • Use padj < alpha (commonly 0.05) for significance.
Contrast specification
  • Format: contrast=["variable", "test_group", "reference_group"]
  • Example: ["condition", "treated", "control"] tests treated relative to control.
LFC shrinkage (optional)
  • ds.lfc_shrink() applies shrinkage to log2FoldChange for more stable ranking/plots.
  • Shrinkage is intended for visualization and prioritization; statistical significance is still based on the (unshrunken) Wald test p-values.
Notes on bundled references/scripts

If your repository includes them, use:

  • references/api_reference.md for parameter/object details.
  • references/workflow_guide.md for extended workflows and troubleshooting.
  • scripts/run_deseq2_analysis.py for a CLI-style batch workflow (counts/metadata/design/contrast/output, optional plots).

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in scientific-skills/Data Analysis/pydeseq2 of aipoch/medical-research-skills.

  • SKILL.md
  • pydeseq2_audit_result_v1.json
  • references/api_reference.md
  • references/workflow_guide.md
  • scripts/run_deseq2_analysis.py

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Pydeseq next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pydeseq compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pydeseq this skillaipoch/medical-research-skills1.9k—~1.8kAutomated safety check: PassMIT
PyDESeq2 Differential Expressiondavila7/claude-code-templates33k11 repos~4kAutomated safety check: PassMIT
Alphagenome Predictionsgenomicsxai/alphagenome-pytorch162—~868Automated safety check: PassApache-2.0
Tooluniverse Epigenomicswu-yc/LabClaw1.1k2 repos~14kAutomated safety check: PassNone
Bio Genome Intervals Bed File BasicsGPTomics/bioSkills1.2k1 repos~4.5kAutomated safety check: PassMIT
Spatial VelocityTianGzlab/OmicsClaw161—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    33k GitHub starsUsed in 11 repos~4k tokens
    Research & ScienceAuto-check passed
  • Alphagenome Predictions

    genomicsxai/alphagenome-pytorch

    Run AlphaGenome-PyTorch to get genomic track predictions — via the agt predict CLI (single locus, BED regions, whole chromosomes, raw FASTA sequences, or per-gene count tables/AnnData), variant…

    162 GitHub stars~868 tokensUpdated 25 days ago
    AI & LLM EngineeringAuto-check passed
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • Handles BED-format genomic intervals (BED3 through BED12, narrowPeak/broadPeak) and the coordinate-system substrate the whole interval category rests on, with bedtools (CLI) and…

    1.2k GitHub starsUsed in 1 repo~4.5k tokens
    Research & ScienceAuto-check passed
  • Spatial Velocity

    TianGzlab/OmicsClaw

    Load when estimating RNA velocity on a spatial AnnData with layers["spliced"] + layers["unspliced"] via scVelo (stochastic / deterministic / dynamical) or veloVI (deep generative).

    161 GitHub stars~1.3k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    390 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    1.9k GitHub stars~2.2k tokensUpdated 23 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    1.9k GitHub stars~1.4k tokensUpdated 23 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    1.9k GitHub stars~3.7k tokensUpdated 23 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    1.9k GitHub stars~1.8k tokensUpdated 23 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    1.9k GitHub stars~1.7k tokensUpdated 23 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    1.9k GitHub stars~1.3k tokensUpdated 23 days ago
    Auto-check passed

Questions about Pydeseq

What does Pydeseq do?

Differential gene expression analysis for bulk RNA-seq count matrices using a DESeq2-like workflow in Python; use when you need Wald tests, FDR correction, and optional LFC shrinkage for…. Pydeseq is an agent skill from aipoch/medical-research-skills. Differential gene expression analysis for bulk RNA-seq count matrices using a DESeq2-like workflow in Python; use when you need Wald tests, FDR correction, and optional LFC shrinkage for condition/batch/covariate designs.

When should I use Pydeseq?

Pydeseq fits situations like: you need Wald tests; optional LFC shrinkage for condition/batch/covariate designs.

How do I install Pydeseq in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill pydeseq -a claude-code`. Or copy the skill folder (scientific-skills/Data Analysis/pydeseq2 in aipoch/medical-research-skills) into .claude/skills/pydeseq in your project. Claude Code loads it when a task matches its description.

How do I install Pydeseq in Codex?

Run `npx skills add aipoch/medical-research-skills --skill pydeseq -a codex`. Or copy the skill folder (scientific-skills/Data Analysis/pydeseq2 in aipoch/medical-research-skills) into .agents/skills/pydeseq in your project. Codex loads it when a task matches its description.

Can I use Pydeseq in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill pydeseq -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pydeseq, .gemini/skills/pydeseq, .github/skills/pydeseq and .opencode/skills/pydeseq in your project.

What does Pydeseq need to run?

Going by SKILL.md and its folder, Pydeseq needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Pydeseq access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Pydeseq safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pydeseq use?

Pydeseq is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pydeseq use?

About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Pydeseq?

Skills that share tags, products or a category with Pydeseq: PyDESeq2 Differential Expression (davila7/claude-code-templates, 33k stars), Alphagenome Predictions (genomicsxai/alphagenome-pytorch, 162 stars), Tooluniverse Epigenomics (wu-yc/LabClaw, 1.1k stars) and Bio Genome Intervals Bed File Basics (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pydeseq?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,937 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.