Agent skill

Bio Differential Expression De Visualization

by GPTomics in GPTomics/bioSkills

Creates DE-specific diagnostic and result visualizations using DESeq2/edgeR built-in functions and lightweight ggplot2 wrappers.

MITAuto-check passedData & Analytics

Install Bio Differential Expression De Visualization

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-differential-expression-de-visualization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-differential-expression-de-visualization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/differential-expression/de-visualization .claude/skills/bio-differential-expression-de-visualization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-differential-expression-de-visualization
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.1k tokens
SKILL.md length
2,016 words
Files
5
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Creates DE-specific diagnostic and result visualizations using DESeq2/edgeR built-in functions and lightweight ggplot2 wrappers.

  • Generating DE diagnostic plots
  • SKILL.md covers Version Compatibility, Scope, The Single Most Important… and Plot Taxonomy, plus 14 more sections
  • Runs R scripts from its folder
  • Choosing VST vs rlog for visualization

What it does

Bio Differential Expression De Visualization is an agent skill from GPTomics/bioSkills. Creates DE-specific diagnostic and result visualizations using DESeq2/edgeR built-in functions and lightweight ggplot2 wrappers. Covers MA plot (with the shrunken-LFC compression effect), volcano (with the apeglm caveat that p-values are unchanged), PCA on VST/rlog (never raw counts), sample distance heatmaps, top-DE-gene heatmaps with the row-scaling trap, dispersion / BCV plot interpretation, p-value histogram diagnostics, plotCounts for individual genes, blind=TRUE vs FALSE rationale, and the n=3 visualization…

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `usage-guide.md`).

It sits in Data & Analytics, covering Statistics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Generating DE diagnostic plots
  • Choosing VST vs rlog for visualization
  • Troubleshooting suspicious plot patterns (shifted MA cloud
  • Batch-dominated PCA

Example prompts

  • “Use the bio-differential-expression-de-visualization skill to create DE-specific diagnostic and result visualizations using DESeq2/edgeR built-in…”
  • “/bio-differential-expression-de-visualization”

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Differential Expression De Visualization loads about 5.1k tokens when it runs. Until then it costs about 203 tokens; SKILL.md has 2,016 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~203
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,016 words, ~5,075 tokens.

Download SKILL.mdSave it as .claude/skills/bio-differential-expression-de-visualization/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
bio-differential-expression-de-visualization
description
Creates DE-specific diagnostic and result visualizations using DESeq2/edgeR built-in functions and lightweight ggplot2 wrappers. Covers MA plot (with the shrunken-LFC compression effect), volcano (with the apeglm caveat that p-values are unchanged), PCA on VST/rlog (never raw counts), sample distance heatmaps, top-DE-gene heatmaps with the row-scaling trap, dispersion / BCV plot interpretation, p-value histogram diagnostics, plotCounts for individual genes, blind=TRUE vs FALSE rationale, and the n=3 visualization stake. Use when generating DE diagnostic plots, choosing VST vs rlog for visualization, troubleshooting suspicious plot patterns (shifted MA cloud, batch-dominated PCA, anti-conservative p-value histogram), or building a standard QC figure panel.
tool_type
r
primary_tool
DESeq2

Version Compatibility

Reference examples tested with: DESeq2 1.42+, edgeR 4.0+, limma 3.58+, ggplot2 3.5+, pheatmap 1.0+, RColorBrewer 1.1+, ggrepel 0.9+, EnhancedVolcano 1.20+, matrixStats 1.2+

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

DE Visualization

"Make the standard DE figure panel" -> Use built-in functions or thin wrappers to produce diagnostic plots (dispersion, p-value histogram, PCA, sample distance) and result plots (MA, volcano, heatmap of top DE genes, per-gene counts), interpreted as diagnostics of the underlying model.

Scope

This skill covers DE-specific built-in plots and immediate wrappers. For richer customization:

  • Custom volcano/MA with apeglm-shrunken LFC and ggrepel labelling -> data-visualization/volcano-and-ma-plots
  • PCA / UMAP / t-SNE customization -> data-visualization/dimensionality-reduction-plots
  • Heatmap customization and ComplexHeatmap recipes -> data-visualization/heatmaps-clustering

The Single Most Important Modern Insight -- A volcano with shrunken LFC compresses the cloud, but the p-values are unchanged

lfcShrink() pulls noisy estimates toward zero. On the volcano, that pulls genes horizontally toward the center. But the y-axis (-log10(pvalue)) is the unshrunken Wald p-value -- shrinkage does NOT recompute p-values (Zhu, Ibrahim, Love 2019 Bioinformatics 35:2084). A naive reader sees fewer extreme dots and concludes "fewer genes are significant". Wrong: the same genes are significant; the effect sizes are smaller and more honest.

Always label the volcano x-axis "shrunken log2 fold change (apeglm)" and note the y-axis comes from the unshrunken Wald test. The whole point of the apeglm volcano is the honest effect-size axis; if a publication shows an unshrunken volcano, it is showing inflated effects from low-count noise.

The MA plot has its own version of this: shrinkage flattens the left side (low-mean, formerly extreme LFC) and barely touches the right (high-mean, well-estimated LFC). That asymmetry is the visual signature of working shrinkage.

Plot Taxonomy

PlotDiagnostic OR resultBuilt-in functionWhat it tests
Dispersion plotDiagnosticplotDispEsts(dds) (DESeq2), plotBCV(y) (edgeR)Mean-dispersion trend fit quality
p-value histogramDiagnosticNone; use ggplot2Null calibration, hidden batch, over-correction
PCA on VST/rlogDiagnostic + resultplotPCA(vsd, intgroup=...) (DESeq2), plotMDS() (edgeR via limma)Sample clustering, batch effects, outliers
Sample distance heatmapDiagnosticpheatmap on dist(t(assay(vsd)))Within-group consistency, sample swaps
MA plotDiagnostic + resultplotMA(res) (DESeq2), plotMD(qlf) (edgeR)Normalization sanity, LFC vs mean
VolcanoResultggplot2 wrapper; EnhancedVolcanoTop-effect, top-significance gene story
Top-DE heatmapResultpheatmap on assay(vsd)[sig_genes,]Per-gene pattern across conditions
plotCounts per geneResultplotCounts(dds, gene, intgroup)Per-gene biology

Decision Tree by Scenario

ScenarioRecommended approach
PCA for unbiased QCvst(dds, blind = TRUE); ask "do samples group as expected without design influence?"
PCA for results figurevst(dds, blind = FALSE); design is settled, accept its influence on dispersion
n < 30, library sizes vary >4xrlog(dds, blind = FALSE) instead of vst
n > 30vst(); rlog impractical
VolcanoPlot shrunken LFC on x, unshrunken p-value on y; label both axes
Sample distance heatmapvst(blind = TRUE); tells if a sample is the wrong group regardless of design
Top-DE heatmap, want to see PATTERNscale = 'row' (z-score per gene)
Top-DE heatmap, want to see ABSOLUTE LEVELscale = 'none' on assay(vsd); otherwise weak signal looks strong
Top-variable-gene selectionmatrixStats::rowMads(assay(vsd)) instead of rowVars -- MAD is outlier-robust
n = 3, top genes in volcanoNote Schurch 2016 finding: 20-40% of true positives missed; treat as exploratory
Many groups, comparing DE setsUpSet plot (Lex 2014); Venn drowns above 3 sets

Dispersion Diagnostic (Run This First)

Goal: Verify the dispersion-mean trend was fit acceptably before trusting any results.

Approach: plotDispEsts(dds) (DESeq2) or plotBCV(y) (edgeR) shows gene-wise (black/blue), fitted trend (red), and final shrunken (blue) dispersions vs mean.

r
plotDispEsts(dds)

plotBCV(y)
PatternMeaningAction
Cloud follows trend; final shrunken estimates pulled toward red curveHealthy fitProceed
Red trend nowhere near the gene-wise cloudParametric trend failedDESeq(dds, fitType = 'local') or fitType = 'mean'
Many gene-wise dispersions FAR ABOVE the trendOutlier or unmodeled batch genesInspect rather than trust QL F-test alone
Final estimates much lower than gene-wise everywhereExcessive shrinkage; sample too small or trend too flatCheck useEM, robust hyperparameter setting
BCV decreases monotonically with meanCorrect in edgeRDefault trend

A plot inspected before trusting results is worth a hundred lines of statistical safeguards.

P-value Histogram (Run This Second)

Goal: Detect model misspecification or hidden batch before reporting any gene list.

Approach: Histogram of raw p-values; under a correctly specified null, uniform with a spike near zero.

r
library(ggplot2)
ggplot(res_df, aes(x = pvalue)) +
    geom_histogram(bins = 50, fill = 'steelblue', color = 'white') +
    labs(x = 'P-value', y = 'Frequency', title = 'P-value distribution') +
    theme_bw()
ShapeMeaningAction
Uniform + spike at 0Correctly specifiedProceed
U-shape (spikes at 0 AND 1)Anti-conservative; hidden batch or unmodeled covariateAdd the missing covariate; re-fit
Depleted near 0, spike near 1Conservative; over-modeled or wrong dispersionSimplify model; check dispersion plot
Spike only at p = 1Discrete artifact from very-low-count genesPre-filter more aggressively

MA Plot (LFC vs Mean)

Goal: Inspect the relationship between LFC and mean expression for normalization correctness and shrinkage effect.

Approach: plotMA (DESeq2) or plotMD (edgeR). Always pick ylim deliberately; default can flatten the signal.

r
plotMA(res, ylim = c(-5, 5), main = 'MA plot (unshrunken)')

res_apeglm <- lfcShrink(dds, coef = 'condition_treated_vs_control', type = 'apeglm')
plotMA(res_apeglm, ylim = c(-5, 5), main = 'MA plot (apeglm-shrunken)')

plotMD(qlf, main = 'edgeR MD plot')
abline(h = c(-1, 1), col = 'blue', lty = 2)
PatternMeaning
Symmetric cloud centered at LFC = 0Correct normalization
Cloud median clearly above or below 0Normalization failed (TMM/RLE assumption violated) -- see normalization skill
Funnel widening at low meanExpected (low counts noisier)
Dramatic up/down asymmetryPossibly real (large biological perturbation), possibly normalization failure -- cross-check
Discrete horizontal bands at low meanLow-count artifacts; pre-filter more aggressively

The apeglm-shrunken MA visually flattens the left side; the post-shrinkage cloud should be tighter at low means.

Volcano with Shrunken LFC

Goal: Show effect size vs significance with honest fold changes.

Approach: Use a built-in renderer (EnhancedVolcano for quick publication-quality output) on shrunken LFCs. Always plot shrunken LFC; always set max.overlaps = Inf when labeling >10 genes -- the ggrepel default (10) silently drops labels. EnhancedVolcano accepts max.overlaps directly in 1.12+; version 1.10-1.11 has the older maxoverlapsConnectors argument (default 15); for either, falling back to options(ggrepel.max.overlaps = Inf) at the top of the script also works. For full ggplot2 customization (color schemes, faceting, label-set engineering), see data-visualization/volcano-and-ma-plots.

r
library(EnhancedVolcano)

res_apeglm <- lfcShrink(dds, coef = 'condition_treated_vs_control', type = 'apeglm')

EnhancedVolcano(res_apeglm,
    lab = rownames(res_apeglm),
    x = 'log2FoldChange', y = 'pvalue',
    pCutoff = 0.05, FCcutoff = 1,
    title = 'Treatment vs Control',
    subtitle = 'Shrunken LFC (apeglm); unshrunken Wald p',
    max.overlaps = Inf)

PCA on VST/rlog (Never on Raw Counts)

Goal: Show sample clustering by condition; detect batch effects, swaps, outliers.

Approach: Variance-stabilize first (VST or rlog), THEN PCA. Raw counts make PC1 = library size; log(counts+1) makes PC1 = mean expression. Neither carries biological signal until variance is stabilized.

r
vsd <- vst(dds, blind = FALSE)
plotPCA(vsd, intgroup = c('condition', 'batch'))

pca_df <- plotPCA(vsd, intgroup = c('condition', 'batch'), returnData = TRUE)
percentVar <- round(100 * attr(pca_df, 'percentVar'))

library(ggplot2)
ggplot(pca_df, aes(PC1, PC2, color = condition, shape = batch)) +
    geom_point(size = 4) +
    xlab(paste0('PC1: ', percentVar[1], '% variance')) +
    ylab(paste0('PC2: ', percentVar[2], '% variance')) +
    theme_bw()

library(limma)
plotMDS(cpm(y, log = TRUE), col = as.numeric(group), pch = 16)

blind=TRUE (default for vst()) re-estimates dispersions ignoring the design -- appropriate for unbiased QC ("are samples consistent independent of design?"). blind=FALSE uses the fitted dispersions -- appropriate for downstream visualization where the design is settled. Modern DESeq2 vignette recommends blind=FALSE for any plot after the model is fit.

PCA patternInterpretationAction
Clear separation by condition on PC1 or PC2Strong biological signalProceed
Separation by batch, not conditionBatch effect dominatesInclude batch in design; DO NOT subtract before DE (see batch-correction Nygaard 2016)
One sample far from its groupOutlier or swapCheck library QC; sex check; somalier
Condition signal on PC3+, not PC1-PC2Subtle effectMay still find DE; review dispersion plot
Two distinct sample clusters not explained by metadataHidden covariateInvestigate processing date, lane, machine

Sample Distance Heatmap (for QC)

r
library(pheatmap)
vsd <- vst(dds, blind = TRUE)
sd <- dist(t(assay(vsd)))
mat <- as.matrix(sd)
ann <- data.frame(condition = colData(dds)$condition,
                  row.names = colnames(dds))
pheatmap(mat, annotation_col = ann, annotation_row = ann,
         clustering_distance_rows = sd, clustering_distance_cols = sd,
         color = colorRampPalette(c('white', 'steelblue'))(100),
         main = 'Sample distance (vst blind)')

The diagonal should be dark; within-group samples should cluster. A within-group sample distant from its peers is a candidate for sample swap.

Show full SKILL.md (840 more words)Show less

Top-DE Heatmap and the Row-Scaling Trap

Goal: Show expression patterns of significant genes across samples for results figure.

Approach: Use vst(blind=FALSE), select top genes (by adjusted p-value or MAD-robust variance), choose scaling deliberately.

r
library(pheatmap)

sig <- rownames(subset(res, padj < 0.01))[1:50]
vsd <- vst(dds, blind = FALSE)
mat <- assay(vsd)[sig, ]

mat_scaled <- t(scale(t(mat)))

ann_col <- data.frame(condition = colData(dds)$condition,
                      batch     = colData(dds)$batch,
                      row.names = colnames(mat))

pheatmap(mat_scaled, annotation_col = ann_col,
         show_rownames = FALSE,
         clustering_distance_rows = 'correlation',
         clustering_distance_cols = 'correlation',
         color = colorRampPalette(c('blue', 'white', 'red'))(100),
         main = 'Top 50 DE genes (z-scored per gene)')

scale='row' (z-score per gene) is the conventional choice for "show me patterns". It DESTROYS absolute expression level information -- a gene at 5-7 with mean 6 looks identical to a gene at 10-1000. For pattern detection: correct. For QC heatmaps showing batch shifts: WRONG -- use scale='none' on assay(vsd).

Top-variable-gene selection robustness:

r
library(matrixStats)
vars_mad <- rowMads(assay(vsd))
top500 <- order(vars_mad, decreasing = TRUE)[1:500]

rowMads (median absolute deviation) is outlier-robust; rowVars is dominated by single-outlier-sample genes. For exploratory PCA of "top variable genes", MAD selection avoids artifacts.

Per-gene Plot

r
plotCounts(dds, gene = 'GENE_NAME', intgroup = 'condition')

d <- plotCounts(dds, gene = 'GENE_NAME', intgroup = c('condition','batch'),
                returnData = TRUE)
library(ggplot2)
ggplot(d, aes(x = condition, y = count, color = batch)) +
    geom_jitter(width = 0.1, size = 3) +
    scale_y_log10() +
    ggtitle('GENE_NAME') +
    theme_bw()

With n=3, the boxplot is misleading (3 points per box). Prefer geom_jitter over geom_boxplot at small n.

UpSet for Multi-set Comparisons

For >3 DE gene sets (e.g., contrasts treated_drugA, treated_drugB, treated_drugC each vs control), Venn diagrams become unreadable. UpSet (Lex et al. 2014 IEEE Trans Vis Comput Graph 20:1983) scales:

r
library(UpSetR)
upset(fromList(list(drugA = sig_drugA, drugB = sig_drugB, drugC = sig_drugC)))

Per-Method Failure Modes

Volcano with unshrunken LFC -- inflated story

Trigger: ggplot(res_df, aes(x=log2FoldChange, ...)) without lfcShrink(); extreme dots at the corners are low-count genes.

Mechanism: Unshrunken MLE LFCs are dominated by very-low-count genes whose log ratios are noisy. The visual top-left and top-right corners look impressive but are artifacts.

Symptom: Top genes by abs(LFC) are obscure low-count genes; reviewer asks "why are these the top hits?"

Fix: res_apeglm <- lfcShrink(dds, coef=..., type='apeglm'); plot from res_apeglm. Label axis "shrunken log2 fold change (apeglm)".

ggrepel max.overlaps silently drops labels

Trigger: geom_text_repel(data = top30, aes(label = gene)); only 10 labels render.

Mechanism: Default max.overlaps = 10; warning printed but easily missed in a knitr/Quarto render.

Symptom: Reviewer asks "where is gene X?"; it was in top30 but did not render.

Fix: geom_text_repel(..., max.overlaps = Inf) or options(ggrepel.max.overlaps = Inf) at top of script.

PCA shows batch, not condition

Trigger: plotPCA(vsd, intgroup='batch') cleanly separates batches; intgroup='condition' does not separate.

Mechanism: Batch variance exceeds condition variance.

Symptom: Treatment effect looks weak; DE p-values inflated if batch not in design.

Fix: Include batch in design (design = ~ batch + condition). DO NOT use removeBatchEffect then re-do DE on corrected counts (Nygaard 2016 cardinal sin -- see batch-correction). For VISUALIZATION only, removeBatchEffect is OK.

Heatmap row-scaling hid a sample-level shift

Trigger: QC heatmap with scale='row' looks consistent within group; downstream PCA shows clear sample outlier.

Mechanism: z-score per gene removes per-sample additive shifts. A sample that's globally inflated 1.5x looks identical to peers after row scaling.

Symptom: "The heatmap looked fine but PCA shows a problem."

Fix: For QC heatmaps, use scale = 'none' on assay(vsd) directly. For result heatmaps after QC is clean, scale = 'row' is the appropriate choice for pattern emphasis.

Top-N-by-rowVars dominated by single-outlier-sample genes

Trigger: "Top 500 variable genes" PCA shows a striped pattern, one or two samples driving the spread.

Mechanism: rowVars is squared-deviation; one outlier sample of one gene inflates that gene's "variance" massively.

Symptom: Top variable gene list includes many genes where N-1 samples are flat and one sample is extreme.

Fix: matrixStats::rowMads() for MAD-based selection; or genefilter::rowQ().

Common errors

Error / symptomCauseFix
plotPCA reports only 2 PCsDESeq2 plotPCA is hard-coded to PC1/PC2Use prcomp(t(assay(vsd))) and plot any pair
PCA cloud collapses to one pointForgot to log-transform; raw counts plottedvst(dds) first
All MA-plot points redalpha set too high or sig-flag bugVerify alpha; check padj vs pvalue in flag
pheatmap complains "infinite values"NA / Inf in scaled matrix; gene with zero varianceRemove zero-variance rows before scaling
Volcano axis labels obscuredDefault ggplot theme too compacttheme_bw(base_size = 14)
plotCounts says gene not foundWrong ID type (symbol vs Ensembl)Match rownames(dds) exactly
vst() errors with very low gene count post-filterDefault nsub=1000 exceeds available genesLower nsub (e.g., vst(dds, nsub=500))

References

  • Anders S, Huber W. 2010. Differential expression analysis for sequence count data. Genome Biol 11(10):R106. doi:10.1186/gb-2010-11-10-r106
  • Love MI, Huber W, Anders S. 2014. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15(12):550. doi:10.1186/s13059-014-0550-8
  • Zhu A, Ibrahim JG, Love MI. 2019. Heavy-tailed prior distributions for sequence count data: removing the noise and preserving large differences. Bioinformatics 35(12):2084-2092. doi:10.1093/bioinformatics/bty895
  • Robinson MD, McCarthy DJ, Smyth GK. 2010. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics 26(1):139-140. doi:10.1093/bioinformatics/btp616
  • Lex A, Gehlenborg N, Strobelt H, Vuillemot R, Pfister H. 2014. UpSet: Visualization of Intersecting Sets. IEEE Trans Vis Comput Graph 20(12):1983-1992. doi:10.1109/TVCG.2014.2346248
  • Schurch NJ et al. 2016. How many biological replicates are needed in an RNA-seq experiment and which differential expression tool should you use? RNA 22(6):839-851. doi:10.1261/rna.053959.115
  • Nygaard V, Rødland EA, Hovig E. 2016. Methods that remove batch effects while retaining group differences may lead to exaggerated confidence in downstream analyses. Biostatistics 17(1):29-39. doi:10.1093/biostatistics/kxv027
  • deseq2-basics - Generates the dds / res objects plotted here; vst/rlog choice
  • edger-basics - Generates y / qlf for plotMD, plotBCV, plotMDS
  • de-results - p-value histogram, padj=NA diagnosis informs what to plot
  • batch-correction - removeBatchEffect for visualization only (never as DE input)
  • expression-matrix/normalization - VST vs rlog vs log-CPM mechanics
  • data-visualization/volcano-and-ma-plots - Full custom volcano/MA with apeglm + ggrepel
  • data-visualization/dimensionality-reduction-plots - PCA, UMAP, t-SNE customization
  • data-visualization/heatmaps-clustering - pheatmap and ComplexHeatmap recipes
  • data-visualization/upset-plots - UpSet plot customization

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in differential-expression/de-visualization of GPTomics/bioSkills.

  • SKILL.md
  • examples/heatmap.R
  • examples/pca_plot.R
  • examples/volcano_plot.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Differential Expression De Visualization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Differential Expression De Visualization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Differential Expression De Visualization this skillGPTomics/bioSkills1.2k1 repos~5.1kAutomated safety check: PassMIT
Sandbox Benchvercel/next.js143k—~4.1kAutomated safety check: PassMIT
Statistical Analysisspacering-net/codeg3.8k4 repos~5kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.6k16 repos~4.9kAutomated safety check: PassBSD-3-Clause
Statistical Powerspacering-net/codeg3.8k2 repos~3.6kAutomated safety check: NotesMIT
AI Daily DigestvigorX777/ai-daily-digest1.6k—~1.3kAutomated safety check: PassNone

Similar skills

  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Statistical Analysis

    spacering-net/codeg

    Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting.

    3.8k GitHub starsUsed in 4 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 16 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Power

    spacering-net/codeg

    Sample-size and statistical power calculations for planning studies.

    3.8k GitHub starsUsed in 2 repos~3.6k tokens
    Data & AnalyticsAuto-check: notes
  • AI Daily Digest

    vigorX777/ai-daily-digest

    Fetches RSS feeds from 90 top Hacker News blogs (curated by Karpathy), uses AI to score and filter articles, and generates a daily digest in Markdown with Chinese-translated titles, category…

    1.6k GitHub stars~1.3k tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Agent Session Monitor

    higress-group/higress

    Real-time agent conversation monitoring - monitors Higress access logs, aggregates conversations by session, tracks token usage.

    9.5k GitHub stars~3.3k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Differential Expression De Visualization

What does Bio Differential Expression De Visualization do?

Creates DE-specific diagnostic and result visualizations using DESeq2/edgeR built-in functions and lightweight ggplot2 wrappers. Bio Differential Expression De Visualization is an agent skill from GPTomics/bioSkills. Creates DE-specific diagnostic and result visualizations using DESeq2/edgeR built-in functions and lightweight ggplot2 wrappers.

When should I use Bio Differential Expression De Visualization?

Bio Differential Expression De Visualization fits situations like: generating DE diagnostic plots; choosing VST vs rlog for visualization; troubleshooting suspicious plot patterns (shifted MA cloud; batch-dominated PCA.

How do I install Bio Differential Expression De Visualization in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-de-visualization -a claude-code`. Or copy the skill folder (differential-expression/de-visualization in GPTomics/bioSkills) into .claude/skills/bio-differential-expression-de-visualization in your project. Claude Code loads it when a task matches its description.

How do I install Bio Differential Expression De Visualization in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-de-visualization -a codex`. Or copy the skill folder (differential-expression/de-visualization in GPTomics/bioSkills) into .agents/skills/bio-differential-expression-de-visualization in your project. Codex loads it when a task matches its description.

Can I use Bio Differential Expression De Visualization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-differential-expression-de-visualization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-differential-expression-de-visualization, .gemini/skills/bio-differential-expression-de-visualization, .github/skills/bio-differential-expression-de-visualization and .opencode/skills/bio-differential-expression-de-visualization in your project.

What does Bio Differential Expression De Visualization need to run?

Going by SKILL.md and its folder, Bio Differential Expression De Visualization needs R for the scripts in its folder.

Does Bio Differential Expression De Visualization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Differential Expression De Visualization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Differential Expression De Visualization use?

Bio Differential Expression De Visualization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Differential Expression De Visualization use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Differential Expression De Visualization?

Skills that share tags, products or a category with Bio Differential Expression De Visualization: Sandbox Bench (vercel/next.js, 143k stars), Statistical Analysis (spacering-net/codeg, 3.8k stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.6k stars) and Statistical Power (spacering-net/codeg, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Differential Expression De Visualization?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.