Agent skill

Bio Differential Splicing

by GPTomics in GPTomics/bioSkills

Detects differential alternative splicing between conditions using rMATS-turbo (binomial LRT on junction counts), leafcutter (Dirichlet-multinomial GLM on intron clusters), MAJIQ V3 deltapsi/HET…

MITAuto-check passedResearch & Science

Install Bio Differential Splicing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-differential-splicing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-differential-splicing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alternative-splicing/differential-splicing .claude/skills/bio-differential-splicing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-differential-splicing
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.1k tokens
SKILL.md length
2,345 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Detects differential alternative splicing between conditions using rMATS-turbo (binomial LRT on junction counts), leafcutter (Dirichlet-multinomial GLM on intron clusters), MAJIQ V3 deltapsi/HET…

  • Works in 3 steps: Stratification: run rMATS within each… → PSI residuals (logit-transformed): PSI… → Switch to leafcutter (R function accepts…
  • Comparing splicing patterns between treatment groups
  • SKILL.md covers Version Compatibility, Statistical Model Taxonomy, Decision Tree by Experimental… and rMATS-turbo Differential…, plus 16 more sections
  • Runs R and Shell scripts from its folder; calls python, conda and pip

What it does

Bio Differential Splicing is an agent skill from GPTomics/bioSkills. Detects differential alternative splicing between conditions using rMATS-turbo (binomial LRT on junction counts), leafcutter (Dirichlet-multinomial GLM on intron clusters), MAJIQ V3 deltapsi/HET (Bayesian posterior on LSVs), SUPPA2 (empirical-null on TPM-derived PSI), or Shiba (junction-imbalance-corrected, 2025 SOTA at low coverage). Reports FDR-corrected significance and delta PSI effect sizes. Tools differ in statistical model, annotation dependence, calibration regime, and replicate-count requirements. Use…

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/diff_splicing_rmats.sh` and `usage-guide.md`).

It sits in Research & Science, covering Performance reviews and Test coverage. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Comparing splicing patterns between treatment groups
  • Tasks that involve Performance reviews
  • Tasks that involve Test coverage

Example prompts

  • “Use the bio-differential-splicing skill to detect differential alternative splicing between conditions using rMATS-turbo (binomial LRT on junction…”
  • “/bio-differential-splicing”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Stratification: run rMATS within each batch separately and meta-analyze.
  2. PSI residuals (logit-transformed): PSI is bounded [0,1]; raw linear regression near the boundaries is biased. Logit-transform first…
  3. Switch to leafcutter (R function accepts confounders matrix; CLI accepts confounders as additional columns in the groups file).

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • conda
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • sika-zheng-lab.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Differential Splicing loads about 6.1k tokens when it runs. Until then it costs about 157 tokens; SKILL.md has 2,345 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~157
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,345 words, ~6,101 tokens.

Download SKILL.mdSave it as .claude/skills/bio-differential-splicing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-differential-splicing
description
Detects differential alternative splicing between conditions using rMATS-turbo (binomial LRT on junction counts), leafcutter (Dirichlet-multinomial GLM on intron clusters), MAJIQ V3 deltapsi/HET (Bayesian posterior on LSVs), SUPPA2 (empirical-null on TPM-derived PSI), or Shiba (junction-imbalance-corrected, 2025 SOTA at low coverage). Reports FDR-corrected significance and delta PSI effect sizes. Tools differ in statistical model, annotation dependence, calibration regime, and replicate-count requirements. Use when comparing splicing patterns between treatment groups, tissues, or disease states.
tool_type
mixed
primary_tool
rMATS-turbo

Version Compatibility

Reference examples tested with: rMATS-turbo 4.3+, SUPPA2 2.4+, leafcutter 0.2.9+, MAJIQ 3.0+, Shiba 0.5+, STAR 2.7.11+, regtools 1.0+, pandas 2.2+, R 4.4+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Differential Splicing

Detect splicing changes between conditions. Tool choice is a decision about statistical model, annotation dependence, and calibration regime under the specific experimental design — not a preference. Wrong tool for the design produces uncalibrated FDR or systematic effect-size bias.

Statistical Model Taxonomy

ToolModelTest statisticMin reps per groupCalibration regimeFails when
rMATS-turboBinomial counts with hierarchical PSI varianceLRT on |ΔPSI| > cutoff (default 0.0001)n>=3Well-calibrated at n>=3 with adequate junction readsJunction read imbalance; very low coverage; uncorrected for confounders
leafcutterDirichlet-multinomial GLM at cluster levelLRT on group factorn>=2 (n>=3 preferred)Strong at n>=3; novel-junction-friendlyUndersampled clusters (DM dispersion unstable); cluster topology arbitrariness
MAJIQ deltapsiBeta-binomial bootstrap -> posterior over PSI per LSVP(|ΔPSI| > T) threshold (T=0.2)n>=3Replicate-structured n=3 vs n=3Cohorts where between-sample variability dominates between-group
MAJIQ HETSame model, heterogeneity-awarePer-LSV permutation-based testn>=10n>=10 vs n>=10 cohort designsTightly-controlled small replicate experiments
SUPPA2 (empirical)Empirical null from between-replicate ΔPSIECDF on |ΔPSI| conditioned on TPMn>=4n>=4 vs n>=4 with paired-end deep sequencingn<=3 vs n<=3 (sparse null collapses)
SUPPA2 (classical)Wilcoxon rank-sum on PSI distributionsWilcoxon p-valuen>=2Small samples; non-parametric backupCassette events with tight PSI distributions
Shiba (2025)Beta-binomial with explicit junction-imbalance correctionLRTn>=2n=2-3 vs n=2-3Established benchmarks limited (new tool)
LeafcutterMDDirichlet-multinomial outlier modePer-sample p-valuen=1 vs cohort >=20Single-patient vs cohortToo few controls (<20)
FRASER 2.0Beta-binomial autoencoder on Intron Jaccard IndexPer-sample p-value with delta cutoffn=1 vs cohort >=20n>=20 control cohort, single-patient querySee outlier-splicing-detection for this regime

The first decision is which regime the design falls into: between-group with replicates, heterogeneous cohort, or single-sample-vs-cohort. Within each regime, tool choice is much smaller (1-2 options).

Comprehensive 2023-2026 benchmarks: Olofsson 2023 Biochem Biophys Res Commun; Tran 2025 WIREs RNA; Kubota 2025 NAR. Methodology evolves — verify benchmarks and tool docs before reporting. Default 2026 recommendation: run two complementary tools (rMATS + leafcutter) and require concordance for high-confidence calls.

Decision Tree by Experimental Design

ScenarioRecommended toolWhyThreshold
Standard n=3 vs n=3, GENCODE-annotatedrMATS-turbo + leafcutter (concordance)Two algorithmic families; concordant hits = high-confidenceFDR<0.05, |ΔPSI|>0.10
n=2 vs n=2 small pilotShibaJunction-imbalance correction matters most at low coverageFDR<0.10, |ΔPSI|>0.10
n=10+ vs n=10+ heterogeneous (clinical, GTEx-style)MAJIQ V3 HETHET designed for between-sample heterogeneityP(|ΔPSI|>0.2)>0.95
Single rare-disease patient vs panel of n>=20FRASER 2.0 (see outlier-splicing-detection)Outlier detection statistical model is fundamentally differentpadj<0.05, |delta-jaccard|>=0.1
Time-course / multi-condition designCustom DEXSeq or limma on PSI matrixrMATS/leafcutter primarily 2-groupFDR<0.05 on time:group interaction
Paired tumor-normalrMATS with --paired-statsPaired test reduces inter-patient varianceFDR<0.05, paired |ΔPSI|>0.10
Cancer with spliceosomal mutation (SF3B1, U2AF1)leafcutter or MAJIQ denovoCryptic events not in annotationFDR<0.05; check 3'ss shifts in IGV
TDP-43 loss / ALS post-mortemleafcutter denovoCryptic exons not in annotationFDR<0.05; expect UNC13A, STMN2
Non-model organism without GENCODE-grade annotationleafcutterAnnotation-freeFDR<0.05, |ΔPSI|>0.10
Long-read availablerMATS-long, FLAIR diffSpliceSee long-read-splicingTool-specific

rMATS-turbo Differential Analysis

Goal: Detect statistically significant differential splicing between two groups from BAMs.

Approach: Run rMATS-turbo without --statoff, then filter by FDR + ΔPSI + per-replicate coverage.

bash
rmats.py \
    --b1 condition1_bams.txt \
    --b2 condition2_bams.txt \
    --gtf annotation.gtf \
    -t paired \
    --readLength 150 \
    --variable-read-length \
    --libType fr-firststrand \
    --nthread 8 \
    --od rmats_output \
    --tmp rmats_tmp \
    --novelSS \
    --cstat 0.05

--cstat 0.05 tests |ΔPSI| > 0.05; raise to 0.10 for stricter discovery. --novelSS enables novel-junction discovery (recommended with STAR 2-pass). For paired designs, add --paired-stats.

python
import pandas as pd
import numpy as np

se = pd.read_csv('rmats_output/SE.MATS.JC.txt', sep='\t')

def min_per_rep(s):
    return s.str.split(',').apply(lambda x: min(int(v) for v in x))

se['min_inc'] = min_per_rep(se['IJC_SAMPLE_1']).combine(min_per_rep(se['IJC_SAMPLE_2']), min)
se['min_skip'] = min_per_rep(se['SJC_SAMPLE_1']).combine(min_per_rep(se['SJC_SAMPLE_2']), min)

significant = se[
    (se['FDR'] < 0.05) &
    (se['IncLevelDifference'].abs() > 0.10) &
    ((se['min_inc'] + se['min_skip']) >= 10)
].copy()

significant['score'] = -np.log10(significant['FDR']) * significant['IncLevelDifference'].abs()
top = significant.nlargest(50, 'score')

leafcutter Differential Intron Usage

Goal: Detect differential intron-cluster usage annotation-free, capturing novel junctions and complex multi-junction events.

Approach: Extract junctions with regtools, cluster introns by shared splice sites, run cluster-level Dirichlet-multinomial test.

bash
for bam in *.bam; do
    regtools junctions extract -a 8 -m 50 -s XS "$bam" -o "${bam%.bam}.junc"
done
ls *.junc > juncfiles.txt

python leafcutter_cluster_regtools.py \
    -j juncfiles.txt \
    -o leafcutter \
    -m 50 \
    -l 500000
r
library(leafcutter)

groups <- data.frame(
    sample = c('s1', 's2', 's3', 's4', 's5', 's6'),
    group = c('control', 'control', 'control', 'treatment', 'treatment', 'treatment')
)
write.table(groups, 'groups.txt', sep = '\t', quote = FALSE, row.names = FALSE, col.names = FALSE)

system('leafcutter_ds.R --num_threads 4 --exon_file gencode_exons.txt.gz \
    leafcutter_perind_numers.counts.gz groups.txt -o ds_results')

cluster_sig <- read.table('ds_results_cluster_significance.txt', header = TRUE, sep = '\t')
intron_effects <- read.table('ds_results_effect_sizes.txt', header = TRUE, sep = '\t')

sig_clusters <- subset(cluster_sig, p.adjust < 0.05)

LeafCutter2 (Buen Abad Najar 2025 bioRxiv) extends leafcutter with NMD-aware classification of unproductive splicing — useful when AS-NMD coupling is the question.

MAJIQ V3 Differential Analysis

Goal: Detect differential LSVs with full posterior distributions over ΔPSI; ideal for complex multi-junction events and heterogeneous cohorts.

Approach: Build splice graph -> compute coverage per group -> run deltapsi (replicate-structured) or heterogen (cohort-style).

bash
majiq build annotation.gff3 -c settings.ini -j 8 -o build_output

majiq deltapsi \
    -grp1 build_output/ctrl1.majiq build_output/ctrl2.majiq build_output/ctrl3.majiq \
    -grp2 build_output/trt1.majiq build_output/trt2.majiq build_output/trt3.majiq \
    -n control treatment \
    -o deltapsi_output \
    --minreads 10 --minpos 3 \
    -j 8

majiq heterogen \
    -grp1 build_output/het_ctrl{1..20}.majiq \
    -grp2 build_output/het_trt{1..20}.majiq \
    -n control treatment \
    -o heterogen_output \
    -j 8

voila view -p 5000 -j 8 build_output/splicegraph.zarr deltapsi_output/control_treatment.deltapsi.voila -o voila_html

MAJIQ V3 (Aicher, Slaff, Jewell, Barash bioRxiv 2024; public release 2025) uses Zarr storage (splicegraph.zarr); V2's SQLite splicegraph is deprecated. MAJIQ reports posterior probability P(|ΔPSI| > 0.2); thresholds are interpreted differently from FDR. Use HET for n>=10 vs n>=10 cohort designs (clinical, GTEx-style); deltapsi for tightly controlled n=3 vs n=3.

SUPPA2 Differential Analysis

Goal: Quick differential splicing from existing transcript quantifications, useful as a sanity check or pilot.

Approach: Generate per-condition PSI files from Salmon TPM, then run diffSplice with empirical or classical p-values.

bash
suppa.py generateEvents -i annotation.gtf -o events -f ioe -e SE SS MX RI

for ev in SE A5 A3 MX RI; do
    suppa.py psiPerEvent -i events_${ev}_strict.ioe -e ctrl_tpm.tsv -o ctrl_${ev}
    suppa.py psiPerEvent -i events_${ev}_strict.ioe -e trt_tpm.tsv -o trt_${ev}

    suppa.py diffSplice \
        -m empirical \
        -gc \
        -i events_${ev}_strict.ioe \
        -p ctrl_${ev}.psi trt_${ev}.psi \
        -e ctrl_tpm.tsv trt_tpm.tsv \
        -o diff_${ev}
done

For n<=3 designs, switch -m classical (Wilcoxon). Empirical null requires sufficient between-replicate observations to construct.

Shiba for Low-Coverage / Few-Replicate Designs

Goal: Detect differential splicing with explicit junction-imbalance correction — addresses a known false-positive source for rMATS-style methods.

Approach: Shiba is a Snakemake-based pipeline configured via YAML. Install via bioconda, write a config file describing groups + BAMs, then run with snakemake.

bash
conda install -c bioconda shiba

# Edit config.yaml with reference GTF, BAM groups, output dir, thresholds
# Then run the Snakemake workflow:
snakemake -s snakeshiba.smk \
    --configfile config.yaml \
    --cores 8 \
    --use-singularity \
    --singularity-args "--bind $HOME:$HOME"

Shiba (Kubota 2025 NAR) reportedly outperforms rMATS at n=2 vs n=2 by correcting differential mappability between inclusion and skipping junctions; community calibration still emerging. See https://sika-zheng-lab.github.io/Shiba/ for the full config.yaml schema.

Per-Tool Failure Modes

rMATS: Confounder-Blind LRT

Trigger: Sequencing batch, RIN, library prep date, or sex correlates with the comparison of interest.

Mechanism: rMATS' default LRT does not natively accept covariates the way DESeq2 does; the --paired-stats flag handles paired designs but not arbitrary covariates.

Symptom: Many "significant" hits driven by batch rather than biological condition; PCA on PSI matrix shows samples clustering by batch rather than group.

Fix: Either (a) include batch as a stratification (run rMATS within each batch), (b) regress PSI matrix against batch in R, then test residuals, or (c) switch to leafcutter which accepts a confounders argument in differential_splicing.

leafcutter: Cluster Mis-Topology

Trigger: A cluster spans a complex topology (cassette + alternative donor in same cluster).

Mechanism: Cluster-level p-value reports "something in this cluster differs" but doesn't indicate which intron drove the change; downstream analysis needs per-intron effect sizes.

Symptom: Significant cluster, multiple introns with different ΔPSI directions, ambiguous biological interpretation.

Fix: Inspect cluster in leafviz; report per-intron effect sizes from ds_results_effect_sizes.txt; map cluster topology to canonical SE/A5SS/A3SS via flanking exon coordinates manually.

MAJIQ HET: Power vs Type-1 Tradeoff

Trigger: HET module on n=5-10 cohorts (between regimes).

Mechanism: HET assumes between-sample variability dominates; for moderate-replicate designs (n=5-10), HET is conservative and deltapsi is more powerful.

Symptom: HET reports few hits in n=5-10 designs; deltapsi on the same data reports many.

Fix: Use deltapsi for n=3-5; reserve HET for n>=10 with explicit cohort heterogeneity.

SUPPA2 Empirical: Sparse Null at Low Replicate

Trigger: n<=3 vs n<=3 with --method empirical.

Mechanism: Empirical null is ECDF of |ΔPSI| from between-replicate comparisons within each group, binned by transcript expression. Few replicates -> few null observations -> wide confidence on null distribution.

Symptom: Inflated FDR (15-30%); "significant" hits don't replicate or validate.

Fix: Use -m classical (Wilcoxon) for n<=3 vs n<=3; or switch tool entirely (leafcutter, Shiba).

Reconciliation: When Tools Disagree

The two most common short-read tools answer slightly different questions: rMATS classifies on annotated event templates; leafcutter classifies on observed cluster usage. Disagreement is informative.

PatternLikely causeAction
rMATS sig, leafcutter not sigrMATS junction imbalance OR rMATS event hits annotation that leafcutter clustered differentlyInspect locus in IGV; check Shiba on the same locus
leafcutter sig, rMATS not sigNovel junction not in rMATS annotation; rMATS --novelSS may have missed itVerify --novelSS was on; rerun if not
Both sig, opposite ΔPSI directionEvent class mismatch (rMATS calls SE positive, leafcutter sees A5SS shift in same cluster)Manually map cluster topology to event class
Both sig, same directionHigh-confidence callReport; cross-validate with sashimi-plot
All tools null but biology suggests changeUnderpowered design or wrong regimeIncrease replicates; check whether outlier-splicing-detection regime applies

Operational rule: for high-confidence reporting, require concordant detection in two tools from different algorithmic families (event-based + cluster-based, or LSV + isoform-based). Document both calls and any explainable disagreements.

rMATS Output Columns Reference

ColumnMeaning
IJC_SAMPLE_1 / SJC_SAMPLE_1Comma-delimited inclusion / skipping junction counts per replicate, group 1
IJC_SAMPLE_2 / SJC_SAMPLE_2Same for group 2
IncFormLen / SkipFormLenEffective lengths normalizing PSI for differential mapping opportunity
upstreamES/EE, downstreamES/EEFlanking exon coordinates (genomic order; strand-agnostic in column meaning)
exonStart_0base / exonEndCassette exon coordinates (0-based half-open)
PValueLRT p-value of |ΔPSI| > cutoff
FDRBH-adjusted PValue within event class
IncLevel1, IncLevel2Comma-delimited per-replicate PSI values
IncLevelDifferencemean(IncLevel1) - mean(IncLevel2); sign matches --b1 - --b2 order
Show full SKILL.md (919 more words)Show less

Replicate Count and Power

DesignRecommended toolsExpected power for ΔPSI=0.2
n=2 vs n=2leafcutter or Shiba; avoid SUPPA2Marginal; many real effects missed
n=3 vs n=3rMATS-turbo + leafcutterAdequate at moderate coverage; standard
n=5 vs n=5rMATS or leafcutter, MAJIQ deltapsiGood; recommended for publication
n=10+ vs n=10+ heterogeneousMAJIQ-HETDesigned for this scale
Single patient vs n=20+ controlsleafcutterMD or FRASER2Outlier regime; see outlier-splicing-detection

For an effect-size of |ΔPSI|=0.10 (typical biological signal), power generally requires n>=4 and >=20 junction reads per replicate. Below this, expect to miss most real changes.

Significance and Effect-Size Thresholds

Stringency|ΔPSI|FDRUse case
Lenient> 0.05< 0.10Discovery, exploratory, hypothesis generation
Standard> 0.10< 0.05Publication; default reporting threshold
Stringent> 0.20< 0.01Validation cohort, follow-up targets

For MAJIQ: posterior probability P(|ΔPSI| > 0.2) >= 0.95 is roughly equivalent to standard stringency. Always document tool, threshold, and rationale.

Biologically meaningful ΔPSI varies by context:

  • A poison exon shift of |ΔPSI|=0.10 can halve functional protein (huge biology, modest number).
  • A stoichiometric isoform shift of |ΔPSI|=0.10 may be physiologically silent.
  • Therapeutic ASO target: SMA nusinersen aims for ΔPSI~+0.30 in SMN2 exon 7.

Confounder Handling

rMATS does not natively accept arbitrary covariates. Workarounds:

  1. Stratification: run rMATS within each batch separately and meta-analyze.
  2. PSI residuals (logit-transformed): PSI is bounded [0,1]; raw linear regression near the boundaries is biased. Logit-transform first, regress on confounders, then test residuals.
  3. Switch to leafcutter (R function accepts confounders matrix; CLI accepts confounders as additional columns in the groups file).
python
import numpy as np
import statsmodels.formula.api as smf

# logit-transform PSI before residualization (PSI is bounded [0,1])
eps = 1e-3
psi['logit_psi'] = np.log((psi['psi'].clip(eps, 1 - eps)) / (1 - psi['psi'].clip(eps, 1 - eps)))
psi['psi_resid'] = smf.ols('logit_psi ~ batch + RIN', data=psi).fit().resid
# then test psi_resid by group via Wilcoxon

leafcutter accepts confounders two ways:

  • R function: differential_splicing(counts, x, confounders=numeric_matrix) accepts a numeric covariate matrix
  • CLI script: leafcutter_ds.R reads confounders from additional columns in the groups file (3rd, 4th, ... columns), NOT from a --confounders flag

MAJIQ does not accept arbitrary confounders; use stratification or switch tool.

Always check confounding before reporting: PCA on PSI matrix; if PC1 separates by batch rather than group, the comparison is confounded.

Multi-Group / Multi-Factor Designs

DesignApproach
3 groups (e.g. drug A, drug B, control)Pairwise rMATS or leafcutter; OR limma/DESeq2 on logit-PSI matrix
Time-course (e.g. 0h, 6h, 24h)DEXSeq on event counts with time as factor; or limma::lmFit on PSI matrix
2x2 factorial (genotype × treatment)DEXSeq with interaction term; rMATS pairwise on interaction subsets
Continuous covariate (dose, age)limma::lmFit on logit-PSI ~ covariate

For complex designs, custom regression on the PSI matrix is more flexible than rMATS/leafcutter pairwise.

Common Errors

ErrorCauseSolution
rMATS: numpy.AxisErrorrMATS version mismatch with numpy >=2.0Pin numpy<2.0 or update rMATS-turbo to >=4.3
leafcutter: zero variance in clusterCluster has all-zero counts in a groupPre-filter with --min_samples_per_intron 5 --min_samples_per_group 3
MAJIQ: out of memoryDefault settings on >50-sample cohortUse --mem-profile flag; chunk samples; consider HET for large cohorts
SUPPA2: no events with sufficient coverageSalmon/kallisto TPM filter too strict upstreamLower upstream TPM threshold; verify event annotations
voila: missing splicegraph.zarr (V3) or splicegraph.sql (V2; deprecated)Forgot to keep build output directoryRe-run majiq build; output must persist for VOILA
regtools: too many open filesMany BAMs in one batchulimit -n 4096 or batch in groups

Result Prioritization

Goal: Rank events by combined statistical and biological significance for follow-up.

Approach: Composite score combining FDR and effect size, then enrich for biology (RBP binding, NMD sensitivity, conservation, disease relevance).

python
import pandas as pd
import numpy as np

sig['score'] = -np.log10(sig['FDR']) * sig['IncLevelDifference'].abs()
sig['exon_length'] = sig['exonEnd'] - sig['exonStart_0base']
sig['nmd_likely'] = (sig['exon_length'] % 3 != 0)
top_events = sig.nlargest(50, 'score')

Cross-reference top hits with:

  • eCLIP/ENCODE RBP target databases (POSTAR3, oRNAment, RBP2GO) -> candidate trans-regulators
  • Disease-specific signatures: SF3B1 cryptic 3'ss for MDS/CLL/UM; TDP-43 cryptic exons (UNC13A, STMN2) for ALS/FTD
  • Conservation: VastDB cross-species PSI for evolutionary support
  • Splice-site predictions: SpliceAI scores for the involved sites (see splice-variant-prediction)

Common Pitfalls

  • Junction read imbalance (cassette exon flanks have unequal mapping opportunity) inflates rMATS false positives; Shiba explicitly corrects this.
  • Comparing tool outputs naively — MAJIQ posteriors and rMATS FDR are different scales; use threshold equivalents (P>0.95 ~ FDR<0.05 in many regimes) but confirm with simulation when reporting.
  • Forgetting NMD direction — increased PSI of a poison exon decreases protein. Always check whether the alternative form is PTC-introducing using ORF-aware annotation.
  • Cryptic splicing in TDP-43 loss / SF3B1-mutant samples — annotation-bound tools (rMATS, SUPPA2) miss these; need leafcutter or MAJIQ with denovo mode.
  • Forgetting strand — wrong --libType halves usable junctions. Confirm with RSeQC infer_experiment.py.
  • Reporting one tool's call as ground truth — discordance between rMATS and leafcutter is informative, not a problem to hide.
  • Skipping confounder check — always run PCA on PSI matrix before final reporting.
  • Using empirical SUPPA2 at n<=3 — calibration collapses; use classical mode or different tool.
  • splicing-quantification - PSI estimation per event; foundational
  • splicing-qc - Run BEFORE differential to verify library, depth, strandedness; avoid downstream surprises
  • isoform-switching - DTU framework with NMD/ORF/domain consequences; complementary to event-level
  • sashimi-plots - Visualize differential events for QC and reporting
  • outlier-splicing-detection - Single-sample-vs-cohort regime (FRASER2/DROP); use when not 2-group
  • splice-variant-prediction - SpliceAI / Pangolin for variant-driven mechanistic explanation of differential events
  • long-read-splicing - Differential analysis from full-length isoforms; use when short-read insufficient
  • read-alignment/star-alignment - STAR 2-pass cohort-style required upstream

References

  • Shen et al 2014 PNAS - rMATS original
  • Wang et al 2024 Nat Protoc - rMATS-turbo
  • Li et al 2018 Nat Genet - leafcutter (Dirichlet-multinomial GLM)
  • Buen Abad Najar et al 2025 bioRxiv - LeafCutter2 (NMD-aware unproductive splicing)
  • Vaquero-Garcia et al 2016 eLife - MAJIQ LSV framework
  • Vaquero-Garcia et al 2023 Nat Commun - MAJIQ-HET heterogeneity module
  • Aicher, Slaff, Jewell, Barash 2024 bioRxiv - MAJIQ V3
  • Trincado et al 2018 Genome Biol - SUPPA2
  • Kubota et al 2025 NAR - Shiba (junction-imbalance correction)
  • Olofsson et al 2023 Biochem Biophys Res Commun 653:31-37 - benchmark across tools
  • Tran et al 2025 WIREs RNA - methodology review
  • Brown et al 2022 Nature - UNC13A cryptic exon (TDP-43 / ALS)
  • Klim et al 2019 Nat Neurosci - STMN2 cryptic splicing (ALS)
  • Darman et al 2015 Cell Rep - SF3B1 cryptic 3'ss

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in alternative-splicing/differential-splicing of GPTomics/bioSkills.

  • SKILL.md
  • examples/diff_splicing_leafcutter.R
  • examples/diff_splicing_rmats.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Differential Splicing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Differential Splicing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Differential Splicing this skillGPTomics/bioSkills1.2k2 repos~6.1kAutomated safety check: PassMIT
Sealeap Taotie Amazon Third Party Tool Data Calibrationxjli360/sealeap-amazon-skills247—~804Automated safety check: PassMIT
Paper From Zeroyunshenwuchuxun/latex-paper-skills266—~1.5kAutomated safety check: PassMIT
Lsp Inspectblackwell-systems/agent-lsp160—~4.2kAutomated safety check: NotesMIT
Jqte Io Cgefranklee16/academic-research-skills2231 repos~419Automated safety check: PassNone
Ier Robustnessbrycewang-stanford/Awesome-Journal-Skills1.2k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Calibrate third-party keyword and sales-estimate tools against each other: run the same keyword and the same high-volume ASIN through several tools, compare search volume, estimated monthly sales…

    247 GitHub stars~804 tokensUpdated 12 days ago
    Business, Finance & HRAuto-check passed
  • Paper From Zero

    yunshenwuchuxun/latex-paper-skills

    Route a fixed research topic into a rigorous paper-generation workflow.

    266 GitHub stars~1.5k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Lsp Inspect

    blackwell-systems/agent-lsp

    Full code quality audit for a file, package, or directory. An agent skill from blackwell-systems/agent-lsp.

    160 GitHub stars~4.2k tokensUpdated today
    DevelopmentAuto-check: notes
  • Jqte Io Cge

    franklee16/academic-research-skills

    A skill your agent uses when a 《数量经济技术经济研究》 (JQTE) manuscript is built on an input-output table, a CGE model, or a structural decomposition (SDA).

    223 GitHub starsUsed in 1 repo~419 tokens
    Research & ScienceAuto-check passed
  • Ier Robustness

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when an International Economic Review (IER) result may be sensitive to specification, sample, functional form, calibration, or inference choices.

    1.2k GitHub stars~2.4k tokensUpdated 12 days ago
    Research & ScienceAuto-check passed
  • Ppsych Revision

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to a Perspectives on Psychological Science (PoPS) editor and reviewer decision letter — coverage gaps, balance/fairness complaints, framework/argument asks…

    1.2k GitHub stars~1.5k tokensUpdated 12 days ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Differential Splicing

What does Bio Differential Splicing do?

Detects differential alternative splicing between conditions using rMATS-turbo (binomial LRT on junction counts), leafcutter (Dirichlet-multinomial GLM on intron clusters), MAJIQ V3 deltapsi/HET…. Bio Differential Splicing is an agent skill from GPTomics/bioSkills. Detects differential alternative splicing between conditions using rMATS-turbo (binomial LRT on junction counts), leafcutter (Dirichlet-multinomial GLM on intron clusters), MAJIQ V3 deltapsi/HET (Bayesian posterior on LSVs), SUPPA2 (empirical-null on TPM-derived PSI), or Shiba (junction-imbalance-corrected, 2025 SOTA at low coverage).

When should I use Bio Differential Splicing?

Bio Differential Splicing fits situations like: comparing splicing patterns between treatment groups; tasks that involve Performance reviews; tasks that involve Test coverage.

How do I install Bio Differential Splicing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-differential-splicing -a claude-code`. Or copy the skill folder (alternative-splicing/differential-splicing in GPTomics/bioSkills) into .claude/skills/bio-differential-splicing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Differential Splicing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-differential-splicing -a codex`. Or copy the skill folder (alternative-splicing/differential-splicing in GPTomics/bioSkills) into .agents/skills/bio-differential-splicing in your project. Codex loads it when a task matches its description.

Can I use Bio Differential Splicing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-differential-splicing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-differential-splicing, .gemini/skills/bio-differential-splicing, .github/skills/bio-differential-splicing and .opencode/skills/bio-differential-splicing in your project.

What does Bio Differential Splicing need to run?

Going by SKILL.md and its folder, Bio Differential Splicing needs R and a shell for the scripts in its folder and the command-line tools its instructions call (python, conda and pip). Our summary lists: Python 3; A Bash shell.

Does Bio Differential Splicing access the network?

SKILL.md names 1 domain. As links in the text: sika-zheng-lab.github.io. This is read from the text; nothing was executed.

Is Bio Differential Splicing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Differential Splicing use?

Bio Differential Splicing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Differential Splicing use?

About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Differential Splicing?

Skills that share tags, products or a category with Bio Differential Splicing: Sealeap Taotie Amazon Third Party Tool Data Calibration (xjli360/sealeap-amazon-skills, 247 stars), Paper From Zero (yunshenwuchuxun/latex-paper-skills, 266 stars), Lsp Inspect (blackwell-systems/agent-lsp, 160 stars) and Jqte Io Cge (franklee16/academic-research-skills, 223 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Differential Splicing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.