Agent skill

Bio Splicing Quantification

by GPTomics in GPTomics/bioSkills

Quantifies alternative splicing as PSI (percent spliced in) from RNA-seq using rMATS-turbo (BAM-based event), SUPPA2 (TPM-based event), MAJIQ V3 (LSV-based Bayesian), leafcutter (annotation-free…

MITAuto-check passedResearch & Science

Install Bio Splicing Quantification

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-splicing-quantification -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-splicing-quantification --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alternative-splicing/splicing-quantification .claude/skills/bio-splicing-quantification && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-splicing-quantification
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.8k tokens
SKILL.md length
3,056 words
Files
3
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Quantifies alternative splicing as PSI (percent spliced in) from RNA-seq using rMATS-turbo (BAM-based event), SUPPA2 (TPM-based event), MAJIQ V3 (LSV-based Bayesian), leafcutter (annotation-free…

  • Works in 3 steps: Canonical RI (cytoplasmic, NMD-substrate… → Detained intron (DI) (Boutz 2015 Genes… → Co-transcriptional unspliced: nascent…
  • Measuring splice-site usage
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Event Taxonomy (Beyond… and Tool Selection Matrix, plus 17 more sections
  • Runs Python scripts from its folder; calls python and pip

What it does

Bio Splicing Quantification is an agent skill from GPTomics/bioSkills. Quantifies alternative splicing as PSI (percent spliced in) from RNA-seq using rMATS-turbo (BAM-based event), SUPPA2 (TPM-based event), MAJIQ V3 (LSV-based Bayesian), leafcutter (annotation-free intron clusters), VAST-TOOLS (cross-species with microexon support), Shiba (junction-imbalance-corrected, 2025 SOTA at low coverage), or IRFinder-S (intron retention coverage-aware). Distinguishes the five canonical event classes (SE, A5SS, A3SS, MXE, RI), special classes (microexons, exitrons, AFE/ALE), intron retention…

Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/quantify_splicing.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Database schema design. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Measuring splice-site usage
  • Isoform inclusion ratios from short-read RNA-seq

Example prompts

  • “Use the bio-splicing-quantification skill to quantify alternative splicing as PSI (percent spliced in) from RNA-seq using rMATS-turbo (BAM-based…”
  • “/bio-splicing-quantification”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Canonical RI (cytoplasmic, NMD-substrate often): mature polyadenylated mRNA carries the intron; usually PTC-bearing and NMD-targeted…
  2. Detained intron (DI) (Boutz 2015 Genes Dev): nuclear-localized, mature transcripts retaining a specific intron; a regulated reservoir…
  3. Co-transcriptional unspliced: nascent pre-mRNA captured before splicing complete; not a regulated state.

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Splicing Quantification loads about 6.8k tokens when it runs. Until then it costs about 181 tokens; SKILL.md has 3,056 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~181
When it runs · the whole SKILL.md, loaded when a task matches
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,056 words, ~6,827 tokens.

Download SKILL.mdSave it as .claude/skills/bio-splicing-quantification/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-splicing-quantification
description
Quantifies alternative splicing as PSI (percent spliced in) from RNA-seq using rMATS-turbo (BAM-based event), SUPPA2 (TPM-based event), MAJIQ V3 (LSV-based Bayesian), leafcutter (annotation-free intron clusters), VAST-TOOLS (cross-species with microexon support), Shiba (junction-imbalance-corrected, 2025 SOTA at low coverage), or IRFinder-S (intron retention coverage-aware). Distinguishes the five canonical event classes (SE, A5SS, A3SS, MXE, RI), special classes (microexons, exitrons, AFE/ALE), intron retention subtypes (canonical RI vs detained introns), and applies effective-length normalization. Use when measuring splice-site usage or isoform inclusion ratios from short-read RNA-seq.
tool_type
mixed
primary_tool
rMATS-turbo

Version Compatibility

Reference examples tested with: rMATS-turbo 4.3+, SUPPA2 2.4+, leafcutter 0.2.9+, MAJIQ 3.0+, IRFinder-S 2.0+, kallisto 0.50+, Salmon 1.10+, pandas 2.2+, STAR 2.7.11+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Splicing Quantification

Quantify alternative splicing events as PSI (percent spliced in) from RNA-seq. PSI = inclusion read evidence / (inclusion + skipping read evidence), normalized for differential mapping opportunity between isoforms. The choice of quantification unit (event, intron cluster, LSV, transcript) determines which biological questions can be answered and which failure modes apply.

Algorithmic Taxonomy

FamilyUnitReference toolsFails when
Event-basedPre-defined SE/A5SS/A3SS/MXE/RI events from annotationrMATS-turbo, SUPPA2, VAST-TOOLSEvent isn't in annotation; complex multi-junction events split arbitrarily; AFE/ALE confounded with splicing
LSV-basedLocal Splice Variations at single source/target nodesMAJIQ V3Memory-constrained environments; cohorts smaller than ~3 reps; non-academic users (license)
Junction-clusterAnnotation-free intron clusters by shared splice sitesleafcutter, leafcutter2Undersampled clusters lose power; topology biologically uninterpretable for novel events
Splice-graphGraph nodes (non-overlapping exonic regions)Whippet, ShibaWhippet maintenance status uncertain since ~2022; complex multi-exon graphs
Coverage-aware (IR)Intron body coverage + flanking junctionsIRFinder-S, S-IRFindeR, iREADConfounded by overlapping exons, repeats, low mappability regions
Isoform-basedTranscript abundance via EMSalmon/kallisto + tximportSalmon EM uncertainty propagates; many similar isoforms (TTN, MAPT) become indistinguishable

The agent's first decision is which family the question requires, not which tool. Switching family is the response to a tool failing within a family — switching tool within a family rarely fixes systematic blind spots.

Event Taxonomy (Beyond Standard SE/A5SS/A3SS/MXE/RI)

ClassCodeBiologyDetection caveat
Skipped exon (cassette)SEDefault cassette-exon ASMost common; well-handled by all tools
Alternative 5' splice siteA5SSAlternative donor; intron 5' end variesSign convention tool-specific (see below)
Alternative 3' splice siteA3SSAlternative acceptor; intron 3' end variesSensitive to BPS cancer mutations (SF3B1) — cryptic 3'ss ~10-30nt upstream
Mutually exclusive exonsMXETwo exons paired, one includedTool implementations vary; verify which "form 1" is yours
Retained intronRIWhole intron retained in mature mRNAJunction-only quant systematically underdetects; needs IRFinder-S
Microexon(sub-SE)3-27 nt; neural-enriched, SRRM4-regulatedMissed by default aligner anchor lengths (>=20-30 nt); needs VAST-TOOLS, MicroExonator, or long-read
Exitronn/aIntronic region within an annotated CDS exonMis-classified as A5SS/A3SS by most tools; use ExitronFinder, ScanExitron
Alternative first exon (AFE)AFEAlternative TSS / promoter usePromoter-driven, NOT spliceosomal; confirm with FANTOM CAGE before reporting as splicing
Alternative last exon (ALE)ALEAlternative cleavage/polyadenylationAPA-driven, NOT spliceosomal; confirm with 3'-end-seq
Detained intron (DI)(RI subtype)Nuclear-retained on mature mRNA, regulated by Clk kinasesDistinct from cytoplasmic NMD-targeted RI (Boutz 2015 Genes Dev); requires fractionation to confirm
Recursive splicing(intron subtype)Long introns >50kb spliced via internal ratchet pointsSibley 2015 Nature; only detectable with long-read or nascent RNA-seq

Tool-agnostic taxonomy reference: Wang 2008 Nature; Vaquero-Garcia 2016 eLife (LSV); Tapial 2017 Genome Res (VastDB).

Tool Selection Matrix

ToolBest forInputStrengthsFails when
rMATS-turboStandard SE/A5SS/A3SS/MXE/RI in annotated organism, n>=3BAM + GTFFast, well-calibrated at n>=3, novel SS supportJunction read imbalance; novel multi-junction events; underdetects RI and microexons
SUPPA2Quick PSI from existing TPM, pilot analysisSalmon/kallisto TPM + GTFNo alignment; fastestAnnotation-bound; high FDR (15-30%) at n<=2 vs n<=2
MAJIQ V3Complex events, heterogeneous cohorts (HET)BAM + GFF3Bayesian posterior PSI, complete LSV semanticsHigh memory (~50+ GB on cohorts); academic license; complex LSV interpretation needs care
leafcutterNovel junction discovery, sQTL, low-memory environmentsBAM (regtools junctions)Annotation-free; ~400 MB memory; SOTA for unannotated organismsSensitive to read depth; cluster topology arbitrary for complex multi-junction events
Shiba (2025)Low-coverage / few-replicate designs; junction imbalance correctionBAM + GTFBest calibration in own benchmark at n=2 vs n=2New (2025); limited community calibration
VAST-TOOLSCross-species comparative AS, microexonsFASTQVastDB orthology, ExOrthist co-toolLimited to species in VastDB
WhippetLaptop-scale exploratoryFASTQFast splice-graph-based PSIReduced active development since ~2022; underperforms on complex topologies
IRFinder-SIntron retention specificallyFASTQCoverage + junction integration; CNN-based artifact filteringIR-only; not for cassette events
S-IRFindeRReplicate-stable IR ratioBAMStable IR ratio metricIR-only; less integrated than IRFinder-S

Methodology evolves; verify benchmarks (Olofsson 2023 Biochem Biophys Res Commun; Kubota 2025 NAR; Tran 2025 WIREs RNA) and tool docs before committing. Default 2026 recommendation: run rMATS-turbo + leafcutter and reconcile; add MAJIQ V3 for complex events / heterogeneous cohorts; switch to Shiba for n=2 vs n=2.

PSI Definition and Effective Length Normalization

For a cassette exon, naive PSI ignores that the inclusion isoform contains more positions where a junction read can map than the skipping isoform. rMATS reports IncFormLen (= 2*(read_length - anchor) for the two flanking junctions, plus exon body bases) and SkipFormLen (= read_length - anchor) and computes:

PSI = (IJC / IncFormLen) / (IJC / IncFormLen + SJC / SkipFormLen)

where IJC = inclusion junction counts, SJC = skipping junction counts. Skipping this normalization biases PSI by ~10-30% depending on read length and exon size. SUPPA2 derives PSI from transcript TPMs, which the upstream Salmon/kallisto already accounts for. leafcutter operates on intron usage proportions within a cluster (a different statistic).

For long-read data, every read carries full isoform identity — effective-length normalization becomes unnecessary because each read counts as one isoform.

Sign Conventions for Alternative Splice Sites

ToolA5SS interpretationA3SS interpretation"Inclusion" direction
rMATS"long" form = donor downstream of alternative donor"long" form = acceptor upstream of alternative acceptorPSI > 0 = more long form
SUPPA2Same as rMATS (long = inclusion of additional exon body)SamePSI > 0 = more long form
VAST-TOOLSEncoded in event ID (_D1 vs _D2)Encoded in event ID (_A1 vs _A2)Document the chosen reference
MAJIQPer-junction within LSV; explicit donor/acceptor naming in VOILASamePSI per junction in the LSV

Always record which alternative form ΔPSI > 0 corresponds to in publication-grade reporting. Confusion is the most common reviewer comment for AS papers.

rMATS-turbo Workflow

Goal: Quantify SE/A5SS/A3SS/MXE/RI events from BAMs aligned with STAR 2-pass.

Approach: Group BAMs by condition, run rMATS with --statoff for quantification only, then parse JC.txt files for per-replicate PSI.

bash
rmats.py \
    --b1 condition1_bams.txt \
    --b2 condition2_bams.txt \
    --gtf annotation.gtf \
    -t paired \
    --readLength 150 \
    --variable-read-length \
    --libType fr-firststrand \
    --nthread 8 \
    --od rmats_output \
    --tmp rmats_tmp \
    --novelSS \
    --statoff

Key flags: --novelSS discovers junctions absent from the GTF (recommended with STAR 2-pass output). --variable-read-length allows mixed read lengths in the cohort. --libType fr-firststrand matches Illumina TruSeq stranded; verify with RSeQC infer_experiment.py. --statoff is for quantification-only runs; omit for differential testing.

python
import pandas as pd

se_jc = pd.read_csv('rmats_output/SE.MATS.JC.txt', sep='\t')

inc_cols = [c for c in se_jc.columns if c.startswith('IncLevel')]
se_jc['mean_PSI'] = se_jc[inc_cols].mean(axis=1)

per_rep_inc = se_jc['IJC_SAMPLE_1'].str.split(',').apply(lambda x: list(map(int, x)))
per_rep_skip = se_jc['SJC_SAMPLE_1'].str.split(',').apply(lambda x: list(map(int, x)))
min_inc = per_rep_inc.apply(min)
min_skip = per_rep_skip.apply(min)

reliable = se_jc[(min_inc + min_skip) >= 20]
JC vs JCEC files

SE.MATS.JC.txt uses only junction-spanning reads. SE.MATS.JCEC.txt adds reads contained within the alternative exon body as inclusion evidence.

  • Prefer JC for clean cassette-exon analysis when reads spanning junctions are sufficient.
  • Use JCEC when alternative exons are short (<50nt) and junction-spanning reads are scarce.
  • Avoid JCEC when intron retention overlaps the alternative exon body — exon-body reads may come from retained introns, not inclusion isoform.

SUPPA2 Workflow

Goal: Compute event PSI from transcript TPM without alignment; useful when Salmon/kallisto TPMs already exist.

Approach: Generate IOE event definitions from GTF, then aggregate TPMs of transcripts including/excluding each event.

bash
suppa.py generateEvents -i annotation.gtf -o events -f ioe -e SE SS MX RI FL

# -e takes {SE,SS,MX,RI,FL}; SS emits A5+A3 files, FL emits AF+AL files
for ev in SE A5 A3 MX RI AF AL; do
    suppa.py psiPerEvent -i events_${ev}_strict.ioe -e transcript_tpm.tsv -o psi_${ev}
done

SUPPA2 is annotation-bound: events absent from the GTF cannot be quantified. Whether an event is detected depends entirely on which transcripts the upstream Salmon/kallisto index contains. Use GENCODE comprehensive over basic when SUPPA2 detection sensitivity matters.

MAJIQ V3 Workflow

Goal: Quantify LSVs with Bayesian posterior PSI distributions; ideal for complex multi-junction events that don't fit canonical event types.

Approach: Build a splice graph from BAMs + GFF3, compute per-junction coverage with bootstrap, then run majiq psi for posterior PSI per LSV.

bash
majiq build annotation.gff3 -c settings.ini -j 8 -o build_output
majiq psi build_output/sample1.majiq build_output/sample2.majiq -j 4 -o psi_output -n condition_psi
voila view -p 5000 -j 8 build_output/splicegraph.zarr psi_output/condition_psi.psi.voila -o voila_output

MAJIQ V3 (Aicher, Slaff, Jewell, Barash bioRxiv 2024; public release 2025) replaced V2's SQLite splicegraph (splicegraph.sql) with Zarr storage (splicegraph.zarr); the .sql is deprecated. V3 is substantially faster than V2 via a rewritten xarray/zarr implementation with parallelized coverage calculations. LSV output includes posterior mean PSI plus the full posterior distribution; this enables threshold-based testing (e.g. P(|ΔPSI| > 0.2)) rather than frequentist p-values.

leafcutter Junction Quantification

Goal: Detect junctions and intron clusters annotation-free for downstream cluster-level usage.

Approach: Extract junctions per BAM with regtools, write filenames into a list, then cluster introns sharing splice sites.

bash
for bam in *.bam; do
    regtools junctions extract -a 8 -m 50 -s XS "$bam" -o "${bam%.bam}.junc"
done
ls *.junc > juncfiles.txt

python leafcutter_cluster_regtools.py \
    -j juncfiles.txt \
    -o leafcutter \
    -m 50 \
    -l 500000

-a 8 = 8nt anchor minimum (raise to 12 for stricter; lower to 6 for microexon-friendly). -m 50 = minimum junction reads per cluster. -l 500000 = max intron length (relevant for long brain-gene introns; raise for genes like DSCAM, ROBO2, ANK3).

Per-Tool Failure Modes

rMATS-turbo: Junction Read Imbalance

Trigger: A cassette exon's flanking exons have unequal read mapping opportunity (very short upstream exon, repeat-overlapping flanks, or low-mappability regions).

Mechanism: rMATS' binomial model treats inclusion vs skipping junctions as having equal mappability. When mappability differs between the two junction types, the PSI estimate is biased.

Symptom: "Significant" rMATS calls with no concordant change in leafcutter or MAJIQ at the same locus; ΔPSI direction inconsistent with sashimi-plot intuition.

Fix: Run Shiba (Kubota 2025 NAR) which corrects junction-imbalance, or filter rMATS hits requiring concordant detection in leafcutter.

SUPPA2: Sparse Empirical Null at Low Replicate Count

Trigger: n=2 vs n=2 (or n=3 vs n=2) design with --method empirical.

Mechanism: SUPPA2's empirical null is constructed from between-replicate ΔPSI distributions binned by transcript expression. With few replicates, the binned null is sparse and conservative-looking but actually under-calibrated.

Symptom: Inflated FDR (15-30% in benchmarks); many "significant" hits don't replicate or validate.

Fix: Switch to leafcutter or Shiba for n<=3 designs; or use --method classical (Wilcoxon) for very low replicate count; reconcile against orthogonal tool.

MAJIQ V3: Complex LSV Interpretation

Trigger: A gene with 4+ alternative splice sites at one node (e.g. one source, multiple acceptors).

Mechanism: A complete LSV at a single node lists all observed junctions; PSI is per-junction within the LSV, not "PSI of one event."

Symptom: Reporting "PSI of the gene" doesn't make sense; per-junction PSIs sum to 1 across the LSV but no single number represents the gene.

Fix: Use VOILA to visualize the LSV graph and identify which junction(s) shifted; for cassette-style reporting, derive equivalent PSI from sum of inclusion-junctions / total junctions in the LSV.

leafcutter: Cluster Topology Arbitrariness

Trigger: A cluster has 4+ introns sharing splice sites with non-canonical topology (e.g. mixed cassette + alternative donor + IR).

Mechanism: leafcutter clusters introns by shared splice sites; complex topologies don't map onto SE/A5SS/A3SS taxonomy and cluster-level "ΔPSI" hides which intron drove the change.

Symptom: Significant cluster-level p-value but multiple introns showing different effect-size directions.

Fix: Inspect the cluster in leafviz; report per-intron effect sizes (effect_sizes.txt); for canonical event reporting, map to SE/A5SS/A3SS via flanking exon coordinates manually.

Reconciliation: When rMATS and leafcutter Disagree

The two most common short-read tools answer slightly different questions: rMATS classifies on annotated event templates; leafcutter classifies on observed cluster usage. Disagreement is informative.

PatternLikely causeAction
rMATS sig, leafcutter not sigrMATS junction read imbalance OR rMATS event hits an annotation that leafcutter clustered differentlyInspect locus in IGV; check Shiba
leafcutter sig, rMATS not sigNovel junction not in rMATS annotation; rMATS --novelSS may have missed itCheck --novelSS was on; rerun if not
Both sig, opposite ΔPSI directionEvent class mismatch (e.g. rMATS calls SE positive but leafcutter sees A5SS shift in same cluster)Manually map cluster topology to event class
Both sig, same directionHigh-confidence callReport; cross-validate with sashimi-plot

Operational rule: for high-confidence reporting, require concordant detection in two tools from different algorithmic families (event-based + cluster-based, or LSV + isoform-based).

Show full SKILL.md (1,194 more words)Show less

Intron Retention: Canonical vs Detained vs Co-Transcriptional Unspliced

Three biologically distinct states all called "IR" by generic tools:

  1. Canonical RI (cytoplasmic, NMD-substrate often): mature polyadenylated mRNA carries the intron; usually PTC-bearing and NMD-targeted, sometimes encoding an alternative protein.
  2. Detained intron (DI) (Boutz 2015 Genes Dev): nuclear-localized, mature transcripts retaining a specific intron; a regulated reservoir released into translation upon signaling.
  3. Co-transcriptional unspliced: nascent pre-mRNA captured before splicing complete; not a regulated state.

Library prep determines which state(s) are visible:

  • Poly(A) selection: enriches (1), depletes (2)/(3)
  • rRNA depletion (cytoplasmic): captures (1)
  • rRNA depletion (whole cell or nuclear): captures all three

To distinguish DI from canonical RI: subcellular fractionation (nuclear vs cytoplasmic RNA-seq), or NMD inhibitor (cycloheximide, NMDi-14) treatment — canonical RI mRNA increases under NMD inhibition; DI does not.

bash
IRFinder FastQ -r REF/ -d ir_output sample.fastq

IRFinder-S (Lorenzi 2021 Genome Biol) uses CNN-based filtering of true IR vs noise; current SOTA for IR analysis. iREAD and S-IRFindeR (Broseus & Ritchie 2020 bioRxiv) are alternatives.

Microexon Detection

Microexons (3-27 nt, neural-enriched, SRRM4-regulated; Irimia 2014 Cell) are missed by default short-read aligners requiring 20-30 nt anchors. Options:

ApproachToolNotes
Curated database lookupVAST-TOOLS + VastDBCross-species, microexon-aware (Tapial 2017 Genome Res)
De novo discoveryMicroExonator (Parada 2021 Genome Biol)Snakemake pipeline
Tune the upstream alignerSTAR --alignSJoverhangMin 6 --alignSJDBoverhangMin 1 --outFilterMismatchNoverReadLmax 0.04rMATS itself cannot recover microexons that STAR didn't pass through; lower DB-junction overhang to 1 (trusts annotated microexon coords) and combine with strict mismatch filter. Typical AS pipelines use STAR 8/3 which is too strict for microexons
Long-read sequencingPacBio Iso-Seq, ONTSolves the problem entirely; reads span microexons fully

For brain / neural tissue or autism-spectrum studies, microexon analysis must be explicit — default short-read pipelines underdetect them by ~70%.

Quality Thresholds

MetricThresholdSource / Rationale
Junction reads per replicate>=10-20 (per-replicate minimum)Empirical PSI variance becomes <0.05 above this; below, PSI becomes a coin flip
PSI dynamic rangemean PSI 0.05-0.95Outside is near-constitutive; rMATS, SUPPA2 default filters drop these
Missing values<50% of samplesHigher missingness indicates low expression — re-test with subset
Read length>=75nt PE preferred; >=100nt for microexonsShorter single-end reads bias junction detection toward shorter exons (convention)
LibraryrRNA depletion for IR analysis; poly(A) acceptable for cassettepoly(A) selection loses pre-mRNA/intronic signal (convention)
STAR 2-passCohort-style preferred over per-sample basicVeeneman 2016 Bioinformatics: >=94% novel junction recovery
MAJIQ minreads / minpos--minreads 10 --minpos 3Default; lower for low-coverage
leafcutter -m50 reads per clusterHigher for rare events; lower for sQTL discovery
Anchor length>=8 nt for short-readBelow this, false-positive junctions dominate (CIGAR-N noise)

Decision Tree by Scenario

ScenarioRecommended tool(s)Why
Standard cassette analysis, n>=3, GENCODE-annotatedrMATS-turbo + leafcutter (concordance)Default workflow; complementary algorithmic families
Non-model organism, no GENCODE-grade annotationleafcutter + de novo discoveryAnnotation-free
Heterogeneous cohort, n>=10 vs n>=10 (clinical, GTEx-style)MAJIQ V3 with HET moduleHET designed for between-sample variability dominance
Low coverage / few replicates (n=2 vs n=2)ShibaJunction-imbalance correction; SOTA at low coverage in 2025 benchmarks
Cross-species comparative (vertebrate panel)VAST-TOOLS + VastDBOrthology-aware events; ExOrthist co-tool
TPM-only available (no BAMs)SUPPA2Annotation-bound but fast
Microexon focus (neural / ASD)VAST-TOOLS or MicroExonatorDefault tools systematically miss microexons
Intron retention focusIRFinder-S (rRNA-depleted library)Coverage-aware; CNN artifact filter
Detained introns specificallyIRFinder-S + nuclear/cytoplasmic fractionationRequired to separate DI from cytoplasmic RI
Long reads availablerMATS-long, FLAIR, IsoQuantFull-isoform resolution; see long-read-splicing
Single-cell (full-length plate)MARVEL, BRIE2See single-cell-splicing
Single-cell (10X 3')Likely don't attempt; consider Sierra for APA10X 3' chemistry insufficient for AS

Common Errors

ErrorCauseSolution
error: GTF gene_id parsing (rMATS)rMATS expects GENCODE-style gene_id; some Ensembl GTFs use different attribute ordergffread input.gff3 -T -o standardized.gtf
KeyError: 'IJC_SAMPLE_1' (rMATS parsing)Output column missing; sometimes occurs when --statoff combined with novel events on older versionsUpdate rMATS-turbo to >=4.3.x; re-run
MAJIQ: too few reads at junctionDefault --minreads 10 --minpos 3 filters out the locusLower thresholds for low-coverage data; document filtering
leafcutter: dispersion estimation failedCluster has all-zero counts in one groupPre-filter with leafcutter_ds.R --min_samples_per_group 3 --min_samples_per_intron 5
SUPPA2: empirical p computed on N=4 nullsInsufficient replicates for empirical modeSwitch --method classical (Wilcoxon) for very low replicate count
regtools: invalid CIGARNon-BAM-spec read in inputFilter with samtools view -h -F 0x100 -F 0x800 (drop secondary/supplementary)
STAR: too many SJs (in pass 2)Cohort SJ.out.tab too largeFilter to junctions seen in >=3 samples or with >=3 unique reads before merging

Output Interpretation

PSI ranges 0 to 1: 1 = always included, 0 = always skipped, 0.5 = balanced. Sign of IncLevelDifference matches --b1 minus --b2 group order — always document which is which in publications.

NMD direction matters: an increase in PSI of a poison exon (PTC-introducing) decreases functional protein due to NMD. Always check whether the alternative form is PTC-bearing using ORF-aware annotation (IsoformSwitchAnalyzeR consequences, or manual stop-codon distance check vs last exon-exon junction).

Disease signatures:

  • SF3B1 mutations (MDS, CLL, uveal melanoma): cryptic 3'ss ~10-30 nt upstream of canonical (Darman 2015 Cell Rep). Look for clustered A3SS hits.
  • U2AF1 mutations (lung adeno, MDS): altered preferences at 3'ss -3 position; cassette-exon shifts.
  • TDP-43 loss (ALS/FTD): de novo cryptic exons in UNC13A, STMN2, ATG4B (Brown 2022 Nature; Klim 2019 Nat Neurosci) — annotation-free tools required (leafcutter denovo).

Common Pitfalls

  • Treating AFE/ALE as splicing — these are typically promoter-driven (AFE) or APA-driven (ALE), not spliceosomal. Confirm with FANTOM CAGE or 3'-end-seq.
  • Confusing detained introns with NMD-targeted RI — both call as "IR" but have opposite biological fates.
  • Using poly(A) libraries for IR analysis — biases toward mature transcripts, depletes pre-mRNA.
  • Single-end short reads — junction-spanning reads need >=8nt overhang on both sides; biases toward shorter exons.
  • Quoting "PSI of the gene" from MAJIQ LSV output — only per-junction PSI within an LSV is meaningful.
  • Skipping STAR 2-pass — loses ~14% of novel junctions; matters for any non-canonical organism or condition.
  • Trusting rMATS calls without --novelSS when STAR 2-pass found new junctions — rMATS will only quantify pre-annotated events.
  • differential-splicing - Compare PSI between conditions; use the same upstream alignment but switch to with-stat tools
  • splicing-qc - Run BEFORE quantification to verify library, depth, strandedness, alignment quality
  • isoform-switching - DTU framework with NMD/ORF/domain consequences; complementary to event-level PSI
  • sashimi-plots - Visualize specific events for QC and reporting; concordance check across tools
  • splice-variant-prediction - SpliceAI/Pangolin for variant impact predictions to test against PSI changes
  • long-read-splicing - Full-isoform PSI without anchor-length limits; preferred for microexons and complex isoforms
  • read-alignment/star-alignment - STAR 2-pass cohort-style alignment is required upstream
  • rna-quantification/alignment-free-quant - Salmon/kallisto TPM is required for SUPPA2

References

  • Wang et al 2008 Nature - AS event taxonomy
  • Vaquero-Garcia et al 2016 eLife - MAJIQ LSV framework
  • Trincado et al 2018 Genome Biol - SUPPA2
  • Li et al 2018 Nat Genet - leafcutter
  • Tapial et al 2017 Genome Res - VAST-TOOLS / VastDB
  • Wang et al 2024 Nat Protoc - rMATS-turbo
  • Aicher, Slaff, Jewell, Barash 2024 bioRxiv - MAJIQ V3
  • Kubota et al 2025 NAR - Shiba
  • Lorenzi et al 2021 Genome Biol - IRFinder-S
  • Boutz et al 2015 Genes Dev - detained introns
  • Irimia et al 2014 Cell - SRRM4 microexons
  • Darman et al 2015 Cell Rep - SF3B1 cryptic 3'ss
  • Olofsson et al 2023 Biochem Biophys Res Commun 653:31-37 - benchmark
  • Tran et al 2025 WIREs RNA - methodology review
  • Brown et al 2022 Nature - UNC13A cryptic exon (TDP-43)
  • Klim et al 2019 Nat Neurosci - STMN2 cryptic splicing
  • Veeneman et al 2016 Bioinformatics - STAR 2-pass benchmark

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in alternative-splicing/splicing-quantification of GPTomics/bioSkills.

  • SKILL.md
  • examples/quantify_splicing.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Splicing Quantification next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Splicing Quantification compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Splicing Quantification this skillGPTomics/bioSkills1.2k2 repos~6.8kAutomated safety check: PassMIT
Tooluniverse Rnaseq Deseq2wu-yc/LabClaw1.1k2 repos~4.5kAutomated safety check: PassNone
Tooluniverse Metabolomics Analysiswu-yc/LabClaw1.1k2 repos~5.9kAutomated safety check: PassNone
Bio Single Cell PreprocessingFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2.4kAutomated safety check: PassNone
Bio Geo Datamajiayu000/claude-skill-registry6663 repos~4.4kAutomated safety check: PassMIT
Gene Protein Expression Matrix Normalizationaipoch/medical-research-skills2k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Production-ready RNA-seq differential expression analysis using PyDESeq2.

    1.1k GitHub starsUsed in 2 repos~4.5k tokens
    Research & ScienceAuto-check passed
  • Analyze metabolomics data including metabolite identification, quantification, pathway analysis, and metabolic flux.

    1.1k GitHub starsUsed in 2 repos~5.9k tokens
    Research & ScienceAuto-check passed
  • Bio Single Cell Preprocessing

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control, filtering, and normalization for single-cell RNA-seq using Seurat (R) and Scanpy (Python).

    3.1k GitHub starsUsed in 1 repo~2.4k tokens
    Research & ScienceAuto-check passed
  • Bio Geo Data

    majiayu000/claude-skill-registry

    Query and download from NCBI Gene Expression Omnibus (GEO) and EMBL-EBI's BioStudies/ArrayExpress mirror.

    666 GitHub starsUsed in 3 repos~4.4k tokens
    Research & ScienceAuto-check passed
  • Gene Protein Expression Matrix Normalization

    aipoch/medical-research-skills

    A skill your agent uses when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory…

    2k GitHub stars~1.5k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed
  • Bio Crispr Screens Screen Qc

    majiayu000/claude-skill-registry

    Quality control for pooled CRISPR screens covering library representation, Gini index, log-skew, replicate Pearson and Spearman concordance, essentialome precision-recall AUC against CEGv2 (Hart…

    666 GitHub starsUsed in 3 repos~5.9k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Splicing Quantification

What does Bio Splicing Quantification do?

Quantifies alternative splicing as PSI (percent spliced in) from RNA-seq using rMATS-turbo (BAM-based event), SUPPA2 (TPM-based event), MAJIQ V3 (LSV-based Bayesian), leafcutter (annotation-free…. Bio Splicing Quantification is an agent skill from GPTomics/bioSkills. Quantifies alternative splicing as PSI (percent spliced in) from RNA-seq using rMATS-turbo (BAM-based event), SUPPA2 (TPM-based event), MAJIQ V3 (LSV-based Bayesian), leafcutter (annotation-free intron clusters), VAST-TOOLS (cross-species with microexon support), Shiba (junction-imbalance-corrected, 2025 SOTA at low coverage), or IRFinder-S (intron retention coverage-aware).

When should I use Bio Splicing Quantification?

Bio Splicing Quantification fits situations like: measuring splice-site usage; isoform inclusion ratios from short-read RNA-seq.

How do I install Bio Splicing Quantification in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-splicing-quantification -a claude-code`. Or copy the skill folder (alternative-splicing/splicing-quantification in GPTomics/bioSkills) into .claude/skills/bio-splicing-quantification in your project. Claude Code loads it when a task matches its description.

How do I install Bio Splicing Quantification in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-splicing-quantification -a codex`. Or copy the skill folder (alternative-splicing/splicing-quantification in GPTomics/bioSkills) into .agents/skills/bio-splicing-quantification in your project. Codex loads it when a task matches its description.

Can I use Bio Splicing Quantification in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-splicing-quantification -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-splicing-quantification, .gemini/skills/bio-splicing-quantification, .github/skills/bio-splicing-quantification and .opencode/skills/bio-splicing-quantification in your project.

What does Bio Splicing Quantification need to run?

Going by SKILL.md and its folder, Bio Splicing Quantification needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3.

Does Bio Splicing Quantification access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Splicing Quantification safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Splicing Quantification use?

Bio Splicing Quantification is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Splicing Quantification use?

About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Splicing Quantification?

Skills that share tags, products or a category with Bio Splicing Quantification: Tooluniverse Rnaseq Deseq2 (wu-yc/LabClaw, 1.1k stars), Tooluniverse Metabolomics Analysis (wu-yc/LabClaw, 1.1k stars), Bio Single Cell Preprocessing (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars) and Bio Geo Data (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Splicing Quantification?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.