Agent skill

Bio Single Cell Splicing

by GPTomics in GPTomics/bioSkills

Analyzes alternative splicing at single-cell resolution. An agent skill from GPTomics/bioSkills.

MITAuto-check passedResearch & Science

Install Bio Single Cell Splicing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-splicing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-single-cell-splicing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alternative-splicing/single-cell-splicing .claude/skills/bio-single-cell-splicing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-single-cell-splicing
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.5k tokens
SKILL.md length
2,535 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Analyzes alternative splicing at single-cell resolution. An agent skill from GPTomics/bioSkills.

  • Works in 3 steps: 3' enrichment: median fragment <1 kb… → Short R2 (~91 nt): each read straddles… → PCR concatemers and TSO artifacts:…
  • Analyzing isoform usage in scRNA-seq
  • SKILL.md covers Version Compatibility, The 10X 3' Problem (Quantified), Decision: Does the Chemistry… and Tool Selection Matrix, plus 17 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Single Cell Splicing is an agent skill from GPTomics/bioSkills. Analyzes alternative splicing at single-cell resolution. The first decision is library chemistry — 10X 3' is fundamentally limited (RT primes from poly-A, R2 falls in 3' UTR, <0.1 junction read per cell per AS event). Plate-based full-length methods (Smart-seq3, FLASH-seq, VASA-seq, STORM-seq) and single-cell long-read (MAS-Iso-seq, scISOr-Seq2) are the chemistries that give per-cell isoform structure. Tools include MARVEL (R, Smart-seq integrated), BRIE2 (Bayesian PSI with regulatory features and ELBOgain test)…

Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/sc_splicing_brie2.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Analyzing isoform usage in scRNA-seq
  • Identifying cell-type-specific splicing
  • Determining whether scRNA-seq chemistry supports splicing analysis at all

Example prompts

  • “is fundamentally limited (RT primes from poly-A, R2 falls in 3”
  • “Use the bio-single-cell-splicing skill to analyz alternative splicing at single-cell resolution. An agent skill from GPTomics/bioSkills”
  • “/bio-single-cell-splicing”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. 3' enrichment: median fragment <1 kb from poly(A); >70% of unique reads fall within 3' UTR.
  2. Short R2 (~91 nt): each read straddles at most one junction; usually none, because R2 lands in 3' UTR.
  3. PCR concatemers and TSO artifacts: pollute junction detection; UMI collapse is gene-level, not isoform-level.

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Single Cell Splicing loads about 6.5k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 2,535 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~225
When it runs · the whole SKILL.md, loaded when a task matches
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,535 words, ~6,530 tokens.

Download SKILL.mdSave it as .claude/skills/bio-single-cell-splicing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-single-cell-splicing
description
Analyzes alternative splicing at single-cell resolution. The first decision is library chemistry — 10X 3' is fundamentally limited (RT primes from poly-A, R2 falls in 3' UTR, <0.1 junction read per cell per AS event). Plate-based full-length methods (Smart-seq3, FLASH-seq, VASA-seq, STORM-seq) and single-cell long-read (MAS-Iso-seq, scISOr-Seq2) are the chemistries that give per-cell isoform structure. Tools include MARVEL (R, Smart-seq integrated), BRIE2 (Bayesian PSI with regulatory features and ELBO_gain test), scQuint (junction-cluster, plate-based; not for 10X), SpliZ (annotation-free Z-score), Psix (graph-smoothness regulated AS), and Sierra (alternative polyadenylation, often confused with AS). Use when analyzing isoform usage in scRNA-seq, identifying cell-type-specific splicing, or determining whether scRNA-seq chemistry supports splicing analysis at all.
tool_type
python
primary_tool
MARVEL

Version Compatibility

Reference examples tested with: MARVEL 2.0+, BRIE2 0.2.4+, scQuint 0.1+, SpliZ 0.0.1+, Sierra 1.0+, Psix 0.1+, anndata 0.10+, scanpy 1.10+, pandas 2.2+, scipy 1.13+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Single-Cell Splicing Analysis

The fundamental decision is chemistry, not tool. Most droplet 3' scRNA-seq cannot support transcriptome-wide splicing inference because reverse transcription primes from the poly(A) tail and most reads land in the 3' UTR — far from CDS-region splicing events. Plate-based full-length methods and single-cell long-read sequencing are the chemistries that give per-cell isoform structure across the gene body.

The 10X 3' Problem (Quantified)

Three compounding mechanisms make 10X Chromium 3' (v3.1, GEM-X, v4) hostile to splicing:

  1. 3' enrichment: median fragment <1 kb from poly(A); >70% of unique reads fall within 3' UTR.
  2. Short R2 (~91 nt): each read straddles at most one junction; usually none, because R2 lands in 3' UTR.
  3. PCR concatemers and TSO artifacts: pollute junction detection; UMI collapse is gene-level, not isoform-level.

Quantitative estimate: Only a small fraction of cassette exons sit close enough to the polyA site to be sampled by 3' chemistry (empirical estimates from APA/3'-end atlases — see Tian & Manley 2017 Nat Rev Mol Cell Biol for the 3' UTR isoform landscape). Effective junction read yield from 10X 3' is <0.1 per cell per AS event — vs the 5-10 needed for stable per-cell PSI. Most splicing analyses on 10X 3' data report artifacts.

The 5' kit (10X 5' GEX) does not solve this — it shifts capture from 3' UTR to 5' UTR / TSS-proximal regions. Marginal improvement; not a transcriptome-wide solution. Note that V(D)J recovery requires the 10X Chromium Single Cell Immune Profiling kit (with TCR/BCR-specific enrichment), not 5' GEX alone — postdocs designing immune-repertoire experiments must use the dedicated V(D)J kit.

Decision: Does the Chemistry Support Splicing Analysis?

ChemistrySplicing analysis viable?Best alternative if no
10X 3' (Chromium v3, GEM-X, v4, Flex)No (transcriptome-wide); maybe near-3'-end eventsSierra for APA
10X 5' GEXLimited; near-5'-end events onlySierra for alternative TSS; switch to MAS-Iso-seq
Smart-seq2Yes (full transcript)MARVEL or BRIE2
Smart-seq3 / Smart-seq3xpressYes + UMI molecule countingMARVEL or BRIE2
FLASH-seqYes (faster, cheaper Smart-seq3)MARVEL or BRIE2
VASA-seqYes + total RNA (incl. nascent, IR)MARVEL with IR analysis
STORM-seqYes + total RNA + ribodepletionMARVEL with IR analysis
MAS-Iso-seq + 10X 5' (PacBio Kinnex)Yes — full isoforms per cellFLAMES, scNanoGPS, IsoQuant, see long-read-splicing
scISOr-Seq2 (PacBio + 10X)Yes — full isoforms with cell-typingFLAMES, IsoQuant
ONT direct cDNA scRNAYesFLAMES
ONT direct RNA scRNAYes + native modificationsFLAMES

Tool Selection Matrix

ToolBest forInputStrengthsFails when
MARVELSmart-seq plate-based and (v2+) 10X droplet unified workflowPlate or droplet BAMs + SeuratSE/A5SS/A3SS/MXE/RI/AFE/ALE; modality classification; native Seurat integration; v2 droplet supportR-only
BRIE2Plate-based with regulatory feature priorPlate BAM + GFF3 eventsBayesian variational PSI + ELBO_gain test; principled uncertainty; CLI-driven (brie-count, brie-quant)TensorFlow dependency; slow at scale
scQuintPlate-based annotation-free junction-cluster quantification (validated on Smart-seq2)STAR junctions across cellsCluster-level junction usage; latent DirichletAuthors recommend AGAINST use on 10X 3'/5' data (3'-bias confounds); plate-based only
SpliZAnnotation-free discovery of cell-state-associated splicingSTAR-aligned BAMsPer-gene Z-score; no event database neededAnnotation-free = power tradeoff
PsixRegulated AS along trajectoriesPSI matrix + kNN graphTests graph smoothness; robust to dropoutNeeds cell-state graph upstream
SierraAPA in 10X 3' (NOT splicing)10X BAM + GTFPeak-calling 3' ends; DEXSeq DTU on UTR isoformsAPA only; not for cassette exons
pseudobulk leafcutter / rMATSBetween-cell-type differential splicingAggregated BAMsBulk-level statistical powerLoses within-cluster heterogeneity
MAS-Iso-seq + FLAMESFull-length single-cell isoforms10X 5' + PacBio KinnexFull isoforms per cell at scaleCost; complex pipeline

Decision Tree by Goal

GoalRecommended approach
"Will my 10X 3' data support splicing?"No transcriptome-wide; consider Sierra for APA. Note: scQuint authors recommend against use on 10X data
Cassette exon analysis in cell types from Smart-seq2MARVEL with ComputePSI + AssignModality + CompareValues
Discover cell-state-associated splicing without an event databaseSpliZ
Test regulated AS along developmental pseudotimePsix
Per-cell PSI with uncertainty in low-coverage cellsBRIE2
Differential splicing between two well-defined cell typesPseudobulk leafcutter or rMATS on aggregated BAMs
APA (alternative polyadenylation, often confused with AS)Sierra
Full-length single-cell isoforms at scaleMAS-Iso-seq + FLAMES (long-read)
Microexons (3-27 nt)Long-read or aligner with low overhang (uLTRA, deSALT)
snRNA-seq (nuclei) — IR questionLibrary captures nuclear RNA enriched for incomplete splicing — interpret IR cautiously

MARVEL Plate-Based Workflow

Goal: Run a unified workflow from STAR junctions to cell-type-specific splicing calls.

Approach: Build a wide splice-junction count matrix (rows = junctions keyed by coord.intron, columns = cells), assemble per-event feature tables, then construct MARVEL object with named slots (SpliceJunction, SplicePheno, SpliceFeature, IntronCounts, GeneFeature, Exp, GTF). Quantify PSI per event class, classify modality, test differential splicing.

r
library(MARVEL); library(Seurat); library(data.table)

seurat_obj <- readRDS('cells.rds')

# Build wide SJ matrix: first column 'coord.intron' (e.g. 'chr1:100007082:100022621'),
# subsequent columns are per-cell sample IDs with junction counts as values.
# This is constructed from STAR SJ.out.tab files (one per cell) merged on intron coord.
sj_files <- list.files('star_pass2/', pattern='SJ.out.tab$', full.names=TRUE)
sj_long <- rbindlist(lapply(sj_files, function(f) {
    d <- fread(f, sep='\t', header=FALSE,
               col.names=c('chr','start','end','strand','motif','annot','unique','multi','overhang'))
    d$coord.intron <- paste(d$chr, d$start, d$end, sep=':')
    d$sample <- gsub('_SJ.out.tab$', '', basename(f))
    d[, .(coord.intron, sample, unique)]
}))
sj <- dcast(sj_long, coord.intron ~ sample, value.var='unique', fill=0)

# SpliceFeature is a NAMED LIST keyed by event class
df.feature.list <- list(
    SE   = read.table('events_SE.txt',   header=TRUE, sep='\t'),
    A5SS = read.table('events_A5SS.txt', header=TRUE, sep='\t'),
    A3SS = read.table('events_A3SS.txt', header=TRUE, sep='\t'),
    MXE  = read.table('events_MXE.txt',  header=TRUE, sep='\t'),
    RI   = read.table('events_RI.txt',   header=TRUE, sep='\t')
)

# SplicePheno: per-cell metadata; sample.id column maps to SpliceJunction column names
df.pheno <- seurat_obj@meta.data
df.pheno$sample.id <- rownames(df.pheno)

marvel <- CreateMarvelObject(
    SpliceJunction = sj,
    SplicePheno    = df.pheno,
    SpliceFeature  = df.feature.list,
    GeneFeature    = read.table('gene_features.tsv', header=TRUE, sep='\t'),
    Exp            = read.table('tpm.tsv', header=TRUE, sep='\t', row.names=1),
    GTF            = rtracklayer::import('annotation.gtf')
)

marvel <- ComputePSI(marvel, CoverageThreshold=10, EventType='SE')
marvel <- AssignModality(marvel, EventType='SE')
marvel <- CompareValues(
    marvel,
    cell.group.g1 = neurons, cell.group.g2 = glia,
    method = 'wilcox', n.cells = 25, psi.delta = 0.1
)

For 10X droplet data, MARVEL v2+ provides CreateMarvelObject.10x() and AnnotateSJ.10x() constructors. Verify the exact API via ?CreateMarvelObject.10x in installed MARVEL.

MARVEL classifies events into modalities (Song 2017 Mol Cell): included (PSI1), excluded (PSI0), bimodal (mixture at 0/1), middle (peaked ~0.5), multimodal. Bimodality usually reflects mixed cell states or stochastic monoallelic-like bursting. Mid-modality (peaked at 0.5) can be technical (mixed cells in a droplet) — confirm with full-length data.

BRIE2 Bayesian PSI

Goal: Estimate per-cell PSI with informative regulatory-feature prior; test cell-state association via likelihood-ratio testing on covariate effects.

Approach: BRIE2 is a CLI-driven workflow (brie-count for read counting, brie-quant for variational inference + LRT). Prepare a GFF3 of splicing events, count cell-barcoded junction reads, then fit the model with covariate testing.

bash
# 1. Count splicing events per cell
brie-count \
    -a splicing_events.gff3 \
    -S sample_list.tsv \
    -o brie_counts/ \
    -p 16

# 2. Fit BRIE2 with LRT against the cell-type covariate
brie-quant \
    -i brie_counts/brie_count.h5ad \
    -c cell_metadata.tsv \
    -o brie_quant.h5ad \
    --interceptMode gene \
    --LRTindex All \
    --testBase null \
    --MCsize 3 \
    --batchSize 1000000 \
    -p 16

--interceptMode gene fits a gene-specific intercept (recommended); --LRTindex All tests all covariates; --testBase null uses the null model as the LRT reference. Verify exact flag set via brie-quant -h in installed BRIE2.

python
import scanpy as sc

adata_splice = sc.read_h5ad('brie_quant.h5ad')
# Per-event covariate effects, ELBO values, and LRT statistics live in
# adata_splice.varm and adata_splice.var; column names depend on BRIE2 version.
# Inspect with: print(adata_splice); print(adata_splice.varm.keys())
# Per-event significance is typically derived from LRT delta-ELBO.

BRIE2 (Huang & Sanguinetti 2021 Genome Biol) uses a sequence-derived feature prior (exon length, GC content, splice site strength, motif counts) to regularize PSI estimates in low-coverage cells. The LRT-based covariate test answers "is this event associated with cell state?" without requiring per-cell PSI accuracy. Threshold the delta-ELBO at ~3 (analogous to log-Bayes-factor); confirm against version-specific output keys via the brie-tutorials repo.

SpliZ for Annotation-Free Discovery

Goal: Identify splicing-defined cell populations without an event database.

Approach: Compute per-gene splicing Z-score across cells; test for cell-state association via permutation.

bash
# SpliZ is a Nextflow pipeline (not a standalone CLI). Configure inputs in a .config
# file (dataname, input_file, libraryType, grouping_level_1/2) - either SICILIAN
# output (SICILIAN=true) or BAMs via a samplesheet CSV + metadata + GTF (SICILIAN=false).
nextflow run salzmanlab/spliz -r main -latest -c spliz.config

SpliZ (Olivieri 2022 Nat Methods) is robust to dropout because it pools junction information across the gene; particularly useful for discovering splicing diversity in heterogeneous tumor samples.

Psix for Regulated AS Along Trajectories

Goal: Detect AS that varies coherently with cell state along a developmental trajectory, robust to dropout.

Approach: Score whether observed PSI is smooth on the cell-cell kNN graph from expression-space embedding.

python
import psix
import scanpy as sc

adata = sc.read_h5ad('cells.h5ad')
sc.pp.neighbors(adata, n_neighbors=30, use_rep='X_pca')

psix_obj = psix.Psix(adata, psi_matrix_path='psi_matrix.tsv')
psix_obj.run_psix()

regulated = psix_obj.psix_results.query('psix_score > 1.5 and pvalue < 0.05')

Psix (Buen Abad Najar 2022 Genome Res 32:1385) is the principled alternative to imputing PSI: do not impute (it obliterates heterogeneity); test for graph smoothness instead.

Sierra for APA (Not Splicing)

Goal: Detect alternative polyadenylation in 10X 3' data — frequently confounded with AS.

Approach: Peak-call read pile-ups at 3' ends, then DEXSeq-style DTU on 3' UTR isoforms.

r
library(Sierra)

peak_file <- FindPeaks(
    output.file = 'peaks.txt',
    gtf.file = 'annotation.gtf',
    bam.file = 'possorted_genome_bam.bam'
)

counts <- CountPeaks(
    peak.sites.file = 'peaks.txt',
    gtf.file = 'annotation.gtf',
    bamfile = 'possorted_genome_bam.bam',
    whitelist.file = 'barcodes.tsv'
)

# CountPeaks returns a peak x cell matrix; annotate it and build a peak Seurat
# object before differential-usage testing.
peak.annotations <- AnnotatePeaksFromGTF(
    peak.sites.file = 'peaks.txt',
    gtf.file = 'annotation.gtf',
    output.file = 'peak_annotations.txt'
)

peaks.seurat <- NewPeakSeurat(
    peak.data = counts,
    annot.info = peak.annotations,
    cell.idents = cell_identities
)

apa_results <- DUTest(peaks.seurat, population.1 = ctrl_cells, population.2 = trt_cells)

If only 10X 3' data is available, this is often what is actually wanted. Distinct UTRs change miRNA targeting, RBP binding, and stability — biologically meaningful but not splicing.

Pseudobulk for Statistical Power

Goal: Recover bulk-level statistical power for differential splicing between cell types.

Approach: Sum junction counts across cells of the same cluster, then run leafcutter / rMATS on aggregated counts.

python
import pandas as pd
import numpy as np

def pseudobulk_junctions(junction_counts, cell_metadata, groupby='cell_type'):
    out = {}
    for group, cells in cell_metadata.groupby(groupby).groups.items():
        mask = junction_counts.columns.isin(cells)
        out[group] = junction_counts.loc[:, mask].sum(axis=1)
    return pd.DataFrame(out)

Use pseudobulk for differential splicing between well-defined cell types; use per-cell methods for within-population heterogeneity (graded splicing along pseudotime, bimodal cell-state mixtures).

Single-Cell Long-Read = Future of Single-Cell Splicing

In 2024-2026, full-length single-cell long-read sequencing has become practical and is the recommended chemistry for splicing-focused single-cell experiments:

  • MAS-Iso-seq / PacBio Kinnex: concatenated full-length cDNA arrays, ~16x throughput vs plain Iso-Seq, compatible with 10X 5' libraries (Al'Khafaji 2024 Nat Biotech)
  • scISOr-Seq2: hybrid 10X + PacBio for cell typing + isoform structure (Joglekar et al 2024 Nat Neurosci 27:1051-1063, single-cell long-read brain isoform mapping)
  • ONT direct cDNA + 10X: lower cost, similar information content
  • FLAMES: barcode demultiplexing + isoform quantification + SNV calling for ONT scRNA (Tian 2021 Genome Biol 22:310)

For splicing-specific full-length single-cell analysis, see long-read-splicing skill.

Per-Tool Failure Modes

MARVEL: SpliceJunction Matrix Format

Trigger: Building the SpliceJunction matrix from STAR SJ.out.tab incorrectly (e.g. long-format instead of wide).

Mechanism: MARVEL plate-based CreateMarvelObject(SpliceJunction = ...) expects a wide matrix with first column coord.intron (formatted chr:start:end) and subsequent columns being per-cell sample IDs with integer junction counts. Long-format data.frames or missing coord.intron column cause runtime errors.

Symptom: "no coord.intron column found" errors; or empty PSI tables despite junction reads being present.

Fix: Verify wide-matrix structure; ensure SJ.out.tabs are merged on the chr:start:end key with cells as columns. Use data.table::dcast for the long->wide reshape.

BRIE2: TensorFlow Memory

Trigger: Large cohort (>10k cells) with deep coverage.

Mechanism: Variational inference loads full count matrix; TensorFlow allocates GPU memory aggressively.

Symptom: OOM kills; training stalls.

Fix: Reduce --batchSize from default (500000) to 100000 or 50000; train per-chromosome batch; use CPU mode for very small cohorts. Note flag is camelCase --batchSize, not --batch_size.

Show full SKILL.md (1,020 more words)Show less
scQuint: 3' Data Sparsity

Trigger: Running scQuint on 10X 3' v3 data hoping for splicing signal.

Mechanism: scQuint's latent Dirichlet model needs junction counts; 10X 3' yields too few junction reads to fit the model robustly.

Symptom: All cells assign to one cluster; no informative splicing signal.

Fix: Pivot to APA analysis with Sierra; or upgrade chemistry to MAS-Iso-seq.

Psix: Missing kNN Graph

Trigger: Running Psix without precomputed cell-cell graph.

Mechanism: Psix tests PSI smoothness on a pre-existing cell-cell graph; without one, no smoothness statistic.

Symptom: Empty results or error about missing connectivities.

Fix: Run sc.pp.neighbors(adata) before Psix; ensure connectivities is in adata.obsp.

Sierra: Annotation Gaps

Trigger: GTF missing 3'UTR annotations.

Mechanism: Sierra peak-calls within annotated 3'UTRs; missing annotations mean missed peaks.

Symptom: Few peaks detected; gene-level coverage but no APA calls.

Fix: Use comprehensive GENCODE annotation; or run de-novo peak calling first.

Reconciliation: When Single-Cell Tools Disagree

PatternLikely causeAction
MARVEL sig, BRIE2 notPer-cell PSI noise (BRIE2 conservative); MARVEL pseudobulk-likeTrust MARVEL for cell-type comparisons; BRIE2 for within-cluster
BRIE2 sig, MARVEL notCell-state effect smoother than cell-type boundaryTest along trajectory with Psix
SpliZ sig, MARVEL notAnnotation-free SpliZ catches novel eventsInvestigate junction structure manually
Sierra sig, MARVEL notSierra is APA, MARVEL is splicing — different biologyDistinguish in interpretation
Pseudobulk sig, per-cell notPower issue; effect averaged out per-cellReport at cluster level, not per-cell

Quantitative Concepts Unique to Single-Cell

Per-cell PSI vs pseudobulk PSI:

  • Per-cell PSI: meaningful only when junction coverage exceeds ~10-20 reads per cell per event (plate-based or long-read).
  • Pseudobulk PSI: aggregate, recovers bulk-level statistical power, discards within-cluster heterogeneity.

Modality detection in PSI distributions (Song 2017 Mol Cell):

ModalityPSI distributionBiology
IncludedPeaked at 1Constitutive inclusion
ExcludedPeaked at 0Constitutive skipping
BimodalMixture at 0 and 1Mixed cell states or monoallelic-like bursting
MiddlePeaked ~0.5Often technical (well-contamination, doublets, or low-coverage shrinkage to prior); confirm with full-length
MultimodalMultiple peaksComplex regulation; deserves follow-up

Beta-binomial vs binomial models: with sparse counts, binomial PSI is overdispersed. Beta-binomial models (BRIE2; leafcutter2 as Dirichlet-multinomial cluster-level) handle this. For very sparse droplet data, even beta-binomial fits poorly per cell — collapse to pseudobulk.

Imputation pitfalls: naive imputation (MAGIC, scImpute, ALRA) of expression matrices is not appropriate for PSI: imputing missing junction counts averages over neighboring cells and obliterates the very heterogeneity under study. Psix's approach — testing smoothness of observed PSI on the kNN graph — is the principled alternative.

Cell-Type-Specific Splicing Biology

SystemEventRegulator
Neural microexons3-27 nt exons enriched in brainSRRM4/nSR100 (Irimia 2014 Cell); SRRM3 in retina/photoreceptors (Ciampi 2022 PNAS)
Neural differentiationPTBP1 -> PTBP2 switchmiR-124 represses PTBP1; derepresses neural exons (Boutz 2007 Genes Dev)
T-cell activationCD45 RA -> ROhnRNP-L, ESRP-mediated
ErythropoiesisEPB41 exon 16Splicing factor switching during maturation
Cardiac developmentTTN N2BA -> N2BMBNL1/CELF1 antagonism
EMTFGFR2 IIIb -> IIIc, ENAH exon 11aESRP1/2 loss in mesenchymal state (Warzecha 2009 Mol Cell)
Activated T cellCD45 isoform shiftMultiple SR/hnRNP regulators

Quality Thresholds

MetricRecommendation
Cells per event with reads>=50 (per-cell PSI); >=200 cells per cluster (pseudobulk)
Junction reads per event per cell>=5 with coverage; <=1 = unreliable
PSI variance for cell-type call<0.1 within cluster, >0.2 between clusters
Libraryfull-length plate or long-read for transcriptome-wide; 3' for APA only
Doublet filteringRequired before splicing analysis (DoubletFinder, Scrublet)
Cells per cluster (pseudobulk)>=100 ideal; >=50 minimum
nuclear vs whole-cellsnRNA-seq enriches IR; treat with caution

Common Errors

ErrorCauseSolution
MARVEL: ComputePSI returns emptySTAR SJ.out.tab missing strand infoRe-run STAR with --outSJtype Standard
brie.tl.fit: NaN lossInsufficient junction reads per cellFilter cells with min_reads=20; raise threshold
scQuint: convergence not reachedLDA model fit on too-few junctionsAggregate by chromosome; or switch chemistry
Psix: missing connectivitiesNeighbors graph not computedRun sc.pp.neighbors(adata) first
Sierra: no peaks calledGTF missing 3'UTR annotationsUse comprehensive GENCODE; or de-novo peak-call
MARVEL: ggplot errorSeurat version mismatchMatch MARVEL and Seurat versions
FLAMES: barcode rescue failedShort-read 10X output not in expected directoryVerify cellranger output structure

Common Pitfalls

  • Treating 10X 3' splicing analysis as legitimate — the chemistry doesn't support it. Use Sierra for APA or upgrade to MAS-Iso-seq.
  • Imputing PSI matrices — destroys the heterogeneity to be detected. Use Psix or BRIE2 instead.
  • Per-cell PSI on droplet data — typically too sparse for stable estimates. Use pseudobulk first, then drill down to per-cell.
  • Confusing APA with splicing — Sierra results look like AS but are 3' UTR isoforms. Different machinery, different biology.
  • snRNA-seq IR signal misinterpreted as splicing dysregulation — nuclear RNA is enriched for incompletely spliced transcripts; baseline IR is high.
  • Trusting per-cell PSI from BRIE2 without ELBO_gain test — BRIE2's per-cell point estimates are noisy; the principled output is the ELBO_gain cell-state-association statistic.
  • Microexon analysis with default short-read aligners — anchors >=20 nt miss most microexons; use VAST-TOOLS, MicroExonator, or long-read.
  • Skipping doublet filtering before splicing — doublets create artificial PSI mid-modality.
  • single-cell/preprocessing - QC and normalization (must run before splicing)
  • single-cell/clustering - Cell type annotation prerequisite
  • single-cell/doublet-detection - Doublet filtering critical for splicing
  • single-cell/data-io - h5ad / Seurat I/O
  • splicing-quantification - Bulk RNA-seq comparison context
  • long-read-splicing - Full-isoform analysis from MAS-Iso-seq, scISOr-Seq2; future of single-cell splicing

References

  • Huang & Sanguinetti 2021 Genome Biol - BRIE2
  • Wen et al 2023 Nucleic Acids Research 51:e29 - MARVEL
  • Benegas, Fischer & Song 2022 eLife - scQuint (annotation-free single-cell splicing analysis, validated on Smart-seq2)
  • Olivieri et al 2022 Nat Methods - SpliZ
  • Buen Abad Najar et al 2022 Genome Research 32:1385 - Psix
  • Patrick et al 2020 Genome Biol - Sierra
  • Song et al 2017 Mol Cell - splicing modality classification
  • Picelli et al 2014 Nat Protoc - Smart-seq2
  • Hagemann-Jensen et al 2020 Nat Biotech - Smart-seq3
  • Hagemann-Jensen et al 2022 Nat Biotech - Smart-seq3xpress
  • Hahaut et al 2022 Nat Biotech - FLASH-seq
  • Salmen et al 2022 Nat Biotech - VASA-seq
  • Johnson et al 2022 bioRxiv 10.1101/2022.03.14.484332 - STORM-seq (preprint)
  • Al'Khafaji et al 2024 Nat Biotech - MAS-Iso-seq / Kinnex
  • Tian et al 2021 Genome Biology 22:310 - FLAMES
  • Joglekar et al 2024 Nat Neurosci 27:1051-1063 - scISOr-Seq2 single-cell brain isoform mapping
  • Irimia et al 2014 Cell - neural microexons / SRRM4
  • Ciampi et al 2022 PNAS 119:e2117090119 - SRRM3-dependent photoreceptor microexons
  • Boutz et al 2007 Genes Dev - PTBP1/PTBP2 neural switch
  • Tian & Manley 2017 Nat Rev Mol Cell Biol - alternative polyadenylation and 3' UTR isoforms

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in alternative-splicing/single-cell-splicing of GPTomics/bioSkills.

  • SKILL.md
  • examples/sc_splicing_brie2.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Single Cell Splicing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Single Cell Splicing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Single Cell Splicing this skillGPTomics/bioSkills1.2k2 repos~6.5kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Single Cell Splicing

What does Bio Single Cell Splicing do?

Analyzes alternative splicing at single-cell resolution. An agent skill from GPTomics/bioSkills. Bio Single Cell Splicing is an agent skill from GPTomics/bioSkills. Analyzes alternative splicing at single-cell resolution.

When should I use Bio Single Cell Splicing?

Bio Single Cell Splicing fits situations like: analyzing isoform usage in scRNA-seq; identifying cell-type-specific splicing; determining whether scRNA-seq chemistry supports splicing analysis at all.

How do I install Bio Single Cell Splicing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-splicing -a claude-code`. Or copy the skill folder (alternative-splicing/single-cell-splicing in GPTomics/bioSkills) into .claude/skills/bio-single-cell-splicing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Single Cell Splicing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-splicing -a codex`. Or copy the skill folder (alternative-splicing/single-cell-splicing in GPTomics/bioSkills) into .agents/skills/bio-single-cell-splicing in your project. Codex loads it when a task matches its description.

Can I use Bio Single Cell Splicing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-single-cell-splicing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-single-cell-splicing, .gemini/skills/bio-single-cell-splicing, .github/skills/bio-single-cell-splicing and .opencode/skills/bio-single-cell-splicing in your project.

What does Bio Single Cell Splicing need to run?

Going by SKILL.md and its folder, Bio Single Cell Splicing needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Single Cell Splicing access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Single Cell Splicing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Single Cell Splicing use?

Bio Single Cell Splicing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Single Cell Splicing use?

About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Single Cell Splicing?

Skills that share tags, products or a category with Bio Single Cell Splicing: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Single Cell Splicing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.