Agent skill

Bio Splice Variant Prediction

by GPTomics in GPTomics/bioSkills

Predicts whether a DNA variant alters mRNA splicing using sequence-based deep-learning tools — SpliceAI (10kb context dilated CNN, clinical default), Pangolin (multi-tissue), MMSplice (modular…

MITAuto-check passedAI & LLM Engineering

Install Bio Splice Variant Prediction

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-splice-variant-prediction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-splice-variant-prediction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alternative-splicing/splice-variant-prediction .claude/skills/bio-splice-variant-prediction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-splice-variant-prediction
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.4k tokens
SKILL.md length
2,481 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Predicts whether a DNA variant alters mRNA splicing using sequence-based deep-learning tools — SpliceAI (10kb context dilated CNN, clinical default), Pangolin (multi-tissue), MMSplice (modular…

  • Interpreting splice impact of clinical variants
  • SKILL.md covers Version Compatibility, Predictor Taxonomy, Tool Selection Matrix and Decision Tree by Use Case, plus 17 more sections
  • Runs Python scripts from its folder; calls pip and python; reaches kidsneuro.shinyapps.io
  • Prioritizing VUS

What it does

Bio Splice Variant Prediction is an agent skill from GPTomics/bioSkills. Predicts whether a DNA variant alters mRNA splicing using sequence-based deep-learning tools — SpliceAI (10kb context dilated CNN, clinical default), Pangolin (multi-tissue), MMSplice (modular per-region CNN with calibrated ΔPSI), SpliceTransformer/TrASPr (tissue-aware transformers), SpliceVault (empirical 300K-RNA lookup of likely mis-splicing outcomes), CADD-Splice (composite score). Applies the ClinGen SVI 2023 framework for ACMG/AMP variant interpretation (PVS1, PP3, BP4 evidence codes), HGVS splicing…

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/spliceai_clingen_classify.py` and `usage-guide.md`).

It sits in AI & LLM Engineering, covering Deep learning. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Interpreting splice impact of clinical variants
  • Prioritizing VUS
  • Identifying deep-intronic pathogenic variants

Example prompts

  • “Use the bio-splice-variant-prediction skill to predict whether a DNA variant alters mRNA splicing using sequence-based deep-learning tools —…”
  • “/bio-splice-variant-prediction”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • kidsneuro.shinyapps.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Splice Variant Prediction loads about 6.4k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 2,481 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~225
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,481 words, ~6,417 tokens.

Download SKILL.mdSave it as .claude/skills/bio-splice-variant-prediction/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-splice-variant-prediction
description
Predicts whether a DNA variant alters mRNA splicing using sequence-based deep-learning tools — SpliceAI (10kb context dilated CNN, clinical default), Pangolin (multi-tissue), MMSplice (modular per-region CNN with calibrated ΔPSI), SpliceTransformer/TrASPr (tissue-aware transformers), SpliceVault (empirical 300K-RNA lookup of likely mis-splicing outcomes), CADD-Splice (composite score). Applies the ClinGen SVI 2023 framework for ACMG/AMP variant interpretation (PVS1, PP3, BP4 evidence codes), HGVS splicing nomenclature (c.123+1G>A, c.123-3T>G, r.spl?), extended-window scoring for deep-intronic pseudoexons, tissue-specific predictions, branchpoint variant detection (BPHunter, LaBranchoR), and splice-switching ASO design. Use when interpreting splice impact of clinical variants, prioritizing VUS, identifying deep-intronic pathogenic variants, or designing ASOs.
tool_type
python
primary_tool
SpliceAI

Version Compatibility

Reference examples tested with: SpliceAI 1.3+, Pangolin 1.0+, MMSplice 2.4+, pyensembl 2.3+, pysam 0.22+, pandas 2.2+, gffutils 0.13+, tensorflow 2.15+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Splice Variant Prediction

Predict whether a DNA variant alters mRNA splicing. Distinct from "variant pathogenicity" generally: a variant can be a strong splice disruptor without being pathogenic for the gene's standard mechanism, or pathogenic for reasons orthogonal to splicing. Splice prediction asks specifically: does this variant change splice-site usage?

Predictor Taxonomy

FamilyArchitectureOutputFails when
Context-aware CNN10 kb dilated ResNetPer-position donor/acceptor probabilityLong-range (>5 kb) regulatory effects; tissue-specific events
Tissue-aware CNN/transformerSame arch + multi-tissue trainingPer-tissue ΔPSITissue not in training set; novel cell types
Modular per-region CNNSeparate sub-models for 5'ss/3'ss/exon/intronCalibrated quantitative ΔPSIAtypical events; complex multi-junction effects
Foundation transformerPretrained on broad genomic contextSplice probability or ΔPSINew tools; less battle-tested
Empirical lookupPublic RNA-seq event databaseTop-N most likely mis-splicing outcomesVariant types not represented in training cohorts
Composite scoreBlend of multiple predictorsSingle scaled scoreWhen component predictors disagree internally

Tool Selection Matrix

ToolBest forOutputWhen to useFails when
SpliceAIClinical screening; canonical splice site disruptionDelta score 0-1Default for ACMG variant classificationTissue-specific events; deep-intronic with default 50nt window
PangolinTissue-aware predictionsPer-tissue ΔPSIWhen disease tissue is known (brain, heart, liver, testis)Tissue not in 4-tissue training set
MMSpliceQuantitative ΔPSIΔlogit_psiResearch where calibrated effect-size mattersAtypical events outside cassette-exon model
SpliceTransformer2024+ benchmark improvementsTissue-specific ΔPSIWhen transformer foundation models outperform CNN on benchmark variant setsNew (2024); limited clinical adoption
TrASPrMulti-transformer, 2024-2025Tissue-specific PSI/ΔPSIStrong on tissue-specific test setsNew; verify before clinical use
SpliceVaultEmpirical mis-splicing outcomeTop-N events at the affected splice sitePredicting consequence (skip vs cryptic) of canonical-disrupting variantsVariants not represented in 300K-RNA training
CADD-SpliceSingle composite scoreScaled C-scoreClinical pipelines wanting one numberWhen knowing which sub-component drove the score is needed

Methodology evolves; verify benchmarks (Smith & Kitzman 2023 Genome Biol 24:294; You et al 2024 Nat Commun) and ClinGen SVI splicing recommendations before reporting clinical interpretations. Concordance across SpliceAI + Pangolin + MMSplice is gold-standard evidence; discordance flags need RNA validation.

Decision Tree by Use Case

Use caseRecommended approach
Clinical variant report (single variant, ACMG classification)SpliceAI default 50nt + ClinGen SVI 2023 thresholds
Tissue-specific clinical question (brain disease, cardiomyopathy)SpliceAI + Pangolin (tissue-matched)
Unsolved Mendelian case (suspect deep-intronic)SpliceAI extended window (-D 500-2000) + SpliceVault
VUS panel screeningSpliceAI + Pangolin + MMSplice concordance scoring
Predict consequence of canonical-disrupting variantSpliceVault top-N empirical events
Branchpoint variant suspectedBPHunter (branchpoint screen) — SpliceAI is weak here
Splice-switching ASO design (target ESE/ESS occlusion)SpliceAI on masked sequence + RNAfold accessibility
Validate predicted splice change in patientRNA-seq + FRASER2 (see outlier-splicing-detection)
Pseudoexon prediction in deep intronSpliceAI extended window + CI-SpliceAI; require RNA validation

ClinGen SVI 2023 Framework

The ClinGen Sequence Variant Interpretation (SVI) splicing subgroup (Walker 2023 Am J Hum Genet) extended the ACMG/AMP 2015 framework with explicit splice-prediction rules.

Evidence codeThresholdNotes
PP3 (supporting pathogenic)SpliceAI delta >= 0.20ClinGen SVI: apply at supporting weight (not standalone)
BP4 (supporting benign)SpliceAI delta <= 0.10ClinGen SVI: apply at supporting weight
PVS1 (very strong null)Canonical +/-1, +/-2 site disruption with predicted LoF + NMDRequires gene where LoF is established mechanism (Abou Tayoun 2018 Hum Mutat PVS1 decision tree)
PS3 / BS3 (functional)RNA evidence (RT-PCR, RNA-seq, minigene)Supersedes computational evidence

Operational rules: Computational evidence (PP3/BP4) is supporting, not standalone. ClinGen SVI 2023 recommends applying predictive splice PP3/BP4 at supporting weight only; higher SpliceAI cutoffs (0.5, 0.8) increase precision but are the tool's own tiers (Jaganathan 2019), NOT ClinGen-endorsed evidence-strength upgrades — reaching moderate/strong requires functional/RNA evidence (PS3/BS3), not a higher SpliceAI score alone. Splicing variants benefit from concordance across SpliceAI + Pangolin + MMSplice. RNA validation supersedes prediction. Always log SpliceAI version, distance window, and reference transcript. SpliceAI alone is not sufficient for PVS1; canonical site disruption requires gene-level LoF context.

SpliceAI Workflow

Goal: Annotate VCF variants with per-variant delta scores for splice-site change.

Approach: Run spliceai CLI with reference genome and annotation; parse INFO field for delta scores. SpliceAI is human-only (-A grch37 or -A grch38); the model was trained on GENCODE human and does not directly transfer to mouse, fly, or other species. For mouse, retrained variants exist (e.g. mouseSpliceAI); for other species, use Pangolin (4 species: human, mouse, rat, rhesus macaque) or accept that prediction will be unreliable.

bash
spliceai \
    -I input.vcf \
    -O output.vcf \
    -R GRCh38.primary_assembly.genome.fa \
    -A grch38 \
    -D 50 \
    -M 0

-D 50 = distance window in nt around variant (default 50). For deep-intronic variants suspected of creating pseudoexons, raise to 500-2000:

bash
spliceai -I input.vcf -O output_extended.vcf -R genome.fa -A grch38 -D 500 -M 1

-M 0 (default) returns raw scores; -M 1 masks splice gains at annotated sites and losses at unannotated sites (cleaner for clinical use). Output INFO format: SpliceAI=ALLELE|SYMBOL|DS_AG|DS_AL|DS_DG|DS_DL|DP_AG|DP_AL|DP_DG|DP_DL. Delta score = max(DS_AG, DS_AL, DS_DG, DS_DL).

python
import pandas as pd
import re

def parse_spliceai_vcf(vcf_path):
    rows = []
    with open(vcf_path) as f:
        for line in f:
            if line.startswith('#'):
                continue
            fields = line.strip().split('\t')
            info = fields[7]
            m = re.search(r'SpliceAI=([^;]+)', info)
            if not m:
                continue
            for ann in m.group(1).split(','):
                parts = ann.split('|')
                allele, symbol = parts[0], parts[1]
                ds = [float(p) if p != '.' else 0 for p in parts[2:6]]
                dp = parts[6:10]
                rows.append({
                    'chrom': fields[0], 'pos': int(fields[1]),
                    'ref': fields[3], 'alt': allele,
                    'gene': symbol,
                    'DS_AG': ds[0], 'DS_AL': ds[1],
                    'DS_DG': ds[2], 'DS_DL': ds[3],
                    'delta_max': max(ds),
                })
    return pd.DataFrame(rows)

df = parse_spliceai_vcf('output.vcf')
df['acmg_evidence'] = pd.cut(
    df['delta_max'],
    bins=[-0.01, 0.10, 0.20, 0.50, 0.80, 1.01],
    # ClinGen SVI applies splice PP3/BP4 at supporting weight; 0.5/0.8 are SpliceAI
    # precision tiers (Jaganathan 2019), NOT ACMG evidence-strength upgrades
    labels=['BP4', 'inconclusive', 'PP3_supporting', 'PP3_supporting_prec0.5', 'PP3_supporting_prec0.8']
)

DS labels: AG = acceptor gain, AL = acceptor loss, DG = donor gain, DL = donor loss.

Pangolin for Tissue-Specific Prediction

Goal: Get tissue-specific splice impact predictions when disease tissue is known.

Approach: Run Pangolin CLI with VCF + reference + gffutils annotation database.

bash
python -c "import gffutils; gffutils.create_db('gencode.v45.annotation.gff3', 'gencode.db', force=True)"

pangolin \
    input.vcf \
    GRCh38.primary_assembly.genome.fa \
    gencode.db \
    pangolin_output \
    -d 500 \
    -m True \
    -s 0.2

-m True masks splice gains at annotated sites and losses at unannotated sites (recommended for clinical use). -s 0.2 outputs all sites with predicted change >= cutoff.

Pangolin output is a VCF with per-tissue predictions across the 4 tissues used at training: brain, heart, liver, testis (Zeng & Li 2022 Genome Biol). The model outputs per-species per-tissue predictions but extrapolates poorly to tissues outside this set. Use the tissue closest to disease-relevant context. For tissues not in the 4-tissue training set, fall back to SpliceAI — Pangolin extrapolates poorly to unseen tissues.

SpliceVault for Empirical Mis-Splicing Outcomes

Goal: Predict the type of mis-splicing (exon skipping vs cryptic site activation) given a canonical-disrupting variant.

Approach: Query SpliceVault's database of empirical mis-splicing events from public RNA-seq.

python
import requests

# Web API: https://kidsneuro.shinyapps.io/splicevault/
# Or use the R/Python package at github.com/kidsneuro-lab/SpliceVault

# Example: NM_000546.6:c.673-2A>G (TP53)
# Returns top-N most likely mis-splicing events: exon skipping, cryptic 3'ss usage, etc.

SpliceVault (Dawes 2023 Nat Genet) showed that the Top-4 events at any splice site predict variant-associated mis-splicing with ~92% sensitivity overall (96% of exon-skipping and 86% of cryptic-activation events) — a striking regularity that makes consequence prediction tractable. Use SpliceVault when the question is not "will splicing change?" but "what specific aberrant splicing will occur?".

MMSplice for Calibrated ΔPSI

Goal: Predict quantitative ΔPSI (not just probability of disruption) for cassette exons.

Approach: Score variant impact on each splicing region (5'ss, 3'ss, exon, intron-3'/5') and combine.

python
from mmsplice.vcf_dataloader import SplicingVCFDataloader
from mmsplice import MMSplice, predict_save

dl = SplicingVCFDataloader(
    gtf='gencode.v45.basic.gtf',
    fasta_file='GRCh38.fa',
    vcf_file='input.vcf'
)

model = MMSplice()
predict_save(model, dl, 'mmsplice_predictions.csv', pathogenicity=True)

MMSplice (Cheng 2019 Genome Biol) reports Δlogit_psi per variant. Useful when calibrated effect sizes matter (research) more than probability of disruption (clinical screening). Companion MTSplice (Cheng 2021 Genome Biol) adds tissue-specific Δψ predictions.

HGVS Splicing Nomenclature

Following den Dunnen 2016 Hum Mutat:

NotationMeaning
c.123+1G>A+1 of intron downstream of exon ending at cDNA position 123 (canonical 5'ss G)
c.123+5G>A+5 position of donor (consensus region)
c.124-1G>A-1 of acceptor (canonical AG)
c.124-3T>G-3 of acceptor (Py-tract / BPS region)
c.124-50A>GDeep-intronic; may activate cryptic site
r.123_456delRNA-level deletion (predicted exon skipping)
r.spl?Unknown splice consequence
r.0?No detectable RNA
p.0?Unknown protein consequence
p.(=)No predicted protein change (silent)

Validation tools: VariantValidator (Freeman 2018 Hum Mutat), Mutalyzer 2 (Lefter et al 2021 Bioinformatics 37:2811-2817).

Extended-Window Scoring for Deep-Intronic Variants

SpliceAI's default precomputed scores use a 50-nt window, missing variants that create pseudoexons in deep intronic regions. For unsolved Mendelian cases:

bash
# Recompute with extended window
spliceai -I input.vcf -O output_2kb.vcf -R genome.fa -A grch38 -D 2000

# Or use CI-SpliceAI (Strauch 2022 PLoS One), SpliceAI retrained on curated GENCODE splice sites
WindowTradeoff
-D 50 (default)Fast; captures canonical-site disruption; misses deep-intronic
-D 500Captures most pseudoexon-creating deep-intronic variants
-D 2000Maximum sensitivity; some false positives at large distances

Pseudoexon creation in deep introns explains a substantial fraction of unsolved Mendelian disease alleles in current cohorts (estimates 5-15% across studies; specific quantitative range will vary by cohort and panel — verify against current literature). Disease examples: CFTR 3849+10kbC>T, USH2A c.7595-2144A>G, CEP290 c.2991+1655A>G (LCA10).

Concordance Across Predictors

python
import pandas as pd

merged = (spliceai_df
    .merge(pangolin_df, on=['chrom', 'pos', 'alt'], suffixes=('_sai', '_pang'))
    .merge(mmsplice_df, on=['chrom', 'pos', 'alt'])
)

merged['concordance'] = (
    (merged['delta_max_sai'] >= 0.2).astype(int) +
    (merged['pangolin_score'].abs() >= 0.2).astype(int) +
    (merged['delta_logit_psi'].abs() >= 1.0).astype(int)
)

merged['interpretation'] = merged['concordance'].map({
    0: 'concordant_benign',
    1: 'discordant_low_evidence',
    2: 'concordant_evidence',
    3: 'high_concordance_pathogenic'
})
ConcordanceInterpretationAction
3/3 above thresholdHigh confidencePP3 (supporting); strong candidate for RNA validation (PS3)
2/3 aboveConcordant evidencePP3 (supporting)
1/3 aboveDiscordantReport inconclusive; flag for RNA validation
0/3 aboveConcordant benignBP4 (supporting)

Discordance is the most informative pattern — variants where one model sees impact and others don't are high priority for RNA validation.

Branchpoint Variant Detection

All current tools are weak at branchpoint variants because the BPS motif (yUnAy) has low information content. Specific branchpoint tools:

ToolMethodNotes
BPPMixture model (BP motif + polypyrimidine tract)Zhang 2017 Bioinformatics 33:3166
LaBranchoRBidirectional LSTMPaggi & Bejerano 2018 RNA 24:1647
SVM-BPfinderSVM on conservation+sequenceCorvelo 2010 PLoS Comput Biol
BPHunterGenome-wide branchpoint screen against an aggregated experimental (lariat/RNA-seq) + computational BP databaseZhang 2022 PNAS

Branchpoint variants are under-recognized in clinical pipelines; SpliceAI captures only some because branchpoint motifs have low information content. Recommendation: when SpliceAI delta is borderline (0.1-0.3) for a variant in the BPS region (-18 to -40 from 3'ss), run BPHunter as supplement.

Show full SKILL.md (1,013 more words)Show less

Splice-Switching ASO Design

Goal: Design antisense oligonucleotides to modulate splicing therapeutically (e.g. SMA ISS-N1, DMD exon skipping).

Approach: Use SpliceAI to predict impact of binding-site occlusion; check accessibility (RNAfold); avoid SR/hnRNP off-target motifs.

python
# Conceptual workflow - actual design uses ASO synthesis platforms
# 1. Identify target ESE/ESS/ISE/ISS region from MaxEntScan + SpliceAI scan
# 2. Design candidate 18-22 nt ASOs spanning the regulatory element
# 3. For each ASO, simulate splice-site occlusion impact via SpliceAI on the masked sequence
# 4. Filter for RNA accessibility (avoid stable hairpins) using RNAfold
# 5. Whole-transcriptome SpliceAI scan for off-target binding (>=17/20 nt match)
# 6. Avoid TLR9 immunostimulatory CpG motifs

# Chemistry choices:
# - 2'-MOE-PS: nusinersen-like (CNS, intrathecal)
# - PMO: DMD ASOs (systemic IV)
# - GalNAc-conjugated: hepatic targeting

Approved precedents: nusinersen (SMA ISS-N1 occlusion, exon 7 inclusion); risdiplam (small-molecule SMN2 splicing modulator); eteplirsen/golodirsen/casimersen/viltolarsen (DMD exon skipping). Design references: Hua 2008 AJHG; Roberts et al 2023 Nat Rev Drug Discov 22:917 (DMD therapeutic approaches).

Per-Tool Failure Modes

SpliceAI: 50nt Window Limitation

Trigger: Variant deep in an intron (>50 nt from canonical splice site).

Mechanism: Default precomputed scores use ±50 nt window; the model is trained on this context but pre-stored scores limit lookups.

Symptom: Known pathogenic deep-intronic variant scores low (<0.2); no pseudoexon detected.

Fix: Re-run with -D 500 or -D 2000; or try CI-SpliceAI (SpliceAI retrained on curated GENCODE splice sites) as a second predictor.

SpliceAI: Tissue Agnosticism

Trigger: Variant in a tissue-specific gene (NEFM in neurons, MAPT brain, DMD muscle isoforms).

Mechanism: SpliceAI is trained on aggregate GENCODE annotation; tissue-specific events with weak constitutive use score low.

Symptom: Tissue-specific pathogenic variant has low SpliceAI delta; functional impact still observed in target tissue.

Fix: Use Pangolin for tissue-aware prediction; or SpliceTransformer; require RNA validation in disease-relevant tissue.

Pangolin: Out-of-Training Tissue

Trigger: Disease tissue not represented in Pangolin's 4-species, 4-tissue (Cardoso-Moreira 2019 developmental) training set.

Mechanism: Pangolin extrapolates poorly to tissues outside training distribution.

Symptom: Pangolin score uncalibrated for queried tissue; doesn't agree with patient RNA-seq from that tissue.

Fix: Fall back to SpliceAI for tissues not in Pangolin training; or run patient RNA-seq directly.

MMSplice: Atypical Events

Trigger: Variant affecting a non-cassette event (MXE, complex multi-junction, AFE/ALE).

Mechanism: MMSplice modular model is trained primarily on cassette exon events.

Symptom: MMSplice ΔPSI doesn't match other predictors or empirical data for non-cassette events.

Fix: Use SpliceAI for non-cassette events; restrict MMSplice to cassette exon contexts.

CADD-Splice: Loss of Component Information

Trigger: Wanting to know which sub-component drove a high CADD-Splice score.

Mechanism: CADD-Splice combines SpliceAI + MMSplice + CADD into a single C-score; sub-component contributions are abstracted.

Symptom: "High CADD-Splice score but unclear why."

Fix: Run SpliceAI and MMSplice separately to see which contributed.

Branchpoint Variants: Low Information Motif

Trigger: Variant in the BPS region (-18 to -40 from 3'ss).

Mechanism: BPS motif (yUnAy) has low information content; CNNs struggle to learn the consensus.

Symptom: Confirmed BPS variant scores SpliceAI delta <0.2 despite functional disruption.

Fix: Use BPHunter (Zhang 2022 PNAS) for genome-wide branchpoint screening; require RNA validation.

Population Database Lookup

DatabaseUse for
gnomAD v4Allele frequency; SpliceAI annotations integrated
ClinVarExisting classifications; SpliceAI integrated since 2020
SpliceVarDBCurated splice variants with experimental RNA validation
dbNSFP4Pre-computed splice scores aggregated
Recount3Tissue-specific PSI lookups from public RNA-seq
GTEx sQTL v8Tissue-specific splicing QTLs across 49 tissues
MaveDBSplice MAVE results (e.g. BRCA1 saturation; Findlay 2018 Nature)

Always check ClinVar first for existing classifications; cross-reference with gnomAD for population frequency before committing to PP3/PP4.

Common Errors

ErrorCauseSolution
spliceai: tensorflow not foundTensorFlow not installedpip install tensorflow>=2.0 separately
spliceai: chrom not in referenceVCF chrom name mismatch (chr1 vs 1)bcftools annotate --rename-chrs chr_map.txt
pangolin: no annotations found for variantgffutils db doesn't contain queried geneRebuild gffutils db with comprehensive GENCODE GFF3
mmsplice: variant outside any cassette eventMMSplice model assumes cassette contextUse SpliceAI for non-cassette events
SpliceVault: variant not foundVariant outside common splice sites in 300K-RNA databaseUse SpliceAI for prediction (no empirical baseline available)
VariantValidator: invalid HGVSWrong reference transcript or buildSpecify NM_*.* version explicitly

Common Pitfalls

  • Using SpliceAI score alone for clinical reporting — must combine with concordant predictors and ideally RNA validation; ClinGen SVI requires this for non-canonical positions.
  • 50nt window for deep intronic variants — pseudoexon-creating variants 100-2000 nt deep are systematically missed.
  • Tissue-agnostic prediction for tissue-specific genes — use Pangolin or SpliceTransformer when tissue context matters (NEFM, MAPT, DMD isoforms).
  • Branchpoint variants — all current predictors are weak here. Use BPHunter for branchpoint screening.
  • Forgetting NMD direction — confirmed splice disruption needs NMD-status check. Last-exon PTCs escape NMD and can be dominant-negative or gain-of-function.
  • In-silico-only PVS1 application — PVS1 for non-canonical positions requires functional or strong computational evidence; SpliceAI alone is supporting (PP3), not very strong.
  • Trusting LLMs for variant interpretation — use as orchestrators on top of SpliceAI/VariantValidator/ClinVar; all clinical-grade calls require human expert sign-off.
  • Skipping HGVS validation — invalid HGVS leads to silent reference-transcript mismatches; always run VariantValidator first.

Quality Thresholds

MetricRecommendationSource
Default SpliceAI window-D 50 (clinical screening)Jaganathan 2019
Deep-intronic SpliceAI window-D 500-2000 (unsolved Mendelian)Convention (verify current literature)
ACMG PP3 (supporting)SpliceAI delta >= 0.2Walker 2023 AJHG (apply at supporting weight)
ACMG BP4 (supporting)SpliceAI delta <= 0.1Walker 2023 AJHG
SpliceAI higher-precision cutoffs0.5 / 0.8 raise precision, NOT ACMG strengthJaganathan 2019 (not ClinGen graded tiers)
Off-target ASO match<=16/20 nt to any non-target transcriptDesign convention
Concordance for high-confidence2/3 predictors above PP3 thresholdPragmatic
  • splicing-qc - MaxEntScan + library QC for confirming predicted impact
  • splicing-quantification - Empirical PSI from RNA-seq to validate predictions
  • outlier-splicing-detection - FRASER2/DROP for RNA-seq confirmation in clinical samples
  • variant-calling/clinical-interpretation - Broader ACMG/AMP variant interpretation framework
  • variant-calling/variant-annotation - VEP plugin integration for SpliceAI

References

  • Jaganathan et al 2019 Cell - SpliceAI
  • Zeng & Li 2022 Genome Biol - Pangolin
  • Cheng et al 2019 Genome Biol - MMSplice
  • Cheng et al 2021 Genome Biol - MTSplice (tissue MMSplice)
  • You et al 2024 Nat Commun 15:9129 - SpliceTransformer
  • Strauch et al 2022 PLoS One 17:e0269159 - CI-SpliceAI extended window
  • Smith & Kitzman 2023 Genome Biol 24:294 - SpliceAI/Pangolin MPSA benchmark
  • Rentzsch et al 2021 Genome Med - CADD-Splice
  • Dawes et al 2023 Nat Genet - SpliceVault
  • Walker et al 2023 Am J Hum Genet - ClinGen SVI splicing recommendations
  • Riepe et al 2021 Hum Mutat 42:799 - SpliceAI in clinical pipelines (Riepe TV et al)
  • Abou Tayoun et al 2018 Hum Mutat - PVS1 decision tree
  • Richards et al 2015 Genet Med - ACMG/AMP framework
  • den Dunnen et al 2016 Hum Mutat - HGVS standard
  • Zhang et al 2022 PNAS (PMID 36306325) - BPHunter for branchpoints
  • Hua et al 2008 AJHG - ISS-N1 / nusinersen mechanism
  • Roberts et al 2023 Nat Rev Drug Discov 22:917-934 - DMD therapeutic approaches (exon-skipping ASOs)
  • Findlay et al 2018 Nature - BRCA1 saturation genome editing (MAVE)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in alternative-splicing/splice-variant-prediction of GPTomics/bioSkills.

  • SKILL.md
  • examples/spliceai_clingen_classify.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Splice Variant Prediction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Splice Variant Prediction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Splice Variant Prediction this skillGPTomics/bioSkills1.2k2 repos~6.4kAutomated safety check: PassMIT
Sparse Autoencoder Training with SAELensOrchestra-Research/AI-Research-SKILLs13k5 repos~3.2kAutomated safety check: PassMIT
Alphagenome Predictionsgenomicsxai/alphagenome-pytorch162—~868Automated safety check: PassApache-2.0
TransformerLens InterpretabilityOrchestra-Research/AI-Research-SKILLs13k3 repos~3kAutomated safety check: PassMIT
pyvene Causal InterventionsOrchestra-Research/AI-Research-SKILLs13k2 repos~3.5kAutomated safety check: PassMIT
ML Training RecipesOrchestra-Research/AI-Research-SKILLs13k1 repos~2.8kAutomated safety check: PassMIT

Similar skills

  • Sparse Autoencoder Training with SAELens

    Orchestra-Research/AI-Research-SKILLs

    Guides training and analyzing sparse autoencoders with SAELens to break neural network activations into interpretable features, including superposition and monosemanticity studies.

    13k GitHub starsUsed in 5 repos~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Alphagenome Predictions

    genomicsxai/alphagenome-pytorch

    Run AlphaGenome-PyTorch to get genomic track predictions — via the agt predict CLI (single locus, BED regions, whole chromosomes, raw FASTA sequences, or per-gene count tables/AnnData), variant…

    162 GitHub stars~868 tokensUpdated 25 days ago
    AI & LLM EngineeringAuto-check passed
  • TransformerLens Interpretability

    Orchestra-Research/AI-Research-SKILLs

    Guides mechanistic interpretability work with TransformerLens: loading models, caching activations, using HookPoints, activation patching and attention-pattern analysis.

    13k GitHub starsUsed in 3 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • pyvene Causal Interventions

    Orchestra-Research/AI-Research-SKILLs

    Guides causal experiments on PyTorch models with pyvene, such as causal tracing, activation patching and interchange intervention training, to test how a model works.

    13k GitHub starsUsed in 2 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • ML Training Recipes

    Orchestra-Research/AI-Research-SKILLs

    PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM.

    13k GitHub starsUsed in 1 repo~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Cosmos Policy Evaluation

    Orchestra-Research/AI-Research-SKILLs

    Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling.

    13k GitHub stars~3.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Splice Variant Prediction

What does Bio Splice Variant Prediction do?

Predicts whether a DNA variant alters mRNA splicing using sequence-based deep-learning tools — SpliceAI (10kb context dilated CNN, clinical default), Pangolin (multi-tissue), MMSplice (modular…. Bio Splice Variant Prediction is an agent skill from GPTomics/bioSkills. Predicts whether a DNA variant alters mRNA splicing using sequence-based deep-learning tools — SpliceAI (10kb context dilated CNN, clinical default), Pangolin (multi-tissue), MMSplice (modular per-region CNN with calibrated ΔPSI), SpliceTransformer/TrASPr (tissue-aware transformers), SpliceVault (empirical 300K-RNA lookup of likely mis-splicing outcomes), CADD-Splice (composite score).

When should I use Bio Splice Variant Prediction?

Bio Splice Variant Prediction fits situations like: interpreting splice impact of clinical variants; prioritizing VUS; identifying deep-intronic pathogenic variants.

How do I install Bio Splice Variant Prediction in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-splice-variant-prediction -a claude-code`. Or copy the skill folder (alternative-splicing/splice-variant-prediction in GPTomics/bioSkills) into .claude/skills/bio-splice-variant-prediction in your project. Claude Code loads it when a task matches its description.

How do I install Bio Splice Variant Prediction in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-splice-variant-prediction -a codex`. Or copy the skill folder (alternative-splicing/splice-variant-prediction in GPTomics/bioSkills) into .agents/skills/bio-splice-variant-prediction in your project. Codex loads it when a task matches its description.

Can I use Bio Splice Variant Prediction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-splice-variant-prediction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-splice-variant-prediction, .gemini/skills/bio-splice-variant-prediction, .github/skills/bio-splice-variant-prediction and .opencode/skills/bio-splice-variant-prediction in your project.

What does Bio Splice Variant Prediction need to run?

Going by SKILL.md and its folder, Bio Splice Variant Prediction needs Python for the scripts in its folder and the command-line tools its instructions call (pip and python). Our summary lists: Python 3.

Does Bio Splice Variant Prediction access the network?

SKILL.md names 1 domain. In commands or code: kidsneuro.shinyapps.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Splice Variant Prediction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Splice Variant Prediction use?

Bio Splice Variant Prediction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Splice Variant Prediction use?

About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Splice Variant Prediction?

Skills that share tags, products or a category with Bio Splice Variant Prediction: Sparse Autoencoder Training with SAELens (Orchestra-Research/AI-Research-SKILLs, 13k stars), Alphagenome Predictions (genomicsxai/alphagenome-pytorch, 162 stars), TransformerLens Interpretability (Orchestra-Research/AI-Research-SKILLs, 13k stars) and pyvene Causal Interventions (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Splice Variant Prediction?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.