Agent skill

Bio Splicing Qc

by GPTomics in GPTomics/bioSkills

Assesses RNA-seq data quality specifically for alternative splicing analysis.

MITAuto-check passedResearch & Science

Install Bio Splicing Qc

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-splicing-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-splicing-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alternative-splicing/splicing-qc .claude/skills/bio-splicing-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-splicing-qc
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.2k tokens
SKILL.md length
2,265 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Assesses RNA-seq data quality specifically for alternative splicing analysis.

  • Evaluating data suitability for splicing analysis
  • SKILL.md covers Version Compatibility, QC Layer Taxonomy, Decision Tree by Question and Experimental Design Audit…, plus 16 more sections
  • Runs Python scripts from its folder; calls pip
  • Troubleshooting low event detection

What it does

Bio Splicing Qc is an agent skill from GPTomics/bioSkills. Assesses RNA-seq data quality specifically for alternative splicing analysis. QC layers include experimental design audit (library prep, read length, depth, replicates), STAR 2-pass cohort-style alignment, junction saturation curves and discovery plateau detection, novel-vs-known junction ratio diagnostics, junction-overhang distribution, splice-site strength scoring (MaxEntScan intrinsic + SpliceAI context-aware), strandedness verification, GENCODE basic vs comprehensive choice, and rRNA contamination screening…

Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/splicing_qc.py` and `usage-guide.md`).

It sits in Research & Science, covering Experimental design, Creative writing and fiction and Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Evaluating data suitability for splicing analysis
  • Troubleshooting low event detection
  • Designing sequencing experiments where AS is a primary endpoint

Example prompts

  • “Use the bio-splicing-qc skill to assess RNA-seq data quality specifically for alternative splicing analysis”
  • “/bio-splicing-qc”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Splicing Qc loads about 6.2k tokens when it runs. Until then it costs about 227 tokens; SKILL.md has 2,265 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~227
When it runs · the whole SKILL.md, loaded when a task matches
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,265 words, ~6,212 tokens.

Download SKILL.mdSave it as .claude/skills/bio-splicing-qc/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-splicing-qc
description
Assesses RNA-seq data quality specifically for alternative splicing analysis. QC layers include experimental design audit (library prep, read length, depth, replicates), STAR 2-pass cohort-style alignment, junction saturation curves and discovery plateau detection, novel-vs-known junction ratio diagnostics, junction-overhang distribution, splice-site strength scoring (MaxEntScan intrinsic + SpliceAI context-aware), strandedness verification, GENCODE basic vs comprehensive choice, and rRNA contamination screening. Splicing analysis is more demanding than DGE on read length, depth, library prep, alignment strategy, and annotation choice — failures silently bias PSI estimates and inflate novel-junction false positives. Use when evaluating data suitability for splicing analysis, troubleshooting low event detection, or designing sequencing experiments where AS is a primary endpoint.
tool_type
python
primary_tool
RSeQC

Version Compatibility

Reference examples tested with: RSeQC 5.0+, STAR 2.7.11+, samtools 1.19+, pysam 0.22+, regtools 1.0+, maxentpy 0.0.1+, spliceai 1.3+, matplotlib 3.8+, pandas 2.2+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Splicing-Specific Quality Control

Splicing analysis is more demanding than DGE on read length, depth, library prep, alignment strategy, and annotation choice. Failures in any of these silently bias PSI estimates and inflate novel-junction false positives. The decision sequence is: experimental design -> library prep -> alignment strategy -> annotation -> diagnostic metrics. Each layer's failure mode is distinct.

QC Layer Taxonomy

LayerTargetToolFails when
Experimental designRead length, depth, replicates, library typePre-sequencing review<PE 75nt; n<3 vs n<3; <30M reads/sample
Library preppoly(A) vs rRNA depletionPre-sequencing reviewpoly(A) library used for IR analysis
AlignmentSTAR 2-pass cohort-styleSTAR1-pass loses 14% novel junctions; per-sample 2-pass introduces inconsistency
Junction discoverySaturation, noveltyRSeQC junction_saturation, junction_annotationCurve still rising = under-sequenced; novel% >40% suggests biology or artifact
Strand specificityLibrary protocol consistencyRSeQC infer_experimentWrong --libType halves usable junctions
Splice site strengthCryptic vs canonicalMaxEntScan, SpliceAIWeak splice sites (MaxEnt<5) may indicate cryptic, regulated, or annotation error
Junction overhangRead-junction support qualitypysam CIGAR parsingOverhang <8nt = high false-positive rate
ContaminationrRNA, adaptersfastq_screen>20% rRNA in "depleted" library = failed depletion
AnnotationGENCODE basic vs comprehensiveAnnotation choiceBasic for canonical events; comprehensive for DTU

Decision Tree by Question

QuestionRecommended QC
Will my planned RNA-seq design support AS analysis?Pre-sequencing audit: library type, read length, depth, replicates
Is my data suitable for cassette exon analysis?Junction saturation + known/novel ratio + read length
Why does my AS analysis call so few events?Saturation curve, depth, library type, alignment 2-pass
Why does my AS analysis call so many novel junctions?Annotation completeness + novel% + biology check (TDP-43, SF3B1)
Are my SpliceAI predictions calibrated for my tissue?MaxEntScan + SpliceAI concordance for known sites
Did STAR 2-pass actually run cohort-style?Verify SJ.out.tab merging across samples
Is intron retention detectable in my data?Library type (must be rRNA-depleted); strand-specific
Are my microexons detectable?Read length >=100; aligner anchor settings; consider VAST-TOOLS

Experimental Design Audit (Before Sequencing)

DecisionFor splicing analysisRationale
Library preprRNA depletion (Ribo-Zero, RiboCop)poly(A) selection loses pre-mRNA, nascent transcripts, and detained introns; for IR analysis rRNA depletion is mandatory (convention)
Read lengthPE 100-150 nt (PE 150 strongly preferred)Junction-spanning reads need >=8 nt overhang on each exon; shorter single-end reads bias junction detection toward shorter exons (convention)
PairingPaired-endSingle-end loses fragment-level disambiguation of junctions
Depth50-100M reads/sampleDGE-grade 30M misses low-PSI events; 100M for low-abundance event discovery
StrandednessStranded library (Illumina TruSeq stranded)Distinguishes overlapping antisense; some tools double-count unstranded junctions
Replicatesn>=3 per conditionn=2 vs n=2 has poor calibration in most tools (especially SUPPA2)
AnnotationGENCODE basic for canonical, comprehensive for DTU/discoverybasic = high-confidence; comprehensive includes putative — affects FDR control
MicroexonsPE 100+ with --alignSJoverhangMin 8; VAST-TOOLSDefault aligners miss 3-27nt exons
Long-intron genes (TTN, brain)Increased --alignIntronMaxDefault 1Mb may miss >1Mb introns

STAR 2-Pass Alignment

Goal: Maximize novel-junction sensitivity for downstream AS analysis.

Approach: Run STAR once per sample to discover novel junctions (pass 1), merge novel junctions across cohort, then re-align with the augmented junction set (pass 2). Cohort-style 2-pass beats per-sample basic 2-pass for differential splicing because all samples use the same junction reference.

bash
# Pass 1: per-sample
STAR --runMode alignReads \
    --runThreadN 8 \
    --genomeDir genome_index \
    --sjdbGTFfile gencode.v45.basic.gtf \
    --sjdbOverhang 149 \
    --readFilesIn sample_R1.fq.gz sample_R2.fq.gz \
    --readFilesCommand zcat \
    --outSAMtype BAM SortedByCoordinate \
    --outFileNamePrefix pass1_${sample}_ \
    --outSJtype Standard \
    --outFilterMultimapNmax 20 \
    --alignSJoverhangMin 8 \
    --alignSJDBoverhangMin 1
bash
# Cohort-style 2-pass: collect all SJ.out.tab from pass 1
cat pass1_*_SJ.out.tab | awk '$5 > 0 && $7 >= 3' | sort -u > cohort_novel_SJ.tab

# Pass 2: re-align with augmented junctions
STAR --runMode alignReads \
    --runThreadN 8 \
    --genomeDir genome_index \
    --sjdbGTFfile gencode.v45.basic.gtf \
    --sjdbFileChrStartEnd cohort_novel_SJ.tab \
    --sjdbOverhang 149 \
    --readFilesIn sample_R1.fq.gz sample_R2.fq.gz \
    --readFilesCommand zcat \
    --outSAMtype BAM SortedByCoordinate \
    --outFileNamePrefix pass2_${sample}_ \
    --outSJtype Standard \
    --twopassMode None \
    --quantMode GeneCounts \
    --alignSJoverhangMin 8 \
    --alignSJDBoverhangMin 3
ApproachNovel-junction recoveryCohort consistency
1-pass with annotation~80-86% (depends on GENCODE completeness)High (annotation-based)
Per-sample basic 2-pass (--twopassMode Basic)>=94%Variable (each sample has its own junction set)
Cohort-style 2-pass (manual merge)>=94%High (shared junction reference)

Per-sample 2-pass (--twopassMode Basic) is simpler but produces inconsistent junction sets across samples; for differential splicing the cohort-style version is preferred (Veeneman 2016 Bioinformatics).

The pass-1 filter awk '$5 > 0 && $7 >= 3' keeps junctions with strand info AND >=3 unique reads — adjust threshold to balance discovery vs noise.

Junction Saturation

Goal: Determine whether sequencing depth is sufficient for comprehensive splicing detection.

Approach: Run RSeQC junction saturation; check whether the discovery curve plateaus.

bash
junction_saturation.py \
    -i sample.bam \
    -r gencode_v45.bed \
    -o sample_junc_sat \
    -m 50
python
import subprocess
import pandas as pd

samples = ['s1.bam', 's2.bam', 's3.bam']
for sample in samples:
    subprocess.run([
        'junction_saturation.py',
        '-i', sample,
        '-r', 'gencode_v45.bed',
        '-o', sample.replace('.bam', '_junc_sat')
    ], check=True)

The output *.junctionSaturation_plot.r plots known + novel junctions vs subsampled reads.

Plateau detection rule: if from 80% to 100% of reads, the junction count rises by <2%, consider it plateaued. Still rising means more sequencing would yield more junctions.

For AS analysis, plateau on the known junction curve is the requirement; novel-junction curves often don't plateau even at deep coverage (which is biologically informative — novel junctions are inherently rarer events).

Novel-vs-Known Junction Ratio

Goal: Detect annotation/mapping issues or biologically interesting cryptic splicing.

Approach: Classify junctions with RSeQC and compute the novel:known ratio.

bash
junction_annotation.py -i sample.bam -r gencode_v45.bed -o sample_junc_annot
python
import pandas as pd

# RSeQC .junction.xls has a header: chrom, intron_st(0-based), intron_end(1-based), read_count, annotation
junc = pd.read_csv('sample_junc_annot.junction.xls', sep='\t')
total = junc['read_count'].sum()

by_class = junc.groupby('annotation')['read_count'].sum()
known_frac = by_class.get('annotated', 0) / total
novel_frac = (by_class.get('partial_novel', 0) + by_class.get('complete_novel', 0)) / total

print(f'known: {known_frac:.1%}, novel: {novel_frac:.1%}')
Known fractionStatusInterpretation
>=80%HealthyComprehensive annotation, good alignment
60-80%AcceptableCheck annotation completeness or organism
<60%Suspect or interestingMapping artifacts, contamination, OR biologically informative

High novel-junction rate may be biology, not artifact:

  • TDP-43 loss (ALS/FTD post-mortem brain): cryptic exon de-repression in UNC13A, STMN2, ATG4B (Brown 2022 Nature; Klim 2019 Nat Neurosci)
  • SF3B1-mutant cancer (MDS, CLL, uveal melanoma): cryptic 3'ss ~10-30nt upstream of canonical (Darman 2015 Cell Rep)
  • Non-model organism: GENCODE-grade annotation unavailable; novel junctions reflect annotation gaps not biology
  • Microbial / viral contamination: reads aligning to host but with unusual junctions

If novel% >40%, drill down: check organism, check spliceosomal mutation status, check known disease signatures.

Junction Read Overhang and Coverage

Goal: Profile per-junction read counts and overhang distribution to identify weakly-supported events.

Approach: Parse CIGAR for N (intron) operations; tally per-junction reads and minimum exon overhangs.

python
import pysam
from collections import defaultdict

def junction_stats(bam_path):
    bam = pysam.AlignmentFile(bam_path, 'rb')
    counts = defaultdict(int)
    min_overhang = defaultdict(lambda: float('inf'))

    for read in bam.fetch():
        if read.is_unmapped or read.is_secondary:
            continue
        ref_pos = read.reference_start
        cumulative_query = 0
        cigar = read.cigartuples
        for i, (op, length) in enumerate(cigar):
            if op == 3:
                left_match = sum(l for o, l in cigar[:i] if o in (0, 7, 8))
                right_match = sum(l for o, l in cigar[i+1:] if o in (0, 7, 8))
                overhang = min(left_match, right_match)
                key = (read.reference_name, ref_pos, ref_pos + length)
                counts[key] += 1
                min_overhang[key] = min(min_overhang[key], overhang)
            if op in (0, 2, 3, 7, 8):
                ref_pos += length

    bam.close()
    return counts, dict(min_overhang)

counts, overhang = junction_stats('sample.bam')
print(f'total junctions: {len(counts)}')
print(f'>= 10 reads: {sum(1 for c in counts.values() if c >= 10)}')
print(f'overhang >= 8 nt: {sum(1 for k, c in counts.items() if overhang[k] >= 8)}')

Junction reads with overhang <8 nt are common false positives, especially for novel sites. Most callers default to >=8 nt anchor for this reason. Microexon-aware aligners use overhang as low as 6 nt with explicit configuration.

Splice Site Strength (MaxEntScan and SpliceAI)

Goal: Score donor and acceptor splice sites to flag weak / cryptic sites and to predict variant impact on splicing.

Approach: Use MaxEntScan (sequence information content) and SpliceAI (context-aware deep-learning) — they answer different questions.

python
from maxentpy.maxent import score5, score3

donor = 'CAGGTAAGT'
acceptor = 'TTTTTTTTTTTTTTTTTTTTCAG'
print(f"5'ss MaxEnt: {score5(donor):.2f}")
print(f"3'ss MaxEnt: {score3(acceptor):.2f}")
ScoreInterpretationSource
5'ss MaxEnt > 8Strong donorYeo & Burge 2004 J Comput Biol
5'ss MaxEnt 5-8Moderate
5'ss MaxEnt < 5Weak / cryptic
3'ss MaxEnt > 8Strong acceptor
3'ss MaxEnt < 5Weak / cryptic
SpliceAI delta >= 0.2PP3 (applied at supporting weight); BP4 at <= 0.1Walker 2023 AJHG (ClinGen SVI 2023)
SpliceAI delta 0.5 / 0.8Higher-precision cutoffs (SpliceAI recommended/high-precision tiers)Jaganathan 2019 Cell — NOT ClinGen graded evidence-strength upgrades

MaxEntScan vs SpliceAI:

  • MaxEntScan scores sequence information content (intrinsic strength). Captures position-wise dependencies at the consensus.
  • SpliceAI predicts in-vivo usage probability given full pre-mRNA context (10 kb window).
  • A position with high MaxEnt but low SpliceAI is intrinsically strong but contextually silenced (chromatin, trans factors).
  • A position with low MaxEnt but high SpliceAI is intrinsically weak but contextually used (enhancer-driven, e.g. weak donors stabilized by ESEs).
  • Report both for variant interpretation; for variant impact see splice-variant-prediction.

Picard CollectRnaSeqMetrics and Gene-Body Coverage

Goal: Get integrated RNA-seq QC including intronic / exonic / intergenic mapping rates and gene-body coverage uniformity.

Approach: Run picard CollectRnaSeqMetrics for mapping distribution; RSeQC geneBody_coverage.py for 5'-3' bias.

bash
picard CollectRnaSeqMetrics \
    I=sample.bam \
    O=sample.rna_metrics.txt \
    REF_FLAT=refFlat.txt \
    STRAND_SPECIFICITY=SECOND_READ_TRANSCRIPTION_STRAND \
    RIBOSOMAL_INTERVALS=rRNA_intervals.interval_list

# Strandedness conversion (foot-gun):
# Reverse-stranded (Illumina TruSeq Stranded; NEB Ultra II Directional — both dUTP):
#   rMATS  --libType fr-firststrand
#   featureCounts -s 2
#   Picard STRAND_SPECIFICITY=SECOND_READ_TRANSCRIPTION_STRAND
# Forward-stranded (Lexogen QuantSeq FWD, certain ligation-based kits):
#   rMATS  --libType fr-secondstrand
#   featureCounts -s 1
#   Picard STRAND_SPECIFICITY=FIRST_READ_TRANSCRIPTION_STRAND
# STAR has no library-strand flag; pass --outSAMstrandField intronMotif
# (works for any library) so downstream tools can read XS tags.

geneBody_coverage.py \
    -i sample.bam \
    -r gencode_v45.bed \
    -o sample_geneBody
MetricHealthyConcerning
PCT_CODING_BASES>=50%<30% (suggests degradation or mis-priming)
PCT_UTR_BASES20-40%>>50% (3' bias)
PCT_INTRONIC_BASES<30% (poly(A)); <60% (rRNA-depleted)>50% (poly(A)) suggests pre-mRNA contamination
PCT_INTERGENIC_BASES<10%>20% (genomic DNA contamination)
MEDIAN_5PRIME_TO_3PRIME_BIAS0.7-1.3>2 or <0.5 (severe degradation)
Gene body coverage curveFlatStrong 3' skew = RIN low or library mis-prep

3' bias (degraded RNA) directly reduces splicing-event detection because junction reads scatter across the gene body; with 3' bias they concentrate near the 3' end and miss CDS junctions.

Strandedness Verification

bash
infer_experiment.py -i sample.bam -r gencode_v45.bed -s 200000

Output reports the fraction of reads consistent with each library type:

Output patternLibrary typerMATS --libType
~50% / ~50%Unstrandedfr-unstranded
>=90% "++ , --"Forward-strandedfr-secondstrand
>=90% "+- , -+"Reverse-stranded (Illumina TruSeq stranded)fr-firststrand

Wrong strand setting halves usable junction reads — always verify before quantification. RSeQC infer_experiment.py is fast and authoritative.

Annotation Choice

GENCODE levelContentsUse for
BasicHigh-confidence canonical isoformsStandard rMATS, leafcutter, SUPPA2
ComprehensiveAll transcripts including putative/predictedDTU pipelines (DRIMSeq+DEXSeq, satuRn), isoform discovery
RefSeqNCBI curatedLess complete than GENCODE; legacy use
EnsemblSame content as GENCODE in vertebratesDifferent attribute conventions

Comprehensive captures more biology but inflates DTU multiple-testing burden and includes annotation noise. For event-level (rMATS) AS, basic is usually adequate; for transcript-level DTU (DRIMSeq, satuRn), comprehensive may be necessary to capture rare isoforms.

Show full SKILL.md (873 more words)Show less

rRNA Contamination Check

bash
fastq_screen --conf fastq_screen.conf --threads 8 sample_R1.fq.gz

Or post-alignment:

bash
samtools view -c sample.bam | awk '{print "total:",$0}'
samtools view -c -L rRNA_intervals.bed sample.bam | awk '{print "rRNA:",$0}'
rRNA fractionLibrary typeStatus
>=20%"depleted"Failed depletion; redo
5-20%"depleted"Acceptable; some rRNA leakage
<5%poly(A)Healthy
<5%"depleted"Excellent depletion
1-3%poly(A)Suggests RNA degradation

5% rRNA in a poly(A) library suggests degraded RNA; >20% in a "depleted" library indicates failed depletion.

Per-Tool Failure Modes

RSeQC junction_saturation: Subsampling Behavior

Trigger: Running on extremely deep BAM (>200M reads).

Mechanism: RSeQC subsamples at 5%, 10%, ..., 100%; with very deep BAMs, the early subsamples are still tens of millions of reads, masking saturation behavior.

Symptom: Curve appears flat throughout; uninformative.

Fix: Subsample BAM with samtools view -s 0.1 before running junction_saturation; or use -s flag to set custom step intervals.

STAR 2-Pass: Per-Sample Inconsistency

Trigger: Using --twopassMode Basic on differential splicing cohorts.

Mechanism: Per-sample 2-pass means each sample has its own SJ.out.tab; samples may differ in which novel junctions they re-align against.

Symptom: Inconsistent novel junction calls across replicates; rMATS --novelSS differential calls don't replicate.

Fix: Switch to cohort-style 2-pass (collect all pass-1 SJ.out.tabs, merge, re-align all samples with merged set).

MaxEntScan: Out-of-Range Sequences

Trigger: Sequences with N bases or wrong length.

Mechanism: score5 expects exactly 9 nt (3 exon + 6 intron); score3 expects 23 nt (20 intron + 3 exon).

Symptom: ValueError or silently incorrect score.

Fix: Pre-validate sequence length and N-content; use a wrapper that returns NaN for invalid inputs.

SpliceAI: TensorFlow Memory

Trigger: Running spliceai on large VCF without GPU.

Mechanism: TensorFlow CPU mode is slow; default batch size may exceed memory.

Symptom: OOM kill; very slow runtime (hours per chromosome).

Fix: Use -D 50 for screening (fastest); split VCF by chromosome; use GPU when available.

infer_experiment.py: Sample Size

Trigger: Running on very low-coverage region or small subsample (-s).

Mechanism: Default sample size is 200,000 reads; with low coverage, this isn't met.

Symptom: "0 of 200000 reads" output; cannot infer strand.

Fix: Lower -s to actual available reads; or use -q 30 to filter by quality.

Common Errors

ErrorCauseSolution
STAR: SJDBoverhang differs from genomeIndex built with different overhang than current runRebuild index with --sjdbOverhang matching read length - 1
RSeQC: BED format errorAnnotation BED has wrong column orderConvert with awk or bedtools
MaxEntScan: invalid sequence character NN in inputFilter or replace; document
samtools view: missing indexBAM not indexedsamtools index sample.bam
STAR: too many SJs in cohort mergeCohort SJ.out.tab too large after mergeFilter to junctions in >=3 samples or with >=3 unique reads
regtools: invalid CIGARNon-spec read in BAMFilter with samtools view -h -F 0x100 -F 0x800

Quality Thresholds

MetricGoodAcceptablePoorSource
Read length (PE)150 nt100 nt<75 ntconvention
Sequencing depth>=100M50-100M<30MDGE-grade insufficient
Junction saturationPlateau (<2% growth in last 20%)Near plateauStill risingRSeQC convention
Known-junction fraction>=80%60-80%<60% (suspect or interesting)RSeQC convention
Junctions >=10 reads>=50%30-50%<30%rMATS reliability cutoff
5'ss / 3'ss MaxEnt>85-8<5Yeo & Burge 2004
Strandedness>90% one direction70-90%<70%RSeQC convention
rRNA in depleted library<5%5-20%>20%convention
2-pass STARCohort-stylePer-sample basic1-pass onlyVeeneman 2016 Bioinformatics

Troubleshooting Low Event Detection

IssuePossible causesSolutions
Few events calledLow depth; short reads; SE; wrong strandIncrease depth; use PE150; verify libType
High novel junctionsAnnotation gaps; mapping artifacts; biology (TDP-43, SF3B1)Update annotation; check 2-pass; consider biology
Low IR detectionpoly(A) libraryUse rRNA depletion
Microexons missingDefault aligner anchors too longVAST-TOOLS, MicroExonator, or long-read
Many weak splice sitesCryptic splicingValidate with MaxEnt + SpliceAI; consider RNA-seq from secondary tissue
FDR uncalibrated at low nn=2 vs n=2Use leafcutter or Shiba; avoid SUPPA2 alone
PSI variance high across replicatesLibrary prep / RIN inconsistencyCheck RIN; consider RNA degradation
Sashimi plot mismatch with PSIJunction-imbalance bias in rMATSRun Shiba; or filter by overhang distribution

Common Pitfalls

  • Skipping STAR 2-pass — loses ~14% of novel junctions; matters for any non-canonical organism or condition.
  • Per-sample 2-pass instead of cohort-style — produces inconsistent junction sets; differential splicing calls don't replicate.
  • poly(A) library for IR analysis — biases toward mature transcripts; depletes pre-mRNA / nascent / detained intron signal.
  • PE 50nt single-end — junction-spanning reads need >=8nt overhang on both sides; biases toward shorter exons.
  • Wrong --libType — halves usable junctions; always verify with infer_experiment.py.
  • Using basic GENCODE for DTU — basic excludes putative/rare isoforms; DTU pipelines may underdetect.
  • Using MaxEntScan alone for variant interpretation — misses context-dependent regulation; pair with SpliceAI.
  • Treating high novel% as artifact reflexively — could be biology (TDP-43, SF3B1, non-model organism); investigate.
  • splicing-quantification - PSI estimation after QC passes
  • read-alignment/star-alignment - STAR 2-pass detail and parameter tuning
  • read-qc/quality-reports - General sequencing QC (FastQC, MultiQC)
  • read-qc/contamination-screening - rRNA / adapter / cross-species contamination
  • splice-variant-prediction - SpliceAI / Pangolin for variant impact
  • long-read-splicing - When short-read QC is fundamentally limiting (microexons, complex isoforms)
  • differential-splicing - Downstream tool that requires QC pass

References

  • Yeo & Burge 2004 J Comput Biol - MaxEntScan
  • Jaganathan et al 2019 Cell - SpliceAI
  • Walker et al 2023 Am J Hum Genet - ClinGen SVI splicing thresholds
  • Veeneman et al 2016 Bioinformatics - STAR 2-pass benchmark
  • Brown et al 2022 Nature - cryptic exons in TDP-43 loss
  • Klim et al 2019 Nat Neurosci - STMN2 cryptic splicing in ALS
  • Darman et al 2015 Cell Rep - SF3B1 cryptic 3'ss
  • Wang et al 2024 Nat Protoc - rMATS-turbo
  • Dobin et al 2013 Bioinformatics - STAR aligner

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in alternative-splicing/splicing-qc of GPTomics/bioSkills.

  • SKILL.md
  • examples/splicing_qc.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Splicing Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Splicing Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Splicing Qc this skillGPTomics/bioSkills1.2k2 repos~6.2kAutomated safety check: PassMIT
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Aviv RegevK-Dense-AI/mimeographs129—~1.4kAutomated safety check: PassMIT
Rare Disease RnaseqClawBio/ClawBio1.2k1 repos~1.2kAutomated safety check: PassMIT
Bio Proteomics Data ImportFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~1.2kAutomated safety check: PassNone
Knn Imputationaipoch/medical-research-skills1.9k—~2.5kAutomated safety check: PassMIT

Similar skills

  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Aviv Regev

    K-Dense-AI/mimeographs

    Applies the computational biology and AI-driven reasoning of Aviv Regev (computational biologist, Genentech, single-cell genomics).

    129 GitHub stars~1.4k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Rare Disease Rnaseq

    ClawBio/ClawBio

    Blood RNA-seq expression-outlier detection for rare-disease diagnostics.

    1.2k GitHub starsUsed in 1 repo~1.2k tokens
    Research & ScienceAuto-check passed
  • Bio Proteomics Data Import

    FreedomIntelligence/OpenClaw-Medical-Skills

    Load and parse mass spectrometry data formats including mzML, mzXML, and quantification tool outputs like MaxQuant proteinGroups.txt.

    3.1k GitHub starsUsed in 1 repo~1.2k tokens
    Research & ScienceAuto-check passed
  • Knn Imputation

    aipoch/medical-research-skills

    A skill your agent uses when filtering genes with high missingness and then imputing missing values in a bulk expression matrix with group-aware KNN through DMwR2, where donor samples are restricted…

    1.9k GitHub stars~2.5k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Popv Cell Annotation

    jaechang-hits/SciAgent-Skills

    Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via…

    374 GitHub starsUsed in 2 repos~6.9k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Splicing Qc

What does Bio Splicing Qc do?

Assesses RNA-seq data quality specifically for alternative splicing analysis. Bio Splicing Qc is an agent skill from GPTomics/bioSkills. Assesses RNA-seq data quality specifically for alternative splicing analysis.

When should I use Bio Splicing Qc?

Bio Splicing Qc fits situations like: evaluating data suitability for splicing analysis; troubleshooting low event detection; designing sequencing experiments where AS is a primary endpoint.

How do I install Bio Splicing Qc in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-splicing-qc -a claude-code`. Or copy the skill folder (alternative-splicing/splicing-qc in GPTomics/bioSkills) into .claude/skills/bio-splicing-qc in your project. Claude Code loads it when a task matches its description.

How do I install Bio Splicing Qc in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-splicing-qc -a codex`. Or copy the skill folder (alternative-splicing/splicing-qc in GPTomics/bioSkills) into .agents/skills/bio-splicing-qc in your project. Codex loads it when a task matches its description.

Can I use Bio Splicing Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-splicing-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-splicing-qc, .gemini/skills/bio-splicing-qc, .github/skills/bio-splicing-qc and .opencode/skills/bio-splicing-qc in your project.

What does Bio Splicing Qc need to run?

Going by SKILL.md and its folder, Bio Splicing Qc needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Splicing Qc access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Splicing Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Splicing Qc use?

Bio Splicing Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Splicing Qc use?

About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Splicing Qc?

Skills that share tags, products or a category with Bio Splicing Qc: Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars), Aviv Regev (K-Dense-AI/mimeographs, 129 stars), Rare Disease Rnaseq (ClawBio/ClawBio, 1.2k stars) and Bio Proteomics Data Import (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Splicing Qc?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.