Agent skill

Bio Outlier Splicing Detection

by GPTomics in GPTomics/bioSkills

Detects aberrant splicing in single rare-disease patients vs a control panel using FRASER 2.0 (Bioconductor; Beta-binomial autoencoder on Intron Jaccard Index, default delta cutoff 0.1, q…

MITAuto-check passedResearch & Science

Install Bio Outlier Splicing Detection

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-outlier-splicing-detection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-outlier-splicing-detection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alternative-splicing/outlier-splicing-detection .claude/skills/bio-outlier-splicing-detection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-outlier-splicing-detection
GitHub stars
1.2k
Used in
2 other repos
Token cost
~5.1k tokens
SKILL.md length
1,985 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Detects aberrant splicing in single rare-disease patients vs a control panel using FRASER 2.0 (Bioconductor; Beta-binomial autoencoder on Intron Jaccard Index, default delta cutoff 0.1, q…

  • Applying RNA-seq to undiagnosed Mendelian disease
  • SKILL.md covers Version Compatibility, Tool Taxonomy, Decision Tree by Diagnostic… and When to Use Outlier vs…, plus 16 more sections
  • Runs R scripts from its folder; calls python, mamba and conda
  • Validating predicted splice variants in clinical samples

What it does

Bio Outlier Splicing Detection is an agent skill from GPTomics/bioSkills. Detects aberrant splicing in single rare-disease patients vs a control panel using FRASER 2.0 (Bioconductor; Beta-binomial autoencoder on Intron Jaccard Index, default delta cutoff 0.1, q hyperparameter), OUTRIDER (gene-level outlier expression via autoencoder denoising), LeafcutterMD (Dirichlet-multinomial outlier mode of LeafCutter for annotation-free junctions), and DROP (Snakemake pipeline integrating FRASER2 + OUTRIDER + monoallelic expression for clinical diagnostics). The statistical model is fundamentally…

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics, Data cleaning and Reproducible research. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Applying RNA-seq to undiagnosed Mendelian disease
  • Validating predicted splice variants in clinical samples
  • Detecting cryptic splicing in disease tissue

Example prompts

  • “Use the bio-outlier-splicing-detection skill to detect aberrant splicing in single rare-disease patients vs a control panel using FRASER 2.0…”
  • “/bio-outlier-splicing-detection”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • mamba
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Outlier Splicing Detection loads about 5.1k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 1,985 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~224
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,985 words, ~5,115 tokens.

Download SKILL.mdSave it as .claude/skills/bio-outlier-splicing-detection/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-outlier-splicing-detection
description
Detects aberrant splicing in single rare-disease patients vs a control panel using FRASER 2.0 (Bioconductor; Beta-binomial autoencoder on Intron Jaccard Index, default delta cutoff 0.1, q hyperparameter), OUTRIDER (gene-level outlier expression via autoencoder denoising), LeafcutterMD (Dirichlet-multinomial outlier mode of LeafCutter for annotation-free junctions), and DROP (Snakemake pipeline integrating FRASER2 + OUTRIDER + monoallelic expression for clinical diagnostics). The statistical model is fundamentally different from differential splicing — single-sample-vs-cohort outlier detection rather than two-group comparison. Standard tool in EU rare-disease (Solve-RD) and NIH UDN programs. Use when applying RNA-seq to undiagnosed Mendelian disease, validating predicted splice variants in clinical samples, or detecting cryptic splicing in disease tissue.
tool_type
r
primary_tool
FRASER

Version Compatibility

Reference examples tested with: FRASER 2.0 (>=1.99.0), OUTRIDER 1.20+, LeafcutterMD via leafcutter 0.2.9+, DROP 1.4+, R 4.4+, BiocManager 1.30+

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Outlier Splicing Detection

For clinical RNA-seq diagnostics in rare disease, the question is not "what differs between groups?" but "what is aberrant in this single patient relative to a panel of unaffected samples?". The statistical framework is single-sample-vs-cohort outlier detection, fundamentally different from two-group differential splicing. Tools in this space are designed for clinical Mendelian diagnostic settings.

Tool Taxonomy

ToolStatisticTest targetFails when
FRASER 2.0Beta-binomial autoencoder on Intron Jaccard IndexSplicing outliers (per-sample, per-junction)Cohort <20 samples; tissue mismatch
OUTRIDERAutoencoder-denoised expression Z-scoreGene-level expression outliers (LoF, monoallelic)Cohort <20 samples
LeafcutterMDDirichlet-multinomial outlier modeAnnotation-free intron usageBeta-binomial fits poorly OR few controls
DROPSnakemake pipelineAll of above + monoallelic expressionPipeline complexity for small projects

Core reference: FRASER 2.0 for splicing outliers, OUTRIDER for expression outliers, DROP to combine. Standard tool in EU rare-disease programs (Solve-RD) and NIH UDN.

Decision Tree by Diagnostic Scenario

ScenarioRecommended approach
Single rare-disease patient + panel of n>=50 controlsFRASER 2.0 (Intron Jaccard Index)
Single patient + small panel (n=20-50)FRASER 2.0 with auxiliary GTEx controls; tune q carefully
Patient + cohort <20Insufficient for outlier detection; consider differential or recruit more samples
Outlier expression suspected (loss of function, monoallelic)OUTRIDER on same cohort
Annotation-free outlier (cryptic exon, novel junction)LeafcutterMD
Integrated diagnostic pipeline (splicing + expression + MAE)DROP
TDP-43 ALS post-mortem brain (cryptic exons)FRASER 2.0; expect UNC13A, STMN2, ATG4B
SF3B1-mutant cancer sampleFRASER 2.0 with cohort-matched RNA-seq; expect cryptic 3'ss
Familial dysautonomia (ELP1)FRASER 2.0 in fibroblast/iPSC; CNS tissue gives strongest signal
Stargardt deep-intronic ABCA4FRASER 2.0 in retina-relevant tissue
Solid tumor splicing biomarkerDifferential splicing (n>=10 vs cohort) — see differential-splicing skill
RNA validation of SpliceAI hitFRASER 2.0 + cross-reference with predicted variant location

When to Use Outlier vs Differential

Outlier regime (this skill):

  • Single patient or small case series vs control panel
  • Question: "What is aberrant in this patient?"
  • Statistical model: single-sample p-value vs cohort distribution

Differential regime (differential-splicing skill):

  • Two well-defined groups, n>=3 each
  • Question: "What differs between groups?"
  • Statistical model: two-group LRT or related

If n>=10 patients with a shared phenotype are available, prefer differential (more power); if single patient or heterogeneous case series, use outlier.

FRASER 2.0 Workflow

Goal: Detect aberrant splicing in patient samples vs cohort using the Intron Jaccard Index.

Approach: Count split reads per junction, compute Intron Jaccard Index per intron, fit a Beta-binomial autoencoder to estimate expected values, then flag outliers by p-value and delta.

r
library(FRASER); library(BiocParallel)

bam_files <- list.files('bams/', pattern='.bam$', full.names=TRUE)
sample_table <- data.frame(
    sampleID = gsub('.bam', '', basename(bam_files)),
    bamFile = bam_files,
    pairedEnd = TRUE
)

settings <- FraserDataSet(
    colData = sample_table,
    workingDir = 'fraser_workdir',
    name = 'rare_disease_cohort'
)

settings <- countRNAData(settings, BPPARAM = MulticoreParam(8))
fds <- calculatePSIValues(settings)

fds <- filterExpressionAndVariability(
    fds,
    minDeltaPsi = 0.0,
    minExpressionInOneSample = 20,
    quantile = 0.05,
    quantileMinExpression = 1
)

fitMetrics(fds) <- 'jaccard'
currentType(fds) <- 'jaccard'  # canonical setter for active metric in FRASER 2.0
fds <- FRASER(
    fds,
    q = c(jaccard = 10),
    BPPARAM = MulticoreParam(8)
)

results <- results(
    fds,
    psiType = 'jaccard',
    padjCutoff = 0.05,
    deltaPsiCutoff = 0.1
)

patient_results <- results[results$sampleID == 'PATIENT_001', ]
patient_results <- patient_results[order(patient_results$padjust), ]

FRASER 2.0 changes vs FRASER 1.x:

  • Default psiType changed from three metrics (psi5, psi3, theta) to single Intron Jaccard Index
  • Default deltaPsiCutoff dropped from 0.3 to 0.1
  • Pseudocount and filtering parameter optimization
  • Bioconductor package version >=1.99.0 == FRASER 2.0

q = 10 is the autoencoder dimension hyperparameter. Tune via estimateBestQ(fds, type='jaccard') for cohort-specific optimum — too low: confounders not removed; too high: real signal absorbed.

OUTRIDER for Gene-Level Outlier Expression

Goal: Detect genes with aberrantly high or low expression in patient samples.

Approach: Autoencoder denoising of expression matrix; outliers identified by Z-score and adjusted p-value.

r
library(OUTRIDER); library(BiocParallel)

countTable <- read.table('counts.tsv', header=TRUE, row.names=1)
ods <- OutriderDataSet(countData = countTable)

ods <- filterExpression(ods, minCounts=TRUE, filterGenes=TRUE)
# OUTRIDER's estimateBestQ returns a scalar q (unlike FRASER's, which returns the object)
q_best <- estimateBestQ(ods)
ods <- OUTRIDER(ods, q = q_best, BPPARAM = MulticoreParam(8))

res <- results(ods, padjCutoff = 0.05, zScoreCutoff = 0)
patient_outliers <- res[res$sampleID == 'PATIENT_001', ]

OUTRIDER (Brechtmann 2018 Am J Hum Genet) catches loss-of-function alleles producing transcript collapse, monoallelic effects, and tissue-inappropriate expression — complements splice outlier detection.

LeafcutterMD for Annotation-Free Outlier Intron Usage

Goal: Detect outlier intron usage relative to a control panel without annotation dependence.

Approach: Run LeafcutterMD (LeafCutter's Dirichlet-multinomial outlier mode for Mendelian disease) against the control panel.

bash
for bam in *.bam; do
    regtools junctions extract -a 8 -m 50 -s XS "$bam" -o "${bam%.bam}.junc"
done

ls *.junc > juncfiles.txt
python leafcutter_cluster_regtools.py -j juncfiles.txt -o leafcutter -m 50 -l 500000

leafcutterMD.R \
    --num_threads 4 \
    --output_prefix patient_outlier \
    leafcutter_perind_numers.counts.gz

LeafcutterMD (Jenkinson 2020 Bioinformatics) reports per-sample p-values per intron-cluster; useful when FRASER's Beta-binomial model fits poorly or when novel-junction sensitivity matters.

DROP Pipeline (Integrated Workflow)

Goal: Run FRASER2 + OUTRIDER + monoallelic expression in a unified Snakemake pipeline for clinical diagnostics.

Approach: DROP is distributed via bioconda (not PyPI). Install in a dedicated environment, then configure with patient + control sample sheets; pipeline handles QC, alignment, counting, autoencoding, and reporting.

bash
# Install via bioconda (DROP is not on PyPI)
mamba create -n drop_env -c conda-forge -c bioconda drop --override-channels
conda activate drop_env

drop init my_diagnostic_run
cd my_diagnostic_run

# Edit config.yaml:
#  - sample_table: samples.tsv (patient + controls)
#  - aberrantSplicing: enabled
#  - aberrantExpression: enabled
#  - mae: enabled (monoallelic expression)

snakemake --cores 16 --use-conda

DROP (Yepez 2021 Nat Protocols) is the standard tool in EU rare-disease genome+RNA-seq programs (Solve-RD) and the NIH UDN. v1.4+ uses FRASER 2.0. The MAE module uses a custom z-score test on heterozygous SNPs from RNA-seq (allele-specific expression) — useful for catching dominant-negative or monoallelic LoF that splicing/expression outliers miss. Cohort >=30 samples recommended for confident outlier detection.

Variant + Outlier Integration

Goal: Connect a candidate splice-altering DNA variant to RNA-level confirmation.

Approach: Cross-reference SpliceAI hits with FRASER2 outliers in the same sample.

r
library(dplyr)

variants <- read.table('spliceai_hits.tsv', header=TRUE, sep='\t')
fraser_hits <- read.table('fraser_results.tsv', header=TRUE, sep='\t')

confirmed <- variants %>%
    filter(delta_max >= 0.2) %>%
    inner_join(
        fraser_hits %>% filter(sampleID == 'PATIENT_001', padjust < 0.05),
        by = c('chrom' = 'seqnames'),
        relationship = 'many-to-many'
    ) %>%
    filter(abs(pos - start) < 1000 | abs(pos - end) < 1000)

A SpliceAI hit + concordant FRASER2 outlier in the patient = strong PS3 functional evidence in the ACMG framework. This integration is the highest-value clinical pipeline step — converts a computational PP3 to functional PS3.

Cohort Size and Power

Cohort sizePowerComment
n < 20MarginalHigh FDR; consider GTEx tissue-matched controls as auxiliary
n = 20-50AcceptableFRASER autoencoder can fit; tune q carefully
n >= 50RecommendedStandard clinical diagnostic cohort size
n >= 100OptimalTissue-matched and batch-matched gives best calibration

GTEx-derived tissue-matched controls can supplement small in-house cohorts but introduce batch effects; use only when in-house n < 30 and document the pooling strategy.

Tissue Choice for Mendelian RNA-seq

TissueProsConsGenes captured
Whole blood (PAXgene)Easy, standardGlobin contamination; many disease genes silent~70-80% of clinical genes
Fibroblast (skin biopsy)Reasonable expressionRequires culture; senescence variability~75-85%
Muscle biopsyBest for muscular dystrophyInvasive~85-90% for muscle disorders
iPSC-derived neuron / cardiomyocyteDisease-relevant tissueCost, variability~95% if differentiation works
Urine sedimentNon-invasiveLow yield~50-60%

For UDN-style cases: blood first, then fibroblast if blood lacks expression of candidate gene. Critical: a negative blood RNA-seq doesn't rule out a candidate gene that's silent in blood — verify gene expression with GTEx tissue panel before committing to the tissue.

Hyperparameter Tuning

r
# useOHT=FALSE runs the injection-based q grid so plotEncDimSearch has a curve to show;
# useOHT=TRUE (default) is the fast deterministic OHT estimate but produces no search table to plot.
fds <- estimateBestQ(fds, type='jaccard', useOHT=FALSE, q_param=c(2, 5, 10, 15, 20))
plotEncDimSearch(fds, type='jaccard')

The encoding dimension q should be where the loss curve plateaus. Too low: confounders not removed; too high: real signal absorbed.

For typical 50-100 sample cohorts, q=8-15 is the usual operating range (DROP / FRASER workflow convention; no single primary citation — verify with plotEncDimSearch on the actual cohort). For very small cohorts (n=20-30), q=5-8 is typical.

Per-Tool Failure Modes

FRASER 2.0: Q Hyperparameter Mistuning

Trigger: Default q=10 used without tuning; or wrong q for cohort size.

Mechanism: Q is the autoencoder bottleneck dimension; too small -> confounders leak into outlier signal; too large -> real biological signal absorbed by autoencoder.

Symptom: Either no significant outliers (q too high) or many spurious calls clustering by batch (q too low).

Fix: Run estimateBestQ(fds, type='jaccard', useOHT=TRUE) (Optimal Hard Thresholding default; very fast); use the returned bestQ(fds) value. For exhaustive search, pass useOHT=FALSE, q_param=c(2, 5, 10, 15) and inspect plotEncDimSearch for the plateau.

FRASER 2.0: Tissue Mismatch

Trigger: Patient sample from different tissue than majority of controls.

Mechanism: FRASER autoencoder learns tissue-specific expression patterns; tissue-mismatched patient appears as global outlier.

Symptom: Hundreds of "significant" outliers in patient; not biologically interpretable.

Fix: Strict tissue matching; if controls are mixed-tissue, use only controls from patient's tissue.

Show full SKILL.md (813 more words)Show less
OUTRIDER: Few Controls

Trigger: Cohort <20 samples.

Mechanism: Autoencoder needs sufficient samples to learn expression covariance; fails to fit at very small cohort sizes.

Symptom: Convergence warnings; uncalibrated p-values.

Fix: Pool with GTEx auxiliary controls; or use simpler outlier methods (z-score on log-CPM).

LeafcutterMD: Cluster Count Limits

Trigger: Very few clusters in patient sample (low coverage or filtered out).

Mechanism: LeafcutterMD fits a Dirichlet-multinomial (Beta-binomial per-intron) model over cluster counts; few observations -> unstable fit -> unreliable p-values.

Symptom: Inflated or deflated p-values; few significant calls.

Fix: Increase coverage; relax filtering (-m 10 instead of 50); or switch to FRASER2.

DROP: Snakemake Pipeline Failures

Trigger: Missing dependencies or incompatible R/Bioconductor versions.

Mechanism: DROP orchestrates many tools; version mismatches cascade through pipeline.

Symptom: Snakemake step fails partway through; cryptic R errors.

Fix: Use --use-conda flag for environment isolation; pin versions in environment yamls.

Reconciliation: When Outlier Tools Disagree

PatternLikely causeAction
FRASER2 sig, OUTRIDER notSplicing change without expression collapseStandard splicing outlier; report
OUTRIDER sig, FRASER2 notExpression LoF without splicing changeLikely promoter / regulatory; not splicing
Both sig at same geneLoF allele triggering NMD on splicing-disrupted transcriptStrong combined evidence; expect downstream
LeafcutterMD sig, FRASER2 notNovel cryptic event not in annotationHigh-priority novel finding; investigate
All tools null but biology suggests changeUnderpowered cohort or wrong tissueVerify gene expression in tissue; recruit larger cohort

Disease-Specific Expectations

ConditionExpected outlier signatureTissue
ALS / FTD (TDP-43 loss)Cryptic exons in UNC13A, STMN2, ATG4BPost-mortem brain ONLY
SF3B1-mutant MDS / CLL / uveal melanomaAberrant 3'ss ~10-30nt upstream of canonicalBone marrow / tumor tissue
Spinal muscular atrophy (untreated SMN2)SMN exon 7 skippingFibroblast / iPSC-MN
Familial dysautonomia (ELP1 c.2204+6T>C)ELP1 exon 20 skipping (>=99% in CNS, partial elsewhere)iPSC-neuron > fibroblast > blood
Deep-intronic CFTR / USH2A / CEP290Pseudoexon inclusionCognate disease tissue (lung / retina)
Duchenne muscular dystrophy (DMD)Out-of-frame exon skipping patternMuscle biopsy
Stargardt (ABCA4) deep-intronicPseudoexon in retinaRetinal organoid / iPSC-RPE

For each, the gene must be expressed in the queried tissue. Verify with GTEx before assuming negative result rules out the gene.

Common Errors

ErrorCauseSolution
FRASER: cohort too smalln<10Pool with auxiliary controls; or recruit more patients
FRASER: countRNAData failed on chromosome XBAM index missing or corruptedRe-index BAMs; check samtools idxstats
estimateBestQ: convergence not reachedDefault q range insufficientExpand q_param=c(2,5,10,15,20); or use useOHT=TRUE for the fast deterministic alternative
OUTRIDER: encoding-dim search slowfindEncodingDim grids many q valuesUse estimateBestQ(ods) for a fast single-q estimate
DROP: snakemake job failed at FRASERDROP-FRASER version mismatchUpdate DROP to latest; verify FRASER 2.0 compatibility
LeafcutterMD: insufficient clustersCluster filter too strictLower -m minimum cluster reads
Variant integration: chrom format mismatchVCF uses 1, FRASER uses chr1 (or vice versa)Standardize with bcftools annotate --rename-chrs

Common Pitfalls

  • Using bulk differential-splicing tools for n=1 vs cohort — rMATS, leafcutter (regular), SUPPA2 are not designed for this. Use FRASER2 / LeafcutterMD.
  • Ignoring tissue choice — clinical gene expression varies dramatically across blood / fibroblast / muscle. A negative blood RNA-seq doesn't rule out a candidate gene that's silent in blood.
  • Forgetting batch effects — combining in-house and external (GTEx) controls introduces sequencing batch confounding; use ComBat or include batch as covariate.
  • Skipping the variant + outlier integration — RNA-only outlier without DNA variant suggests cellular state or technical artifact; DNA-only prediction without RNA confirmation is supporting only (PP3, not PS3).
  • Treating all FRASER2 outliers as pathogenic — many are benign tissue-specific variation. Filter against gnomAD splice constraint and disease gene panels.
  • Q hyperparameter not tuned — default q=10 works for ~50-100 sample cohorts; tune for outliers.
  • Wrong default delta cutoff for FRASER 1.x vs 2.0 — 1.x default 0.3, 2.0 default 0.1; document which version.

Quality Thresholds

MetricRecommendationSource
Cohort sizen>=50 (ideal); n>=20 (minimum)Solve-RD / UDN convention
FRASER 2.0 padj< 0.05Standard
FRASER 2.0 delta Jaccard>= 0.1 (default in v2.0)Scheller 2023 AJHG
OUTRIDER padj< 0.05Brechtmann 2018 AJHG
OUTRIDER zScoreabs >= 2Conservative
Sequencing depth>=50M PE reads/sampleStandard for AS analysis
Tissue match between patient and controlsRequiredCritical for FRASER2 calibration
Batch matchStrongly recommendedReduces autoencoder confounding
  • splice-variant-prediction - SpliceAI / Pangolin for in-silico prediction; integration target
  • differential-splicing - When testing multiple patients vs controls (>=10 vs cohort)
  • splicing-qc - Library / depth / tissue prerequisites
  • variant-calling/clinical-interpretation - ACMG/AMP framework integration
  • workflows/clinical-trial-pipeline - Trial-grade RNA-seq diagnostics

References

  • Mertes et al 2021 Nat Commun - FRASER 1.x
  • Scheller et al 2023 Am J Hum Genet - FRASER 2.0 (Intron Jaccard Index)
  • Brechtmann et al 2018 Am J Hum Genet - OUTRIDER
  • Jenkinson et al 2020 Bioinformatics - LeafcutterMD
  • Yepez et al 2021 Nat Protocols - DROP pipeline
  • Cummings et al 2017 Sci Transl Med - RNA-seq for muscular dystrophy diagnostics
  • Kremer et al 2017 Nat Commun - RNA-seq for mitochondrial disease
  • Brown et al 2022 Nature - UNC13A cryptic exon (TDP-43 / ALS)
  • Klim et al 2019 Nat Neurosci - STMN2 cryptic splicing (ALS)
  • Darman et al 2015 Cell Rep - SF3B1 cryptic 3'ss
  • Walker et al 2023 Am J Hum Genet - ClinGen SVI splicing recommendations

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in alternative-splicing/outlier-splicing-detection of GPTomics/bioSkills.

  • SKILL.md
  • examples/fraser2_rare_disease.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Outlier Splicing Detection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Outlier Splicing Detection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Outlier Splicing Detection this skillGPTomics/bioSkills1.2k2 repos~5.1kAutomated safety check: PassMIT
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
LaminDB Biological Data Managementdavila7/claude-code-templates33k12 repos~3.6kAutomated safety check: PassMIT
Bio Proteomics Data ImportFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~1.2kAutomated safety check: PassNone
Knn Imputationaipoch/medical-research-skills1.9k—~2.5kAutomated safety check: PassMIT
AI Scientist EvaluatorBioTender-max/awesome-bio-agent-skills200—~2.4kAutomated safety check: PassCustom licence

Similar skills

  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    33k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Bio Proteomics Data Import

    FreedomIntelligence/OpenClaw-Medical-Skills

    Load and parse mass spectrometry data formats including mzML, mzXML, and quantification tool outputs like MaxQuant proteinGroups.txt.

    3.1k GitHub starsUsed in 1 repo~1.2k tokens
    Research & ScienceAuto-check passed
  • Knn Imputation

    aipoch/medical-research-skills

    A skill your agent uses when filtering genes with high missingness and then imputing missing values in a bulk expression matrix with group-aware KNN through DMwR2, where donor samples are restricted…

    1.9k GitHub stars~2.5k tokensUpdated 23 days ago
    Research & ScienceAuto-check passed
  • AI Scientist Evaluator

    BioTender-max/awesome-bio-agent-skills

    Critically review, score, compare, and rank one or more AI scientist outputs for biology, bioinformatics, computational life science, or adjacent research tasks.

    200 GitHub stars~2.4k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Latchbio Integration

    davila7/claude-code-templates

    Latch platform for bioinformatics workflows. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 11 repos~2.4k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Outlier Splicing Detection

What does Bio Outlier Splicing Detection do?

Detects aberrant splicing in single rare-disease patients vs a control panel using FRASER 2.0 (Bioconductor; Beta-binomial autoencoder on Intron Jaccard Index, default delta cutoff 0.1, q…. Bio Outlier Splicing Detection is an agent skill from GPTomics/bioSkills.1, q hyperparameter), OUTRIDER (gene-level outlier expression via autoencoder denoising), LeafcutterMD (Dirichlet-multinomial outlier mode of LeafCutter for annotation-free junctions), and DROP (Snakemake pipeline integrating FRASER2 + OUTRIDER + monoallelic expression for clinical diagnostics).

When should I use Bio Outlier Splicing Detection?

Bio Outlier Splicing Detection fits situations like: applying RNA-seq to undiagnosed Mendelian disease; validating predicted splice variants in clinical samples; detecting cryptic splicing in disease tissue.

How do I install Bio Outlier Splicing Detection in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-outlier-splicing-detection -a claude-code`. Or copy the skill folder (alternative-splicing/outlier-splicing-detection in GPTomics/bioSkills) into .claude/skills/bio-outlier-splicing-detection in your project. Claude Code loads it when a task matches its description.

How do I install Bio Outlier Splicing Detection in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-outlier-splicing-detection -a codex`. Or copy the skill folder (alternative-splicing/outlier-splicing-detection in GPTomics/bioSkills) into .agents/skills/bio-outlier-splicing-detection in your project. Codex loads it when a task matches its description.

Can I use Bio Outlier Splicing Detection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-outlier-splicing-detection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-outlier-splicing-detection, .gemini/skills/bio-outlier-splicing-detection, .github/skills/bio-outlier-splicing-detection and .opencode/skills/bio-outlier-splicing-detection in your project.

What does Bio Outlier Splicing Detection need to run?

Going by SKILL.md and its folder, Bio Outlier Splicing Detection needs R for the scripts in its folder and the command-line tools its instructions call (python, mamba and conda). Our summary lists: Python 3.

Does Bio Outlier Splicing Detection access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Outlier Splicing Detection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Outlier Splicing Detection use?

Bio Outlier Splicing Detection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Outlier Splicing Detection use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Outlier Splicing Detection?

Skills that share tags, products or a category with Bio Outlier Splicing Detection: CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars), LaminDB Biological Data Management (davila7/claude-code-templates, 33k stars), Bio Proteomics Data Import (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars) and Knn Imputation (aipoch/medical-research-skills, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Outlier Splicing Detection?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.