Agent skill

Bio Workflows Clip Pipeline

by GPTomics in GPTomics/bioSkills

End-to-end CLIP-seq pipeline from FASTQ to ENCODE-compliant binding sites, single-nucleotide crosslink maps, annotation, motifs, and (optionally) differential binding.

MITAuto-check passedResearch & Science

Install Bio Workflows Clip Pipeline

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-workflows-clip-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-workflows-clip-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workflows/clip-pipeline .claude/skills/bio-workflows-clip-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-workflows-clip-pipeline
GitHub stars
1.2k
Used in
2 other repos
Token cost
~5.1k tokens
SKILL.md length
1,343 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

End-to-end CLIP-seq pipeline from FASTQ to ENCODE-compliant binding sites, single-nucleotide crosslink maps, annotation, motifs, and (optionally) differential binding.

  • Works in 10 steps: Quality Control of Raw FASTQ → Preprocessing (Protocol-Specific) → Alignment (ENCODE STAR Block) → …
  • Running the full Yeo lab eCLIP / iCLIP / iCLIP2 / iCLIP3 / irCLIP / PAR-CLIP analysis with SMInput control
  • SKILL.md covers Version Compatibility, The governing principle, Pipeline Overview and CLIP Variant Selection, plus 15 more sections
  • Runs Shell scripts from its folder; calls python and pip

What it does

Bio Workflows Clip Pipeline is an agent skill from GPTomics/bioSkills. End-to-end CLIP-seq pipeline from FASTQ to ENCODE-compliant binding sites, single-nucleotide crosslink maps, annotation, motifs, and (optionally) differential binding. Use when running the full Yeo lab eCLIP / iCLIP / iCLIP2 / iCLIP3 / irCLIP / PAR-CLIP analysis with SMInput control, protocol-specific UMI extraction, ENCODE STAR parameters, CLIPper or Skipper peak calling with stringent log2 FC and -log10 p thresholds, IDR rescue and self-consistency QC, and downstream motif registration with mCross or PEKA.

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/clip_full_pipeline.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and End-to-end testing. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Running the full Yeo lab eCLIP / iCLIP / iCLIP2 / iCLIP3 / irCLIP / PAR-CLIP analysis with SMInput control
  • Protocol-specific UMI extraction
  • ENCODE STAR parameters
  • Skipper peak calling with stringent log2 FC and -log10 p thresholds

Example prompts

  • “/bio-workflows-clip-pipeline”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Quality Control of Raw FASTQ
  2. Preprocessing (Protocol-Specific)
  3. Alignment (ENCODE STAR Block)
  4. QC (Five Gates)
  5. Peak Calling
  6. Single-Nucleotide Crosslink-Site Detection
  7. IDR Across Replicates
  8. Binding-Site Annotation
  9. Motif Analysis (De Novo + CL-Registered)
  10. Differential Binding (Optional, Across Conditions)

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Workflows Clip Pipeline loads about 5.1k tokens when it runs. Until then it costs about 135 tokens; SKILL.md has 1,343 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,343 words, ~5,080 tokens.

Download SKILL.mdSave it as .claude/skills/bio-workflows-clip-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-workflows-clip-pipeline
description
End-to-end CLIP-seq pipeline from FASTQ to ENCODE-compliant binding sites, single-nucleotide crosslink maps, annotation, motifs, and (optionally) differential binding. Use when running the full Yeo lab eCLIP / iCLIP / iCLIP2 / iCLIP3 / irCLIP / PAR-CLIP analysis with SMInput control, protocol-specific UMI extraction, ENCODE STAR parameters, CLIPper or Skipper peak calling with stringent log2 FC and -log10 p thresholds, IDR rescue and self-consistency QC, and downstream motif registration with mCross or PEKA.
workflow
true
depends_on
clip-seq/clip-preprocessing, clip-seq/clip-alignment, clip-seq/clip-qc, clip-seq/clip-peak-calling, clip-seq/crosslink-site-detection…
qc_checkpoints
preprocessing_retention, alignment_rate, library_complexity, frip, idr
tool_type
mixed
primary_tool
CLIPper

Version Compatibility

Reference examples tested with: umi_tools 1.1.5+, cutadapt 4.6+, fastp 0.23+, STAR 2.7.11b+, samtools 1.19+, bedtools 2.31+, CLIPper 2.0+, Skipper (commit 2023.05+), PureCLIP 1.3.1+, HOMER 4.11+, ChIPseeker 1.40+, preseq 3.2+, picard 3.1+, idr 2.0.4+, MultiQC 1.21+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws unexpected errors, introspect the installed tool and adapt the example rather than retrying.

CLIP-seq End-to-End Pipeline

"Analyze my CLIP-seq data from raw FASTQ to ENCODE-compliant binding sites" -> Orchestrate protocol-specific UMI extraction, 3'-only adapter trimming (preserving the R2 5' truncation = crosslink site -1), ENCODE STAR alignment, UMI-based deduplication, library complexity QC, peak calling against SMInput with stringent thresholds (log2 FC >= 3 AND -log10 p >= 3), single-nucleotide crosslink-site detection, ChIPseeker annotation with CLIP-appropriate tssRegion, motif discovery with GC-matched background and CL-position registration, and optional differential binding between conditions.

This is a workflow skill: it owns the chaining decisions and hand-offs, not the internals of any one step.

The governing principle

A CLIP callset is decided at four seams, not inside the peak caller.

  1. The protocol variant is the master commitment made once at the top. It fixes the UMI pattern, the STAR mismatch ceiling, and the crosslink signal (truncation vs PAR-CLIP T->C vs STAMP C->U edit). The CLIP Variant Selection table below is that decision; everything downstream inherits it.
  2. Every IP is normalized against its SMInput with ENCODE-stringent thresholds (log2 FC >= 3 AND -log10 p >= 3). Peaks called without SMInput normalization are enrichment-uncontrolled and dominated by abundance.
  3. The R2 5' end IS the crosslink site (-1), so preprocessing and alignment must PRESERVE it. Trim 3'-only and permissively (-q 6, never -g on R1), and align with STAR --alignEndsType EndToEnd - soft-clipping or aggressive 5' trimming destroys the truncation base and with it single-nucleotide resolution.
  4. 40-70% PCR duplication is BY DESIGN; low duplication signals a FAILED IP, not a clean library. The IP enriches a small molecule pool, so the real quality metric is the UNIQUE-fragment count after UMI dedup - never the raw duplication rate.

Pipeline Overview

FASTQ + SMInput
  -> [clip-preprocessing]    UMI extract + 3' adapter trim (-q 6 -m 18) + two-pass for eCLIP
  -> [clip-alignment]        STAR ENCODE block (alignEndsType EndToEnd, mismatch 0.04 or 0.07 for PAR-CLIP) + UMI dedup
  -> [clip-qc]               preseq, FRiP, IDR rescue + self-consistency, read distribution
  -> [clip-peak-calling]     CLIPper + SMInput log2 norm (stringent: log2 FC >= 3, -log10 p >= 3) OR Skipper (substantially more sites)
  -> [crosslink-site-detection] PureCLIP or CTK CITS for single-nt CL positions
  -> [binding-site-annotation] ChIPseeker (tssRegion=c(-100,100), level=transcript) + RBP-Maps for splicing factors
  -> [clip-motif-analysis]   HOMER + mCross (registered) + RBNS Kd cross-check
  -> [differential-clip]     DEWSeq window-level NB with type:condition interaction (optional)

CLIP Variant Selection

VariantWhen to useUMI patternSTAR mismatch ceilingDetection signal
eCLIP (Van Nostrand 2016)ENCODE comparability; SMInput available10 nt R10.04R2 5' truncation
iCLIP / iCLIP2 / iCLIP3Single-end; high motif specificityNNNXXXXNN (3+4+2; demux first)0.04R1 5' truncation
irCLIP / FLASHNon-radioactive; fastProtocol-specific0.04Truncation
PAR-CLIPPhotoactivatable nucleoside (4SU); HEK293/K5624 nt typical0.07 (raised for T->C)T->C transitions
miCLIP / miCLIP2m6A modificationiCLIP-style0.04Truncation + C->T at m6A
STAMP / scSTAMPAntibody-free; in vivo or single-cellNA (no UV)0.04 (RNA-seq mode)C->U editing (RBP-APOBEC1 fusion)
chimeric eCLIP / miR-eCLIPDirect miRNA-target pairs10 nt R10.04Chimeric reads

Step 1: Quality Control of Raw FASTQ

bash
# Initial QC
fastqc raw_R1.fq.gz raw_R2.fq.gz -o qc/raw/

# Inspect first 12 bases of 100 reads to verify UMI pattern matches the prep
zcat raw_R1.fq.gz | awk 'NR%4==2' | head -100 | cut -c1-12 | sort | uniq -c | sort -rn | head
# Random barcode positions show ~25% per base; library barcodes are fixed

Step 2: Preprocessing (Protocol-Specific)

Goal: Convert raw CLIP FASTQ into UMI-deduplicated, alignment-ready FASTQ while preserving the R2 5' end (= crosslink site -1) that drives single-nucleotide resolution downstream.

Approach: Use the protocol-matched UMI pattern (10 nt eCLIP, NNNXXXXNN iCLIP, 4 nt PAR-CLIP), run umi_tools extract to move random barcodes to read names, then apply cutadapt with 3'-only adapter trimming at -q 6 -m 18 (permissive 5' to protect the truncation base). eCLIP uses two-pass trimming to remove read-through inline adapters from R2 5' only; iCLIP and PAR-CLIP use single-pass.

bash
# eCLIP: 10 nt UMI on R1; two-pass adapter trim for read-through
# See clip-seq/clip-preprocessing for protocol-specific patterns
umi_tools extract \
    --bc-pattern=NNNNNNNNNN \
    --stdin=raw_R1.fq.gz --read2-in=raw_R2.fq.gz \
    --stdout=R1.umi.fq.gz --read2-out=R2.umi.fq.gz \
    --log=qc/umi_extract.log

# Pass 1: 3' adapter on both reads
# -q 6 is intentionally permissive; aggressive trimming destroys R2 5' = CL site -1
cutadapt \
    -a AGATCGGAAGAGCACACGTCT \
    -A AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT \
    --quality-base 33 -q 6 -m 18 \
    -j 8 \
    -o R1.p1.fq.gz -p R2.p1.fq.gz \
    R1.umi.fq.gz R2.umi.fq.gz \
    > qc/cutadapt_pass1.log 2>&1

# Pass 2: strip read-through 5' adapter from R2 only (NEVER -g on R1)
cutadapt \
    -G GATCGTCGGACTGTAGAACTCTGAAC \
    --quality-base 33 -q 6 -m 18 \
    -j 8 \
    -o R1.trim.fq.gz -p R2.trim.fq.gz \
    R1.p1.fq.gz R2.p1.fq.gz \
    >> qc/cutadapt_pass2.log 2>&1

For PAR-CLIP: same UMI extraction but downstream alignment raises --outFilterMismatchNoverReadLmax from 0.04 to 0.07 (the T->C signature would otherwise be filtered as sequencing error). See clip-seq/clip-preprocessing for full per-protocol guidance.

Step 3: Alignment (ENCODE STAR Block)

bash
# ENCODE eCLIP convention. Sacred: --alignEndsType EndToEnd (soft-clip would destroy truncation = CL site -1)
STAR --runMode alignReads \
    --runThreadN 16 \
    --genomeDir /path/to/STAR_hg38_index \
    --genomeLoad NoSharedMemory \
    --readFilesIn R1.trim.fq.gz R2.trim.fq.gz \
    --readFilesCommand zcat \
    --outFilterType BySJout \
    --outFilterMultimapNmax 1 \
    --alignEndsType EndToEnd \
    --outFilterMismatchNoverReadLmax 0.04 \
    --outFilterScoreMinOverLread 0.66 \
    --outFilterMatchNminOverLread 0.66 \
    --outSAMtype BAM SortedByCoordinate \
    --outSAMattributes All \
    --outFileNamePrefix sample_

samtools index sample_Aligned.sortedByCoord.out.bam

# MAPQ >= 10 (255 = unique in STAR; lower = multi-mapper)
samtools view -b -q 10 sample_Aligned.sortedByCoord.out.bam > sample_q10.bam
samtools index sample_q10.bam

# UMI dedup. ENCODE convention: --method=unique
umi_tools dedup \
    --stdin=sample_q10.bam \
    --stdout=sample_dedup.bam \
    --method=unique \
    --paired \
    --log=qc/dedup.log
samtools index sample_dedup.bam

For PAR-CLIP: change --outFilterMismatchNoverReadLmax 0.04 to 0.07. For repeat-binding RBPs (MATR3, ZFP36, FUS at LINE-1, HNRNPK at SINEs): change --outFilterMultimapNmax 1 to 100 and add --outSAMmultNmax -1, then run CLAM downstream for EM-based multi-mapper assignment. See clip-seq/clip-alignment for full guidance.

Step 4: QC (Five Gates)

bash
# Gate 1: preprocessing retention (cutadapt log, target >= 70%)
grep -E "passing filters|Pairs written" qc/cutadapt_pass1.log

# Gate 2: alignment rate (STAR Log.final.out, target >= 60% eCLIP, 70% iCLIP)
grep "Uniquely mapped reads %" sample_Log.final.out

# Gate 3: library complexity (preseq, target >= 1M unique at sequenced depth)
preseq lc_extrap -B -P sample_q10.bam -o qc/preseq.txt

# Gate 4: FRiP (after peak calling; target >= 0.005 narrow-binding RBP)
# Gate 5: IDR replicate reproducibility (after peak calling; target rescue and self-consistency < 2)

# Aggregate all QC into a single MultiQC report
multiqc qc/ -o qc/multiqc/

CLIP libraries have 40-70% PCR duplication BY DESIGN (the IP enriches a small molecule pool). Low duplication usually means failed IP, not a good library. The unique-fragment count after UMI dedup is the actual quality metric. See clip-seq/clip-qc for full five-gate diagnostic.

Step 5: Peak Calling

bash
# CLIPper (ENCODE canonical) + SMInput log2 normalization
clipper \
    -b sample_dedup.bam \
    -s GRCh38 \
    -o peaks/sample.clipper.bed \
    --FDR 0.05 \
    --save-pickle \
    --processors 8   # super-local p-values are hard-coded ON in current CLIPper; the --superlocal flag was removed

# ENCODE stringent: log2(IP/SMInput) >= 3 AND -log10 p >= 3
# (Yeo lab eclip-pipeline scripts implement the normalization; see clip-seq/clip-peak-calling)
python overlap_peakfi_with_bam_PE.py \
    peaks/sample.clipper.bed \
    sample_dedup.bam sminput_dedup.bam \
    sample_dedup.bam.readnum.txt sminput_dedup.bam.readnum.txt \
    peaks/sample.normed.bed

python compress_l2foldenrpeakfi_for_replicate_overlapping_bedformat.py \
    peaks/sample.normed.bed \
    peaks/sample.compressed.bed

# Stringent filter
awk 'BEGIN{FS=OFS="\t"} $5 >= 3 && $6 >= 3' peaks/sample.compressed.bed > peaks/sample.stringent.bed

For maximum sensitivity (substantially more sites than CLIPper for mRNA-binding RBPs), use the Skipper Snakemake workflow with the same SMInput control. Mandatory for FASTKD2 / mt-RBPs which CLIPper misses on chrM. See clip-seq/clip-peak-calling for the full caller taxonomy.

bash
# PureCLIP: HMM jointly modeling enrichment + truncation + CL motif.
# -iv learns HMM parameters on a CHROMOSOME SUBSET (semicolon-delimited) to cut memory/runtime
# (per PureCLIP docs); it is NOT a BED. To limit the callset to expressed regions, pre-filter the input BAM.
pureclip \
    -i sample_dedup.bam -bai sample_dedup.bam.bai \
    -g genome.fa \
    -ibam sminput_dedup.bam -ibai sminput_dedup.bam.bai \
    -o crosslinks/sample.sites.bed \
    -or crosslinks/sample.regions.bed \
    -nt 8 -dm 8 \
    -iv 'chr1;chr2;chr3;'

Single-nt CL sites feed mCross motif registration and allele-specific binding analyses. They are NOT a replacement for the broad peak list; complementary outputs. See clip-seq/crosslink-site-detection.

Step 7: IDR Across Replicates

bash
# Sort each replicate's compressed BED by signal (log2 FC, column 5)
sort -k5,5gr peaks/rep1.compressed.bed > peaks/rep1.sorted.bed
sort -k5,5gr peaks/rep2.compressed.bed > peaks/rep2.sorted.bed

# True replicates threshold 0.05
idr --samples peaks/rep1.sorted.bed peaks/rep2.sorted.bed \
    --input-file-type bed --rank 5 \
    --output-file qc/idr.true.out \
    --idr-threshold 0.05 \
    --plot --log-output-file qc/idr.log

# ENCODE rule: rescue + self-consistency ratios both < 2 to pass
# Pseudo-replicate IDR (split BAM in half) at threshold 0.10

Step 8: Binding-Site Annotation

r
# CLIP-appropriate ChIPseeker (tssRegion tight; level=transcript)
library(ChIPseeker)
library(TxDb.Hsapiens.UCSC.hg38.knownGene)
txdb <- TxDb.Hsapiens.UCSC.hg38.knownGene

peaks <- readPeakFile('peaks/sample.stringent.bed')
anno <- annotatePeak(
    peaks,
    TxDb = txdb,
    level = 'transcript',
    tssRegion = c(-100, 100),
    genomicAnnotationPriority = c('Promoter','5UTR','3UTR','Exon','Intron','Downstream','Intergenic')
)
plotAnnoPie(anno)

Default ChIPseeker tssRegion=c(-3000, 3000) over-extends for CLIP (would label 30-50% peaks as "Promoter"). Splicing factors additionally need RBP-Maps (Yeo lab) for the 1400 nt cassette-exon regulatory metagene. See clip-seq/binding-site-annotation.

Show full SKILL.md (549 more words)Show less

Step 9: Motif Analysis (De Novo + CL-Registered)

bash
# Extract peak sequences (strand-preserving)
bedtools getfasta -fi genome.fa -bed peaks/sample.stringent.bed -s -fo motifs/peaks.fa

# GC-matched 3' UTR background (NOT auto-shuffled, which biases to AU)
bedtools shuffle -i peaks/sample.stringent.bed -g chrom.sizes \
    -incl expressed_3utr.bed -seed 42 > motifs/background.bed
bedtools getfasta -fi genome.fa -bed motifs/background.bed -s -fo motifs/background.fa

# HOMER de novo
findMotifs.pl motifs/peaks.fa fasta motifs/homer \
    -rna -len 5,6,7,8 -p 8 -fasta motifs/background.fa

# mCross for CL-position-registered motif. mCross.pl takes a POSITIONAL FASTA of sequences
# pre-extracted/registered around the CL sites and an output stem (not a BED + genome + -i/-g/-k/-o):
#   bedtools slop -i crosslinks/sample.sites.bed -g genome.sizes -b 10 | bedtools getfasta -fi genome.fa -bed - -s > motifs/peakseqs.fa
mCross.pl motifs/peakseqs.fa motifs/mcross   # see clip-seq/clip-motif-analysis for options

UV254 crosslinking has a strong U bias at CL sites; naive logos centered on CL positions are U-enriched even for non-U-binding RBPs. mCross corrects this by registering motif relative to the CL offset. See clip-seq/clip-motif-analysis.

Step 10: Differential Binding (Optional, Across Conditions)

r
# DEWSeq window-level NB with the interaction-term design
# The interaction `~ type + condition + type:condition` tests whether IP/SMInput ratio shifts;
# naive `~ condition` confounds binding with expression changes.
library(DEWSeq)
counts <- read.table('counts/merged.tsv', sep='\t', header=TRUE, row.names=1)
colData <- data.frame(
    type = relevel(factor(c('ip','ip','ip','ip','sminput','sminput','sminput','sminput')), ref='sminput'),
    condition = relevel(factor(c('treat','treat','ctrl','ctrl','treat','treat','ctrl','ctrl')), ref='ctrl')
)
dds <- DESeqDataSetFromSlidingWindows(
    countData=counts, colData=colData,
    annotObj='annotation.txt',   # htseq-clip TAB annotation table (named columns), NOT a plain BED
    design = ~ type + condition + type:condition
)
dds <- DESeq(dds)
# with sminput/ctrl as the references, the interaction coefficient is typeip.conditiontreat
res <- results(dds, name='typeip.conditiontreat')

See clip-seq/differential-clip for full DEWSeq workflow and the htseq-clip preprocessing required upstream.

Quality Checkpoints

StepMetricENCODE target
PreprocessingRetention after adapter trim>= 70%
AlignmentUnique mapping rate>= 60% (eCLIP); >= 70% (iCLIP)
Complexitypreseq predicted unique at 100M reads>= 10M (good); >= 1M (minimum acceptable)
Peak callingFRiP (narrow-binding RBP)>= 0.005
Peak callingStringent peaks log2(IP/SMI)>= 3
Peak callingStringent peaks -log10 p>= 3
IDRRescue ratio< 2
IDRSelf-consistency ratio< 2
AnnotationTop RBP-class match expectationY (HuR -> 3' UTR; PTBP1 -> intron; FASTKD2 -> chrM)

Per-Variant Adjustments

  • PAR-CLIP: Raise STAR --outFilterMismatchNoverReadLmax from 0.04 to 0.07; downstream use PARalyzer or CTK CIMS substitution T->C
  • iCLIP / iCLIP2 multiplexed: Demultiplex by inline library barcode (NNNXXXXNN) BEFORE umi_tools extract
  • HITS-CLIP: Use deletion-tolerant aligner (BWA-aln); downstream CTK CIMS deletion mode
  • Repeat-binding RBPs: STAR --outFilterMultimapNmax 100 --outSAMmultNmax -1 + CLAM EM rescue
  • m6A profiling: Switch to clip-seq/m6a-clip (miCLIP2 + m6Aboost or GLORI)
  • Antibody unavailable: Switch to clip-seq/stamp-antibody-free (STAMP or TRIBE)
  • miRNA targets: Switch to clip-seq/ago-clip-mirna-targets (chimeric eCLIP / miR-eCLIP)
  • Variant-effect prediction: Use clip-seq/clip-deep-learning (RBPNet or RNAProt)

Common Errors

SymptomCauseFix
Peaks everywhere, dominated by abundant transcriptsNo SMInput normalizationNormalize IP against SMInput; keep log2 FC >= 3 AND -log10 p >= 3
Single-nucleotide resolution lostSoft-clipping or aggressive 5' trim destroyed the R2 truncation baseSTAR --alignEndsType EndToEnd; 3'-only -q 6 trim; never -g on R1
PAR-CLIP T->C signal missingMismatch ceiling 0.04 filtered the transitions as errorRaise --outFilterMismatchNoverReadLmax to 0.07
"Low-complexity" library discardedJudged on raw duplication (40-70% is normal for CLIP)Use the unique-fragment count after UMI dedup as the quality metric
Motif logo is all-U even for a non-U-binding RBPNaive CL-centered logo + UV U-biasmCross CL-registered motif + GC-matched (not shuffled) background
30-50% of peaks labeled "Promoter"Default tssRegion=c(-3000,3000) over-extends for CLIPTight tssRegion=c(-100,100), level='transcript'
Differential binding confounded with expression~ condition design~ type + condition + type:condition interaction (DEWSeq)

References

  • Van Nostrand EL, Pratt GA, Shishkin AA, et al (2016) Robust transcriptome-wide discovery of RNA-binding protein binding sites with enhanced CLIP (eCLIP). Nature Methods 13:508-514. DOI 10.1038/nmeth.3810.
  • Van Nostrand EL, Freese P, Pratt GA, et al (2020) A large-scale binding and functional map of human RNA-binding proteins. Nature 583:711-719. DOI 10.1038/s41586-020-2077-3. (ENCODE RBP; SMInput + IDR practice.)
  • Krakau S, Richard H, Marsico A (2017) PureCLIP: capturing target-specific protein-RNA interaction footprints from single-nucleotide CLIP-seq data. Genome Biology 18:240. DOI 10.1186/s13059-017-1364-2.
  • Li Q, Brown JB, Huang H, Bickel PJ (2011) Measuring reproducibility of high-throughput experiments. Annals of Applied Statistics 5:1752-1779. DOI 10.1214/11-AOAS466. (IDR.)
  • clip-seq/clip-preprocessing - UMI extraction and adapter trimming details
  • clip-seq/clip-alignment - STAR ENCODE block + multi-mapper rescue
  • clip-seq/clip-qc - Five-gate QC framework
  • clip-seq/clip-peak-calling - CLIPper / Skipper / PureCLIP / CTK taxonomy
  • clip-seq/crosslink-site-detection - Single-nt CL detection by chemistry
  • clip-seq/binding-site-annotation - ChIPseeker + RBP-Maps
  • clip-seq/clip-motif-analysis - HOMER + mCross + RBNS validation
  • clip-seq/differential-clip - DEWSeq + Flipper for cross-condition
  • clip-seq/m6a-clip - miCLIP2 / GLORI / DART for m6A modifications
  • clip-seq/stamp-antibody-free - STAMP / TRIBE for antibody-free profiling
  • clip-seq/ago-clip-mirna-targets - chimeric eCLIP for direct miRNA-target pairs
  • clip-seq/clip-deep-learning - RBPNet / RNAProt for variant-effect prediction
  • read-qc/quality-reports - FastQC / MultiQC upstream QC
  • reporting/automated-qc-reports - MultiQC aggregates the per-tool QC into one report; gating stays in the five-gate framework, not MultiQC
  • alternative-splicing/differential-splicing - Cassette exon tables for RBP-Maps

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in workflows/clip-pipeline of GPTomics/bioSkills.

  • SKILL.md
  • examples/clip_full_pipeline.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Workflows Clip Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Workflows Clip Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Workflows Clip Pipeline this skillGPTomics/bioSkills1.2k2 repos~5.1kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Workflows Clip Pipeline

What does Bio Workflows Clip Pipeline do?

End-to-end CLIP-seq pipeline from FASTQ to ENCODE-compliant binding sites, single-nucleotide crosslink maps, annotation, motifs, and (optionally) differential binding. Bio Workflows Clip Pipeline is an agent skill from GPTomics/bioSkills. End-to-end CLIP-seq pipeline from FASTQ to ENCODE-compliant binding sites, single-nucleotide crosslink maps, annotation, motifs, and (optionally) differential binding.

When should I use Bio Workflows Clip Pipeline?

Bio Workflows Clip Pipeline fits situations like: running the full Yeo lab eCLIP / iCLIP / iCLIP2 / iCLIP3 / irCLIP / PAR-CLIP analysis with SMInput control; protocol-specific UMI extraction; ENCODE STAR parameters; skipper peak calling with stringent log2 FC and -log10 p thresholds.

How do I install Bio Workflows Clip Pipeline in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-workflows-clip-pipeline -a claude-code`. Or copy the skill folder (workflows/clip-pipeline in GPTomics/bioSkills) into .claude/skills/bio-workflows-clip-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Bio Workflows Clip Pipeline in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-workflows-clip-pipeline -a codex`. Or copy the skill folder (workflows/clip-pipeline in GPTomics/bioSkills) into .agents/skills/bio-workflows-clip-pipeline in your project. Codex loads it when a task matches its description.

Can I use Bio Workflows Clip Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-workflows-clip-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-workflows-clip-pipeline, .gemini/skills/bio-workflows-clip-pipeline, .github/skills/bio-workflows-clip-pipeline and .opencode/skills/bio-workflows-clip-pipeline in your project.

What does Bio Workflows Clip Pipeline need to run?

Going by SKILL.md and its folder, Bio Workflows Clip Pipeline needs a shell for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3; A Bash shell.

Does Bio Workflows Clip Pipeline access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Workflows Clip Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Workflows Clip Pipeline use?

Bio Workflows Clip Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Workflows Clip Pipeline use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Workflows Clip Pipeline?

Skills that share tags, products or a category with Bio Workflows Clip Pipeline: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Workflows Clip Pipeline?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.