Agent skill

Bio Clip Seq Clip Qc

by GPTomics in GPTomics/bioSkills

Comprehensive quality control for CLIP-seq libraries (eCLIP, iCLIP, iCLIP2, PAR-CLIP) covering library complexity (preseq), FRiP, IDR replicate reproducibility, read-distribution metagene, SMInput…

MITAuto-check passedResearch & Science

Install Bio Clip Seq Clip Qc

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clip-seq-clip-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clip-seq-clip-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clip-seq/clip-qc .claude/skills/bio-clip-seq-clip-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clip-seq-clip-qc
GitHub stars
1.2k
Used in
2 other repos
Token cost
~5.4k tokens
SKILL.md length
2,022 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Comprehensive quality control for CLIP-seq libraries (eCLIP, iCLIP, iCLIP2, PAR-CLIP) covering library complexity (preseq), FRiP, IDR replicate reproducibility, read-distribution metagene, SMInput…

  • Assessing whether a CLIP library passed
  • SKILL.md covers Version Compatibility, QC Stage Hierarchy, Library Complexity with preseq and FRiP (Fraction Reads in Peaks), plus 13 more sections
  • Runs Shell scripts from its folder; calls pip
  • Deciding lenient vs stringent peak thresholds

What it does

Bio Clip Seq Clip Qc is an agent skill from GPTomics/bioSkills. Comprehensive quality control for CLIP-seq libraries (eCLIP, iCLIP, iCLIP2, PAR-CLIP) covering library complexity (preseq), FRiP, IDR replicate reproducibility, read-distribution metagene, SMInput vs IgG control rationale, rRNA / snoRNA contamination, fragment-length distribution, and ENCODE-compliance thresholds. Use when assessing whether a CLIP library passed, deciding lenient vs stringent peak thresholds, comparing replicates with IDR rescue and self-consistency ratios, or distinguishing failed IP from…

Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/clip_qc.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Reproducible research. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Assessing whether a CLIP library passed
  • Deciding lenient vs stringent peak thresholds
  • Comparing replicates with IDR rescue and self-consistency ratios
  • Distinguishing failed IP from over-amplified library

Example prompts

  • “/bio-clip-seq-clip-qc”

Requirements

  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clip Seq Clip Qc loads about 5.4k tokens when it runs. Until then it costs about 139 tokens; SKILL.md has 2,022 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~139
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,022 words, ~5,362 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clip-seq-clip-qc/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clip-seq-clip-qc
description
Comprehensive quality control for CLIP-seq libraries (eCLIP, iCLIP, iCLIP2, PAR-CLIP) covering library complexity (preseq), FRiP, IDR replicate reproducibility, read-distribution metagene, SMInput vs IgG control rationale, rRNA / snoRNA contamination, fragment-length distribution, and ENCODE-compliance thresholds. Use when assessing whether a CLIP library passed, deciding lenient vs stringent peak thresholds, comparing replicates with IDR rescue and self-consistency ratios, or distinguishing failed IP from over-amplified library.
tool_type
mixed
primary_tool
preseq

Version Compatibility

Reference examples tested with: preseq 3.2+, picard 3.1+, samtools 1.19+, bedtools 2.31+, deeptools 3.5+, idr 2.0.4+, MultiQC 1.21+, RSeQC 5.0+, pysam 0.22+, fastp 0.23+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws unexpected errors, introspect the installed binary and adapt the example to match the actual CLI rather than retrying.

CLIP-seq Quality Control

"Did my CLIP library pass?" -> Assess preprocessing retention, alignment rate, library complexity, replicate reproducibility (IDR), fraction reads in peaks (FRiP), read-distribution metagene, rRNA/snoRNA contamination, fragment-length distribution, and SMInput vs IP enrichment. ENCODE eCLIP compliance is the canonical bar: >= 1M unique fragments per replicate, IDR rescue and self-consistency ratios both < 2, FRiP >= 0.005 (narrow-binding), library complexity rising linearly with depth on preseq lc_extrap. A library can fail at any of these stages, and the failure mode determines whether the data is salvageable.

  • CLI (library complexity, primary QC): preseq lc_extrap -B -P aligned.bam -o complexity.txt
  • CLI (FRiP after peak calling): bedtools intersect -c -s -a peaks.bed -b dedup.bam | awk '{s+=$NF} END{print s}' then divide by total reads
  • CLI (IDR, ENCODE convention): idr --samples rep1.sorted.bed rep2.sorted.bed --input-file-type bed --rank 5 --output-file idr.out --idr-threshold 0.05 --plot
  • CLI (read distribution / metagene): RSeQC read_distribution.py -i dedup.bam -r gencode.v38.bed + geneBody_coverage.py -i dedup.bam -r housekeeping.bed -o gb
  • CLI (rRNA contamination check): samtools idxstats dedup.bam | awk '$1 ~ /rRNA|45S|18S|28S/ { sum+=$3 } END {print sum}'
  • CLI (consolidated report): multiqc <run_dir> aggregates FastQC + cutadapt + STAR + umi_tools + preseq + samtools stats

The ENCODE eCLIP standards (encodeproject.org/eclip) define: >= 2 biological replicates with >= 1M unique fragments each (or saturated peak detection); IDR rescue and self-consistency ratios both < 2; narrow-binding RBPs FRiP >= 0.005. CLIP libraries should have 40-70% PCR duplication BY DESIGN - the IP enriches a small molecule pool, so high pre-dedup duplication is normal; low duplication suggests failed IP.

QC Stage Hierarchy

CLIP QC progresses through five gates; failure at an earlier gate makes later gates meaningless.

GateMetricToolENCODE thresholdFailure interpretation
1. Preprocessing retention% reads retained after UMI + adapter trimcutadapt log>= 70%Adapter pattern wrong; degraded RNA
2. Alignment rate% reads aligned to genome (unique)STAR Log.final.out>= 60% for eCLIP, 70% for iCLIP/PAR-CLIPWrong genome; rRNA pre-map missing
3. Library complexityPredicted unique fragments at sequenced depthpreseq lc_extrap>= 1M uniqueOver-amplified or under-input library
4. IP enrichmentlog2(IP/SMInput) at expected sites; FRiPbedtools + idrFRiP >= 0.005; log2 >= 3 at top peaksFailed antibody / antibody not IP-grade
5. ReproducibilityIDR rescue + self-consistency ratiosidrboth < 2Biological variation too high; or low complexity

A library failing at Gate 3 (complexity) cannot be rescued analytically; gates 4-5 fail downstream of complexity by construction.

Library Complexity with preseq

Goal: Determine whether the CLIP library captured enough independent molecules to support genome-wide peak calling (ENCODE: >= 1M unique fragments per replicate).

Approach: Run preseq lc_extrap on the PRE-dedup BAM (preseq counts PCR duplicates to extrapolate); also compute picard ESTIMATED_LIBRARY_SIZE at sequenced depth. Flag any library predicted to plateau below 3M unique fragments at infinite depth.

bash
# After alignment, BEFORE UMI dedup (preseq counts PCR duplicates)
preseq lc_extrap \
    -B -P \
    -o sample_complexity.txt \
    sample_aligned.bam

# Output columns:
# TOTAL_READS  EXPECTED_DISTINCT  LOWER_0.95CI  UPPER_0.95CI
# At 100M reads, EXPECTED_DISTINCT:
#   >= 10M  = excellent complexity
#   3-10M   = acceptable; restrict to high-expression transcripts
#   < 3M    = library failed; cannot rescue analytically

# picard direct estimate at current depth
picard EstimateLibraryComplexity \
    I=sample_aligned.bam \
    O=picard_complexity.txt
# ESTIMATED_LIBRARY_SIZE > 5M = healthy CLIP library

A linear plateau on preseq's curve at low depth indicates over-amplification; a curve still climbing at sequenced depth means more sequencing would yield more unique fragments.

FRiP (Fraction Reads in Peaks)

FRiP measures how much of the IP signal falls into the called peak set. ENCODE eCLIP narrow-binding RBP minimum: FRiP >= 0.005. Atypical-binding RBPs (rare-transcript binders like TROVE2 on Y RNAs) are exempt.

bash
# Reads in peaks (use stringent peaks: log2 FC >= 3, -log10 p >= 3)
reads_in_peaks=$(bedtools intersect -c -s -a peaks.stringent.bed -b dedup.bam | awk '{s+=$NF} END {print s}')
total_reads=$(samtools view -c -F 4 dedup.bam)
frip=$(echo "scale=4; $reads_in_peaks / $total_reads" | bc)
echo "FRiP: $frip"

# Per-region FRiP breakdown
for region in three_utr exon intron; do
    rip=$(bedtools intersect -c -s -a peaks_${region}.bed -b dedup.bam | awk '{s+=$NF} END {print s}')
    echo "${region}: $(echo "scale=4; $rip / $total_reads" | bc)"
done
RBP classExpected FRiP (ENCODE eCLIP)
Splicing factors (PTBP1, U2AF2)0.01 - 0.10
3' UTR mRNA stability (HuR, PUM2)0.02 - 0.20
Translation factors (EIF3J)0.01 - 0.05
Repeat binders (MATR3)0.05 - 0.30 (high; concentrated in repeats)
Mitochondrial (FASTKD2)0.05 - 0.40 (very high; chrM is small)
snoRNA binders (DKC1)0.10 - 0.50 (high; snoRNA is rare)
Failed IP (any RBP)< 0.005

IDR for CLIP Reproducibility

IDR (Li et al 2011) measures peak-rank reproducibility across replicates. ENCODE eCLIP convention applies IDR identically to ChIP-seq, using CLIPper + SMInput log2 FC + -log10 p as the ranking signal. The CLIP-specific consideration: rank by signalValue (log2 FC) or p-value, NOT by score column (CLIPper score is sparse and tied).

bash
# Sort each replicate's peaks by signal
sort -k5,5gr rep1.compressed.bed > rep1.sorted
sort -k5,5gr rep2.compressed.bed > rep2.sorted

# True-replicates IDR (threshold 0.05)
idr --samples rep1.sorted rep2.sorted \
    --input-file-type bed --rank 5 \
    --output-file idr_true.out \
    --idr-threshold 0.05 \
    --plot --log-output-file idr.log

# Pseudo-replicates from each individual replicate (split BAM in half)
samtools view -b -h -s 1.5 rep1.dedup.bam > rep1.psr1.bam   # seed 1, fraction 0.5
samtools view -b -h -s 2.5 rep1.dedup.bam > rep1.psr2.bam   # seed 2 (different)
# Re-run peak calling on each pseudoreplicate, then IDR at threshold 0.10

ENCODE consistency rules for eCLIP:

  • Nt = peaks passing IDR on true replicates
  • Nself = peaks passing IDR on pseudo-replicates of each rep
  • Library passes if: max(Nt, Nself) / min(Nt, Nself) <= 2
  • If both ratios > 2: library rejected

SMInput vs IgG Control: Which?

The eCLIP design uses a size-matched input (SMInput) from the SAME lysate, treated identically (UV, IP buffer, RNase, ligation, IP, but with NO antibody addition - just bead-only control or a non-specific control IP). This is fundamentally different from IgG controls and from RNA-seq.

ControlWhat it measuresProsCons
SMInputBackground from non-specific binding + ligation/RT/gel biases at same sizeCaptures all CLIP-specific biases; ENCODE standardRequires same-day prep; cannot use a previous IgG library
IgG-IPNon-specific antibody binding to the same RBP-naive lysateDirect nonspecificity measureYields very low (~3-10x less reads); high PCR dup; hard to normalize
Empty beads (mock)Bead surface non-specificity onlyCleanest baselineMisses real CLIP background (IP buffer + ligation step bias)
RNA-seq (matched cell type)Transcript abundanceEasy to obtainMisses CLIP-specific biases entirely; not a real CLIP control
Total RNA / nuclear RNACellular RNA distributionEasySame as RNA-seq above

Consensus (ENCODE / Hentze / Yeo): SMInput. The bead-only IgG/mock alternatives are ill-suited for quantification because their library yields are 5-10x lower, dominated by PCR duplicates, and produce sparse read-density tracks. RNA-seq cannot replace SMInput because it does not capture the non-specific binding that occurs during IP.

bash
# SMInput preparation - same as IP except no antibody added during incubation
# Same UV dose, same lysate aliquot, same RNase, same library prep
# Critical: same SDS-PAGE size cut from membrane as the IP

# Check SMInput vs IP enrichment. A ratio of whole-library totals only measures relative sequencing
# depth; normalize the in-peak fraction in each library instead:
#   log2( (IP_in_peaks/IP_total) / (SMI_in_peaks/SMI_total) )
# > 0.5 indicates the IP concentrates reads into peaks beyond SMInput (good)
# ~ 0 indicates failed IP (SMInput == IP)

Read Distribution Metagene

bash
# RSeQC read_distribution.py shows fractional read placement
read_distribution.py -i dedup.bam -r gencode.v38.bed > read_dist.txt

# Output reports:
#   CDS_Exons, 5'UTR_Exons, 3'UTR_Exons, Introns, TES_down_10kb, TES_down_1kb,
#   TSS_up_10kb, TSS_up_1kb, Intergenic_region
# CLIP-seq biology-specific patterns:
#   Splicing factor: > 60% Introns + 5' UTR exons (containing 5' splice sites)
#   3' UTR factor: > 50% 3'UTR_Exons
#   m6A reader: 3'UTR_Exons + Stop_codon region
#   Failed IP: matches RNA-seq distribution (no enrichment)

# GeneBody coverage for 5' vs 3' bias
geneBody_coverage.py \
    -i dedup.bam \
    -r housekeeping.bed \
    -o sample_gb
# Flat curve = no positional bias (normal for most RBPs)
# 3' end bias = polyA-dependent enrichment (suspect)
# 5' end bias = nascent / TSS-proximal (only sensible for some RBPs)

Fragment-Length Distribution (Paired-End)

bash
# Paired-end fragment-length distribution
samtools view -f 2 dedup.bam | awk '$9 > 0 && $9 < 500 {print $9}' | sort -n | uniq -c > fragment_lengths.txt

# CLIP normal range: 20-75 nt insert (rises from short trimmed reads to ~75 nt)
# Wider range (20-200) seen in high-RNase / long-fragment protocols
# Narrow peak at one length (e.g., all reads at 30 nt) = over-trimmed
# Bimodal at 30 nt and 150 nt = library preparation artifact

Pre-Map rRNA / snoRNA Contamination Check

bash
# rRNA reads as fraction of total
total=$(samtools view -c -F 4 dedup.bam)
rRNA=$(samtools idxstats dedup.bam | awk '$1 ~ /rRNA|45S|18S|28S|5_8S/ { sum+=$3 } END {print sum}')
rrna_frac=$(echo "scale=4; $rRNA / $total" | bc)
echo "rRNA fraction: $rrna_frac"

# eCLIP without pre-map: rRNA 5-30% normal
# After pre-map: < 2% expected
# > 30% indicates rRNA dominance - IP captured ribosomes preferentially
# Some RBPs (RPL/RPS) expected to bind ribosomes; verify against RBP biology

Antibody Validation Sanity Check

The antibody is the single most common point of failure in CLIP. Even commercial "IP-grade" antibodies fail at 30-50% rates in practice. Cross-check IP enrichment against expected biology:

python
import pandas as pd

# Load peak file
peaks = pd.read_csv('peaks.stringent.bed', sep='\t', header=None,
                    names=['chr','start','end','name','log2fc','strand'])

# Filter top 100 peaks
top_peaks = peaks.nlargest(100, 'log2fc')

# Check enrichment in expected biology
# 1. If RBP is splicing factor: top peaks should be intronic or splice-site flanking
# 2. If RBP is HuR: > 70% top peaks should be 3' UTR
# 3. If RBP is FASTKD2: top peaks should be chrM
# 4. If RBP is FUS: GUGGU motif should appear in top motif enrichment

# Quick check
chrom_dist = top_peaks['chr'].value_counts(normalize=True)
print(chrom_dist.head(10))
# Healthy RBP: chromosomes represented in proportion to expression
# Failed IP: top chromosome chrM (mt artifact) or chr21/22 (housekeeping bias) > 20%

GO term sanity: top-peak genes should enrich for the expected biology (e.g., HuR -> immune / inflammation / mRNA stability terms; PTBP1 -> RNA splicing terms). If the top GO term is "cellular metabolism" or "translation" for a splicing factor, the IP failed.

Per-Stage Failure Modes

Gate 1: Preprocessing retention < 70%

Trigger: cutadapt log shows > 30% reads filtered (too short or no adapter).

Mechanism: Adapter pattern wrong; or RNA was degraded before fragmentation (all reads adapter-only).

Symptom: Catastrophic loss at the adapter trim step.

Fix: Verify adapter sequence in library prep documentation; check library quality control (Bioanalyzer / TapeStation) for RIN value; re-prep if RIN < 7.

Gate 2: Alignment rate < 60%

Trigger: STAR Log.final.out reports Uniquely mapped reads % < 60% for eCLIP, < 70% for iCLIP.

Mechanism: rRNA contamination dominating (no pre-map); wrong species genome; or chimeric library (contamination).

Symptom: Low alignment rate; samtools idxstats shows most reads as rRNA-aligned.

Fix: Run bowtie2 pre-map to rRNA + repeats index; verify genome species; check for sample swap with samtools view -h sample.bam | grep '^@SQ'.

Gate 3: Library complexity < 1M unique

Trigger: preseq lc_extrap predicts plateau < 1M unique at sequenced depth; or picard reports ESTIMATED_LIBRARY_SIZE < 1M.

Mechanism: Over-amplification (> 25 PCR cycles); low input (< 5M cells); failed IP capturing only a few molecules.

Symptom: preseq curve flattens early; UMI clusters have median size > 8.

Fix: No analytic rescue. Re-prep with more input cells (10-20M for eCLIP), fewer PCR cycles (14-18 for eCLIP, 16-20 for iCLIP2). If forced to use this library, restrict downstream analysis to high-expression transcripts and acknowledge the limitation.

Show full SKILL.md (818 more words)Show less
Gate 4: FRiP < 0.005 for narrow-binding RBP

Trigger: FRiP value below ENCODE threshold for the RBP class.

Mechanism: IP failed to enrich specific binding sites; antibody cross-reactive or not IP-grade.

Symptom: Top peaks dominated by abundant transcripts (GAPDH, ACTB); GO enrichment generic; motif analysis returns AU-rich background.

Fix: Re-test antibody on a knockdown lysate (siRNA against the RBP) - if WB signal does not decrease, antibody is non-specific; switch antibody. ENCODE-validated antibodies are at encodeproject.org/biosamples.

Gate 4 variant: IP/SMInput global log2 ~ 0

Trigger: Whole-genome log2(IP/SMInput) is near 0 instead of positive.

Mechanism: IP and SMInput are equivalent - the antibody is not enriching any RBP-bound RNA.

Symptom: Per-peak log2 FC scaled around 0; no obvious enrichment at known motif sites.

Fix: Same as above - antibody failure. Consider tagged-RBP system (Halo-CLIP, GoldCLIP, or knock-in endogenous tag).

Gate 5: IDR fails - rescue/self-consistency > 2

Trigger: ENCODE IDR test reports rescue ratio > 2 OR self-consistency > 2.

Mechanism: Replicates inconsistent. Biological variation high; OR one replicate has lower library complexity; OR one replicate failed IP.

Symptom: Per-replicate peak counts differ > 2x; IDR plot shows the two replicates have very different peak rank distributions.

Fix: Run preseq and FRiP on each replicate independently; identify the failing one; re-sequence to higher depth or re-prep. ENCODE-style: down-sample both replicates to common depth, re-test IDR.

High PCR duplication confused with low complexity

Trigger: Saw 60% PCR duplication and panicked.

Mechanism: CLIP libraries have 40-70% PCR duplication BY DESIGN - the IP enriches a small pool. UMI dedup recovers the unique molecules underneath.

Symptom: "60% duplication" headline. Actual unique count is what matters.

Fix: Confirm UMI dedup has been applied and unique fragment count >= 1M. The duplication rate metric is meaningless without UMI context for CLIP.

Decision Tree by Failure

SymptomLikely failureAction
> 50% reads "too short" in cutadaptAdapter wrong or RNA degradedVerify adapter; check RIN
Most reads align to rRNANo pre-map; or IP captured ribosomesbowtie2 pre-map; or accept if RBP is ribosomal
Unique frag count < 1MOver-amplified or under-inputNo rescue; re-prep
FRiP < 0.005 (narrow RBP)Failed IPSwitch antibody
IP/SMInput global log2 ~ 0Antibody non-specificSwitch antibody or use tagged-RBP
IDR rescue > 2Replicate inconsistencyIdentify failing replicate; re-sequence
Top GO terms genericFailed IP or contaminationCheck antibody on KD lysate WB
chrM peaks abundant for non-mt-RBPMitochondrial contaminationFilter chrM or accept if mt-RBP
Fragment length all 30 ntOver-trimmedLoosen quality trim from -q 20 to -q 6
Strand-specific peaks lostbedtools without -sAdd -s strand flag

Standard QC Report (MultiQC)

bash
# Aggregate all QC tools into a single report
multiqc \
    fastqc_output/ \
    cutadapt_logs/ \
    star_logs/ \
    umi_tools_logs/ \
    preseq_output/ \
    samtools_stats/ \
    -o multiqc_report/

MultiQC reads logs from FastQC, cutadapt, STAR, umi_tools, preseq, samtools stats, picard, and others, and produces a single HTML report. For CLIP-seq, additionally run RSeQC read_distribution.py and document FRiP / IDR results manually.

ENCODE-Style Comprehensive QC Table

MetricThresholdSource
Biological replicates>= 2ENCODE eCLIP
Read length>= 50 nt (PE)ENCODE
Unique fragments per replicate>= 1M (or saturated)ENCODE
Preprocessing retention>= 70%Practitioner consensus
Genome alignment rate (unique)>= 60% eCLIP, 70% iCLIPPractitioner consensus
Library complexity (preseq)>= 10M expected unique at 100MENCODE
FRiP (narrow-binding RBP)>= 0.005ENCODE
FRiP (atypical-binding RBP)ExemptENCODE
log2(IP/SMInput) at stringent peaks>= 3ENCODE
-log10 p at stringent peaks>= 3ENCODE
IDR rescue ratio< 2ENCODE
IDR self-consistency ratio< 2ENCODE
rRNA fraction post pre-map< 2%Practitioner consensus
Read distribution match to RBP classYSanity check
Top motif consistent with literatureYSanity check

Common Errors

Error / symptomCauseSolution
preseq "all reads have duplicates"Pre-deduplicated BAM passed to preseqRe-run on PRE-dedup BAM
FRiP value > 1.0Peak set overlaps reads counted twiceUse unique-reads BAM; verify peak BED is non-redundant
IDR plot emptyRanking column wrong; all peaks tied at same valueRank by --rank 5 (log2 FC); sort BED first
MultiQC missing modulesLogs not in expected directory structureVerify log paths; rerun multiqc with explicit -d
Read distribution: > 50% intergenic for mRNA RBPTxDb missing transcripts; rRNA contaminationUpdate GENCODE; pre-map rRNA
GeneBody coverage 3' biasedpolyA-selected library; not CLIPVerify library prep was random hexamer
Fragment-length distribution narrow at 30 ntAdapter aggressively trimmed at 5'Loosen trim parameters
Antibody KD WB shows no decreaseAntibody non-specificSwitch antibody; use ENCODE-validated

References

  • Van Nostrand EL et al 2016 Nat Methods 13:508 (eCLIP, ENCODE QC standards)
  • Li Q et al 2011 Ann Appl Stat 5:1752 (IDR framework)
  • Landt SG et al 2012 Genome Res 22:1813 (ChIP/CLIP QC guidelines, IDR Nself rule)
  • Daley T & Smith AD 2013 Nat Methods 10:325 (preseq library complexity)
  • Wang L et al 2012 Bioinformatics 28:2184 (RSeQC)
  • Ewels P et al 2016 Bioinformatics 32:3047 (MultiQC)
  • Smith T et al 2017 Genome Res 27:491 (UMI-tools)
  • ENCODE eCLIP Data Standards (encodeproject.org/eclip) - canonical thresholds
  • Van Nostrand EL et al 2020 Nature 583:711 (ENCODE 150 RBP QC patterns)
  • clip-seq/clip-preprocessing - Gate 1 preprocessing logs
  • clip-seq/clip-alignment - Gate 2 alignment metrics
  • clip-seq/clip-peak-calling - Gates 4-5 peak QC + FRiP
  • clip-seq/differential-clip - QC for differential analyses
  • read-qc/quality-reports - General FastQC / MultiQC
  • read-qc/contamination-screening - Cross-species contamination
  • chip-seq/chipseq-qc - DNA-protein QC analogue

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clip-seq/clip-qc of GPTomics/bioSkills.

  • SKILL.md
  • examples/clip_qc.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clip Seq Clip Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clip Seq Clip Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clip Seq Clip Qc this skillGPTomics/bioSkills1.2k2 repos~5.4kAutomated safety check: PassMIT
LaminDB Biological Data Managementdavila7/claude-code-templates33k12 repos~3.6kAutomated safety check: PassMIT
AI Scientist EvaluatorBioTender-max/awesome-bio-agent-skills200—~2.4kAutomated safety check: PassCustom licence
Latchbio Integrationdavila7/claude-code-templates33k11 repos~2.4kAutomated safety check: PassMIT
Remote Compute Sshaipoch/open-science5.5k—~5.7kAutomated safety check: PassApache-2.0
Bio OrchestratorClawBio/ClawBio1.2k3 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • LaminDB Biological Data Management

    davila7/claude-code-templates

    Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.

    33k GitHub starsUsed in 12 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • AI Scientist Evaluator

    BioTender-max/awesome-bio-agent-skills

    Critically review, score, compare, and rank one or more AI scientist outputs for biology, bioinformatics, computational life science, or adjacent research tasks.

    200 GitHub stars~2.4k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Latchbio Integration

    davila7/claude-code-templates

    Latch platform for bioinformatics workflows. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 11 repos~2.4k tokens
    Research & ScienceAuto-check passed
  • Remote Compute Ssh

    aipoch/open-science

    Evaluate and use SSH Remote Compute before choosing where to run GPU, high-memory, parallel, batch, model-inference, bioinformatics, or other long-running scientific work; supports short remote…

    5.5k GitHub stars~5.7k tokensUpdated today
    Research & ScienceAuto-check passed
  • Bio Orchestrator

    ClawBio/ClawBio

    Meta-agent that routes bioinformatics requests to specialised sub-skills.

    1.2k GitHub starsUsed in 3 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Latchbio Integration

    K-Dense-AI/scientific-agent-skills

    Builds, registers, debugs, and operates bioinformatics workflows on Latch using the Python SDK, CLI, Latch Data and Registry, Nextflow, Snakemake, programmatic execution, and Latch MCP.

    48k GitHub starsUsed in 1 repo~2.5k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Clip Seq Clip Qc

What does Bio Clip Seq Clip Qc do?

Comprehensive quality control for CLIP-seq libraries (eCLIP, iCLIP, iCLIP2, PAR-CLIP) covering library complexity (preseq), FRiP, IDR replicate reproducibility, read-distribution metagene, SMInput…. Bio Clip Seq Clip Qc is an agent skill from GPTomics/bioSkills. Comprehensive quality control for CLIP-seq libraries (eCLIP, iCLIP, iCLIP2, PAR-CLIP) covering library complexity (preseq), FRiP, IDR replicate reproducibility, read-distribution metagene, SMInput vs IgG control rationale, rRNA / snoRNA contamination, fragment-length distribution, and ENCODE-compliance thresholds.

When should I use Bio Clip Seq Clip Qc?

Bio Clip Seq Clip Qc fits situations like: assessing whether a CLIP library passed; deciding lenient vs stringent peak thresholds; comparing replicates with IDR rescue and self-consistency ratios; distinguishing failed IP from over-amplified library.

How do I install Bio Clip Seq Clip Qc in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-clip-qc -a claude-code`. Or copy the skill folder (clip-seq/clip-qc in GPTomics/bioSkills) into .claude/skills/bio-clip-seq-clip-qc in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clip Seq Clip Qc in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-clip-qc -a codex`. Or copy the skill folder (clip-seq/clip-qc in GPTomics/bioSkills) into .agents/skills/bio-clip-seq-clip-qc in your project. Codex loads it when a task matches its description.

Can I use Bio Clip Seq Clip Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-clip-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clip-seq-clip-qc, .gemini/skills/bio-clip-seq-clip-qc, .github/skills/bio-clip-seq-clip-qc and .opencode/skills/bio-clip-seq-clip-qc in your project.

What does Bio Clip Seq Clip Qc need to run?

Going by SKILL.md and its folder, Bio Clip Seq Clip Qc needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.

Does Bio Clip Seq Clip Qc access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Clip Seq Clip Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clip Seq Clip Qc use?

Bio Clip Seq Clip Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clip Seq Clip Qc use?

About 5.4k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clip Seq Clip Qc?

Skills that share tags, products or a category with Bio Clip Seq Clip Qc: LaminDB Biological Data Management (davila7/claude-code-templates, 33k stars), AI Scientist Evaluator (BioTender-max/awesome-bio-agent-skills, 200 stars), Latchbio Integration (davila7/claude-code-templates, 33k stars) and Remote Compute Ssh (aipoch/open-science, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clip Seq Clip Qc?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.