Agent skill

Bio Clip Seq Binding Site Annotation

by GPTomics in GPTomics/bioSkills

Annotate CLIP-seq peaks or crosslink sites to RNA features (5'UTR, CDS, 3'UTR, intron, splice junction, snoRNA, tRNA, ncRNA, repeat elements) with ChIPseeker, RCAS, RBP-Maps (Yeo splicing regulatory…

MITAuto-check passedResearch & Science

Install Bio Clip Seq Binding Site Annotation

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clip-seq-binding-site-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clip-seq-binding-site-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clip-seq/binding-site-annotation .claude/skills/bio-clip-seq-binding-site-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clip-seq-binding-site-annotation
GitHub stars
1.2k
Used in
2 other repos
Token cost
~5.4k tokens
SKILL.md length
2,353 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Annotate CLIP-seq peaks or crosslink sites to RNA features (5'UTR, CDS, 3'UTR, intron, splice junction, snoRNA, tRNA, ncRNA, repeat elements) with ChIPseeker, RCAS, RBP-Maps (Yeo splicing regulatory…

  • Works in 11 steps: 3' UTR (specific to mature mRNA biology) → 5' UTR → CDS (coding sequence) → …
  • Characterizing where in transcripts an RBP binds
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Critical Choice: Annotation… and RBP Class -> Expected Annotation, plus 8 more sections
  • Runs R scripts from its folder; calls git, python and pip; reaches github.com

What it does

Bio Clip Seq Binding Site Annotation is an agent skill from GPTomics/bioSkills. Annotate CLIP-seq peaks or crosslink sites to RNA features (5'UTR, CDS, 3'UTR, intron, splice junction, snoRNA, tRNA, ncRNA, repeat elements) with ChIPseeker, RCAS, RBP-Maps (Yeo splicing regulatory maps), and bedtools, applying feature-priority hierarchies, transcript-context resolution, and metagene aggregation. Use when characterizing where in transcripts an RBP binds, comparing peak distribution across regions, generating splicing-regulatory maps relative to alternative-splicing events, or distinguishing…

Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Characterizing where in transcripts an RBP binds
  • Comparing peak distribution across regions
  • Generating splicing-regulatory maps relative to alternative-splicing events
  • Distinguishing exonic vs intronic vs UTR binding

Example prompts

  • “UTR, CDS, 3”
  • “/bio-clip-seq-binding-site-annotation”

Requirements

  • Python 3

Workflow steps

11 steps, taken from the first numbered list in SKILL.md.

  1. 3' UTR (specific to mature mRNA biology)
  2. 5' UTR
  3. CDS (coding sequence)
  4. Exon (catch-all for non-UTR exonic)
  5. Splice site (within 50 nt of donor/acceptor; key for splicing factors)
  6. Intron
  7. Promoter (within 1 kb of TSS; usually not relevant for RNA-binding)
  8. 5' flanking / 3' flanking (within 1 kb of gene)
  9. ncRNA (lncRNA, snoRNA, snRNA, miRNA host)
  10. Repeat element (Alu, LINE, LTR, SINE) - separate axis from above
  11. Intergenic

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clip Seq Binding Site Annotation loads about 5.4k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 2,353 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,353 words, ~5,353 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clip-seq-binding-site-annotation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clip-seq-binding-site-annotation
description
Annotate CLIP-seq peaks or crosslink sites to RNA features (5'UTR, CDS, 3'UTR, intron, splice junction, snoRNA, tRNA, ncRNA, repeat elements) with ChIPseeker, RCAS, RBP-Maps (Yeo splicing regulatory maps), and bedtools, applying feature-priority hierarchies, transcript-context resolution, and metagene aggregation. Use when characterizing where in transcripts an RBP binds, comparing peak distribution across regions, generating splicing-regulatory maps relative to alternative-splicing events, or distinguishing exonic vs intronic vs UTR binding.
tool_type
mixed
primary_tool
ChIPseeker

Version Compatibility

Reference examples tested with: ChIPseeker 1.40+, RCAS 1.30+, GenomicFeatures 1.56+, GenomicRanges 1.56+, rbp-maps (Yeo github), bedtools 2.31+, pybedtools 0.10+, pyranges 0.0.129+.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws unexpected errors, introspect the installed package and adapt the example to match the actual API rather than retrying. ChIPseeker 1.40+ changed the priority defaults; verify the priority vector before pipelining.

Binding Site Annotation

"Annotate where in transcripts my RBP binds" -> Map CLIP peaks or single-nucleotide crosslink sites to RNA features and report the per-feature distribution. The interpretation is RBP-class-specific: splicing factors (PTBP1, U2AF2, RBFOX) bind intron-exon junctions; mRNA-stability regulators (HuR, PUM2) bind 3' UTRs; translation factors (EIF3J, RPS19) bind 5' UTRs and CDS; and small-ncRNA-binding RBPs (NSUN2 tRNAs, LARP7 7SK, TROVE2 Y-RNAs) bind specific non-coding transcripts. A correct annotation pipeline (a) resolves overlapping features by priority, (b) preserves transcript-isoform context, (c) generates metagene distributions, and (d) flags repeat-element overlap separately.

  • R (peak-level, fast): ChIPseeker::annotatePeak(peaks, TxDb=txdb, level='gene', tssRegion=c(-100,100)) then plotAnnoPie(anno)
  • R (transcript-level with RNA-specific regions): RCAS runReport(queryRegions=peaks, gffData=gencode_gtf, genomeVersion='hg38')
  • CLI (regions only): bedtools intersect -s -wa -wb -a peaks.bed -b features.bed (manual hierarchy)
  • R (splicing-regulatory map for splice factors): RBP-Maps (yeolab/rbp-maps) generates the 1400 nt vectorized cassette-exon map used in Yeo lab ENCODE papers
  • CLI (metagene aggregation): deepTools computeMatrix or RSeQC geneBody_coverage.py for read profiles

The Yeo lab convention for ENCODE eCLIP: ChIPseeker for global region distribution + RBP-Maps for position-resolved splicing context + custom analysis for repeat-element overlap. ChIPseeker alone over-counts intronic binding because TSS-region default extends too far for RNA features.

Algorithmic Taxonomy

ToolInputAnnotation levelFeature priorityOutputStrengthFails when
ChIPseeker (R)Peak BED + TxDbGene-level or transcript-levelPromoter > 5' UTR > 3' UTR > Exon > Intron > IntergenicPer-peak annotation + pie chart + distance-to-TSSMature, widely used, fastTSS-region default c(-3000, 3000) over-extends; promoter category meaningless for RNA
RCAS (R)Peak BED + GFFTranscript-levelCustomizableHTML report + per-region tableRNA-specific design; ncRNA-awareSlow; HTML-heavy; less actively maintained
RBP-Maps (Yeo)eCLIP BAM + alternative splicing event tablesSplice junction-levelNA (positional metagene)Splicing regulatory map (1400nt vector around cassette exon)The standard for splicing-factor CLIP analysisSplicing-only; not for 3' UTR or ncRNA binders
HOMER annotatePeaks.plPeak BEDGene-levelPromoter > UTR > Exon > Intron > IntergenicAnnotation TSVFamiliar from ChIP-seqSame TSS issue as ChIPseeker; less RNA-aware
bedtools intersectPeak BED + feature BEDUser-definedUser-definedPer-feature overlapMost flexibleManual hierarchy logic must be written
pyranges (Python)Peak BED + GTFCustomizableUser-definedDataFrameFast, pythonic, scriptableHand-written priority logic
pybedtoolsPeak BED + GTFCustomizableUser-definediteratorsFlexible, scriptedSame as pyranges
Peakhood (Uhl 2022)Peak BED + transcriptsTranscript contextNAPer-peak transcript contextResolves transcript-isoform ambiguitySpecialized; not a global region tool

Methodology evolves; verify the GTF source (GENCODE preferred over Ensembl for human; both have feature differences) and the priority hierarchy against published RBP literature.

Critical Choice: Annotation Hierarchy

RNA features are nested and overlapping. A 3' UTR peak is ALSO in an exon; an intronic peak near a splice site is ALSO in a transcribed-but-spliced region. Without a priority hierarchy the same peak gets counted multiple times.

Standard CLIP hierarchy (most-specific first):

  1. 3' UTR (specific to mature mRNA biology)
  2. 5' UTR
  3. CDS (coding sequence)
  4. Exon (catch-all for non-UTR exonic)
  5. Splice site (within 50 nt of donor/acceptor; key for splicing factors)
  6. Intron
  7. Promoter (within 1 kb of TSS; usually not relevant for RNA-binding)
  8. 5' flanking / 3' flanking (within 1 kb of gene)
  9. ncRNA (lncRNA, snoRNA, snRNA, miRNA host)
  10. Repeat element (Alu, LINE, LTR, SINE) - separate axis from above
  11. Intergenic

ChIPseeker's default level='gene' reports the FIRST hit by priority. For CLIP, set level='transcript' and explicitly customize tssRegion=c(-100, 100) (default c(-3000, 3000) is for ChIP).

Why CLIP needs the tight TSS window: ChIPseeker was designed for ChIP-seq where the "Promoter" category captures transcription-factor binding within 3 kb upstream of a gene. For CLIP-seq, the binding substrate is RNA - the RBP cannot bind upstream of the TSS because there is no RNA upstream of the TSS. A 6 kb TSS window catches 5'-UTR peaks and unrelated upstream sequence, mislabeling 30-50% of CLIP peaks as "Promoter (2-3kb)" - implausible for an RNA-binding study. Tightening to c(-100, 100) constrains the "Promoter" category to the mRNA 5' end immediate vicinity and pushes deeper-5'-UTR peaks into the "5UTR" category where they belong.

r
library(ChIPseeker)
library(TxDb.Hsapiens.UCSC.hg38.knownGene)
txdb <- TxDb.Hsapiens.UCSC.hg38.knownGene

peaks <- readPeakFile('peaks.stringent.bed')

# CLIP-appropriate annotation
anno <- annotatePeak(
    peaks,
    TxDb = txdb,
    level = 'transcript',     # RNA-isoform-aware
    tssRegion = c(-100, 100), # tight; not a meaningful promoter for RNA binding
    genomicAnnotationPriority = c(
        'Promoter', '5UTR', '3UTR', 'Exon',
        'Intron', 'Downstream', 'Intergenic'
    )
)

print(anno)
plotAnnoPie(anno)
plotAnnoBar(anno)
plotDistToTSS(anno)  # CLIP-specific: dist to TSS is dist to mRNA 5' end

RBP Class -> Expected Annotation

RBP classExpected dominant regionRBP examples
Splicing factorsIntron (within 500 nt of splice junction)PTBP1, U2AF2, RBFOX1/2, SRSF1-9, HNRNPC, MBNL1
3' UTR mRNA stability/decay3' UTRHuR (ELAVL1), PUM2, AUF1, ZFP36, TIA1, NUDT21
5' UTR translation regulation5' UTR, CDS startEIF4A3, EIF3J, eIF2 subunits, LARP1, PCBP1/2
Ribosome-associatedCDSRPL/RPS proteins, RACK1
Nuclear exportCDS, 3' UTRNXF1, ALYREF, SRSF3
m6A reader3' UTR, stop codon regionYTHDF1/2/3, IGF2BP1/2/3
miRNA effector3' UTR (seed-matched sites)AGO1/2/3/4, DICER1, TNRC6A/B/C
snoRNA-targeting RBPsnoRNA, rRNADKC1, FBL, NOP56, NHP2
ncRNA-bindingSpecific ncRNA targetTROVE2 (Y RNA), LARP7 (7SK), SSB (vault)
MitochondrialchrM (mt-mRNA + 12S/16S rRNA)FASTKD2, LRPPRC, TFAM, MTPAP
Histone mRNAHistone 3' end stem-loopSLBP
Repeat-binding (TE)Alu, LINE-1, LTR, SINEMATR3, ZFP36, HNRNPK, HNRNPA2B1, FUS (LINE-1)
LCD/IDR / phase separationDiverse (Alu, intron, 3' UTR)FUS, TDP-43, EWSR1, TAF15

Mismatch between observed and expected suggests: (a) failed IP (cross-reactive antibody binding wrong RBP), (b) artifact (high-expression contamination), (c) genuinely new biology - investigate.

RBP-Maps: Splicing Regulatory Maps

For splicing factors (PTBP1, U2AF2, RBFOX, SRSF1, etc.), the Yeo lab RBP-Maps tool aggregates eCLIP signal in a 350 nt window flanking the upstream, cassette, and downstream exons, producing a 1400 nt vectorized regulatory map. This map shows the position-dependent enrichment of RBP binding relative to alternative splicing events and is THE standard for splicing-CLIP analysis.

bash
# Yeo lab rbp-maps Snakemake workflow
git clone https://github.com/YeoLab/rbp-maps
cd rbp-maps

# Inputs:
#   1. eCLIP BAM (rep1 + rep2)
#   2. Cassette exon BED (from RNA-seq KD experiment; see clip-seq/differential-clip)
#   3. Background "native" exons (constitutively spliced)

python rbpmaps/RBPMaps.py \
    --ip rep1.bam rep2.bam \
    --input sminput.bam \
    --positive cassette_inc.bed \
    --negative cassette_exc.bed \
    --background native_constitutive.bed \
    --output rbpmaps_out/ \
    --normalization input_normalization

Output: 1400 nt vector with eCLIP signal at every base, plotted as a metagene for cassette inclusion vs exclusion vs constitutive. A peak in the upstream-intron region (right of the cassette exon's 5' splice site) indicates intronic regulation; a peak in the cassette exon body indicates exonic regulation.

Per-Tool Failure Modes

ChIPseeker -- TSS-region default over-extends

Trigger: Used ChIPseeker default tssRegion=c(-3000, 3000) for CLIP peaks.

Mechanism: The 6 kb TSS window catches CLIP peaks deep in the 5' UTR or upstream introns and labels them "Promoter (2-3kb)". For RNA biology, the "promoter" category is meaningless - the RBP cannot bind pre-mRNA upstream of the TSS.

Symptom: 30-50% of peaks labeled "Promoter (2-3kb)" - implausible for an RNA-binding study.

Fix: Set tssRegion=c(-100, 100) or c(-50, 50). Remove the "Promoter" category from the priority vector entirely if it gives 0 hits.

ChIPseeker -- Gene-level loses isoform context

Trigger: Used level='gene' (default) on peaks that should be transcript-specific.

Mechanism: Gene-level annotation picks the canonical/longest transcript and assigns the peak to that transcript's features. A peak in an alternatively-spliced exon may be labeled "intron" because the gene-level reference picks an isoform where the exon is skipped.

Symptom: Splicing factor peaks heavily "intronic" when they should be exonic at alternative-cassette exons.

Fix: Use level='transcript'. Match isoforms to RBP-Maps cassette-exon coordinates for splicing factors.

RCAS -- Slow on large peak sets

Trigger: RCAS report on > 100k peaks.

Mechanism: RCAS generates an HTML report with many plots; rendering is slow.

Symptom: Run time > 1 h; out-of-memory; HTML failing to render.

Fix: Sub-sample peaks to top 10k; or use ChIPseeker for the global view and reserve RCAS for per-region deep dives.

RBP-Maps -- Requires SE event table

Trigger: Splicing-factor CLIP with no companion RNA-seq KD or no SE event BED.

Mechanism: RBP-Maps's regulatory metagene only makes sense relative to cassette exons differentially regulated by the RBP. Without the SE event table, the tool falls back to plotting at all native exons (no differential signal).

Symptom: Flat metagene; no position-dependent enrichment.

Fix: Generate the cassette-exon table from companion RNA-seq KD experiments (rMATS / MAJIQ / LeafCutter; see alternative-splicing skills) OR use a published table from ENCODE shRNA RNA-seq.

bedtools -- Manual hierarchy errors

Trigger: Hand-written bedtools intersect chain for region annotation.

Mechanism: Forgetting to enforce priority leads to double-counting: a 3' UTR peak intersects 3UTR.bed AND exon.bed.

Symptom: Per-region counts sum to > total peak count by 20-50%.

Fix: Process in priority order; remove already-annotated peaks from the remaining set with -v before next intersect. Or use ChIPseeker / RCAS which encode the hierarchy internally.

Show full SKILL.md (968 more words)Show less
Repeat-element overlap missed

Trigger: Repeat-binding RBP (MATR3, HNRNPK) annotated without RepeatMasker BED.

Mechanism: Genomic feature BEDs (UTR/CDS/intron) do not flag repeat instances within them. An MATR3 peak in an intronic LINE-1 is labeled "intron" when "LINE-1 within intron" is the biology.

Symptom: Repeat-binding RBP looks like an intronic binder; functional analysis (GO terms) returns generic "RNA processing".

Fix: Pass RepeatMasker BED (UCSC) as a separate axis; annotate each peak with BOTH region AND repeat-class. Cross-check repeat-class fraction against literature: > 15% Alu / LINE / LTR overlap indicates a repeat binder.

Mitochondrial peaks excluded by default

Trigger: Standard TxDb annotation drops chrM transcripts; FASTKD2 / LRPPRC peaks all "intergenic".

Mechanism: Some TxDb objects (older releases) exclude mitochondrial chromosome; ChIPseeker reports those peaks as having no annotation.

Symptom: chrM peaks labeled "Distal Intergenic" or simply missing from the annotation table.

Fix: Verify TxDb includes chrM (genes(txdb, filter=list(tx_chrom='chrM'))); add custom mt-mRNA + mt-rRNA + mt-tRNA BED if missing.

Strand information ignored

Trigger: bedtools intersect -wa -wb without -s flag.

Mechanism: CLIP is strand-specific. Without -s, peaks on one strand can be annotated to features on the other strand - especially for overlapping transcripts (sense/antisense pairs).

Symptom: Antisense transcript peaks contaminate the sense annotation; per-region distribution shifted.

Fix: Always pass -s to bedtools intersect for CLIP. ChIPseeker handles strand automatically when peaks have BED column 6.

Metagene Analysis

Metagene = aggregated profile of CLIP signal at each base of a normalized feature (e.g., 5' UTR, CDS, 3' UTR scaled to common length). Reveals position-dependent binding.

Goal: Generate a position-resolved CLIP-signal profile aggregated across all 3' UTRs (or other features) to identify position-dependent binding modes (e.g., stop-codon-proximal m6A readers, polyA-proximal CPEB).

Approach: Convert dedup BAM to strand-stranded bigWig, then use deepTools computeMatrix scale-regions to align signal across feature instances of variable length, and plotProfile / plotHeatmap for visualization.

bash
# deepTools computeMatrix for metagene
# Step 1: peaks BED -> bedgraph
bedtools genomecov -bg -strand + -ibam dedup.bam > clip_plus.bg
bedtools genomecov -bg -strand - -ibam dedup.bam > clip_minus.bg
# combine to bigwig
bedGraphToBigWig clip_plus.bg chrom.sizes clip_plus.bw

# Step 2: metagene over 3' UTRs
computeMatrix scale-regions \
    -S clip_plus.bw clip_minus.bw \
    -R three_prime_UTRs.bed \
    --regionBodyLength 500 \
    --beforeRegionStartLength 100 \
    --afterRegionStartLength 100 \
    -o matrix.gz

plotProfile -m matrix.gz -o metagene.png --perGroup
plotHeatmap -m matrix.gz -o metagene_heatmap.png
Metagene shapeBiology
Peak at stop codon regionm6A reader (YTHDF1/2/3); IGF2BP1; some 3' UTR factors
Peak in proximal 3' UTR (within 200 nt of stop)PUM2, IGF2BP, miRNA effectors
Peak in distal 3' UTR (within 200 nt of poly-A)NUDT21, CPEB1, polyadenylation regulators
Peak at 5' splice site flanking intronRBFOX, MBNL, HNRNPC
Peak at 3' splice site / branch pointU2AF2, PTBP1
Peak at 5' UTR (within 50 nt of TSS)EIF3, EIF4A3, LARP1
Flat across CDSCytoplasmic ribosome-associated; less position-specific

Decision Tree by Scenario

ScenarioTool + parametersWhy
Global region distribution (3' UTR / CDS / intron)ChIPseeker tssRegion=c(-100,100) level='transcript'Mature, fast, ENCODE-comparable
Splicing factor regulatory mapRBP-Maps (Yeo) with RNA-seq KD cassette tableThe standard for splicing-CLIP
RNA-feature aware (snoRNA, lncRNA-specific)RCAS or custom GTF + bedtoolsRCAS handles ncRNA features
Transcript-isoform context resolvedPeakhood + ChIPseeker level='transcript'Resolves which isoform the peak supports
Metagene 3' UTR / 5' UTRdeepTools computeMatrix + plotProfileStandard metagene visualization
Repeat-element overlap (MATR3, HNRNPK)bedtools intersect with RepeatMaskerRepeat axis is independent of region axis
Mitochondrial RBP (FASTKD2)Custom chrM-aware annotationStandard TxDb may drop chrM
Bulk RBP overviewChIPseeker plotAnnoPie + plotDistToTSSQuick orienting plots
Compare two RBPs' region distributionsChIPseeker annotatePeakList + side-by-side barDirect comparability
miRNA target site (AGO-CLIP)bedtools intersect with 3' UTR + TargetScan seed scanmiRNA biology specific
m6A reader (YTHDF, IGF2BP)ChIPseeker + custom stop-codon-distance plotm6A biology stops-codon-proximal

Reconciliation: When Annotations Disagree

PatternLikely causeAction
ChIPseeker "Promoter" >> 5' UTR for CLIPTSS region too wideTighten tssRegion=c(-100,100)
ChIPseeker gene-level intronic; RBP-Maps shows exonic at cassettesGene-level picks canonical isoform; RBP-Maps uses cassettesTrust RBP-Maps for splicing factors
Many peaks "intergenic" near a known geneTxDb missing the transcriptVerify GENCODE vs Ensembl GTF; add lincRNA / unannotated transcripts
chrM peaks missingTxDb excluded chrMAdd custom mt-mRNA BED
Sum of per-region counts > total peaksHierarchy not enforcedUse ChIPseeker; or bedtools intersect -v chain
Repeat-binding RBP looks "intronic"RepeatMasker axis not addedAnnotate peaks with RepMask intersect separately
Metagene flat for known position-specific RBPWrong feature BED (e.g., gene-body instead of transcript-isoform)Use isoform-aware feature BED
RBP-Maps signal flat at cassette exonsNo RNA-seq KD reference tableGenerate cassettes from companion KD; use ENCODE shRNA tables

Operational rule: Report (a) ChIPseeker global region pie, (b) RBP-Maps metagene for splicing factors, (c) repeat-element fraction separately, (d) transcript-isoform context with Peakhood if ambiguous. Three views together prevent misinterpretation.

Common Errors

Error / symptomCauseSolution
"Promoter" >> 30% of peaks in CLIPTSS region default too widetssRegion=c(-100,100)
Strand-mixing artifactsBedtools intersect without -sAlways pass -s for CLIP
3' UTR fraction < 10% for HuRWrong txdb version (UTR poorly annotated)Use GENCODE v38+ TxDb
chrM peaks lostTxDb excludes mitochondrialAdd mt-transcript BED
RBP-Maps flat metageneNo cassette exon BEDGenerate from RNA-seq KD; or use ENCODE table
Repeat fraction over-countedSame repeat instance counted multiple timesbedtools merge repeat BED first
ChIPseeker level='gene' collapses isoformsDefault level=geneSet level='transcript'
Annotation runs slow for large peak setRCAS HTML generationUse ChIPseeker for global; RCAS for top-K
GO enrichment yields generic termsRepeat-binding RBP annotated as intronicRe-annotate with RepeatMasker axis
Peak in unannotated lincRNA labeled intergenicOlder GENCODE missing the lincRNAUpdate GENCODE; add manual transcript BED

References

  • Yu G et al 2015 Bioinformatics 31:2382 (ChIPseeker)
  • Uyar B, Yusuf D et al 2017 Nucleic Acids Res 45:e91 (RCAS)
  • Yee BA et al 2019 RNA 25:193 (RBP-Maps splicing regulatory maps)
  • Van Nostrand EL et al 2020 Nature 583:711 (ENCODE 150 RBP eCLIP, region distribution analysis)
  • Quinlan AR & Hall IM 2010 Bioinformatics 26:841 (bedtools)
  • Hentze MW et al 2018 Nat Rev Mol Cell Biol 19:327 (RBP function review)
  • Uhl M et al 2022 Bioinformatics 38:1139 (Peakhood transcript-context)
  • ENCODE eCLIP standards (encodeproject.org/eclip) - region distribution conventions
  • clip-seq/clip-peak-calling - Upstream peak calls
  • clip-seq/crosslink-site-detection - Single-nt CL sites for fine-grained metagene
  • clip-seq/clip-motif-analysis - Motif discovery within annotated regions
  • clip-seq/differential-clip - Cross-condition annotated peaks for regulatory maps
  • clip-seq/ago-clip-mirna-targets - 3' UTR seed-matched site annotation
  • clip-seq/m6a-clip - Stop-codon-proximal m6A annotation
  • genome-intervals/gtf-gff-handling - GTF preparation for annotation
  • genome-intervals/interval-arithmetic - bedtools intersect patterns
  • alternative-splicing/differential-splicing - Cassette exon tables for RBP-Maps
  • chip-seq/peak-annotation - DNA-protein annotation analogue

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clip-seq/binding-site-annotation of GPTomics/bioSkills.

  • SKILL.md
  • examples/annotate_peaks.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clip Seq Binding Site Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clip Seq Binding Site Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clip Seq Binding Site Annotation this skillGPTomics/bioSkills1.2k2 repos~5.4kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Clip Seq Binding Site Annotation

What does Bio Clip Seq Binding Site Annotation do?

Annotate CLIP-seq peaks or crosslink sites to RNA features (5'UTR, CDS, 3'UTR, intron, splice junction, snoRNA, tRNA, ncRNA, repeat elements) with ChIPseeker, RCAS, RBP-Maps (Yeo splicing regulatory…. Bio Clip Seq Binding Site Annotation is an agent skill from GPTomics/bioSkills. Annotate CLIP-seq peaks or crosslink sites to RNA features (5'UTR, CDS, 3'UTR, intron, splice junction, snoRNA, tRNA, ncRNA, repeat elements) with ChIPseeker, RCAS, RBP-Maps (Yeo splicing regulatory maps), and bedtools, applying feature-priority hierarchies, transcript-context resolution, and metagene aggregation.

When should I use Bio Clip Seq Binding Site Annotation?

Bio Clip Seq Binding Site Annotation fits situations like: characterizing where in transcripts an RBP binds; comparing peak distribution across regions; generating splicing-regulatory maps relative to alternative-splicing events; distinguishing exonic vs intronic vs UTR binding.

How do I install Bio Clip Seq Binding Site Annotation in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-binding-site-annotation -a claude-code`. Or copy the skill folder (clip-seq/binding-site-annotation in GPTomics/bioSkills) into .claude/skills/bio-clip-seq-binding-site-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clip Seq Binding Site Annotation in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-binding-site-annotation -a codex`. Or copy the skill folder (clip-seq/binding-site-annotation in GPTomics/bioSkills) into .agents/skills/bio-clip-seq-binding-site-annotation in your project. Codex loads it when a task matches its description.

Can I use Bio Clip Seq Binding Site Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-binding-site-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clip-seq-binding-site-annotation, .gemini/skills/bio-clip-seq-binding-site-annotation, .github/skills/bio-clip-seq-binding-site-annotation and .opencode/skills/bio-clip-seq-binding-site-annotation in your project.

What does Bio Clip Seq Binding Site Annotation need to run?

Going by SKILL.md and its folder, Bio Clip Seq Binding Site Annotation needs R for the scripts in its folder and the command-line tools its instructions call (git, python and pip). Our summary lists: Python 3.

Does Bio Clip Seq Binding Site Annotation access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Clip Seq Binding Site Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clip Seq Binding Site Annotation use?

Bio Clip Seq Binding Site Annotation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clip Seq Binding Site Annotation use?

About 5.4k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clip Seq Binding Site Annotation?

Skills that share tags, products or a category with Bio Clip Seq Binding Site Annotation: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clip Seq Binding Site Annotation?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.