Agent skill

Bio Clip Seq Clip Motif Analysis

by GPTomics in GPTomics/bioSkills

Discover RBP binding motifs from CLIP-seq peaks or single-nucleotide crosslink sites using HOMER, MEME/STREME, kpLogo, mCross (CL-position-registered motifs), PEKA (positional k-mer enrichment)…

MITAuto-check passedResearch & Science

Install Bio Clip Seq Clip Motif Analysis

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clip-seq-clip-motif-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clip-seq-clip-motif-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clip-seq/clip-motif-analysis .claude/skills/bio-clip-seq-clip-motif-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clip-seq-clip-motif-analysis
GitHub stars
1.2k
Used in
2 other repos
Token cost
~5.5k tokens
SKILL.md length
2,562 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Discover RBP binding motifs from CLIP-seq peaks or single-nucleotide crosslink sites using HOMER, MEME/STREME, kpLogo, mCross (CL-position-registered motifs), PEKA (positional k-mer enrichment)…

  • Characterizing RBP sequence specificity
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Critical Choice: Peak-Based vs… and Uracil Crosslinking Bias, plus 8 more sections
  • Runs Shell scripts from its folder; calls pip
  • Registering motifs to crosslink positions

What it does

Bio Clip Seq Clip Motif Analysis is an agent skill from GPTomics/bioSkills. Discover RBP binding motifs from CLIP-seq peaks or single-nucleotide crosslink sites using HOMER, MEME/STREME, kpLogo, mCross (CL-position-registered motifs), PEKA (positional k-mer enrichment), RBPamp (affinity), and RNA Bind-n-Seq (RBNS) cross-validation. Use when characterizing RBP sequence specificity, registering motifs to crosslink positions, validating in vivo CLIP motifs against in vitro RBNS Kd, reconciling motif disagreements across tools, or correcting for the uracil crosslinking bias that contaminates…

Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/find_motifs.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Characterizing RBP sequence specificity
  • Registering motifs to crosslink positions
  • Validating in vivo CLIP motifs against in vitro RBNS Kd
  • Reconciling motif disagreements across tools

Example prompts

  • “/bio-clip-seq-clip-motif-analysis”

Requirements

  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clip Seq Clip Motif Analysis loads about 5.5k tokens when it runs. Until then it costs about 143 tokens; SKILL.md has 2,562 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,562 words, ~5,501 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clip-seq-clip-motif-analysis/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clip-seq-clip-motif-analysis
description
Discover RBP binding motifs from CLIP-seq peaks or single-nucleotide crosslink sites using HOMER, MEME/STREME, kpLogo, mCross (CL-position-registered motifs), PEKA (positional k-mer enrichment), RBPamp (affinity), and RNA Bind-n-Seq (RBNS) cross-validation. Use when characterizing RBP sequence specificity, registering motifs to crosslink positions, validating in vivo CLIP motifs against in vitro RBNS Kd, reconciling motif disagreements across tools, or correcting for the uracil crosslinking bias that contaminates raw CLIP motif logos.
tool_type
mixed
primary_tool
HOMER

Version Compatibility

Reference examples tested with: HOMER 4.11+, MEME Suite 5.5+ (STREME, MEME-ChIP, FIMO), bedtools 2.31+, kpLogo 1.1+, mCross v1+, PEKA v1+, RBPamp 0.9+, ggseqlogo 0.1+, biopython 1.83+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures

If code throws unexpected errors, introspect the installed tool and adapt the example to match the actual API rather than retrying.

CLIP-seq Motif Analysis

"Find enriched RNA motifs at my RBP binding sites" -> Discover the in vivo sequence preference of an RNA-binding protein from CLIP-seq peaks or single-nucleotide crosslink sites. The fundamental confound is the uracil bias of UV254 crosslinking: U is the most-crosslinked base, so naive motif logos centered on CL positions are U-enriched even for non-U-binding RBPs. Modern tools (mCross, PEKA) register motifs relative to the CL position and correct for this bias; legacy tools (HOMER, MEME) need careful background selection.

  • CLI (de novo, peak-based, HOMER RNA mode): findMotifs.pl peaks.fa fasta motif_out -rna -len 5,6,7,8 -p 4
  • CLI (de novo, peak-based, MEME-ChIP / STREME): streme --rna --oc streme_out -p peaks.fa -n background.fa --minw 5 --maxw 10
  • CLI (positional, single-nt CL-registered, mCross): mCross -i crosslinks.bed -g genome.fa -k 7 -o mcross_out (jointly models motif + CL position)
  • CLI (positional k-mer, no input control needed, PEKA): peka --peak_file_name peaks.bed --crosslinks_file_name crosslinks.bed --genome_file_name genome.fa --regions_file_name regions.bed --kmer_length 5 --percentile 30 --outpath peka_out (the short flags -i/-x/-g/-r/-k/-p shown in earlier docs are not all stable; use the long forms or check peka --help)
  • CLI (affinity-weighted, RBPamp): rbpamp run -i peaks.fa -k 7 -o rbpamp_out (joint affinity + motif model)
  • CLI (positional logo, kpLogo): kpLogo crosslinks_with_kmer_scores.txt -o kplogo_out (position-specific significance)

ENCODE eCLIP motif convention: extract sequences from peaks (CLIPper + SMInput stringent), use shuffled or expression-matched background, run HOMER -rna -len 5,6,7,8. The Yeo lab eCLIP papers typically report the top 1-3 HOMER hits per RBP and validate with mCross for CL-registered position. For ASB and variant-effect work, mCross is mandatory.

Algorithmic Taxonomy

ToolInputBackgroundCL-position awareOutputStrengthFails when
HOMER findMotifs.plPeak FASTAAuto-shuffled or user-providedNoPWM + de novo + known scanMature, fast, RNA mode (U instead of T)Background generation is heuristic; uracil bias inflates U-content motifs
MEME (MEME Suite)Peak FASTAOptional Markov modelNoPWMStatistically rigorous; broadly citedSlow on large peak sets (> 1000 sequences); single-motif default
STREME (MEME Suite)Peak FASTAShuffled or user-providedNoPWMFast successor to MEME-ChIP; tolerant of large setsSame uracil bias issue as HOMER
MEME-ChIPPeak FASTAMarkov order 1-3NoMultiple motifsWrapper for MEME + CentriMo + Tomtom comparisonDesigned for ChIP; for CLIP use STREME directly
kpLogok-mer scores with positionsNA (significance from kpLogo internals)Yes (positional)Position-specific logoVisualizes position-specific signal; complements PWMRequires k-mer-score input file; not a peak-to-motif tool by itself
mCross (Feng 2019)Single-nucleotide CL sites + genomeGenerated internallyYes (jointly models motif + CL position)Registered PWM + CL offsetThe standard for in vivo RBP motif registration; SRSF1 example showed clustered GGA half-sitesRequires high-quality single-nt CL sites (PureCLIP or CTK CITS)
PEKA (Kuret 2022)Peaks + crosslink BEDLow-count crosslinks within same datasetYes (positional k-mer)Position-enriched k-mers with significanceNo external input required; cross-validates with mCrossLess granular than mCross's jointly-modeled PWM
RBPamp (Jens 2022)Peak FASTA or CL sitesJoint affinity modelNo (affinity-based)Affinity-weighted motif + Kd estimateProvides Kd; compatible with RBNS calibrationSlow; small community; needs validation
RCASPeak BED + GTFAutoNoMotif + region annotationOne-stop for motif + region distributionLess granular than per-tool dedicated runs
MEME FIMO (motif scan)PWM + FASTAMarkov bgNoMatch locations and scoresScan known motifs across genomeNot a de novo tool; complements HOMER known scan
RBPNet (Horlacher et al 2023)Raw sequenceTrained modelYes (sequence-to-signal at single nt)Predicted CL count distributionDeep learning; predicts per-base affinity profileSee clip-seq/clip-deep-learning

Methodology evolves; verify motif comparison databases (CISBP-RNA, ATtRACT, oRNAment, mCrossBase, RBPDB) for known motif validation. RBNS in vitro Kd values (Lambert 2014; Dominguez 2018 for 78 RBPs) are the most reliable orthogonal reference for in vivo CLIP motifs.

Peak-based (HOMER, MEME/STREME, RBPamp): Extract sequences from peak BED, find enriched k-mers / PWMs. Produces classical motif logos. Fast and intuitive. Captures broad binding zones.

Crosslink-site-based (mCross, PEKA, kpLogo): Use single-nucleotide CL positions; analyze sequence in a window around each CL position. Produces motif logos REGISTERED to the CL site. Required for understanding the structural relationship between motif and crosslink.

The fundamental difference: peak-based logos show "what sequence is enriched near binding"; CL-registered logos show "what sequence the RBP contacts at the crosslink position." For most RBPs they agree on the core motif but diverge on flanking context.

Uracil Crosslinking Bias

UV254 crosslinking is U-biased: single-nt CL events are strongly enriched at U residues (Konig 2010; Sugimoto 2012). Consequence: motif logos centered on raw CL positions are U-enriched at the center even for non-U-binding RBPs (e.g., PUM2 binds UGUANAUA but peak-center is shifted; HNRNPK binds C-rich tracts but centers can show spurious U).

Mitigations:

  • Use peak boundaries, not single-nt CL, for HOMER/MEME (peak-level averaging dilutes the U bias).
  • Use mCross or PEKA, which explicitly model the CL position offset from the motif center.
  • Shuffle background should preserve mononucleotide composition (HOMER default does; MEME with --markov-order 1 does).
  • For PAR-CLIP, the bias is at U residues but T->C marks the EXACT CL position; mCross / PEKA register more cleanly than for iCLIP/eCLIP.

For naive HOMER/MEME output, examine the central column of the PWM: if it is U-skewed, suspect CL bias rather than true U preference. Cross-check against published RBNS Kd if available.

Per-Tool Failure Modes

HOMER -- Background mismatch inflates GC-biased motifs

Trigger: HOMER auto-generates a shuffled background that preserves only nucleotide frequency, not GC distribution per region; CLIP peaks come disproportionately from 3' UTRs which are AU-rich.

Mechanism: AU-rich foreground vs GC-shuffled background produces "AU-rich motif" as the top hit for any 3' UTR-binding RBP, regardless of biology.

Symptom: Top HOMER motif looks like the 3' UTR mean composition (AU-rich, no clear specificity); known motif scan (-known) finds the RBP's canonical motif at lower rank than the spurious "AU" motif.

Fix: Provide a GC-matched background: extract sequences from random 3' UTR regions of expressed transcripts matched by length distribution. Use -bg matched_bg.fa flag. Or use STREME with -n matched_bg.fa.

MEME -- Slow on large peak sets

Trigger: MEME run on > 1000 peak sequences.

Mechanism: MEME is O(N^2) in sequence count; classical OOPS/ZOOPS/ANR models do not scale.

Symptom: Runtime > 12h; out-of-memory on large machines.

Fix: Use STREME (MEME Suite >= 5.0); it is the modern fast successor (sub-quadratic). Or restrict to top-1000 peaks by score.

mCross -- Requires single-nt CL sites

Trigger: mCross run on a peak BED instead of crosslink BED.

Mechanism: mCross's model is motif_PWM x CL_position_offset_PMF. Without single-nt CL positions there is no offset axis; the model degenerates.

Symptom: mCross output looks identical to a peak-centered logo; the diagnostic CL-offset histogram is flat.

Fix: Provide CL sites from PureCLIP, CTK CITS, or PARalyzer. The Yeo lab and Zhang lab use the convention of PureCLIP CL sites as mCross input.

PEKA -- Background from same dataset

Trigger: Concerned that PEKA's "low-count crosslinks within same dataset" background is biased.

Mechanism: PEKA does NOT need external input. It uses low-count crosslinks (below a quantile threshold) as the background and high-count crosslinks (above) as the foreground. This is by design (Kuret 2022) and cross-validated against mCross.

Symptom: False positive concern; user wants to verify.

Fix: Run BOTH PEKA and mCross; cross-validate. The PEKA paper reports high concordance with mCross for ENCODE RBPs; treat them as orthogonal confirmation.

RBPamp -- Slow convergence

Trigger: Large peak set (> 5000); RBPamp jointly optimizes affinity + motif iteratively.

Mechanism: Joint optimization is slow; some RBPs have multimodal binding (multiple motifs) that confuse single-PWM RBPamp.

Symptom: Runtime > 24h; output PWM disagrees with HOMER top motif.

Fix: Down-sample to top-2000 peaks; or run HOMER for initial PWM and use it as RBPamp seed.

RBNS comparison -- in vitro vs in vivo divergence

Trigger: Comparing CLIP-derived motif to RBNS Kd-ranked motif (Dominguez 2018) and finding divergence (e.g., CLIP top motif rank 3 in RBNS; RBNS top motif weakly enriched in CLIP).

Mechanism: In vivo binding is shaped by structure accessibility (RNA secondary structure), cooperativity, and competing factors. RBNS measures intrinsic affinity in vitro. They diverge for:

  • Structurally regulated RBPs (PUM2, RBFOX2 favor accessible loops)
  • Cooperative binders (FUS, TDP-43, SR proteins)
  • Co-factor-dependent (SRSF1 paralogs)

Symptom: CLIP motif logo plausible but rank-3 in RBNS; or RBNS top motif missing from CLIP peaks.

Fix: Accept the divergence as informative. Cite both. Use SHAPE-eCLIP or icSHAPE-MaP to test structural-accessibility hypothesis. For variant-effect studies, prefer RBNS-derived PWM (RBPamp Kd) over CLIP motif for absolute affinity prediction.

Known motif scan -- FIMO threshold too lenient

Trigger: Used FIMO with default -thresh 1e-4 and got matches everywhere.

Mechanism: Default threshold is too lenient for in vivo binding; CLIP background sequences contain many low-affinity matches.

Symptom: FIMO match rate > 50% of peaks AND > 20% of random background; no enrichment signal.

Fix: Use -thresh 1e-5 or -thresh 1e-6; or use CentriMo for local enrichment around peak center (testing positional enrichment, not threshold).

Show full SKILL.md (1,040 more words)Show less

Decision Tree by Scenario

ScenarioToolWhy
Quick de novo motif from CLIPper peaks, ENCODE-comparableHOMER -rna -len 5,6,7,8 with GC-matched bgFast, mature, ENCODE convention
Motif registered to crosslink position (for variant effect)mCross on PureCLIP single-nt sitesThe 2019 Feng paper is the canonical reference
Confirm mCross result with independent methodPEKA on same peaks + crosslinksCross-validation; PEKA does not need input control
Affinity-weighted motif with Kd estimateRBPampProvides quantitative Kd; integrates with RBNS comparison
Compare to in vitro RBNS KdRBPamp + Dominguez 2018 / Lambert 2014 tablesRBNS is the orthogonal in vitro standard
RBP with known motif - validation scanFIMO -thresh 1e-5 against CISBP-RNA / ATtRACT PWMLower-confidence motifs need tight threshold
Multiple binding modes (SRSF1 GGA clusters)mCross + manual inspection of CL-offset histogrammCross reveals the multimodal structure
Position-specific logo (5' vs 3' of CL)kpLogo with k-mer scoresPosition-specific significance visualization
AU-rich 3' UTR binders (HuR, AUF1) - validate U bias is realCompare HOMER motif center vs PUM2-style registerRBP-specific tradition
Allele-specific motif (variant effect)mCross PWM + DeepRiPe/RBPNet for variant scoringSee clip-seq/clip-deep-learning
snoRNA / structural RNA - sequence motif less meaningfulSkip motif analysis; emphasize structural context (icSHAPE)snoRNA binders read structure, not linear sequence

Reconciliation: When Motif Tools Disagree

PatternLikely causeAction
HOMER and STREME agree; mCross disagrees on flankingmCross's CL-offset registers a flanking base that's identical in peak; HOMER sees both flanking bases as motifTrust mCross for CL-registered position; HOMER is "what's nearby"
Top HOMER motif AU-rich; RBP knows to bind structureGC-matched background not usedRe-run HOMER with -bg matched_bg.fa
HOMER finds canonical motif at rank 1; RBNS Kd ranks at rank 3In vitro vs in vivo specificity divergenceBoth valid; cite RBNS for absolute affinity, CLIP for in vivo prevalence
mCross PWM has multiple CL-position peaksMultiple binding modes (SRSF1 GGA half-sites; PTBP1 multimer)Treat as biologically meaningful; report both
Skipper window-level motifs differ from CLIPper peak-levelWindow size dilutes signal; peak boundaries sharpenBoth valid; report at the level of the downstream analysis
PAR-CLIP motifs cleaner than iCLIP for same RBPPAR-CLIP T->C registers exactly; iCLIP truncates -1 from CLBoth correct; PAR-CLIP is the sharper single-nt method when 4SU labeling tolerated
HOMER finds different motif on top vs bottom strandStrand information lost in upstream BEDVerify BED column 6 (strand) is preserved; HOMER respects strand

Operational rule for high-confidence motif reporting: Report (a) HOMER top de novo motif with GC-matched background, (b) mCross PWM with CL-position offset histogram, (c) PEKA k-mer enrichment for orthogonal confirmation, (d) FIMO scan against published RBP PWM (CISBP-RNA / ATtRACT) for known-motif validation, (e) RBNS Kd ranking from Dominguez 2018 if available. Three independent methods agreeing on the core motif is the bar.

Workflow: De Novo + Known + CL-Registered

Goal: Produce a publication-grade motif report combining peak-based de novo discovery, CL-position-registered motif (mCross), and known-motif validation.

Approach: Extract sequences with strand preservation, build a GC-matched 3' UTR background, run HOMER + STREME for de novo, mCross for CL-registered, FIMO scan against CISBP-RNA, and report all three views with information-content QC.

bash
# Step 1: Extract peak sequences (use stringent peak set: log2 FC >= 3, -log10 p >= 3)
bedtools getfasta -fi genome.fa -bed peaks.stringent.bed -s -fo peaks.fa

# Step 2: GC-matched background (random regions from expressed transcripts)
# expressed.bed = transcripts with TPM >= 1 in the same cell type
shuffleBed -i peaks.stringent.bed -g chrom.sizes -incl expressed.bed -seed 42 > shuffled.bed
bedtools getfasta -fi genome.fa -bed shuffled.bed -s -fo background.fa

# Step 3: HOMER de novo + known
findMotifs.pl peaks.fa fasta homer_out \
    -rna -len 5,6,7,8 -p 8 \
    -fasta background.fa

# Step 4: STREME for cross-validation
streme --rna --oc streme_out -p peaks.fa -n background.fa --minw 5 --maxw 10

# Step 5: mCross requires single-nt CL sites (PureCLIP output)
# crosslinks.bed = PureCLIP -o sites.bed (single-nt CL positions, see clip-seq/crosslink-site-detection)
mCross \
    -i crosslinks.bed -g genome.fa -k 7 -n 5 -o mcross_out

# Step 6: PEKA orthogonal positional k-mer
peka --peak_file_name peaks.stringent.bed --crosslinks_file_name crosslinks.bed --genome_file_name genome.fa --regions_file_name regions.bed --kmer_length 5 --percentile 30 --outpath peka_out

# Step 7: Compare to RBNS in vitro Kd (Dominguez 2018 supplementary)
# Manual cross-reference to RBNS Kd-ranked top-5 motifs for the target RBP

Quality Checks

python
import numpy as np
import pandas as pd
from Bio import motifs

def parse_homer_motif(motif_file):
    '''Parse HOMER motif PWM and return information content'''
    pwm = []
    with open(motif_file) as f:
        for line in f:
            if line.startswith('>'):
                continue
            pwm.append([float(x) for x in line.strip().split()])
    pwm = np.array(pwm)
    # Information content per position (bits)
    epsilon = 1e-10
    ic = np.sum(pwm * np.log2(pwm / 0.25 + epsilon), axis=1)
    return {
        'length': pwm.shape[0],
        'mean_IC_per_position': np.mean(ic),
        'max_IC': np.max(ic),
        'min_IC': np.min(ic),
        'is_low_complexity': np.mean(ic) < 0.5
    }

# A high-quality CLIP motif has mean IC per position > 1.0 bit
# Mean IC < 0.5 suggests no real motif - check for background mismatch
MetricGoodInvestigate
Mean IC per position (HOMER PWM)> 1.0 bit< 0.5 = no real motif or background mismatch
Motif center U contentmatches RBP biology> 80% U in central position = U-CL bias suspect
FIMO scan rate in peaks vs backgroundpeaks 3-10x bg< 2x = motif too weak; > 50x = motif too narrow
mCross CL-offset histogramsingle mode at offset Nflat = mCross unable to register
RBNS rank of CLIP top motiftop 3> 10 = strong in vivo / in vitro divergence; investigate
Replicate motif overlaptop motif identical in 2/2 repsdiscordance = under-powered or one rep failed

Common Errors

Error / symptomCauseSolution
HOMER top motif is "ATCC" or similar TruSeq-adapter-likeAdapter trim incomplete in preprocessingRe-run cutadapt with correct adapter; verify trimming completed before peak calling
mCross "no crosslinks found" or output identical to peak-centered logoPassed peak BED instead of single-nt crosslink-site BEDmCross requires PureCLIP or CTK CITS single-nt site BED; not a CLIPper peak BED. See clip-seq/crosslink-site-detection
mCross CL-offset histogram is flatStrand information missing OR alignment soft-clipped 5' baseVerify BED column 6 (strand); confirm upstream --alignEndsType EndToEnd
All HOMER motifs are AU-rich, no specificityDefault shuffled background AU-rich for 3' UTR peaksProvide GC-matched background from expressed 3' UTRs
mCross input "no crosslinks found"Provided peak BED instead of single-nt CL BEDPass PureCLIP / CTK CITS single-nt BED instead
MEME "out of memory" or > 12h runtimePeak set too largeSwitch to STREME; or down-sample to top-1000 peaks
FIMO match in nearly every peakThreshold too lenientTighten to 1e-5 or 1e-6
Motif width > 10 nt looks chimericHOMER reported overlapping co-motifs as oneRe-run with -len 5,6,7,8 (force narrower); inspect rank 2-5 motifs
PWM info content < 0.5 bitsBackground mismatch or peaks have no specificityVerify peak set is stringent (log2 FC >= 3); regenerate background
Strand-specific motif different on +/-BED column 6 missing or wrongVerify strand encoding; re-extract with bedtools getfasta -s
Multiple distinct motifs per RBPBiological (multimodal RBPs) OR contaminating multi-RBP CLIPValidate with mCross multimode + compare to literature

References

  • Heinz S et al 2010 Mol Cell 38:576 (HOMER motif discovery)
  • Bailey TL et al 2015 Nucleic Acids Res 43:W39 (MEME Suite)
  • Bailey TL 2021 Bioinformatics 37:2834 (STREME, MEME's fast successor)
  • Wu X & Bartel DP 2017 Nucleic Acids Res 45:W534 (kpLogo positional logo)
  • Feng H et al 2019 Mol Cell 74:1189 (mCross, jointly modeling motif + CL position)
  • Kuret K, Amalietti AG, Jones DM, Capitanchik C, Ule J 2022 Genome Biol 23:191 (PEKA, positional k-mer no input)
  • Jens M et al 2022 bioRxiv 2022.11.08.515616 (RBPamp, affinity-weighted motif; preprint)
  • Lambert N et al 2014 Mol Cell 54:887 (RNA Bind-n-Seq)
  • Dominguez D et al 2018 Mol Cell 70:854 (78-RBP RBNS atlas)
  • Sugimoto Y et al 2012 Genome Biol 13:R67 (CLIP/iCLIP analysis; uridine crosslink preference)
  • Van Nostrand EL et al 2020 Nature 583:711 (ENCODE 150 RBP eCLIP + motifs)
  • Konig J et al 2010 Nat Struct Mol Biol 17:909 (iCLIP, U bias origin)
  • clip-seq/clip-peak-calling - Upstream peak calls feed motif discovery
  • clip-seq/crosslink-site-detection - Single-nt CL sites required for mCross / PEKA
  • clip-seq/binding-site-annotation - Region-level context for motifs
  • clip-seq/clip-deep-learning - Sequence-to-binding deep models (RBPNet, DeepRiPe)
  • clip-seq/ago-clip-mirna-targets - miRNA seed motifs from CLEAR-CLIP chimeras
  • chip-seq/motif-analysis - DNA-protein motif analogue
  • pathway-analysis/go-enrichment - Functional context of motif-bearing genes

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clip-seq/clip-motif-analysis of GPTomics/bioSkills.

  • SKILL.md
  • examples/find_motifs.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clip Seq Clip Motif Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clip Seq Clip Motif Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clip Seq Clip Motif Analysis this skillGPTomics/bioSkills1.2k2 repos~5.5kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Clip Seq Clip Motif Analysis

What does Bio Clip Seq Clip Motif Analysis do?

Discover RBP binding motifs from CLIP-seq peaks or single-nucleotide crosslink sites using HOMER, MEME/STREME, kpLogo, mCross (CL-position-registered motifs), PEKA (positional k-mer enrichment)…. Bio Clip Seq Clip Motif Analysis is an agent skill from GPTomics/bioSkills. Discover RBP binding motifs from CLIP-seq peaks or single-nucleotide crosslink sites using HOMER, MEME/STREME, kpLogo, mCross (CL-position-registered motifs), PEKA (positional k-mer enrichment), RBPamp (affinity), and RNA Bind-n-Seq (RBNS) cross-validation.

When should I use Bio Clip Seq Clip Motif Analysis?

Bio Clip Seq Clip Motif Analysis fits situations like: characterizing RBP sequence specificity; registering motifs to crosslink positions; validating in vivo CLIP motifs against in vitro RBNS Kd; reconciling motif disagreements across tools.

How do I install Bio Clip Seq Clip Motif Analysis in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-clip-motif-analysis -a claude-code`. Or copy the skill folder (clip-seq/clip-motif-analysis in GPTomics/bioSkills) into .claude/skills/bio-clip-seq-clip-motif-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clip Seq Clip Motif Analysis in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-clip-motif-analysis -a codex`. Or copy the skill folder (clip-seq/clip-motif-analysis in GPTomics/bioSkills) into .agents/skills/bio-clip-seq-clip-motif-analysis in your project. Codex loads it when a task matches its description.

Can I use Bio Clip Seq Clip Motif Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-clip-motif-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clip-seq-clip-motif-analysis, .gemini/skills/bio-clip-seq-clip-motif-analysis, .github/skills/bio-clip-seq-clip-motif-analysis and .opencode/skills/bio-clip-seq-clip-motif-analysis in your project.

What does Bio Clip Seq Clip Motif Analysis need to run?

Going by SKILL.md and its folder, Bio Clip Seq Clip Motif Analysis needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.

Does Bio Clip Seq Clip Motif Analysis access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Clip Seq Clip Motif Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clip Seq Clip Motif Analysis use?

Bio Clip Seq Clip Motif Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clip Seq Clip Motif Analysis use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clip Seq Clip Motif Analysis?

Skills that share tags, products or a category with Bio Clip Seq Clip Motif Analysis: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clip Seq Clip Motif Analysis?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.