Agent skill

Bio Clip Seq Crosslink Site Detection

by GPTomics in GPTomics/bioSkills

Detect single-nucleotide crosslink (CL) sites in CLIP-seq data using truncation patterns (iCLIP/eCLIP CITS), crosslink-induced mutations (HITS-CLIP CIMS deletions, PAR-CLIP T-to-C), or…

MITAuto-check passedResearch & Science

Install Bio Clip Seq Crosslink Site Detection

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clip-seq-crosslink-site-detection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clip-seq-crosslink-site-detection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clip-seq/crosslink-site-detection .claude/skills/bio-clip-seq-crosslink-site-detection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clip-seq-crosslink-site-detection
GitHub stars
1.2k
Used in
2 other repos
Token cost
~5k tokens
SKILL.md length
2,444 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Detect single-nucleotide crosslink (CL) sites in CLIP-seq data using truncation patterns (iCLIP/eCLIP CITS), crosslink-induced mutations (HITS-CLIP CIMS deletions, PAR-CLIP T-to-C), or…

  • Single-nucleotide resolution is required for motif registration (mCross)
  • SKILL.md covers Version Compatibility, Crosslink Chemistry by CLIP…, Algorithmic Taxonomy and Critical Choice: Truncation vs…, plus 7 more sections
  • Runs Shell scripts from its folder; calls pip
  • Allele-specific binding (BEAPR)

What it does

Bio Clip Seq Crosslink Site Detection is an agent skill from GPTomics/bioSkills. Detect single-nucleotide crosslink (CL) sites in CLIP-seq data using truncation patterns (iCLIP/eCLIP CITS), crosslink-induced mutations (HITS-CLIP CIMS deletions, PAR-CLIP T-to-C), or HMM/kernel-density methods (PureCLIP, PARalyzer, CTK). Use when single-nucleotide resolution is required for motif registration (mCross), allele-specific binding (BEAPR), variant-effect prediction, or comparing crosslink chemistry across CLIP variants.

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/detect_crosslinks.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Single-nucleotide resolution is required for motif registration (mCross)
  • Allele-specific binding (BEAPR)
  • Variant-effect prediction
  • Comparing crosslink chemistry across CLIP variants

Example prompts

  • “/bio-clip-seq-crosslink-site-detection”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clip Seq Crosslink Site Detection loads about 5k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 2,444 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,444 words, ~4,982 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clip-seq-crosslink-site-detection/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clip-seq-crosslink-site-detection
description
Detect single-nucleotide crosslink (CL) sites in CLIP-seq data using truncation patterns (iCLIP/eCLIP CITS), crosslink-induced mutations (HITS-CLIP CIMS deletions, PAR-CLIP T-to-C), or HMM/kernel-density methods (PureCLIP, PARalyzer, CTK). Use when single-nucleotide resolution is required for motif registration (mCross), allele-specific binding (BEAPR), variant-effect prediction, or comparing crosslink chemistry across CLIP variants.
tool_type
cli
primary_tool
PureCLIP

Version Compatibility

Reference examples tested with: PureCLIP 1.3.1+, CTK 1.1.4+, PARalyzer 1.5+, wavClusteR 2.34+, pyCRAC 1.5+, samtools 1.19+, bedtools 2.31+, pysam 0.22+, R 4.3+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws unexpected errors, introspect the installed tool and adapt the example to match the actual CLI rather than retrying.

"Detect single-nucleotide crosslink sites in my CLIP data" -> Identify the exact base where the protein-RNA UV adduct caused the reverse transcriptase to stop (truncation, in iCLIP/eCLIP), to read through with a mutation (deletion in HITS-CLIP, T->C in PAR-CLIP), or to leave a multi-base signature (PARalyzer kernel density for PAR-CLIP). Single-nucleotide resolution is the foundation of motif registration (mCross), allele-specific binding (BEAPR/ASPRIN), variant-effect prediction, and the most rigorous comparisons across CLIP variants. The detection chemistry differs by CLIP type: iCLIP/eCLIP read 5' end is one nucleotide downstream of the crosslink (CITS); PAR-CLIP reads contain T->C transitions at the crosslink (CIMS substitution); HITS-CLIP reads contain deletions at the crosslink (CIMS deletion).

  • CLI (HMM, all CLIP variants): pureclip -i dedup.bam -bai dedup.bam.bai -g genome.fa -ibam sminput.bam -ibai sminput.bam.bai -o crosslinks.bed -or regions.bed -nt 8 -dm 8
  • CLI (CTK CITS truncations, iCLIP/eCLIP): parseAlignment.pl --map-qual 1 --min-len 18 --mutation-file mut.txt dedup.bam dedup.bed then tag2cluster.pl ... -cs5 5 -m 1 for truncation-cluster
  • CLI (CTK CIMS deletions, HITS-CLIP): getMutationType.pl dedup.bed mut.txt -type del then CIMS.pl dedup.bed mut.txt -big -c -p 0.01 cims.bed
  • CLI (CTK CIMS T->C, PAR-CLIP): getMutationType.pl dedup.bed mut.txt -type sub -nuc t -mut c then CIMS.pl dedup.bed t2c.mut -p 0.001 t2c_cims.bed
  • CLI (PAR-CLIP kernel density): PARalyzer params.ini (parameters file defines read length, min reads per cluster, mutation rate threshold)
  • CLI (PAR-CLIP wavClusteR R): wavClusteR::filterClusters(cl, snps=NULL, filterFC=FALSE) after wavelet clustering

Single-nt CL sites are the input to mCross motif registration (see clip-seq/clip-motif-analysis), to BEAPR/ASPRIN allele-specific binding analyses, and to RBPNet/DeepRiPe deep-learning models (see clip-seq/clip-deep-learning). They are NOT a replacement for broad peak calls; the two outputs are complementary. A common downstream error is passing a CLIPper peak BED to mCross instead of a PureCLIP single-nt BED - mCross requires the single-nt resolution to register motif position relative to the crosslink offset.

The detection method must match the underlying chemistry of how the reverse transcriptase encountered the protein-RNA adduct:

CLIP variantRT behavior at CLDetection signatureSensitivity (CL captured per read)Sequence bias
iCLIP / iCLIP2 / iCLIP3 / eCLIP / irCLIP / FLASHTruncates at adduct5' end of cDNA = CL site - 1~80% of reads truncate; the minority read throughStrong U bias at CL sites
HITS-CLIPReads through with deletion (~8-20% of tags)Single-base deletion in read~8-20% of tags contain a deletion at the CL siteU bias plus sequence-specific deletion rate
PAR-CLIP (4SU)Reads through with T->C (20-50% of T positions in crosslinked reads)T->C transition (or G->A on reverse strand)Per-T rate 20-50%; per-read 1-5 conversionsRestricted to T positions; depends on 4SU incorporation rate
PAR-CLIP (6SG)Reads through with G->AG->A transitionLower rate than 4SU T->CRestricted to G positions
miCLIP / miCLIP2 (m6A)Primarily truncation at the m6A site (RT stops at antibody-trapped m6A); secondary C->T transition observed in a subset of reads at m6A5' end (CL -1); C->T rate is a feature used by m6Aboost ML in addition to truncation, not a stand-alone primary signalMixed; m6Aboost ML integrates truncation + sequence context to scoreDRACH motif context required
STAMPC->T editing on target (not crosslink)C->T editing in mRNA readsNA (editing-based; not crosslink-based)N/A
TRIBEA->I editing on targetA->G in cDNA (I read as G)NAAdjacent to ADAR consensus

Algorithmic Taxonomy

ToolCLIP variantDetection signalStatistical modelOutput resolutionStrengthFails when
PureCLIP (Krakau 2017)iCLIP/eCLIP/PAR-CLIPTruncation + enrichment + CL motifNon-homogeneous HMMSingle-nucleotideThe most comprehensive HMM modelLowest F1 among callers benchmarked in Boyle 2023 on broad-binding RBPs (highly focal)
CTK CITS (Shah 2017)iCLIP/eCLIPTruncation onlyEmpirical FDR vs backgroundSingle-nucleotideSimple, well-validated for iCLIPLess granular than PureCLIP; no HMM smoothing
CTK CIMS deletion (Shah 2017)HITS-CLIPSingle-base deletionsEmpirical FDRSingle-nucleotideStandard for HITS-CLIPRequires deletion-tolerant aligner (BWA -e)
CTK CIMS T->C (Shah 2017)PAR-CLIPT->C substitutionsEmpirical FDRSingle-nucleotideAlternative to PARalyzerLess popular than PARalyzer
PARalyzer (Corcoran 2011)PAR-CLIPT->C transitions, kernel densityKernel density estimationCluster (10-50 nt)Field standard for PAR-CLIPCluster-level, not single-nt; needs careful parameter tuning
wavClusteR (Comoglio 2015)PAR-CLIPT->C with wavelet smoothingWavelet transform + clusteringClusterRobust to sequencing depth variabilityLess single-nt; legacy R package
pyCRAC / kPLogoCRAC (yeast)Read truncation + deletionEmpiricalSingle-nucleotideOriginal CLIP analytics toolYeast-focused; perl/python legacy
Piranha binsAnyCoverage in binsZTNBBin-level (50-200 nt)Not crosslink-specificBin width too coarse for single-nt
HOMER tag2posAny5' end positionNoneRead-end positions onlyQuick truncation site dumpNo statistical filtering
omniCLIPAnyCoverage + variant patternDirichlet-multinomial HMMRegionNot focalToo broad for single-nt
iCount (paths)iCLIPCrosslink site clusterEmpiricalClusterUsed in nf-core/clipseqLess single-nt than PureCLIP

Methodology evolves; PureCLIP 2.0 and CTK 1.2 have minor flag changes; PARalyzer parameters need careful tuning per-RBP. Cross-validate single-nt sites with at least two tools when high-confidence reporting is required.

Critical Choice: Truncation vs Mutation vs HMM

Three orthogonal approaches:

Truncation-based (CITS, PureCLIP) -- Find positions enriched in cDNA 5' ends. The RT enzyme stops at the adduct; the 5' end of the read maps to CL site - 1. Used for iCLIP/eCLIP. Pro: high yield (~80% of reads truncate). Con: U bias of crosslinking inflates U positions.

Mutation-based (CIMS for deletions/substitutions; PARalyzer for T->C) -- Find positions with crosslink-induced mutations. RT reads through the adduct with high mutation rate at the CL position. Used for HITS-CLIP (deletions, ~8-20%) and PAR-CLIP (T->C, 20-50%). Pro: less U bias; PAR-CLIP signal is restricted to T positions. Con: lower yield (only mutated reads count).

HMM-based (PureCLIP, omniCLIP) -- Joint model of enrichment + CL signature + sequence context. State the most-likely posterior probability of CL state per nucleotide. Used across CLIP variants. Pro: integrates multiple signals; explicit input normalization. Con: F1 ~0.2 on bulk RBPs (very focal); single-nt resolution at the cost of breadth.

GoalMethodTool
Single-nt CL map for motif registration (mCross)HMMPureCLIP
Truncation-based ENCODE-compatible iCLIPTruncationCTK CITS
PAR-CLIP T->C signalMutationPARalyzer (cluster) or CTK CIMS sub T->C (single-nt)
HITS-CLIP deletion signalMutationCTK CIMS deletion
Cross-CLIP-variant comparable single-ntHMMPureCLIP (all variants)
Allele-specific CLIP (BEAPR)Truncation + ASBPureCLIP CL sites + WASP-filtered BAM
Variant effect (computational from CL)HMMPureCLIP + downstream DeepRiPe (see clip-deep-learning)
Yeast / CRACTruncation + deletionpyCRAC

Per-Tool Failure Modes

PureCLIP -- HMM convergence on sparse coverage

Trigger: RBP with low IP enrichment; SMInput depth uneven; PureCLIP run on full genome without -iv interval.

Mechanism: PureCLIP HMM convergence is sensitive to coverage depth. On regions with < 5 reads/100 bp the HMM defaults to background state; if half the genome is sparse, the state distributions become bimodal and the entire run takes > 24 h or runs out of memory.

Symptom: PureCLIP "Convergence not reached" warning; or output has 0 CL sites; or runtime > 24 h on standard server.

Fix: Restrict analysis to expressed transcripts: -iv expressed_tx.bed. Or filter the BAM to high-coverage regions first. Test on chr22 first to verify parameters before genome-wide run.

PureCLIP -- Misses broad binding zones

Trigger: RBP with broad binding (PUM2 3' UTRs, SR proteins exonic enhancers); using PureCLIP exclusively.

Mechanism: PureCLIP HMM emits the CL state at very high stringency. The Boyle 2023 Skipper benchmark reports PureCLIP as the most focal caller with low recall on broad-binding RBPs.

Symptom: PureCLIP CL site count is 100x lower than CLIPper / Skipper peak count.

Fix: Use PureCLIP for single-nt CL maps AND CLIPper/Skipper for broad-zone peak calls. They are complementary, not substitutes. PureCLIP -or regions output is broader than -o sites but still focal.

CITS truncation -- Misapplied to PAR-CLIP

Trigger: CTK CITS run on PAR-CLIP data.

Mechanism: PAR-CLIP RT reads through the adduct (T->C mutation), not truncates. The 5' end of PAR-CLIP reads is at the fragment 5' end, not the CL site.

Symptom: CITS returns few sites; the sites that ARE returned do not align with the T->C mutation positions.

Fix: For PAR-CLIP use PARalyzer or CTK CIMS substitution mode (-type sub -nuc t -mut c). CITS is iCLIP/eCLIP only.

CIMS deletion -- Misapplied to iCLIP

Trigger: CTK CIMS run with -type del on iCLIP data.

Mechanism: iCLIP RT truncates, doesn't delete. The deletion rate is < 1% in iCLIP (background level), so CIMS finds nothing significant.

Symptom: CIMS output BED has < 100 sites genome-wide.

Fix: Use CTK CITS for iCLIP (truncation-based) or PureCLIP. CIMS deletion is HITS-CLIP / older PAR-CLIP only.

Show full SKILL.md (998 more words)Show less
PARalyzer -- Parameter sensitivity

Trigger: PARalyzer default parameters on a new RBP; cluster count unexpected.

Mechanism: PARalyzer's kernel-density approach has 5+ parameters: minimum read depth, minimum mutation read count, mutation rate threshold, bandwidth, kernel type. Different combinations produce 5-100x different cluster counts.

Symptom: PARalyzer cluster count 5-100x off from literature reports for the same RBP.

Fix: Use the PARalyzer default parameters (Corcoran 2011), tuned for that RBP class. Validate on a known-binding region with a reference dataset before publication.

Aligner deletion-tolerance for HITS-CLIP

Trigger: HITS-CLIP aligned with STAR or bowtie2 in standard mode; CIMS finds few deletion sites.

Mechanism: Standard STAR/bowtie2 are aggressive about handling deletions; CIMS expects the deletions to be visible as CIGAR D operations in BAM. STAR/bowtie2 may soft-clip or fail to align deletion-bearing reads.

Symptom: Few CIGAR D operations in BAM (samtools view dedup.bam | awk '$6 ~ /D/' | wc -l); CIMS deletion site count low.

Fix: Re-align HITS-CLIP with BWA-aln (Zhang lab / CTK convention) or STAR --scoreDelOpen -1 --scoreDelBase -1 --scoreInsOpen -1 --scoreInsBase -1 to permit deletions. Or use Novoalign which is more deletion-tolerant.

PAR-CLIP T->C mismatch ceiling

Trigger: PAR-CLIP aligned with STAR --outFilterMismatchNoverReadLmax 0.04; T->C reads discarded.

Mechanism: PAR-CLIP per-T conversion rate 20-50% means a 30 nt read with 8 Ts may have 4 T->C events = 13% mismatch rate, exceeding the 4% ceiling.

Symptom: 40-70% read loss at alignment for PAR-CLIP; downstream T->C site detection finds nothing.

Fix: Raise to 0.07 for PAR-CLIP only (see clip-seq/clip-alignment).

Strand mis-assignment for truncation site

Trigger: Single-end CLIP; using read 5' end as CL site without strand consideration.

Mechanism: For plus-strand RNA, the cDNA 5' end is downstream of the CL position. For minus-strand RNA, it is upstream. Without strand-aware coordinate conversion, half of CL sites are mis-positioned by ~1 nt.

Symptom: mCross / motif analysis shows the motif positioned 1 nt off from expected.

Fix: Verify the BED output handles strand correctly. CTK CITS, PureCLIP, and most modern tools do this automatically; custom scripts must add if strand == '-': cl_position = read_end_5p + 1.

Decision Tree by Use Case

ScenarioToolParameters
Single-nt CL map for mCross motif registrationPureCLIP-i dedup.bam -ibam sminput.bam -dm 8
iCLIP/eCLIP CITS truncation sitesCTK CITStag2cluster.pl -cs5 5 -m 1
HITS-CLIP deletion CIMSCTK CIMSCIMS.pl -big -c -p 0.01
PAR-CLIP T->C clustersPARalyzerPer-RBP params tuned from Corcoran 2011 defaults
PAR-CLIP T->C single-ntCTK CIMS substitutionCIMS.pl -type sub -nuc t -mut c -p 0.001
Yeast CRACpyCRACpyCRAC.py with HTP-tagged protein
Variant effect at CL sitesPureCLIP + DeepRiPeSee clip-seq/clip-deep-learning
Allele-specific bindingWASP-filtered BAM + PureCLIPSee clip-seq/clip-alignment for WASP
Cross-CLIP comparisonPureCLIP (consistent across variants)Same parameters
m6A miCLIP2 sitesmiCLIP2-specific (m6Aboost)See clip-seq/m6a-clip
STAMP edit sitesNot crosslink-basedSee clip-seq/stamp-antibody-free

Reconciliation: When Detection Methods Disagree

PatternLikely causeAction
PureCLIP << CTK CITS site countPureCLIP HMM very focal; CITS empirical broaderBoth correct; report which threshold
PARalyzer clusters << CIMS T->C sites for PAR-CLIPPARalyzer is cluster-level; CIMS is single-ntAggregate CIMS to clusters for comparison
PureCLIP CL sites at unexpected positionsStrand mis-assignment OR aligner soft-clipCheck BED strand column; verify --alignEndsType EndToEnd
HITS-CLIP CIMS emptyAligner not deletion-tolerantRe-align with BWA-aln or STAR with adjusted scoring
PAR-CLIP CIMS << expectedReads lost at mismatch ceilingRaise STAR --outFilterMismatchNoverReadLmax to 0.07
eCLIP CITS positions shifted by 1 ntStrand handling off-by-oneVerify R2 5' end position handling
Multiple CITS sites within 5 ntRT stops near each other; same CL eventCluster within 10 nt window post-detection
Motif registers 5 nt off in mCrossCL positions wrong by RT-stop offsetConfirm tool reports "CL - 1" position consistently

Operational rule: For motif registration and ASB, run PureCLIP with SMInput. For ENCODE-comparable truncation-based output, also run CTK CITS. Cross-validate by checking that mCross motif PWM is the same when fed PureCLIP sites vs CTK CITS sites. If they diverge by > 2 nt in motif position, suspect strand or off-by-one issue.

Workflow: PureCLIP -> mCross -> ASB

bash
# Step 1: PureCLIP single-nt sites
pureclip \
    -i sample.dedup.bam -bai sample.dedup.bam.bai \
    -g genome.fa \
    -ibam sminput.dedup.bam -ibai sminput.dedup.bam.bai \
    -o sample.crosslinks.bed \
    -or sample.regions.bed \
    -nt 8 -dm 8 -iv expressed_tx.bed

# Step 2: mCross motif registration
mCross -i sample.crosslinks.bed -g genome.fa -k 7 -n 5 -o mcross_out

# Step 3: Intersect with heterozygous SNPs for allele-specific binding
bedtools intersect -wa -wb -s -a sample.crosslinks.bed -b het_snps.vcf > cl_at_hets.bed

# Step 4: Test allele bias with BEAPR or ASPRIN
# (see allele-specific CLIP literature; not in this skill)

Common Errors

Error / symptomCauseSolution
PureCLIP "Convergence not reached"HMM didn't converge on sparse coverageRestrict to expressed transcripts with -iv expressed.bed
PureCLIP > 24h runtimeFull genome on large libraryTest on chr22 first; add -iv filter
CTK CIMS empty for HITS-CLIPAligner not deletion-tolerantRe-align with BWA-aln
CITS truncation sites not single-ntCIGAR S operations from soft-clipVerify --alignEndsType EndToEnd upstream
PAR-CLIP CIMS emptyReads lost at mismatch ceilingRaise STAR mismatch tolerance
Off-by-one motif positionStrand handling differsVerify tool documents whether output is CL or CL-1
Single-nt motif lost in HOMER on PureCLIP sitesSite BED too narrow; need flankingExtend bedtools slop -b 15 before motif analysis
mCross "no crosslinks found"Passed peak BED instead of CL BEDUse single-nt site output
PureCLIP output looks like uniform distributionR2 5' end trimmed in preprocessingRe-preprocess; -g and --trim_front2 banned for CLIP
Strand-specific motif invertedBED column 6 wrongVerify strand encoding; PureCLIP preserves correctly

References

  • Konig J et al 2010 Nat Struct Mol Biol 17:909 (iCLIP truncation principle)
  • Hafner M et al 2010 Cell 141:129 (PAR-CLIP T->C signature)
  • Granneman S et al 2009 PNAS 106:9613 (CRAC, single-nt CL detection)
  • Corcoran DL et al 2011 Genome Biol 12:R79 (PARalyzer)
  • Comoglio F et al 2015 BMC Bioinformatics 16:32 (wavClusteR)
  • Shah A et al 2017 Bioinformatics 33:566 (CTK / CIMS / CITS)
  • Krakau S et al 2017 Genome Biol 18:240 (PureCLIP HMM)
  • Boyle EA et al 2023 Cell Genomics 3:100317 (Skipper; PureCLIP focality/F1 benchmark)
  • Sugimoto Y et al 2012 Genome Biol 13:R67 (CLIP/iCLIP analysis; uridine crosslink preference)
  • Feng H et al 2019 Mol Cell 74:1189 (mCross requires CL sites)
  • Yang EW et al 2019 Nat Commun 10:1338 (BEAPR allele-specific protein-RNA binding)
  • Van Nostrand EL et al 2020 Nature 583:711 (ENCODE eCLIP single-nt analyses)
  • clip-seq/clip-preprocessing - 5' base preservation critical for truncation detection
  • clip-seq/clip-alignment - End-to-end alignment for crosslink-preserving BAM
  • clip-seq/clip-peak-calling - Peak vs CL site is a complementary distinction
  • clip-seq/clip-motif-analysis - mCross consumes CL sites
  • clip-seq/clip-deep-learning - RBPNet trained on CL count distributions
  • clip-seq/m6a-clip - miCLIP2 has its own CL detection
  • clip-seq/stamp-antibody-free - STAMP uses C->U editing, not crosslink
  • alignment-files/sam-bam-basics - CIGAR string semantics for D operations

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clip-seq/crosslink-site-detection of GPTomics/bioSkills.

  • SKILL.md
  • examples/detect_crosslinks.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clip Seq Crosslink Site Detection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clip Seq Crosslink Site Detection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clip Seq Crosslink Site Detection this skillGPTomics/bioSkills1.2k2 repos~5kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Clip Seq Crosslink Site Detection

What does Bio Clip Seq Crosslink Site Detection do?

Detect single-nucleotide crosslink (CL) sites in CLIP-seq data using truncation patterns (iCLIP/eCLIP CITS), crosslink-induced mutations (HITS-CLIP CIMS deletions, PAR-CLIP T-to-C), or…. Bio Clip Seq Crosslink Site Detection is an agent skill from GPTomics/bioSkills. Detect single-nucleotide crosslink (CL) sites in CLIP-seq data using truncation patterns (iCLIP/eCLIP CITS), crosslink-induced mutations (HITS-CLIP CIMS deletions, PAR-CLIP T-to-C), or HMM/kernel-density methods (PureCLIP, PARalyzer, CTK).

When should I use Bio Clip Seq Crosslink Site Detection?

Bio Clip Seq Crosslink Site Detection fits situations like: single-nucleotide resolution is required for motif registration (mCross); allele-specific binding (BEAPR); variant-effect prediction; comparing crosslink chemistry across CLIP variants.

How do I install Bio Clip Seq Crosslink Site Detection in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-crosslink-site-detection -a claude-code`. Or copy the skill folder (clip-seq/crosslink-site-detection in GPTomics/bioSkills) into .claude/skills/bio-clip-seq-crosslink-site-detection in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clip Seq Crosslink Site Detection in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-crosslink-site-detection -a codex`. Or copy the skill folder (clip-seq/crosslink-site-detection in GPTomics/bioSkills) into .agents/skills/bio-clip-seq-crosslink-site-detection in your project. Codex loads it when a task matches its description.

Can I use Bio Clip Seq Crosslink Site Detection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clip-seq-crosslink-site-detection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clip-seq-crosslink-site-detection, .gemini/skills/bio-clip-seq-crosslink-site-detection, .github/skills/bio-clip-seq-crosslink-site-detection and .opencode/skills/bio-clip-seq-crosslink-site-detection in your project.

What does Bio Clip Seq Crosslink Site Detection need to run?

Going by SKILL.md and its folder, Bio Clip Seq Crosslink Site Detection needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Clip Seq Crosslink Site Detection access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Clip Seq Crosslink Site Detection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clip Seq Crosslink Site Detection use?

Bio Clip Seq Crosslink Site Detection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clip Seq Crosslink Site Detection use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clip Seq Crosslink Site Detection?

Skills that share tags, products or a category with Bio Clip Seq Crosslink Site Detection: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clip Seq Crosslink Site Detection?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.