Workflow Orchestration
AnastasiyaW/codex-claude-code-config
Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).
Aligns and QCs methylated-RNA-immunoprecipitation (MeRIP / m6A-seq) IP and input libraries using STAR or HISAT2 splice-aware mapping, samtools sort/index, IP/input matched-pair tracking…
$ npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-epitranscriptomics-merip-preprocessing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/epitranscriptomics/merip-preprocessing .claude/skills/bio-epitranscriptomics-merip-preprocessing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-epitranscriptomics-merip-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/epitranscriptomics/merip-preprocessing into .claude/skills/bio-epitranscriptomics-merip-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-epitranscriptomics-merip-preprocessing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/epitranscriptomics/merip-preprocessingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-epitranscriptomics-merip-preprocessing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/epitranscriptomics/merip-preprocessing .agents/skills/bio-epitranscriptomics-merip-preprocessing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-epitranscriptomics-merip-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/epitranscriptomics/merip-preprocessing into .agents/skills/bio-epitranscriptomics-merip-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-epitranscriptomics-merip-preprocessing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-epitranscriptomics-merip-preprocessing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/epitranscriptomics/merip-preprocessing .cursor/skills/bio-epitranscriptomics-merip-preprocessing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-epitranscriptomics-merip-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/epitranscriptomics/merip-preprocessing into .cursor/skills/bio-epitranscriptomics-merip-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-epitranscriptomics-merip-preprocessing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path epitranscriptomics/merip-preprocessing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-epitranscriptomics-merip-preprocessing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/epitranscriptomics/merip-preprocessing .gemini/skills/bio-epitranscriptomics-merip-preprocessing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-epitranscriptomics-merip-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/epitranscriptomics/merip-preprocessing into .gemini/skills/bio-epitranscriptomics-merip-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-epitranscriptomics-merip-preprocessing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-epitranscriptomics-merip-preprocessingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/epitranscriptomics/merip-preprocessing .github/skills/bio-epitranscriptomics-merip-preprocessing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-epitranscriptomics-merip-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/epitranscriptomics/merip-preprocessing into .github/skills/bio-epitranscriptomics-merip-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-epitranscriptomics-merip-preprocessing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-epitranscriptomics-merip-preprocessing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/epitranscriptomics/merip-preprocessing .opencode/skills/bio-epitranscriptomics-merip-preprocessing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-epitranscriptomics-merip-preprocessing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/epitranscriptomics/merip-preprocessing into .opencode/skills/bio-epitranscriptomics-merip-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-epitranscriptomics-merip-preprocessing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-epitranscriptomics-merip-preprocessingAligns and QCs methylated-RNA-immunoprecipitation (MeRIP / m6A-seq) IP and input libraries using STAR or HISAT2 splice-aware mapping, samtools sort/index, IP/input matched-pair tracking…
Bio Epitranscriptomics Merip Preprocessing is an agent skill from GPTomics/bioSkills. Aligns and QCs methylated-RNA-immunoprecipitation (MeRIP / m6A-seq) IP and input libraries using STAR or HISAT2 splice-aware mapping, samtools sort/index, IP/input matched-pair tracking, antibody-lot metadata recording, replicate concordance via deepTools multiBamSummary + plotCorrelation, IP enrichment QC via plotFingerprint and per-transcript IP/input ratio distributions, library-complexity saturation curves via PreSeq, and the explicit do-NOT-deduplicate convention for standard non-UMI MeRIP. Use when…
Its SKILL.md is about 8.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/align_merip.sh`, `examples/merip_qc.py` and `usage-guide.md`).
It sits in Business, Finance & HR, covering Bioinformatics and Accounting and bookkeeping. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell and Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Epitranscriptomics Merip Preprocessing loads about 8.5k tokens when it runs. Until then it costs about 264 tokens; SKILL.md has 3,728 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,728 words, ~8,490 tokens.
.claude/skills/bio-epitranscriptomics-merip-preprocessing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: STAR 2.7.11+, HISAT2 2.2.1+, samtools 1.19+, deepTools 3.5+, PreSeq 3.2+, fastp 0.23+, Trim Galore 0.6.10+, Picard 3.1+, MultiQC 1.25+, bowtie2 2.5+, BWA-MEM2 2.2+.
Before using code patterns, verify installed versions match. If versions differ:
<tool> --version then <tool> --help to confirm flagspip show <package> then help(module.function)If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
STAR --outSAMtype accepts BAM SortedByCoordinate since 2.5.x; check the Log.final.out file for input/output statistics. deepTools bamCompare --operation log2 is the modern syntax (older --ratio log2 still works but is being phased out). PreSeq c_curve and lc_extrap have stable interfaces; HISAT2 reports unique vs multi-mapped in the summary log.
"Get my MeRIP IP and input libraries ready for peak calling" -> Trim adapters with MeRIP-appropriate defaults (do NOT trim UMIs unless the library is UMI-MeRIP — most are not), splice-aware-align IP and input to the GENOME (not transcriptome) with STAR / HISAT2, sort and index, evaluate replicate concordance and IP enrichment with deepTools, build a saturation curve per library so peak counts can be honestly compared across libraries, record antibody clone and lot metadata so cross-batch comparison is later auditable, and produce IP-over-Input log2 bigWig tracks for downstream visualisation. Crucially, do NOT deduplicate non-UMI MeRIP — see the failure-modes section.
STAR --runMode alignReads -- splice-aware genome alignment, the field defaulthisat2 -x index -1 R1.fq.gz -2 R2.fq.gz -- graph-based alternative; lighter memory footprintsamtools sort -@ 8 -o sorted.bam in.bam && samtools index sorted.bam -- post-alignment mechanicsdeeptools multiBamSummary bins -b *.bam -o cov.npz + plotCorrelation -- replicate Spearmandeeptools plotFingerprint -b IP.bam Input.bam -o fp.pdf -- IP enrichment QC (ChIP-seq term; transfers cleanly)preseq lc_extrap -B -o curve.txt sorted.bam -- library complexity / saturationdeeptools bamCompare -b1 IP.bam -b2 Input.bam --operation log2 -o IP_over_Input.bw -- downstream-ready coverage trackA MeRIP library sequenced to 20 million unique reads finds substantially fewer peaks than the same biology at 60 million reads. Per-sample peak counts reported without saturation curves (PreSeq c_curve / lc_extrap; Daley & Smith 2013 Nat Methods 10:325) are uninterpretable across studies and often across replicates within a study. Subsample BAMs to a common unique-read depth before peak calling for any cross-condition peak-count comparison, OR report peaks alongside the saturation curve. A corollary: do NOT deduplicate standard MeRIP — the typical MeRIP protocol (Synaptic Systems 202-003 / Abcam ab151230 / NEB EpiMark E1610 antibody pull-down on fragmented poly(A)-selected RNA) has NO unique molecular identifiers, and picard MarkDuplicates on such libraries collapses real biological replicates of high-coverage transcripts (the opposite of what dedup achieves in DNA ChIP-seq). UMI-MeRIP is the only exception — and most MeRIP libraries in print are NOT UMI. McIntyre et al. 2020 Sci Rep 10:6590 demonstrated that replicate-to-replicate peak overlap is ~80% within a single lab but drops to a median 45% between labs using nominally identical conditions; this irreducible technical noise constrains how strongly any single MeRIP study can support biological claims, and the preprocessing pipeline is where the variance is set.
| Tool / step | Mechanism | Output | Strength | Fails when |
|---|---|---|---|---|
| STAR 2.7+ (Dobin 2013 Bioinformatics 29:15) | Two-pass splice-aware alignment with on-the-fly splice junction database | Sorted BAM + splice-junction TSV | Field default; multi-mapper retention configurable; STAR splice-junction DB | Memory-heavy (~30 GB human); slower than HISAT2 |
| HISAT2 2.2.1+ (Kim 2019 Nat Biotechnol 37:907) | Hierarchical graph FM-index; splice-aware | Sorted BAM | ~5x lighter memory than STAR; comparable accuracy | Less mature splice-junction handling for novel introns |
| BWA-MEM2 (Vasimuddin 2019 IPDPS 314) | DNA-style local alignment; NO splice awareness | Sorted BAM | Use ONLY for transcriptome-aligned MeRIP (rare) | Splits reads across exon junctions if used on genome |
| fastp 0.23+ (Chen 2018 Bioinformatics 34:i884) | Streaming adapter detection + quality trim | Trimmed FASTQ + JSON QC | Fast; JSON-readable QC output | UMI handling disabled by default; do NOT pass --umi for standard non-UMI MeRIP (the opposite of the failure direction in some other library types) |
| Trim Galore | Wrapper over cutadapt with paired-end auto-detect | Trimmed FASTQ | Conservative defaults; widely cited | Slower than fastp on large datasets |
| samtools sort / index | BAM coordinate sort + .bai index | Sorted BAM + index | Standard | None at default |
| Picard MarkDuplicates | Identifies PCR duplicates by 5' alignment start | Marked / removed BAM | Standard in DNA / ChIP | The dominant MeRIP convention is to SKIP dedup for non-UMI libraries (collapses real biology at high-coverage transcripts); a minority of pipelines dedup MeRIP — record the choice in metadata |
| deepTools multiBamSummary + plotCorrelation | Per-bin read counts; Spearman / Pearson matrix | Heatmap + clustering | Standard replicate-concordance plot | Bin size sensitive (use 10 kb for transcriptome-genome) |
| deepTools plotFingerprint (Diaz 2012 Stat Appl Genet Mol Biol 11:9) | Cumulative read-fraction vs cumulative-bin-fraction Lorenz curve | PDF + raw counts | Direct IP-vs-input enrichment QC; "good" IP has steep tail | A flat fingerprint = failed IP (or mock IgG) |
| deepTools bamCompare --operation log2 | Per-bin log2 (IP/input) bigWig | bigWig | Ready for downstream visualisation | Pseudocount choice matters at low-coverage bins |
| PreSeq c_curve / lc_extrap (Daley & Smith 2013 Nat Methods 10:325) | Capture-recapture; rational-function extrapolation | Curve TSV | The only honest library-complexity estimate | Requires uniquely-mapped reads to be reliable |
| MultiQC (Ewels 2016 Bioinformatics 32:3047) | Aggregator across tools | HTML report | Consolidates STAR + HISAT2 + samtools + deepTools + PreSeq into one report | None |
| Scenario | Recommended | Why wrong choices fail |
|---|---|---|
| Standard mammalian MeRIP, downstream exomePeak2 | STAR splice-aware -> genome BAM; do NOT deduplicate; build saturation curve | Transcriptome alignment breaks exomePeak2 (expects genome BAM + GTF); dedup collapses biology |
| Downstream m6Anet (ONT direct RNA) | Defer to m6anet-analysis -- alignment is to TRANSCRIPTOME with minimap2 -ax map-ont -uf -k14 | Genome-aligned ONT input breaks m6Anet entirely (signal-level dataprep requires per-transcript coordinates) |
| Limited memory (<16 GB) | HISAT2 instead of STAR | STAR human genome index requires ~30 GB |
| UMI-MeRIP (rare) | Trim UMI to read header (umi_tools / fastp --umi_loc), align, then dedup with umi_tools dedup | Standard picard MarkDuplicates ignores UMI; effective dedup rate wrong |
| Cross-batch comparison (different antibody lots) | Record antibody clone + lot in sample-sheet metadata; include batch factor in downstream design | Pooling cross-batch counts without batch term inflates false-positive differential peaks |
| Spike-in normalisation (NEB EpiMark control oligos) | Align separately to spike-in reference; report IP-spike-in / Input-spike-in ratio per sample; use for absolute normalisation | Most users discard the NEB EpiMark Gluc / Cluc controls; they are the per-sample IP-efficiency QC anchor |
| Cross-library peak-count comparison | Subsample BAMs to common unique-read depth with samtools view -s BEFORE downstream peak calling, OR fit saturation curves and compare at common depth | Raw peak counts are sequencing-depth-dependent and not biologically interpretable |
| Suspect failed IP | Inspect deepTools plotFingerprint AND per-transcript IP/input ratio distribution; failed IP shows shallow Lorenz tail AND median IP/input ~1.0 | Trusting raw peak count alone — failed IPs still produce peaks |
| Viral / contamination-suspect samples | Build a combined host + viral index (and rRNA index) and check unmapped reads | Single-organism indexes hide systematic contamination |
| Aligning to transcriptome (rare; specific downstream tools) | BWA-MEM2 or bowtie2; defer to read-alignment/ | STAR splice-aware on transcriptome causes spurious splice calls inside transcripts |
Methodology evolves; before any high-stakes preprocessing pipeline, web-search "STAR vs HISAT2 MeRIP 2024" and "MeRIP saturation curve preseq" for current consensus parameters.
Goal: Remove sequencing adapters and low-quality 3' ends WITHOUT removing biological signal (UMI-MeRIP must keep UMI sequences; standard MeRIP does not have UMIs and trimming should be minimal).
Approach: Use fastp or Trim Galore with adapter auto-detection; require minimum read length 25-30 nt (shorter reads multi-map and confound exomePeak2); for standard non-UMI MeRIP, do NOT pass --umi flags; preserve random-hexamer-priming artifacts ONLY if downstream pipeline expects them (most do not).
mkdir -p trimmed
for sample in IP_rep1 IP_rep2 IP_rep3 Input_rep1 Input_rep2 Input_rep3; do
fastp \
--in1 raw/${sample}_R1.fastq.gz \
--in2 raw/${sample}_R2.fastq.gz \
--out1 trimmed/${sample}_R1.fq.gz \
--out2 trimmed/${sample}_R2.fq.gz \
--html qc/${sample}_fastp.html \
--json qc/${sample}_fastp.json \
--length_required 25 \
--detect_adapter_for_pe \
--thread 8
doneFor UMI-MeRIP (rare), insert --umi --umi_loc read1 --umi_len 8 BEFORE the alignment step. Default fastp output preserves base quality information needed by downstream variant-aware tools; do NOT pass --disable_quality_filtering for MeRIP libraries.
Goal: Produce coordinate-sorted, indexed GENOME BAM files for each IP and input library with splice-junction-aware mapping, retaining a moderate number of multi-mappers for accurate per-window read counts at multi-isoform loci.
Approach: Build STAR genome index once with the matched GENCODE / Ensembl GTF used downstream; loop IP and input samples with identical parameters; retain up to 20 multi-mappers per read (MeRIP read counts at multi-isoform genes need this); request explicit BAM SortedByCoordinate; emit splice-junction tables for QC.
STAR \
--runMode genomeGenerate \
--genomeDir star_index \
--genomeFastaFiles genome.fa \
--sjdbGTFfile annotation.gtf \
--sjdbOverhang 100 \
--runThreadN 12
mkdir -p aligned
for sample in IP_rep1 IP_rep2 IP_rep3 Input_rep1 Input_rep2 Input_rep3; do
STAR \
--runMode alignReads \
--genomeDir star_index \
--readFilesIn trimmed/${sample}_R1.fq.gz trimmed/${sample}_R2.fq.gz \
--readFilesCommand zcat \
--outSAMtype BAM SortedByCoordinate \
--outFilterMultimapNmax 20 \
--outSAMattributes NH HI AS nM NM MD \
--outFileNamePrefix aligned/${sample}_ \
--runThreadN 12
samtools index -@ 4 aligned/${sample}_Aligned.sortedByCoord.out.bam
done--outFilterMultimapNmax 20 is intentional: MeRIP at rRNA / snoRNA / pseudogene-rich loci needs multi-mapper retention. Reduce to 1 only if downstream analysis explicitly cannot tolerate multi-mappers. --sjdbOverhang should equal (read length - 1) but 100 is the common-enough default for 100-150 bp reads.
Goal: Achieve splice-aware alignment in ~5x less memory than STAR (12-16 GB suffices for human), with comparable accuracy for MeRIP applications.
Approach: Build HISAT2 graph index; align with --dta for downstream-transcript-assembly compatibility; pipe directly to samtools sort.
hisat2-build genome.fa hisat2_index/genome
for sample in IP_rep1 IP_rep2 IP_rep3 Input_rep1 Input_rep2 Input_rep3; do
hisat2 \
-x hisat2_index/genome \
-1 trimmed/${sample}_R1.fq.gz \
-2 trimmed/${sample}_R2.fq.gz \
--dta \
--summary-file qc/${sample}_hisat2.log \
-p 12 | \
samtools sort -@ 8 -o aligned/${sample}.sorted.bam -
samtools index -@ 4 aligned/${sample}.sorted.bam
doneHISAT2 multi-mapper handling is governed by -k; the default reports the primary alignment only. For MeRIP, pass -k 5 if multi-mapper-aware downstream counting is required.
mkdir -p qc
for bam in aligned/*sortedByCoord.out.bam aligned/*sorted.bam; do
name=$(basename ${bam} .bam)
samtools flagstat ${bam} > qc/${name}.flagstat
samtools idxstats ${bam} > qc/${name}.idxstats
doneInspect flagstat for properly-paired rate (>=85% indicates good pairing); inspect idxstats for unexpected chromosome-level read piles (rRNA bleed-through, mitochondrial domination — both are MeRIP red flags).
Goal: Quantify how similar replicate IP libraries are to each other (and likewise input libraries) using a Spearman correlation matrix; flag a divergent replicate before it propagates into peak calling.
Approach: Compute genome-wide per-bin read counts at 10 kb resolution across all IP and input BAMs; convert to a clustered Spearman heatmap with deepTools plotCorrelation.
multiBamSummary bins \
--bamfiles aligned/IP_rep1*.bam aligned/IP_rep2*.bam aligned/IP_rep3*.bam \
aligned/Input_rep1*.bam aligned/Input_rep2*.bam aligned/Input_rep3*.bam \
--binSize 10000 \
--numberOfProcessors 8 \
--outRawCounts qc/raw_bin_counts.tab \
-o qc/cov.npz
plotCorrelation \
--corData qc/cov.npz \
--corMethod spearman \
--skipZeros \
--whatToPlot heatmap \
--colorMap RdYlBu_r \
--plotNumbers \
-o qc/replicate_correlation.pdfIP replicates within a condition should cluster (Spearman >= 0.85 typical); input replicates should cluster with each other; IP and input should NOT cluster together. A failed IP looks like input.
Goal: Confirm IP libraries are enriched (a few transcripts have many reads) and input libraries are uniform (reads spread across transcripts); fail-fast on poor IP before peak calling.
Approach: deepTools plotFingerprint builds a cumulative Lorenz-style curve; a steep tail = signal concentrated in few regions (good IP); a diagonal = uniform coverage (input or failed IP). The framework is from ChIP-seq (Diaz 2012 Stat Appl Genet Mol Biol 11:9) and transfers cleanly to MeRIP.
plotFingerprint \
--bamfiles aligned/IP_rep1*.bam aligned/IP_rep2*.bam aligned/IP_rep3*.bam \
aligned/Input_rep1*.bam aligned/Input_rep2*.bam aligned/Input_rep3*.bam \
--labels IP1 IP2 IP3 In1 In2 In3 \
--numberOfProcessors 8 \
--skipZeros \
--outQualityMetrics qc/fingerprint_metrics.tab \
-o qc/fingerprint.pdfGood MeRIP IP: cumulative-fraction-of-reads vs cumulative-fraction-of-bins curve sits well below the diagonal in the right half (top-X% of bins capture >50% of reads). Input: near-diagonal. The --outQualityMetrics file reports JS distance and synthetic JS distance; the IP-vs-Input JS distance is a single-number IP-quality summary (higher = more concentrated signal).
Goal: Compute per-library complexity so peak counts can be honestly compared across libraries and conditions of different sequencing depth.
Approach: PreSeq c_curve (interpolation up to observed depth) and lc_extrap (extrapolation beyond observed) on the sorted BAM. Daley & Smith 2013 Nat Methods 10:325 capture-recapture model.
mkdir -p complexity
for bam in aligned/*.bam; do
name=$(basename ${bam} .bam)
preseq c_curve -B -o complexity/${name}_c_curve.txt ${bam}
preseq lc_extrap -B -o complexity/${name}_lc_extrap.txt ${bam}
doneInspect: the lc_extrap curve plots distinct molecules vs total reads; a plateau indicates saturation. For cross-condition peak-count comparison: pick a common depth (often 30M unique reads), subsample with samtools view -s 0.<frac> to that depth, THEN call peaks.
Goal: Produce a per-bin log2 (IP / Input) coverage track per replicate, ready for downstream metagene / browser plots.
Approach: deepTools bamCompare with --operation log2; choose a sensible pseudocount to avoid divide-by-zero at low-coverage bins.
mkdir -p tracks
paste -d ' ' \
<(printf '%s\n' IP_rep1 IP_rep2 IP_rep3) \
<(printf '%s\n' Input_rep1 Input_rep2 Input_rep3) | \
while read ip input; do
bamCompare \
-b1 aligned/${ip}_Aligned.sortedByCoord.out.bam \
-b2 aligned/${input}_Aligned.sortedByCoord.out.bam \
--operation log2 \
--pseudocount 1 \
--binSize 25 \
--normalizeUsing CPM \
--numberOfProcessors 8 \
-o tracks/${ip}_over_${input}.bw
done--pseudocount 1 prevents division-by-zero at zero-coverage bins; --binSize 25 is fine-grained enough to preserve peak topology while keeping bigWig files reasonably sized.
Trigger: picard MarkDuplicates REMOVE_DUPLICATES=true invoked on a standard MeRIP BAM that has no UMI.
Mechanism: Standard MeRIP libraries have no unique molecular identifiers. PCR duplicates and biological re-sampling at high-coverage transcripts look identical at the alignment level. Dedup removes both, collapsing real coverage at the most-abundant transcripts to an artificially flat profile. This is the opposite of dedup's intent in DNA ChIP-seq.
Symptom: Coverage at housekeeping mRNAs (e.g., GAPDH, ACTB) drops 5-20x after dedup; downstream peak counts at highly-expressed transcripts collapse; volcano plot of differential peaks shows expression-driven false positives.
Fix: Skip dedup for standard non-UMI MeRIP. If the library is UMI-MeRIP, use umi_tools dedup (Smith 2017 Genome Res 27:491) which respects UMI rather than alignment position alone. Record dedup status in sample-sheet metadata.
Trigger: STAR or bowtie2 aligned to transcriptome FASTA, then BAM passed to exomePeak2 / MeTPeak / MACS3.
Mechanism: exomePeak2 and MeTPeak expect a GENOME BAM plus GTF; they project peaks back to transcript features internally. A transcriptome BAM has reads in per-transcript coordinates which the GTF cannot resolve back to genome coordinates without re-alignment.
Symptom: exomePeak2 throws errors on TxDb-genome consistency; MeTPeak returns zero peaks; MACS3 calls peaks on transcript IDs as if they were chromosomes.
Fix: Align to GENOME with STAR / HISAT2 for downstream MeRIP peak calling. Transcriptome alignment is correct only for m6anet-analysis (ONT DRS) and rare quantification-only downstream tools.
Trigger: A single replicate IP library has IP/input ratio distribution centred at 1.0 across all transcripts (no enrichment); fingerprint Lorenz curve sits at the diagonal.
Mechanism: Failed IP — antibody-RNA binding did not enrich m6A-containing fragments. Causes include antibody-batch defect, insufficient pulldown wash, RNA degradation during IP, or accidental mock IgG IP.
Symptom: plotFingerprint shows IP overlaying input on the Lorenz plot; per-transcript IP/input ratio histogram is centred at 1.0; downstream peak callers find few or no peaks AT THE FAILED REPLICATE while other replicates produce normal counts.
Fix: Identify the failed replicate via plotFingerprint AND IP/input ratio distribution BEFORE peak calling; exclude or re-do. Single failed IP in a 3-replicate design routinely produces "differential" peaks driven entirely by the failure.
Trigger: A multi-condition MeRIP study uses Synaptic Systems 202-003 antibody lot A for the control IPs and lot B for the treatment IPs (because lot A ran out mid-study).
Mechanism: Anti-m6A polyclonals (Synaptic Systems 202-003, Abcam ab151230, NEB EpiMark E1610, Cell Signaling 56593, Active Motif 61755) have batch-to-batch variability in pulldown efficiency and m6A-vs-m6Am cross-reactivity. Pooling lot-A and lot-B counts in a downstream differential model attributes lot-effect to condition.
Symptom: "Differential" peaks at high abundance transcripts; effect sizes track antibody lot rather than condition; reanalysis with lot in the design matrix removes most differential peaks.
Fix: Record antibody_clone and antibody_lot per sample in metadata; include lot as a fixed effect in downstream differential analysis. Within a single study, ideally use a single lot for ALL replicates and ALL conditions.
Trigger: "Condition A has 14,000 peaks; condition B has 22,000 peaks; condition B has more m6A."
Mechanism: Peak count is library-size-dependent. A library at 60M unique reads finds more peaks than 30M. Without rarefaction or saturation correction, peak-count comparisons across libraries are dominated by sequencing depth.
Symptom: Peak counts track total mapped reads more closely than they track biological condition; downstream "biological m6A change" claims do not survive rarefaction-to-common-depth.
Fix: Either rarefy all BAMs to common unique-read depth before peak calling, OR fit saturation curves with PreSeq lc_extrap and compare at matched depth, OR report peak count alongside the saturation curve.
Trigger: Aggressive 5' trimming of the first 6-12 nt to remove "random hexamer priming bias" applied to MeRIP libraries.
Mechanism: Random hexamer priming bias affects the 5' nucleotide composition of reads but does NOT degrade downstream peak-calling accuracy. Over-trimming removes biological signal and shortens reads enough to inflate multi-mapper fraction.
Fix: Standard adapter trimming with --length_required 25 is sufficient; do not 5'-trim for hexamer bias unless downstream tooling explicitly requires unbiased 5' ends (most do not). The bias is a known artifact in the RNA-seq community and is robust to standard analytical pipelines.
| Pattern | Likely cause | Action |
|---|---|---|
| plotFingerprint diagonal but IP/input ratio shows enrichment | Mismatched chromosome naming (chr1 vs 1) between samples | Verify `samtools view -H bam |
| Replicate Spearman 0.95 but plotFingerprint diverges | Replicates correlate in bulk but differ in IP enrichment depth | Check per-sample sequencing depth; reduce to common depth |
| Saturation curve plateaus early but peak count low | Library complexity exhausted (e.g., over-amplified PCR) | Inspect duplicate rate; re-prep library if possible |
| MultiQC reports input has higher mapping rate than IP | IP enriches non-canonical sequences (m6A on intronic RNA, mt-RNA) that map differently | Acceptable if STAR multi-mapper retention is on; verify with idxstats |
| Properly-paired rate < 60% | Insert size distribution off (RNA degradation; library prep failure) | Inspect samtools view -f 0x2 count; re-prep if severe |
HISAT2 reports many discordant pairs | Splice-junction not captured in index | Re-build with --dta and confirm GTF matches genome |
| plotFingerprint synthetic JS distance < 0.5 | Marginal IP enrichment; borderline failed | Inspect per-transcript IP/input ratio distribution; consider exclusion |
| Quantity | Threshold | Source / rationale |
|---|---|---|
| Minimum read length after trimming | 25 nt | Below this, multi-mapping fraction inflates; downstream peak callers lose specificity |
STAR --outFilterMultimapNmax for MeRIP | 20 | Retains multi-isoform mapping; tighten to 1 only when downstream cannot tolerate |
--sjdbOverhang | read length - 1 | STAR convention; 100 is common for 100-150 bp reads |
| Properly-paired rate (samtools flagstat) | >=85% | Below indicates degraded RNA or library-prep failure |
| Replicate Spearman correlation (multiBamSummary 10 kb bins) | >=0.85 (IP-vs-IP within condition) | Below suggests one replicate is anomalous |
| plotFingerprint IP-vs-input JS distance | >=0.5 | Higher indicates better IP enrichment; <0.3 suggests failed IP |
| Saturation curve plateau depth | ~30-60M unique reads typical | Below this, peak calling under-samples; depth depends on cell type / antibody |
| Per-transcript IP/input ratio median | >1.5 (genome-wide median) | Lower suggests failed IP; conditions / cell lines vary |
| Minimum biological replicates | 3 (4-5 preferred) per condition | McIntyre 2020 Sci Rep 10:6590 — N=2 routinely under-powered |
| Dedup status for non-UMI MeRIP | OFF | Standard convention; UMI-MeRIP is the only exception |
| BAM sort order for downstream tools | Coordinate (SortedByCoordinate) | exomePeak2, MeTPeak, MACS3, deepTools all expect coordinate sort |
bamCompare --binSize for downstream metagene | 25 | Fine enough to preserve peak topology; coarser only for whole-chromosome browser views |
bamCompare --pseudocount | 1 | Prevents divide-by-zero at zero-coverage bins; larger values flatten signal |
| Error / symptom | Cause | Solution |
|---|---|---|
| STAR runs out of memory on human genome | --genomeDir build needs ~30 GB RAM | Use HISAT2 (~12 GB) or run STAR on a high-memory node |
samtools index fails with "is not coordinate sorted" | BAM is name-sorted or unsorted | Re-run samtools sort (not sort -n) |
| exomePeak2 errors on TxDb chromosome mismatch | BAM uses chr1, GTF uses 1 (or reverse) | Verify with `samtools view -H bam |
deepTools bamCompare --ratio log2 deprecation warning | Newer deepTools uses --operation log2 | Switch to --operation log2 |
PreSeq lc_extrap rejects with "low complexity" | Library too shallow OR genome too small (BAM under 1M unique reads) | Use c_curve only; or sequence deeper |
| MultiQC misses STAR Log.final.out | STAR output naming non-standard | Re-run with --outFileNamePrefix and rerun MultiQC; check multiqc_config.yaml search patterns |
picard MarkDuplicates collapses all reads to 1 per position | Tiny BAM or single read pair per fragment | Verify BAM has many properly-paired reads; do NOT dedup non-UMI MeRIP regardless |
| Empty fingerprint output | All BAMs have identical bin coverage | Verify BAMs are different files; check multiBamSummary --outRawCounts |
| bigWig file size too large | Bin size too small at deep coverage | Increase --binSize from 25 to 50; bigWig is lossy at large bin sizes |
| Saturation curve never plateaus | Library deeply under-sampled | Sequence deeper OR accept curve does not plateau and report accordingly |
fastp --umi errors on non-UMI library | UMI flag passed but library has no UMI | Drop --umi flag for standard non-UMI MeRIP |
| Pushback | Response |
|---|---|
| "Was deduplication applied?" | No — standard non-UMI MeRIP protocol; PCR duplicate vs biological resampling indistinguishable without UMI; dedup collapses real coverage at high-expression transcripts |
| "What is the IP enrichment QC?" | deepTools plotFingerprint reported per replicate; JS distance >=0.5 vs input |
| "Are the replicates concordant?" | Spearman correlation matrix reported via deepTools plotCorrelation on 10 kb bins; IP-IP within condition >=0.85 |
| "Saturation curve?" | PreSeq lc_extrap per library; libraries rarefied to common depth before downstream peak calling |
| "What antibody clone and lot?" | Recorded per sample in metadata; same lot for all replicates within study |
| "Why STAR instead of HISAT2?" | STAR splice-junction-DB-based vs HISAT2 graph-based; both valid for MeRIP; choice driven by memory budget |
| "How many biological replicates?" | N >=3 per condition (per McIntyre 2020); N=2 is under-powered for differential downstream |
| "Was alignment to genome or transcriptome?" | Genome (required for exomePeak2 / MeTPeak / MACS3 downstream); transcriptome alignment is for m6anet-analysis only |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in epitranscriptomics/merip-preprocessing of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Epitranscriptomics Merip Preprocessing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Epitranscriptomics Merip Preprocessing this skillGPTomics/bioSkills | 1.2k | 1 repos | ~8.5k | Automated safety check: Pass | MIT | |
| Workflow OrchestrationAnastasiyaW/codex-claude-code-config | 154 | — | ~3.8k | Automated safety check: Pass | MIT | |
| EtetoolkitK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.3k | Automated safety check: Notes | GPL-3.0-or-later | |
| Pharmacoeconomic EvaluationLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.2k | Automated safety check: Pass | MIT | |
| En Journal Workflowfranklee16/academic-research-skills | 223 | 1 repos | ~1.5k | Automated safety check: Pass | None | |
| Stata Accounting Researchwentorai/research-plugins | 298 | 1 repos | ~4k | Automated safety check: Pass | MIT |
AnastasiyaW/codex-claude-code-config
Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).
K-Dense-AI/scientific-agent-skills
Analyzes, manipulates, compares, annotates, and visualizes phylogenetic or other hierarchical trees with ETE 4.
LeoYeAI/openclaw-master-skills
This skill provides comprehensive guidance and tools for conducting pharmacoeconomic evaluations including cost-effectiveness analysis (CEA), cost-utility analysis (CUA), cost-benefit analysis…
franklee16/academic-research-skills
A skill your agent uses when deciding which English economics / finance / management / accounting / marketing / operations / information-systems journal skill to invoke next, comparing fit across…
wentorai/research-plugins
STATA code patterns for empirical accounting and finance research
hh-health-AI/healthcare-equity
A skill your agent uses for healthcare reimbursement, clinical catalysts, utilization, epidemiology, provider adoption or economics, procedure exposure, safety, IP/exclusivity, international access…
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Categories
Aligns and QCs methylated-RNA-immunoprecipitation (MeRIP / m6A-seq) IP and input libraries using STAR or HISAT2 splice-aware mapping, samtools sort/index, IP/input matched-pair tracking…. Bio Epitranscriptomics Merip Preprocessing is an agent skill from GPTomics/bioSkills. Aligns and QCs methylated-RNA-immunoprecipitation (MeRIP / m6A-seq) IP and input libraries using STAR or HISAT2 splice-aware mapping, samtools sort/index, IP/input matched-pair tracking, antibody-lot metadata recording, replicate concordance via deepTools multiBamSummary + plotCorrelation, IP enrichment QC via plotFingerprint and per-transcript IP/input ratio distributions, library-complexity saturation curves via PreSeq, and the explicit do-NOT-deduplicate convention for standard non-UMI MeRIP.
Bio Epitranscriptomics Merip Preprocessing fits situations like: preparing paired IP and input BAM files for exomePeak2 / MeTPeak / MACS3 peak calling; evaluating MeRIP replicate concordance and IP enrichment; deciding whether to deduplicate (standard MeRIP typically NOT); choosing genome-vs-transcriptome alignment for downstream peak vs m6Anet workflows.
Run `npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a claude-code`. Or copy the skill folder (epitranscriptomics/merip-preprocessing in GPTomics/bioSkills) into .claude/skills/bio-epitranscriptomics-merip-preprocessing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a codex`. Or copy the skill folder (epitranscriptomics/merip-preprocessing in GPTomics/bioSkills) into .agents/skills/bio-epitranscriptomics-merip-preprocessing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-epitranscriptomics-merip-preprocessing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-epitranscriptomics-merip-preprocessing, .gemini/skills/bio-epitranscriptomics-merip-preprocessing, .github/skills/bio-epitranscriptomics-merip-preprocessing and .opencode/skills/bio-epitranscriptomics-merip-preprocessing in your project.
Going by SKILL.md and its folder, Bio Epitranscriptomics Merip Preprocessing needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Epitranscriptomics Merip Preprocessing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.5k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Epitranscriptomics Merip Preprocessing: Workflow Orchestration (AnastasiyaW/codex-claude-code-config, 154 stars), Etetoolkit (K-Dense-AI/scientific-agent-skills, 48k stars), Pharmacoeconomic Evaluation (LeoYeAI/openclaw-master-skills, 2.2k stars) and En Journal Workflow (franklee16/academic-research-skills, 223 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.