tangermeme Genomic Model Analysis
jmschrei/tangermeme
Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.
Build a differential-ready consensus peakset from per-replicate ATAC-seq peaks using iterative overlap removal, fixed-width re-centering, and majority-rule overlap.
$ npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-atac-seq-consensus-peakset --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/atac-seq/consensus-peakset .claude/skills/bio-atac-seq-consensus-peakset && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-atac-seq-consensus-peakset" agent skill from https://github.com/GPTomics/bioSkills/tree/main/atac-seq/consensus-peakset into .claude/skills/bio-atac-seq-consensus-peakset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-atac-seq-consensus-peakset", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/atac-seq/consensus-peaksetType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-atac-seq-consensus-peakset --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/atac-seq/consensus-peakset .agents/skills/bio-atac-seq-consensus-peakset && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-atac-seq-consensus-peakset" agent skill from https://github.com/GPTomics/bioSkills/tree/main/atac-seq/consensus-peakset into .agents/skills/bio-atac-seq-consensus-peakset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-atac-seq-consensus-peakset", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-atac-seq-consensus-peakset --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/atac-seq/consensus-peakset .cursor/skills/bio-atac-seq-consensus-peakset && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-atac-seq-consensus-peakset" agent skill from https://github.com/GPTomics/bioSkills/tree/main/atac-seq/consensus-peakset into .cursor/skills/bio-atac-seq-consensus-peakset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-atac-seq-consensus-peakset", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path atac-seq/consensus-peakset--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-atac-seq-consensus-peakset --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/atac-seq/consensus-peakset .gemini/skills/bio-atac-seq-consensus-peakset && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-atac-seq-consensus-peakset" agent skill from https://github.com/GPTomics/bioSkills/tree/main/atac-seq/consensus-peakset into .gemini/skills/bio-atac-seq-consensus-peakset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-atac-seq-consensus-peakset", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-atac-seq-consensus-peaksetInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/atac-seq/consensus-peakset .github/skills/bio-atac-seq-consensus-peakset && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-atac-seq-consensus-peakset" agent skill from https://github.com/GPTomics/bioSkills/tree/main/atac-seq/consensus-peakset into .github/skills/bio-atac-seq-consensus-peakset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-atac-seq-consensus-peakset", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-atac-seq-consensus-peakset --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/atac-seq/consensus-peakset .opencode/skills/bio-atac-seq-consensus-peakset && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-atac-seq-consensus-peakset" agent skill from https://github.com/GPTomics/bioSkills/tree/main/atac-seq/consensus-peakset into .opencode/skills/bio-atac-seq-consensus-peakset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-atac-seq-consensus-peakset", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-atac-seq-consensus-peaksetBuild a differential-ready consensus peakset from per-replicate ATAC-seq peaks using iterative overlap removal, fixed-width re-centering, and majority-rule overlap.
Bio Atac Seq Consensus Peakset is an agent skill from GPTomics/bioSkills. Build a differential-ready consensus peakset from per-replicate ATAC-seq peaks using iterative overlap removal, fixed-width re-centering, and majority-rule overlap. Use when generating a stable peak coordinate system for downstream differential accessibility, ML feature engineering, cross-sample comparison, or fixed-width peak counts; covers Corces 2018 iterative overlap (501 bp), DiffBind summit re-centering, and ENCODE consistency rules.
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/iterative_overlap.sh` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics and Machine learning. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
python3pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Atac Seq Consensus Peakset loads about 5k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 1,791 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,791 words, ~5,010 tokens.
.claude/skills/bio-atac-seq-consensus-peakset/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: bedtools 2.31+, samtools 1.19+, BEDOPS 2.4.41+, GenomicRanges 1.54+, DiffBind 3.12+, Subread 2.0.2+ (featureCounts; --countReadPairs requires >= 2.0.2), pybedtools 0.10+.
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturespackageVersion('<pkg>') then ?function_name to verify parameters<tool> --version then <tool> --help to confirm flagsIf code throws unexpected errors, introspect the installed package and adapt rather than retrying.
"Build a single peakset to count reads against for all my samples" -> Combine per-replicate or per-condition peak calls into a non-redundant, fixed-width set of regions. The strategy chosen drives FDR calibration, peak-width fairness, and reproducibility downstream.
bedtools merge (simple union) and bedtools multiinter (per-sample membership columns emitted by default)DiffBind::dba.count(summits=250) (built-in fixed-width)pybedtools for programmatic mergingThe peakset choice is rarely default-correct. Wrong width or wrong overlap rule propagates to every downstream analysis (differential, motif, footprint, ML).
ATAC peaks vary in width across replicates: same regulatory element might be called 200 bp in rep1 and 800 bp in rep2 because of stochastic Tn5 cuts at edges. Counting reads in different-width intervals confounds peak width with biological signal. A fixed-width consensus avoids this.
For ENCODE-style differential analysis: ALL samples must be counted against the SAME peak coordinates; otherwise the count matrix is non-rectangular and statistical models are misspecified.
| Strategy | Implementation | Width | When to use | Fails when |
|---|---|---|---|---|
| Naive union | bedtools merge of all peaks | Variable, tends wide | Quick exploratory; never for differential | Width inflation drives spurious differential |
| Naive intersection | bedtools multiinter requiring all samples | Variable | High-stringency reproducibility | Loses real condition-specific peaks |
| Majority-rule overlap | multiinter requiring >= n/2 samples | Variable | Balance; DiffBind default with minOverlap | Width still varies; counts are width-biased |
| Iterative overlap removal (Corces 2018) | Sort by significance, greedily keep non-overlapping at fixed width | 501 bp fixed | ML features; cross-study comparison; modern ATAC standard | Loses sub-501bp resolution; overweights high-significance peaks |
| Summit-centered fixed width (DiffBind) | dba.count(summits=250) re-centers all peaks on summit +/- 250 bp | 501 bp fixed | Matches the Corces 501 bp convention (note: dba.count default is summits=200 -> 401 bp); integrates with replicate counts | Requires summit info (MACS narrowPeak); broad peaks lose width info |
| IDR-filtered union | Union of IDR-passed peaks across rep pairs | Variable | ENCODE pipeline-compliant; reproducibility-aware | Requires running IDR per pair; computationally heavier |
| Per-condition union, then global union | Each group consensus separately, then merge | Variable | Different cell types / strong condition shift | Same width issues as naive union |
| Width-controlled extension | Extend each peak to median width centered on midpoint | User-set | Quick fixed-width without summit info | Midpoint != summit; can shift biology |
Methodology evolves; verify against current ENCODE 4 ATAC standards (encodeproject.org/atac-seq) and the Corces 2018 iterative overlap algorithm before locking pipelines.
This is the modern standard for fixed-width consensus peaksets used in cross-study comparison and machine-learning feature matrices.
Algorithm:
Goal: Produce a non-overlapping, fixed-width peakset weighted toward strongest evidence.
Approach: Greedy non-overlap on summit-centered fixed-width peaks ranked by signalValue.
#!/bin/bash
# Corces 2018 iterative overlap removal
PEAKS_DIR=peaks/per_sample
GENOME_SIZES=hg38.chrom.sizes
WIDTH_HALF=250 # 501 bp total width
# 1. Pool all peaks, re-center on summit, extend
awk -v w=$WIDTH_HALF 'BEGIN{OFS="\t"}
{summit=$2+$10; print $1, summit-w, summit+w+1, $4, $7, $6}' \
$PEAKS_DIR/*.narrowPeak | \
awk '$2 >= 0' | \
bedtools slop -i - -g $GENOME_SIZES -b 0 | \
sort -k1,1 -k2,2n > pooled_recentered.bed
# 2. Sort by signalValue descending (column 5 in our BED -- which was column 7 of narrowPeak)
sort -k5,5gr pooled_recentered.bed > pooled_by_sig.bed
# 3. Iterative greedy non-overlap (process in significance order)
python3 - <<'EOF'
import sys
kept = []
with open('pooled_by_sig.bed') as f:
for line in f:
chrom, start, end, name, sig, strand = line.strip().split('\t')[:6]
start, end = int(start), int(end)
overlap = any(c == chrom and not (end <= s or start >= e) for c, s, e in kept)
if not overlap:
kept.append((chrom, start, end))
with open('consensus_iterative.bed', 'w') as f:
for c, s, e in sorted(kept):
f.write(f'{c}\t{s}\t{e}\n')
EOF
sort -k1,1 -k2,2n consensus_iterative.bed > consensus_final.bed
echo "Consensus peakset: $(wc -l < consensus_final.bed) fixed-width 501bp peaks"For a faster approximation at scale, bedtools cluster + per-cluster top-significance selection is tempting, but it is NOT equivalent to greedy iterative overlap: cluster groups transitively-overlapping peaks (a chain can span well beyond 501 bp), so keeping one peak per cluster discards non-overlapping peaks the greedy algorithm would retain. Use it only as an approximation.
Trigger: Using bedtools merge of all per-rep peaks; counting reads in merged intervals.
Mechanism: A peak appearing as 200 bp in one rep but 800 bp in another merges to 800 bp in the union. Read count in 800 bp interval is biased high; differential analysis flags it as condition-specific even when it's just width difference.
Symptom: Top differential peaks track peak width (mean width different by 100+ bp between conditions).
Fix: Use fixed-width strategy (Corces iterative or DiffBind summits=250). Never use merged variable-width peaks for differential.
Trigger: Using bedtools multiinter -i ... (which emits per-sample membership columns by default) and requiring all samples.
Mechanism: Peaks present only in one condition fail intersection requirement and are excluded from the consensus, even though they are the biology of interest.
Symptom: Differential analysis returns near-zero significant peaks; the condition-specific peaks were filtered out before testing.
Fix: Use per-condition consensus, then union of consensus. Or majority rule with per-condition minOverlap.
Trigger: dba.count(minOverlap=ceiling(N/2)) where N = total replicates and one condition has fewer reps.
Mechanism: A peak in 2/2 reps of cond1 but 0/3 reps of cond2 is in 2/5 = 40% < 50% -> dropped.
Fix: Compute consensus per condition first (e.g. peak in >= 2/3 reps), then union across conditions. This preserves condition-specific peaks.
Trigger: Extending peak to fixed width using midpoint ((start+end)/2).
Mechanism: Midpoint of the called peak is rarely at the actual summit; particularly for asymmetric peaks (TSS-flanking) or wide broadPeaks.
Fix: Use summit position from narrowPeak column 10 (start + summit_offset). If summit unavailable (broadPeak), use peak start + half median width as approximation.
Trigger: Running IDR for each rep pair, then unioning IDR-passed peaks.
Mechanism: IDR thresholds differ (true reps 0.05; pseudoreps 0.10). Mixing both into a union biases the consensus toward looser pseudorep peaks unless filtering carefully.
Fix: Decide upfront -- use only true-rep IDR (stricter, fewer peaks) OR pseudorep IDR (looser, more peaks). Don't mix.
| Goal | Strategy |
|---|---|
| Standard 2-3 rep DA analysis with DiffBind | DiffBind summits=250, minOverlap=2 |
| Modern ATAC publication / ML features | Corces 2018 iterative overlap (501 bp fixed) |
| ENCODE-compliant differential | IDR-passed per-rep-pair union (true-rep threshold 0.05) |
| Multi-condition with strong biology shift | Per-condition consensus then union |
| Single-cell ATAC pseudobulk | MACS3 per cluster -> iterative overlap across clusters |
| Cross-study peak comparison | Iterative overlap on shared genome build; lift over if needed |
| Quick exploratory | bedtools merge (acknowledge width bias; not for stats) |
| Pattern | Likely cause | Action |
|---|---|---|
| Iterative overlap peakset much smaller than DiffBind summits=250 | Iterative removes overlapping summits within 500 bp; DiffBind keeps them | Both valid; iterative is sparser, DiffBind preserves resolution |
| Per-condition union then merged peak count >> within-condition consensus | Strong condition-specific peaks; biology is real | Use this strategy; do NOT collapse to global majority rule |
| IDR-pass peaks much fewer than majority-rule | IDR is stricter (reproducibility-aware) than overlap counting | Use IDR for high-confidence reporting; majority for exploratory |
| Same biology gives different peak counts on different genome builds | Lift-over noise; build-specific blacklist differences | Match builds; use ENCODE blacklist v2 specific to build |
Operational rule: Document the strategy explicitly in methods. The strategy choice can shift downstream results by 30-50% in peak count and FDR. Reproducibility requires explicit specification.
# pybedtools: re-center on summit and fix width to 501 bp
import pybedtools as pbt
def fix_width_recenter(narrowpeak_path, half_width=250, out_path='consensus.bed'):
"""Re-center each narrowPeak on its summit and fix width to 2 * half_width + 1 bp."""
rows = []
for line in open(narrowpeak_path):
f = line.strip().split('\t')
chrom = f[0]; start = int(f[1]); summit_off = int(f[9])
signal = float(f[6])
summit = start + summit_off
rows.append((chrom, max(0, summit - half_width), summit + half_width + 1,
f[3], signal, f[5]))
rows.sort(key=lambda r: -r[4]) # Sort by signalValue desc
bt = pbt.BedTool([list(map(str, r)) for r in rows])
return bt.saveas(out_path)# For each condition, build a within-condition consensus first
for cond in cond1 cond2; do
bedtools multiinter -i <(sort -k1,1 -k2,2n ${cond}_rep1.narrowPeak) \
<(sort -k1,1 -k2,2n ${cond}_rep2.narrowPeak) \
<(sort -k1,1 -k2,2n ${cond}_rep3.narrowPeak) \
-names rep1 rep2 rep3 | \
awk '$4 >= 2 {print $1"\t"$2"\t"$3}' > ${cond}_consensus.bed
done
# Union across condition consensuses (preserves condition-specific peaks)
cat cond1_consensus.bed cond2_consensus.bed | \
sort -k1,1 -k2,2n | \
bedtools merge -i - > all_conditions_consensus.bedThis is the right pattern when conditions differ enough that a global majority-rule would discard condition-specific peaks.
For Bioconductor pipelines that avoid intermediate BED files. The iterative-overlap step requires a loop; !duplicated(findOverlaps(...)) does NOT implement greedy non-overlap (every peak self-overlaps).
library(GenomicRanges); library(rtracklayer)
peaks <- import('peaks.narrowPeak') # GRanges
# trim() below only clamps ranges when seqlengths are set, so attach them from the genome index
chrom_sizes <- read.table('hg38.chrom.sizes', col.names=c('chrom', 'len'))
seqlengths(peaks) <- setNames(chrom_sizes$len, chrom_sizes$chrom)[seqlevels(peaks)]
peaks_summit <- peaks # column 10 of narrowPeak is summit offset
start(peaks_summit) <- start(peaks) + peaks$peak # move start to the summit position
end(peaks_summit) <- start(peaks_summit) # collapse to 1 bp at the summit before centering
peaks_fixed <- resize(peaks_summit, width=501, fix='center')
peaks_fixed <- trim(peaks_fixed) # clamp to chromosome bounds (requires seqlengths, set above)
# Sort by signalValue (column 7) descending; greedy iterative non-overlap
peaks_sorted <- peaks_fixed[order(-peaks_fixed$signalValue)]
kept <- logical(length(peaks_sorted))
already_used <- GRanges()
for (i in seq_along(peaks_sorted)) {
if (!any(overlapsAny(peaks_sorted[i], already_used))) {
kept[i] <- TRUE
already_used <- c(already_used, peaks_sorted[i])
}
}
consensus <- peaks_sorted[kept]
export(consensus, 'consensus_iterative.bed')For very large peaksets (>500k), use bedtools shell pipeline (faster) or precompute the per-rank reduce in chunks. The greedy loop is the canonical Corces 2018 algorithm; approximations via reduce() or disjoin() are NOT equivalent.
When merging peaks across genome builds (hg19 -> hg38; human -> mouse for cross-species analysis):
# Lift over hg19 peaks to hg38
liftOver published_peaks.hg19.bed hg19ToHg38.over.chain.gz \
published_peaks.hg38.bed published_peaks.unmapped.bed
# Then merge with current hg38 peaks
cat published_peaks.hg38.bed my_peaks.hg38.bed | \
sort -k1,1 -k2,2n | \
bedtools merge -i - > merged_consensus.bedTrigger: Cross-cohort meta-analysis or comparing to published datasets on older builds.
Mechanism: Genome assemblies differ in coordinates, gap regions, alt-haplotypes; chain files (UCSC) provide the position translation but ~3-5% of regions fail liftover (especially complex regions, sex chromosomes, MHC).
Fix: Always document liftOver chain version; report unmapped fraction; for high-stakes cross-build comparisons, re-align rather than liftover.
For human -> mouse comparison, use synteny-based mapping (bnMapper or halLiftover) rather than chain liftOver because conservation is non-trivial.
Trigger: Combined ATAC narrow peaks + H3K27ac (broadPeak) or H3K4me1 in the same workflow.
Mechanism: Fixed 501 bp from ATAC under-counts H3K27ac which spreads across 2-5 kb domains. Counting H3K27ac reads in 501 bp windows misses most signal.
Fix: Build a dual-resolution consensus: 501 bp for ATAC narrow peaks; broadPeak union (2-5 kb) for H3K27ac. Run differential separately on each resolution. Or use a unified per-feature width derived from the broader signal.
Alternative to iterative-overlap: ENCODE-rE2G (atac-seq/enhancer-gene-linking) predicts scored enhancer-gene regulatory links for ENCODE cell types, prioritizing elements by predicted regulatory function rather than just ATAC signal strength (the genome-wide cCRE registry itself is a distinct product, SCREEN/ENCODE cCREs). Its predicted regulatory elements serve as an alternative consensus peakset for ML feature engineering. Trade-off: limited to ENCODE cell types; cell-type-specific elements may not align with the actual peakset of any specific dataset.
# Convert BED to SAF for featureCounts (BED start is 0-based half-open; SAF Start is 1-based inclusive, so +1)
awk 'BEGIN{OFS="\t"; print "GeneID","Chr","Start","End","Strand"}
{print $1"_"$2"_"$3, $1, $2+1, $3, "+"}' consensus.bed > consensus.saf
featureCounts -F SAF -a consensus.saf \
-o consensus_counts.tsv \
-p --countReadPairs \
-T 8 \
sample1.bam sample2.bam sample3.bam-p --countReadPairs is required for paired-end ATAC; -T 8 parallelizes across 8 threads. Output is the count matrix consumed by DESeq2 / edgeR / DiffBind.
Blacklist filtering belongs after consensus construction, before or alongside differential testing.
# ENCODE blacklist v2 (Amemiya 2019)
bedtools intersect -v -a consensus_final.bed -b hg38-blacklist.v2.bed.gz \
> consensus_no_blacklist.bed
echo "After blacklist: $(wc -l < consensus_no_blacklist.bed) peaks"| Error / symptom | Cause | Solution |
|---|---|---|
| Differential analysis flags hundreds of peaks; biology is implausible | Variable-width consensus driving width effects | Switch to fixed-width (iterative or summits=250) |
Negative coordinates in fixed-width output | Peak summit too close to chromosome start | Clip with awk '$2 >= 0' or bedtools slop -g chrom.sizes |
| featureCounts SAF complaint about non-unique GeneID | Duplicate peak names in SAF | Use chr_start_end as GeneID; ensure peakset is non-redundant |
| Iterative overlap output much smaller than expected | Too many peaks within 500 bp; strict non-overlap | Try smaller half-width (e.g. 100) for higher resolution |
multiinter output empty | Peak files not sorted; or wrong chromosome naming | Sort all inputs first; confirm chr name consistency |
| Per-condition union has 2x peaks of any per-condition | Strong condition-specific biology -- expected | Confirm with browser tracks; use this consensus |
| DiffBind takes hours on consensus | All-by-all counting is N samples x M peaks | Use bParallel=TRUE (default DBA$config$RunParallel); limit cores via DBA$config$cores |
summits=250 parameter)© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in atac-seq/consensus-peakset of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Atac Seq Consensus Peakset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Atac Seq Consensus Peakset this skillGPTomics/bioSkills | 1.2k | 2 repos | ~5k | Automated safety check: Pass | MIT | |
| tangermeme Genomic Model Analysisjmschrei/tangermeme | 311 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Gtars Genomic Interval Toolkitdavila7/claude-code-templates | 32k | 12 repos | ~1.9k | Automated safety check: Pass | MIT | |
| AlphagenomeK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~4.4k | Automated safety check: Notes | MIT | |
| External Model Validationaipoch/medical-research-skills | 2k | — | ~3.2k | Automated safety check: Pass | MIT | |
| Popv Cell Annotationjaechang-hits/SciAgent-Skills | 370 | 2 repos | ~6.9k | Automated safety check: Pass | BSD-3-Clause |
jmschrei/tangermeme
Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.
davila7/claude-code-templates
Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.
K-Dense-AI/scientific-agent-skills
Looks up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC…
aipoch/medical-research-skills
A skill your agent uses when validating an existing prognostic risk signature on an external bulk expression cohort with survival outcomes, producing risk scores, Kaplan-Meier curves, risk…
jaechang-hits/SciAgent-Skills
Consensus cell type annotation: runs 10+ algorithms (KNN-Harmony/BBKNN/Scanorama/scVI, CellTypist, ONCLASS, Random Forest, SCANVI, SVM, XGBoost) on a labeled reference and transfers labels via…
majiayu000/claude-skill-registry
Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Categories
Build a differential-ready consensus peakset from per-replicate ATAC-seq peaks using iterative overlap removal, fixed-width re-centering, and majority-rule overlap. Bio Atac Seq Consensus Peakset is an agent skill from GPTomics/bioSkills. Build a differential-ready consensus peakset from per-replicate ATAC-seq peaks using iterative overlap removal, fixed-width re-centering, and majority-rule overlap.
Bio Atac Seq Consensus Peakset fits situations like: generating a stable peak coordinate system for downstream differential accessibility; ML feature engineering; cross-sample comparison; fixed-width peak counts.
Run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a claude-code`. Or copy the skill folder (atac-seq/consensus-peakset in GPTomics/bioSkills) into .claude/skills/bio-atac-seq-consensus-peakset in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a codex`. Or copy the skill folder (atac-seq/consensus-peakset in GPTomics/bioSkills) into .agents/skills/bio-atac-seq-consensus-peakset in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-consensus-peakset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-atac-seq-consensus-peakset, .gemini/skills/bio-atac-seq-consensus-peakset, .github/skills/bio-atac-seq-consensus-peakset and .opencode/skills/bio-atac-seq-consensus-peakset in your project.
Going by SKILL.md and its folder, Bio Atac Seq Consensus Peakset needs a shell for the scripts in its folder and the command-line tools its instructions call (python3 and pip). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Atac Seq Consensus Peakset is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Atac Seq Consensus Peakset: tangermeme Genomic Model Analysis (jmschrei/tangermeme, 311 stars), Gtars Genomic Interval Toolkit (davila7/claude-code-templates, 32k stars), Alphagenome (K-Dense-AI/scientific-agent-skills, 48k stars) and External Model Validation (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.