Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
De novo and known TF motif enrichment in ChIP-seq/ATAC-seq peaks via HOMER.
$ npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills homer-motif-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/homer-motif-analysis .claude/skills/homer-motif-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "homer-motif-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/homer-motif-analysis into .claude/skills/homer-motif-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "homer-motif-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/homer-motif-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills homer-motif-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/homer-motif-analysis .agents/skills/homer-motif-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "homer-motif-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/homer-motif-analysis into .agents/skills/homer-motif-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "homer-motif-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills homer-motif-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/homer-motif-analysis .cursor/skills/homer-motif-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "homer-motif-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/homer-motif-analysis into .cursor/skills/homer-motif-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "homer-motif-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/homer-motif-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills homer-motif-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/homer-motif-analysis .gemini/skills/homer-motif-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "homer-motif-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/homer-motif-analysis into .gemini/skills/homer-motif-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "homer-motif-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills homer-motif-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/homer-motif-analysis .github/skills/homer-motif-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "homer-motif-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/homer-motif-analysis into .github/skills/homer-motif-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "homer-motif-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills homer-motif-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/homer-motif-analysis .opencode/skills/homer-motif-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "homer-motif-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/homer-motif-analysis into .opencode/skills/homer-motif-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "homer-motif-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
homer-motif-analysisDe novo and known TF motif enrichment in ChIP-seq/ATAC-seq peaks via HOMER.
Homer Motif Analysis is an agent skill from jaechang-hits/SciAgent-Skills. De novo and known TF motif enrichment in ChIP-seq/ATAC-seq peaks via HOMER. findMotifsGenome.pl finds over-represented patterns vs background; annotatePeaks.pl assigns context (TSS distance, gene, repeat). Use after MACS3 to identify enriched TFs, annotate peaks with nearest genes, and validate ChIP-seq via the target motif.
Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is GPL-3.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
condapipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
homer.ucsd.edudoi.orggithub.comjaspar.elixir.noFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Homer Motif Analysis loads about 6.4k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 1,131 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its GPL-3.0 licence (© jaechang-hits). 1,131 words, ~6,428 tokens.
.claude/skills/homer-motif-analysis/SKILL.md (or your agent's skills folder).HOMER (Hypergeometric Optimization of Motif EnRichment) is a suite of Perl/C++ tools for analyzing genomic regulatory elements. Its two primary commands are findMotifsGenome.pl, which performs de novo motif discovery and known motif enrichment against JASPAR/HOMER databases, and annotatePeaks.pl, which maps each peak to the nearest gene, distance to TSS, and genomic feature class (promoter, intron, intergenic, repeat). HOMER takes BED-format peak files from MACS3 or similar peak callers and a reference genome assembly as input, and outputs HTML/text reports ranking enriched motifs by p-value and fold enrichment over a matched background.
macs3-peak-calling first to generate the peak BED files that serve as input to HOMERjaspar-database to cross-reference HOMER-discovered motifs with JASPAR IDs and additional TF metadataMEME-CHIP (web or local) when you need a more probabilistic ZOOPS/TCM model or the MEME Suite ecosystemAME (part of MEME Suite) as a faster alternative for known motif scanning without de novo discoveryinstallGenome.pl after HOMER installpandas, matplotlib, seabornCheck before installing: The tool may already be available in the current environment (e.g., inside a
pixi/condaenv). Runcommand -v findMotifsGenome.plfirst and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool viapixi run findMotifsGenome.plrather than barefindMotifsGenome.pl.
# Install HOMER via conda (recommended — handles Perl dependencies)
conda install -c bioconda homer
# Verify installation
findMotifsGenome.pl 2>&1 | head -3
# Usage: findMotifsGenome.pl <peak/BED file> <genome> <output directory> [options]
annotatePeaks.pl 2>&1 | head -3
# Usage: annotatePeaks.pl <peak/BED file> <genome> [options]
# Install reference genomes (downloads 2-way masker + sequence; ~3–10 GB each)
installGenome.pl hg38
installGenome.pl mm10
# Install Python parsing dependencies
pip install pandas matplotlib seabornSettle these with the user before writing any analysis code.
decisions:
- id: D1
param: genomeAssembly
kind: derived
source: upstream
ask: "Which assembly are the peak coordinates on?"
default: "carried from the peak-calling stage"
- id: D2
param: searchWindow
kind: required
source: user
ask: "How wide a window around each peak centre should be searched?"
default: "200 bp for TF ChIP, 150 bp for ATAC"
- id: D3
param: backgroundRegions
kind: required
source: user
ask: "Compare peaks against GC-matched random genomic regions, or against a specific control set?"
default: "auto-generated GC-matched background"
- id: D4
param: repeatMasking
kind: required
source: user
ask: "Mask repetitive sequence, which otherwise dominates the enrichment as false positives?"
default: "masked"
- id: D5
param: denovoDiscovery
kind: required
source: user
ask: "Search for previously undescribed motifs, or only test the known-motif library?"
default: "known motifs only - de novo discovery is roughly ten times slower"
- id: D6
param: motifLengths
kind: optional
source: user
depends_on: [D5]
ask: "Which motif widths should the de novo search try?"
default: "8, 10, 12"
skip_if: "de novo discovery disabled"
- id: D7
param: numDenovoMotifs
kind: optional
source: user
depends_on: [D5]
ask: "How many de novo motifs should be reported?"
default: 25
skip_if: "de novo discovery disabled"
- id: D8
param: knownMotifMismatches
kind: optional
source: user
ask: "How loosely may a sequence match a known motif and still count?"
default: 2
- id: D9
param: threads
kind: never_ask
source: data
reason: "Affects runtime only, not the enrichment"
default: "min(8, available_cores)"D1 is derived rather than asked: peak coordinates are meaningless outside the assembly they were called on, so the answer is whatever the upstream stage used, not a preference. D3 is the decision most often left at its default without thought - a background that does not match the peaks in GC content produces enrichment for GC content rather than for biology.
# Run de novo + known motif enrichment on TF ChIP-seq peaks (hg38, 200 bp window)
findMotifsGenome.pl peaks/tf_chip_summits.bed hg38 motif_output/ \
-size 200 -mask -p 4
# Annotate peaks with nearest genes and genomic features
annotatePeaks.pl peaks/tf_chip_peaks.narrowPeak hg38 > annotated_peaks.txt
echo "Top known motif:"
head -2 motif_output/knownResults.txt | tail -1 | cut -f1-4
echo "Annotated peaks: $(wc -l < annotated_peaks.txt) lines"Install HOMER and download the reference genome sequence required for motif analysis.
# Activate conda environment (or use existing env)
conda create -n homer_env -c bioconda homer python=3.10 -y
conda activate homer_env
# List available genomes
installGenome.pl list
# Install human (hg38) and mouse (mm10) genomes
# Downloads masked genome sequence and annotation files
installGenome.pl hg38
# Output: Installing hg38... Done. (3-5 min, ~3 GB)
installGenome.pl mm10
# Output: Installing mm10... Done. (3-5 min, ~2.5 GB)
# Verify genome is installed
ls ~/.homer/data/genomes/hg38/
# genome.fa chrom.sizes ...
# Check HOMER motif database
ls ~/.homer/data/knownTFs/
# vertebrates.motifs jaspar.motifs ...Prepare a summit-centered BED file from MACS3 output for optimal motif resolution.
# Option A: Use MACS3 summit file directly (already 1 bp summit positions)
# Expand summits to ±100 bp (200 bp total) centered on summit
awk 'BEGIN{OFS="\t"} {
start = ($2 - 100 < 0) ? 0 : $2 - 100;
print $1, start, $2 + 100, $4, $5
}' peaks/tf_chip_summits.bed > peaks/tf_chip_200bp.bed
echo "Summit-centered peaks: $(wc -l < peaks/tf_chip_200bp.bed)"
# Summit-centered peaks: 12453
# Option B: Use narrowPeak file directly (HOMER accepts multi-column BED)
# HOMER uses columns 1-3 (chr, start, end) and centers internally with -size
cp peaks/tf_chip_peaks.narrowPeak peaks/input_peaks.bed
# Option C: Prepare a custom background region file (matched GC content)
# HOMER auto-generates background if not provided, but explicit background
# is recommended when comparing two peak sets
# Use the control peak set or random genomic regions as background:
bedtools shuffle -i peaks/tf_chip_peaks.narrowPeak \
-g ~/.homer/data/genomes/hg38/chrom.sizes \
-excl peaks/tf_chip_peaks.narrowPeak > peaks/background_regions.bed
echo "Background regions: $(wc -l < peaks/background_regions.bed)"
# Background regions: 12453Run findMotifsGenome.pl for de novo motif discovery and known motif enrichment simultaneously.
mkdir -p motif_output/
# Full run: de novo + known motif enrichment
# -size 200: use 200 bp window centered on peak midpoint
# -mask: mask repetitive elements (recommended for clean motifs)
# -p 4: use 4 CPU threads
# -S 25: find top 25 de novo motifs (default)
findMotifsGenome.pl peaks/tf_chip_200bp.bed hg38 motif_output/ \
-size 200 \
-mask \
-p 4 \
-S 25
# Check progress output:
# Reading genome sizes for hg38 ...
# Scanning for motifs...
# Optimizing 25 motifs...
# Done! Output in motif_output/
echo "Known results: $(wc -l < motif_output/knownResults.txt) motifs"
echo "De novo motifs: $(ls motif_output/homerResults/*.motif 2>/dev/null | wc -l) motifs"
# Known results: 392 motifs
# De novo motifs: 25 motifs
# For mouse peaks (mm10)
# findMotifsGenome.pl peaks/atac_peaks.bed mm10 motif_output_mm10/ \
# -size 200 -mask -p 4Scan peaks for occurrences of a specific known motif or skip de novo discovery for speed.
# Skip de novo discovery (faster when you only need known motifs)
findMotifsGenome.pl peaks/tf_chip_200bp.bed hg38 motif_output_known/ \
-size 200 \
-mask \
-p 4 \
-nomotif
echo "Known motif results: $(wc -l < motif_output_known/knownResults.txt)"
# Known motif results: 392
# Find occurrences of a specific motif across peaks (outputs peak-level annotation)
# Extract the motif matrix file for the TF of interest from homerResults/
findMotifsGenome.pl peaks/tf_chip_200bp.bed hg38 motif_scan_out/ \
-size 200 \
-mask \
-find motif_output/homerResults/motif1.motif \
> peaks_with_motif1.txt
echo "Peaks containing motif1: $(wc -l < peaks_with_motif1.txt)"
# Peaks containing motif1: 8941
# Custom background: compare treated vs. control peak sets
findMotifsGenome.pl peaks/treated_peaks.bed hg38 motif_treated_vs_ctrl/ \
-size 200 \
-mask \
-p 4 \
-bg peaks/control_peaks.bedUse annotatePeaks.pl to assign each peak to a genomic feature and nearest gene.
# Annotate peaks with nearest gene and TSS distance
# Outputs a tab-delimited file with genomic context for each peak
annotatePeaks.pl peaks/tf_chip_peaks.narrowPeak hg38 \
> annotated_peaks.txt
echo "Annotated peaks: $(($(wc -l < annotated_peaks.txt) - 1)) peaks"
# Annotated peaks: 12453 peaks
# Preview column headers and first peak
head -2 annotated_peaks.txt | cut -f1-10
# Annotate ATAC-seq peaks (same command, different input)
annotatePeaks.pl peaks/atac_sample_peaks.narrowPeak hg38 \
> annotated_atac.txt
# Generate TSS-distance histogram (for tag density plots)
# annotatePeaks.pl can compute read density around peaks with -d flag
# annotatePeaks.pl tss hg38 -size 4000 -hist 10 \
# -d chip_tagdir/ > tss_histogram.txtRead knownResults.txt and de novo motif files into pandas for downstream analysis.
import pandas as pd
import subprocess
import io
# --- Parse known motif enrichment results ---
# knownResults.txt columns:
# Motif Name | Consensus | P-value | Log P-value | q-value | # Target Seqs w/ motif | % Target | # Bg Seqs w/ motif | % Bg
known_cols = [
"motif_name", "consensus", "pvalue", "log_pvalue",
"qvalue", "n_target_seqs", "pct_target",
"n_bg_seqs", "pct_bg"
]
known = pd.read_csv(
"motif_output/knownResults.txt",
sep="\t", header=0, names=known_cols
)
# Convert string percentages to floats
known["pct_target"] = known["pct_target"].str.rstrip("%").astype(float)
known["pct_bg"] = known["pct_bg"].str.rstrip("%").astype(float)
known["fold_enrichment"] = known["pct_target"] / known["pct_bg"].replace(0, 0.001)
known["-log10_pvalue"] = -known["log_pvalue"] / 2.303 # log10 from natural log
print(f"Total known motifs tested: {len(known)}")
print(f"Significant (p < 1e-5): {(known['pvalue'].astype(float) < 1e-5).sum()}")
print("\nTop 5 enriched motifs:")
print(
known[["motif_name", "pct_target", "pct_bg", "fold_enrichment", "-log10_pvalue"]]
.head(5)
.to_string(index=False)
)
# Total known motifs tested: 391
# Significant (p < 1e-5): 47
# Top 5 enriched motifs:
# motif_name pct_target pct_bg fold_enrichment -log10_pvalue
# CTCF(Zf)/GM12878-CTCF-ChIP-Seq... 78.21 8.32 9.40 312.4
# ZNF143(Zf)/... 42.11 5.10 8.26 188.7
# ...Read annotated peak output, summarize genomic feature distribution, and plot motif enrichments.
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import seaborn as sns
# --- Parse annotated peaks ---
# annotatePeaks.pl output header: PeakID chr start end strand score ...
# Key columns: "Annotation", "Distance to TSS", "Nearest RefSeq", "Gene Name"
annot = pd.read_csv("annotated_peaks.txt", sep="\t", header=0, low_memory=False)
annot.columns = annot.columns.str.strip()
# Simplify annotation categories
def simplify_annotation(ann):
if pd.isna(ann):
return "Other"
ann = str(ann)
if "promoter" in ann.lower():
return "Promoter (<2 kb)"
elif "exon" in ann.lower():
return "Exon"
elif "intron" in ann.lower():
return "Intron"
elif "tts" in ann.lower() or "downstream" in ann.lower():
return "Downstream"
elif "intergenic" in ann.lower():
return "Intergenic"
else:
return "Other"
annot["simple_ann"] = annot["Annotation"].apply(simplify_annotation)
ann_counts = annot["simple_ann"].value_counts()
print("Peak annotation summary:")
print(ann_counts.to_string())
# Peak annotation summary:
# Intron 5832
# Intergenic 3102
# Promoter (<2 kb) 2218
# Exon 891
# Downstream 287
# Other 123
# --- Plot 1: Annotation pie chart ---
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
colors = ["#E41A1C", "#377EB8", "#4DAF4A", "#984EA3", "#FF7F00", "#A65628"]
axes[0].pie(
ann_counts.values,
labels=ann_counts.index,
colors=colors[:len(ann_counts)],
autopct="%1.1f%%",
startangle=140,
textprops={"fontsize": 9}
)
axes[0].set_title("Peak Genomic Annotation\n(n=12,453 peaks)", fontsize=11)
# --- Plot 2: Top known motif enrichments ---
known = pd.read_csv("motif_output/knownResults.txt", sep="\t", header=0)
known.columns = [
"motif_name", "consensus", "pvalue", "log_pvalue",
"qvalue", "n_target_seqs", "pct_target", "n_bg_seqs", "pct_bg"
]
known["pct_target"] = known["pct_target"].str.rstrip("%").astype(float)
known["pct_bg"] = known["pct_bg"].str.rstrip("%").astype(float)
known["-log10_p"] = (-known["log_pvalue"].astype(float)) / 2.303
known["short_name"] = known["motif_name"].str.split("/").str[0]
top_motifs = known.head(15).copy()
sns.barplot(
data=top_motifs,
x="-log10_p",
y="short_name",
palette="Blues_r",
ax=axes[1]
)
axes[1].set_xlabel("-log10(p-value)", fontsize=10)
axes[1].set_ylabel("Motif", fontsize=10)
axes[1].set_title("Top 15 Enriched Known Motifs\n(HOMER/JASPAR database)", fontsize=11)
axes[1].axvline(x=5, color="red", linestyle="--", linewidth=0.8, label="p=1e-5")
axes[1].legend(fontsize=8)
plt.tight_layout()
plt.savefig("homer_summary.png", dpi=150, bbox_inches="tight")
plt.close()
print("Saved: homer_summary.png")| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
-size | 200 | 50–1000 | Window size (bp) centered on peak midpoint for motif search; 200 for TF ChIP-seq, 150 for ATAC-seq |
-mask | off | flag | Mask repetitive elements (N-mask); strongly recommended to reduce false positives |
-p | 1 | 1–64 | Number of parallel CPU threads; use 4–8 for typical datasets |
-S | 25 | 1–100 | Number of de novo motifs to find; 25 is usually sufficient |
-nomotif | off | flag | Skip de novo motif discovery; only perform known motif enrichment (10× faster) |
-bg | auto-generated | BED file path | Custom background region file; auto-background is GC-matched if omitted |
-find | — | .motif file path | Scan peaks for occurrences of a specific motif matrix; outputs per-peak annotation |
-mis | 2 | 0–4 | Maximum mismatches allowed when matching known motifs |
-len | 8,10,12 | comma-separated integers | Motif lengths to try for de novo search; adding 6 finds shorter motifs |
genome | required | hg38, mm10, hg19, dm6, ce11, etc. | Reference genome assembly; must be installed with installGenome.pl |
Run HOMER from a Python script and capture completion status.
import subprocess
import sys
from pathlib import Path
def run_homer_motifs(
peaks_bed: str,
genome: str,
output_dir: str,
size: int = 200,
mask: bool = True,
cpus: int = 4,
n_motifs: int = 25,
denovo: bool = True,
bg_bed: str = None
) -> int:
"""Run findMotifsGenome.pl and return exit code."""
Path(output_dir).mkdir(parents=True, exist_ok=True)
cmd = [
"findMotifsGenome.pl", peaks_bed, genome, output_dir,
"-size", str(size),
"-p", str(cpus),
"-S", str(n_motifs),
]
if mask:
cmd.append("-mask")
if not denovo:
cmd.append("-nomotif")
if bg_bed:
cmd.extend(["-bg", bg_bed])
print(f"Running: {' '.join(cmd)}")
result = subprocess.run(cmd, capture_output=True, text=True)
if result.returncode != 0:
print(f"HOMER stderr:\n{result.stderr[-2000:]}", file=sys.stderr)
return result.returncode
print(f"Done. Output in {output_dir}/")
return 0
# Usage
exit_code = run_homer_motifs(
peaks_bed="peaks/tf_chip_200bp.bed",
genome="hg38",
output_dir="motif_output/",
size=200, mask=True, cpus=4, n_motifs=25
)
print(f"Exit code: {exit_code}")
# Running: findMotifsGenome.pl peaks/tf_chip_200bp.bed hg38 motif_output/ -size 200 -p 4 -S 25 -mask
# Done. Output in motif_output/
# Exit code: 0Run HOMER on multiple ChIP-seq samples in a loop.
#!/bin/bash
# Batch motif analysis for several ChIP-seq experiments (same genome)
GENOME="hg38"
PEAK_DIR="peaks"
MOTIF_DIR="motif_results"
mkdir -p "$MOTIF_DIR"
SAMPLES=(CTCF H3K4me3 FOXA2 RUNX1)
for sample in "${SAMPLES[@]}"; do
echo "=== Processing $sample ==="
peak_file="${PEAK_DIR}/${sample}_summits.bed"
if [ ! -f "$peak_file" ]; then
echo " Skipping $sample: $peak_file not found"
continue
fi
# Center on summit ±100 bp
awk 'BEGIN{OFS="\t"} {s=$2-100<0?0:$2-100; print $1,s,$2+100,$4,$5}' \
"$peak_file" > "${PEAK_DIR}/${sample}_200bp.bed"
findMotifsGenome.pl "${PEAK_DIR}/${sample}_200bp.bed" "$GENOME" \
"${MOTIF_DIR}/${sample}/" \
-size 200 -mask -p 4 -nomotif \
2> "${MOTIF_DIR}/${sample}.log"
n=$(wc -l < "${MOTIF_DIR}/${sample}/knownResults.txt")
echo " $sample: $n motifs tested. Top hit: $(sed -n '2p' "${MOTIF_DIR}/${sample}/knownResults.txt" | cut -f1)"
done
# === Processing CTCF ===
# CTCF: 392 motifs tested. Top hit: CTCF(Zf)/GM12878-CTCF-ChIP-Seq(GSE32465)/Homer
# === Processing FOXA2 ===
# FOXA2: 392 motifs tested. Top hit: Foxa2(Forkhead)/Liver-Foxa2-ChIP-Seq(GSE25694)/HomerLoad de novo motif PWM matrices for downstream comparison or plotting.
import os
import re
import pandas as pd
def parse_homer_motif(motif_file: str) -> dict:
"""Parse a HOMER .motif file into name, log_odds_threshold, and PWM."""
with open(motif_file) as f:
header = f.readline().strip() # >motif_name\tlog_odds\tlog_p-value\t0\tnucs
rows = []
for line in f:
line = line.strip()
if line:
rows.append([float(x) for x in line.split("\t")])
parts = header.lstrip(">").split("\t")
name = parts[0]
log_odds = float(parts[1]) if len(parts) > 1 else 0.0
log_p = float(parts[2]) if len(parts) > 2 else 0.0
pwm = pd.DataFrame(rows, columns=["A", "C", "G", "T"])
return {"name": name, "log_odds": log_odds, "log_p": log_p, "pwm": pwm}
# Load all de novo motifs
motif_dir = "motif_output/homerResults/"
motifs = []
for fn in sorted(os.listdir(motif_dir)):
if fn.endswith(".motif"):
motif = parse_homer_motif(os.path.join(motif_dir, fn))
motifs.append(motif)
print(f"{fn}: {motif['name']} ({len(motif['pwm'])} positions, log_p={motif['log_p']:.1f})")
print(f"\nLoaded {len(motifs)} de novo motifs")
# motif1.motif: CTCF-motif (19 positions, log_p=-8234.1)
# motif2.motif: CTCFL-motif (17 positions, log_p=-3421.7)
# ...
# Loaded 25 de novo motifsCombine peak annotations with RNA-seq DE results to find regulated genes near peaks.
import pandas as pd
# Load annotated peaks
annot = pd.read_csv("annotated_peaks.txt", sep="\t", header=0, low_memory=False)
annot.columns = annot.columns.str.strip()
# Key columns from HOMER annotation
peak_genes = annot[["PeakID (cmd=annotatePeaks.pl peaks.bed hg38)",
"Chr", "Start", "End",
"Annotation", "Distance to TSS",
"Nearest RefSeq", "Gene Name"]].copy()
peak_genes.columns = ["peak_id", "chr", "start", "end",
"annotation", "tss_dist", "refseq", "gene_name"]
# Filter promoter-proximal peaks (within 2 kb of TSS)
promoter_peaks = peak_genes[peak_genes["tss_dist"].abs() < 2000].copy()
print(f"Promoter-proximal peaks (|TSS| < 2kb): {len(promoter_peaks)}")
# Promoter-proximal peaks (|TSS| < 2kb): 2218
# Load DESeq2 results (gene_name, log2FC, padj)
de_results = pd.read_csv("deseq2_results.csv")
de_sig = de_results[de_results["padj"] < 0.05].copy()
# Merge: find DE genes with a nearby peak
merged = promoter_peaks.merge(de_sig, on="gene_name", how="inner")
print(f"DE genes with promoter-proximal peak: {len(merged)}")
print(merged[["gene_name", "tss_dist", "annotation", "log2FC", "padj"]].head())
# DE genes with promoter-proximal peak: 347| Output | Format | Description |
|---|---|---|
motif_output/knownResults.txt | TSV | All known motifs tested: name, p-value, q-value, % target, % background; primary result file |
motif_output/knownResults.html | HTML | Interactive HTML report with motif logos, statistics, and links |
motif_output/homerResults/ | Directory | De novo motifs: motifN.motif (PWM matrix), motifN.logo.png, similar.motifs.txt |
motif_output/homerResults.html | HTML | Interactive HTML report for de novo motifs |
annotated_peaks.txt | TSV | One row per peak: chr, start, end, annotation, TSS distance, nearest gene, RefSeq ID |
peaks_with_motif1.txt | TSV | Per-peak motif occurrence scores (from -find mode) |
| Problem | Cause | Solution |
|---|---|---|
ERROR: Genome not found | Genome not installed via installGenome.pl | Run installGenome.pl hg38; confirm ~/.homer/data/genomes/hg38/ exists |
command not found: findMotifsGenome.pl | HOMER binaries not in $PATH | Add HOMER bin directory to PATH: export PATH=~/.homer/bin:$PATH (or activate conda env) |
| Poor de novo motif quality (GC-rich artifacts) | Unmasked repetitive elements | Always use -mask; also try increasing -size to 300 or reducing to 150 |
| No significant known motifs (all p > 0.01) | Too few peaks or wrong peak size | Ensure ≥500 peaks; check -size matches expected binding footprint (150–200 for TFs) |
| Conda HOMER conflicts (Perl module errors) | HOMER's Perl scripts conflict with conda environment Perl | Install HOMER in a dedicated conda env: conda create -n homer_only -c bioconda homer |
annotatePeaks.pl gives all "Intergenic" | Genome annotation not installed | HOMER annotation requires installGenome.pl which downloads GTF; check ~/.homer/data/genomes/hg38/ contains *.ann files |
| Very slow run time (>2 hr) | Single-threaded or large peak set | Add -p 8; reduce -S to 10 for speed; use -nomotif for known-only runs |
-bg custom background gives poor motifs | Background regions have very different GC content | Use a GC-matched background; run bedtools shuffle with -noOverlapping to generate random same-size regions |
© jaechang-hits, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/homer-motif-analysis of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Homer Motif Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Homer Motif Analysis this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~6.4k | Automated safety check: Pass | GPL-3.0 | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
De novo and known TF motif enrichment in ChIP-seq/ATAC-seq peaks via HOMER. Homer Motif Analysis is an agent skill from jaechang-hits/SciAgent-Skills. De novo and known TF motif enrichment in ChIP-seq/ATAC-seq peaks via HOMER.
Homer Motif Analysis fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/homer-motif-analysis in jaechang-hits/SciAgent-Skills) into .claude/skills/homer-motif-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/homer-motif-analysis in jaechang-hits/SciAgent-Skills) into .agents/skills/homer-motif-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill homer-motif-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/homer-motif-analysis, .gemini/skills/homer-motif-analysis, .github/skills/homer-motif-analysis and .opencode/skills/homer-motif-analysis in your project.
Going by SKILL.md and its folder, Homer Motif Analysis needs the command-line tools its instructions call (conda and pip). Our summary lists: Python 3.
SKILL.md names 4 domains. As links in the text: homer.ucsd.edu, doi.org, github.com and jaspar.elixir.no. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Homer Motif Analysis is published under the GPL-3.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Homer Motif Analysis: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.