Agent skill

Bio Genome Intervals Overlap Significance

by GPTomics in GPTomics/bioSkills

Tests whether two genomic interval sets overlap (colocalize) more than expected by chance using a permutation test against a structured-genome null model.

MITAuto-check passedResearch & Science

Install Bio Genome Intervals Overlap Significance

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-intervals-overlap-significance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-intervals-overlap-significance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-intervals/overlap-significance .claude/skills/bio-genome-intervals-overlap-significance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-intervals-overlap-significance
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.6k tokens
SKILL.md length
2,416 words
Files
5
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Tests whether two genomic interval sets overlap (colocalize) more than expected by chance using a permutation test against a structured-genome null model.

  • Works in 3 steps: The universe/background is the lever… → A correct null preserves three things,… → One ten-minute sanity move catches most…
  • Asking whether peaks/regions are enriched at enhancers/TFBS/features
  • SKILL.md covers Version Compatibility, The Single Most Important…, Method Taxonomy and Decision Tree by Scenario, plus 11 more sections
  • Runs Python, R and Shell scripts from its folder; calls pip

What it does

Bio Genome Intervals Overlap Significance is an agent skill from GPTomics/bioSkills. Tests whether two genomic interval sets overlap (colocalize) more than expected by chance using a permutation test against a structured-genome null model. Covers bedtools fisher (analytic 2x2 screen), bedtools shuffle + jaccard permutation, GAT (isochore/GC-conditioned simulation with FDR), regioneR (flexible permutation, randomizeRegions vs circularRandomizeRegions, localZScore), LOLA (universe-relative Fisher against a region database), and GREAT/rGREAT (regulatory-domain binomial + hypergeometric for…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `examples/permutation_overlap.py`, `examples/shuffle_permutation_test.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Asking whether peaks/regions are enriched at enhancers/TFBS/features
  • Scoring region-set colocalization
  • Region-set enrichment
  • Comparing CNV/SV concordance

Example prompts

  • “Use the bio-genome-intervals-overlap-significance skill to test whether two genomic interval sets overlap (colocalize) more than expected by chance…”
  • “/bio-genome-intervals-overlap-significance”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The universe/background is the lever that moves the whole answer - bigger than the test choice. Across LOLA (userUniverse), GREAT…
  2. A correct null preserves three things, or it manufactures significance: (a) the regions' size distribution (a 50 kb domain hits anything…
  3. One ten-minute sanity move catches most false claims: shuffle the query within the same workspace and re-run the same overlap pipeline. If…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, R and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Intervals Overlap Significance loads about 5.6k tokens when it runs. Until then it costs about 228 tokens; SKILL.md has 2,416 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~228
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,416 words, ~5,620 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-intervals-overlap-significance/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
bio-genome-intervals-overlap-significance
description
Tests whether two genomic interval sets overlap (colocalize) more than expected by chance using a permutation test against a structured-genome null model. Covers bedtools fisher (analytic 2x2 screen), bedtools shuffle + jaccard permutation, GAT (isochore/GC-conditioned simulation with FDR), regioneR (flexible permutation, randomizeRegions vs circularRandomizeRegions, localZScore), LOLA (universe-relative Fisher against a region database), and GREAT/rGREAT (regulatory-domain binomial + hypergeometric for ontology-from-regions). Stresses the universe/background choice, matched background, blacklist exclusion, and multiple-testing control. Use when asking whether peaks/regions are enriched at enhancers/TFBS/features, scoring region-set colocalization or region-set enrichment, comparing CNV/SV concordance, or turning an overlap count into a defensible p-value.
tool_type
mixed
primary_tool
regioneR

Version Compatibility

Reference examples tested with: bedtools 2.31+, pybedtools 0.10+, regioneR 1.36+ (Bioconductor 3.18+), GAT 1.3+, LOLA 1.30+, rGREAT 2.4+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('regioneR') then ?permTest to verify parameters

Bioconductor packages (regioneR, LOLA, rGREAT) are version-pinned to the Bioconductor release, not just the package version - record the Bioconductor release with results. rGREAT 2.x runs GREAT locally with general background handling; rGREAT 1.x only proxied the (whole-genome-default) web server. If code throws an error, introspect the installed package and adapt the example to match the actual API rather than retrying.

Overlap Significance

"My peaks overlap enhancers a lot - is that more than chance?" -> Compare the observed overlap to a null distribution from a structured-genome model, not to a uniform-random expectation, and report a permutation p-value/z-score, not a raw count.

  • CLI: bedtools fisher -a A -b B -g genome.txt (fast screen); gat-run.py --segments=A --annotations=B --workspace=accessible.bed --isochores=gc.bed --num-samples=10000
  • Python: BedTool(A).shuffle(g='genome.txt', excl='blacklist.bed').jaccard(B) looped for a null (pybedtools)
  • R: permTest(A=peaks, B=enhancers, randomize.function=circularRandomizeRegions, evaluate.function=numOverlaps, genome='hg38', mask=blacklist, ntimes=1000) (regioneR)

The Single Most Important Modern Insight -- A Raw Overlap Count Is an Observation in Search of a Null, and the Universe Choice Dominates

"847 of 1,000 peaks overlap enhancers" means nothing until the analysis can say what that number would have been by chance - and "by chance" is almost never "place the regions uniformly at random on the genome." The genome is structured: genes cluster, GC varies in megabase isochores, mappability is uneven, half the genome is repeat/gap, and query regions are drawn from a biased universe (open chromatin, callable space, exons). Two tracks that share nothing but a gene-rich, high-GC, high-mappability habitat overlap far more than uniform-random expectation, and a naive test returns p < 1e-300. The co-localization is real; the interpretation ("functional association") is false. Three load-bearing moves:

  1. The universe/background is the lever that moves the whole answer - bigger than the test choice. Across LOLA (userUniverse), GREAT (background regions), GAT (--workspace), and regioneR (the mask/resampleRegions universe), the most consequential choice is the set of regions the query could have come from. ATAC/ChIP peaks can only be called in accessible chromatin; their honest universe is "all accessible regions," not the genome. Testing against the whole genome merely rediscovers that open chromatin is gene-rich - every gene-associated annotation lights up, none of it specific. The difference between LOLA/GAT/regioneR on the same correct universe is second-order; the difference between a correct universe and a whole-genome universe on the same tool is often p1 vs p1e-200. Care belongs on the background, not on tool selection.

  2. A correct null preserves three things, or it manufactures significance: (a) the regions' size distribution (a 50 kb domain hits anything; a 200 bp peak rarely does - relocate intervals of the observed sizes, do not sprinkle points); (b) an accessible workspace excluding assembly gaps, centromeres, the ENCODE blacklist (Amemiya 2019), and unmappable bins; (c) local structure - GC/isochore, gene density, and clustering (GAT --isochores; regioneR circularRandomizeRegions for autocorrelated regions). The "right" answer is usually less significant than the naive test - that deflation is the methodology working.

  3. One ten-minute sanity move catches most false claims: shuffle the query within the same workspace and re-run the same overlap pipeline. If shuffled regions also overlap the annotation a lot, the "enrichment" is workspace geography, not biology.

Method Taxonomy

ToolCitationNull modelWhen
bedtools fisherQuinlan 2010 Bioinformaticsanalytic 2x2; estimates the unobserved "in-neither" cell from mean interval size + genome size; ignores genome structurefast triage screen only - never the reported result
bedtools shuffle + jaccard/numOverlapsQuinlan 2010 BioinformaticsDIY size-preserving permutation; -incl/-excl make it matchedentry-level permutation in a shell/Python pipeline; full control, more code
GATHeger 2013 Bioinformaticsper-isochore size-preserving simulation in a workspace; GC/composition conditioned; built-in FDR across annotationscomposition-aware enrichment vs many tracks with multiple-testing control
regioneRGel 2016 Bioinformaticsflexible permutation; randomizeRegions (mask) vs circularRandomizeRegions (preserves clustering) vs resampleRegions (real universe); any evaluator; localZScorepublication-grade, R-native, autocorrelated regions, where-in-the-region probing
LOLASheffield & Bock 2016 BioinformaticsNOT a shuffle - Fisher's exact of query vs a region database, relative to a userUniverseranking enrichment of a peak set against ENCODE/Roadmap region collections
GREAT / rGREATMcLean 2010 Nat Biotechnol; Gu & Hubschmann 2023 Bioinformaticsgene regulatory-domain (basal 5 kb up/1 kb down, extend to 1 Mb); binomial-over-regions AND hypergeometric-over-genesGO/ontology enrichment from regions (cis-regulatory), not from a gene list

Decision Tree by Scenario

ScenarioRecommendedWhy
Quick "is this even worth permuting?"bedtools fisherone command, analytic; p~1 means stop; low means permute
Publication-grade colocalization of two region setsregioneR permTest with a maskflexible, reports p + z-score; circular randomization for clustered regions
Need GC/isochore/composition control + FDR over many tracksGAT with --workspace + --isochoresper-isochore sampling removes the GC confounder; built-in qvalue
Enrich a peak set against a region database (TFBS/chromatin)LOLA runLOLA with a correct userUniverseuniverse-relative Fisher; ranked odds ratios + q-values
GO/ontology terms FROM regions (cis-regulatory)rGREAT with a real backgroundregulatory-domain model; require binomial AND hypergeometric
GWAS/eQTL statistical colocalization (shared causal variant)-> causal-genomics/colocalization-analysisDISTINCT problem: coloc/SuSiE on summary stats, not interval overlap
Gene-list (not region) ontology enrichment-> pathway-analysis/go-enrichmentstart from genes; GREAT is the region-based analog
Query regions are autocorrelated/clusteredregioneR circularRandomizeRegionsrotating the set preserves inter-region spacing; uniform randomization inflates significance
CNV/SV concordance between call setsbedtools intersect -f 0.5 -r (50% reciprocal)"same event" convention; one-sided fractions let a giant call swallow a tiny one
Peaks not yet called-> chip-seq/peak-calling, atac-seq/atac-peak-callingthis skill operates on existing interval sets

bedtools fisher - The Analytic Screen (weakest null)

bash
bedtools fisher -a peaks.bed -b enhancers.bed -g genome.txt   # both inputs sorted; genome file required

fisher builds a 2x2 table (in-A/not x in-B/not) and runs Fisher's exact test. The trap: it cannot observe the "in-neither" cell (there is no negative class of intervals that do not exist), so it estimates the table totals from a heuristic on mean interval size and genome size, assuming intervals are independent points uniformly placeable across the genome - exactly the assumption a structured genome violates. The bedtools docs warn it is prone to inflation and advise validating any low p-value by simulation. Treat it as triage: p1 -> stop; p1e-50 -> run GAT or regioneR, because the structure-corrected p could be anywhere from 1e-30 to 0.3.

Permutation Null with bedtools shuffle (entry-level)

Goal: Decide whether two interval sets overlap more than chance, controlling for feature size and the accessible workspace.

Approach: Compute the observed overlap (jaccard or count), then shuffle one set N times within an include-list / outside a blacklist, recompute each time, and locate the observed value in the resulting null distribution.

python
import pybedtools

N_PERMUTATIONS = 1000   # >=1000 gives a stable empirical p down to ~0.001; fewer cannot resolve small p
a = pybedtools.BedTool('peaks.bed')
b = pybedtools.BedTool('enhancers.bed').sort()
observed = a.sort().jaccard(b)['jaccard']
null = [a.shuffle(g='genome.txt', incl='accessible.bed', excl='blacklist.bed', chrom=True).sort().jaccard(b)['jaccard'] for _ in range(N_PERMUTATIONS)]
p = (sum(x >= observed for x in null) + 1) / (N_PERMUTATIONS + 1)   # +1 avoids a p of exactly 0 (Phipson & Smyth 2010)

-incl restricts placement to the accessible workspace (the universe); -excl avoids gaps/blacklist; -chrom keeps each region on its own chromosome (preserves per-chromosome density). Without -incl/-excl this collapses to a uniform-random null - the wrong one.

GAT - Isochore/GC-Conditioned Simulation

Goal: Test enrichment against one or many annotation tracks while conditioning on GC/composition and controlling FDR.

Approach: Provide the query segments, the annotations, an accessible workspace, and an isochore segmentation; GAT samples size-matched segments per isochore and compares observed to sampled overlap, reporting fold, empirical p, and FDR-adjusted q across annotations.

bash
gat-run.py \
  --segments=peaks.bed \
  --annotations=features.bed \
  --workspace=accessible.bed \
  --isochores=gc_bins.bed \
  --num-samples=10000 \
  --counter=nucleotide-overlap \
  --log=gat.log > gat_results.tsv

--isochores subdivides the workspace (GC bins, chromatin state, or mappability) so sampling happens per isochore, preserving GC/composition confounding rather than averaging it away - GAT's signature over a plain shuffle. Set --num-samples (>=10000 for stable small q) and --counter explicitly; do not rely on defaults.

regioneR - Flexible Permutation in R

Goal: Get a publication-grade colocalization p-value and z-score, with a null that preserves the structure the query trivially has.

Approach: Mask the genome to the workspace, choose a randomizer that concedes the right structure (uniform vs clustering-preserving vs real-universe), permute N times scoring overlaps, then probe where the association lives with localZScore.

r
# Reference: regioneR 1.36+ (Bioconductor 3.18+) | Verify API if version differs
library(regioneR)

N_TIMES <- 1000   # permutation count; >=1000 for a stable empirical p (Gel 2016)
peaks <- toGRanges('peaks.bed')
enhancers <- toGRanges('enhancers.bed')
gam <- getGenomeAndMask(genome = 'hg38', mask = toGRanges('blacklist.bed'))

pt <- permTest(A = peaks, B = enhancers,
               randomize.function = circularRandomizeRegions,   # preserves clustering; use randomizeRegions for non-autocorrelated query
               evaluate.function = numOverlaps,
               genome = gam$genome, mask = gam$mask,
               ntimes = N_TIMES, count.once = TRUE)
pt$numOverlaps$pval; pt$numOverlaps$zscore

lz <- localZScore(A = peaks, B = enhancers, pt = pt, window = 10000, step = 500)   # sharp vs diffuse positional association

circularRandomizeRegions rotates the whole set around the genome, preserving inter-region spacing/clustering - the honest null when the query is autocorrelated (CpG islands, TAD-restricted peaks); plain randomizeRegions breaks clustering and inflates significance. resampleRegions draws from a supplied real universe. The mask is the workspace control.

LOLA - Universe-Relative Region-Set Enrichment

r
# Reference: LOLA 1.30+ | Verify API if version differs
library(LOLA)
regionDB <- loadRegionDB('LOLACore/hg38')
userSets <- readBed('peaks.bed')
userUniverse <- readBed('all_called_regions.bed')   # the candidate pool the query was drawn from - NOT the whole genome
res <- runLOLA(userSets, userUniverse, regionDB, cores = 4)   # odds ratio + p + q per reference set, rankable

LOLA is not a shuffle: it tests the query against many reference region sets (ENCODE TFBS, Roadmap chromatin) with a Fisher's exact test relative to userUniverse. The universe is everything - the LOLA vignette recommends either the union of all query sets across the experiment ("regions that were in play") or the assay's full candidate pool (e.g. all tested DHS/called peaks); pick the one that honestly bounds where the query could have come from. A whole-genome universe inflates every enrichment.

Show full SKILL.md (1,011 more words)Show less

rGREAT - Ontology Enrichment From Regions

r
# Reference: rGREAT 2.4+ | Verify API if version differs
library(rGREAT)
res <- great(toGRanges('peaks.bed'), gene_sets = 'GO:BP', tss_source = 'txdb:hg38',
             background = toGRanges('accessible.bed'))   # supply a real background, not the whole-genome default
tb <- getEnrichmentTable(res)   # has Binom + Hyper p/adjp columns

GREAT assigns each gene a regulatory domain (basal 5 kb upstream / 1 kb downstream, extended up to 1 Mb to the next gene's basal domain), maps regions to those domains, then runs a binomial-over-regions test AND a hypergeometric-over-genes test - by design, with opposite biases. Trust a term only if both fire: a single gene with a huge regulatory domain attracts regions by target size and lights up the binomial; the hypergeometric (counting that gene once) calls the bluff. Supply a real background - the binomial assumes regions are independent and uniformly placeable, which clustered ChIP/ATAC peaks violate (Fulcher 2021 demonstrates orders-of-magnitude false-positive inflation from spatial autocorrelation in genomic enrichment analysis - a transferable critique of region-to-gene-category tests).

Per-Method Failure Modes

Whole-genome universe for a biased query

Trigger: testing ATAC/ChIP peaks (or DMRs, or capture-panel regions) against the whole genome. Mechanism: the query could only have come from accessible/assayed space, which is gene-rich; the genome universe credits that geography as enrichment. Symptom: every gene-associated annotation is "significant," none specific, p absurdly small. Fix: set the universe to the callable/candidate pool (LOLA userUniverse, GAT --workspace, regioneR mask/resampleRegions).

bedtools fisher reported as the result

Trigger: putting a bedtools fisher p-value in a figure or reviewer reply. Mechanism: its analytic 2x2 estimates the unobserved cell from a uniform-placement heuristic, ignoring genome structure; prone to inflation. Symptom: spuriously tiny p that evaporates under a competent permutation null. Fix: use fisher only to triage; validate any low p with GAT/regioneR.

Uniform null on clustered query regions

Trigger: randomizeRegions / plain shuffle on autocorrelated regions (CpG islands, tandem families, TAD-restricted peaks). Mechanism: uniform placement destroys the clustering the observed data has, so permuted overlap is too low. Symptom: inflated significance vs a structure-preserving null. Fix: circularRandomizeRegions (regioneR) or per-chromosome/per-class shuffling.

No blacklist / no workspace exclusion

Trigger: shuffling across the full genome including gaps, centromeres, and the ENCODE blacklist. Mechanism: the null places regions where reads/peaks could never occur, lowering expected overlap. Symptom: enrichment that is really mappability artifact. Fix: exclude the ENCODE blacklist (Amemiya 2019) and assembly gaps via -excl/mask before any test.

Ignoring multiple testing across tracks/terms

Trigger: testing a query against many annotation tracks or many GO terms and reading raw p-values. Mechanism: dozens-to-thousands of tests inflate the family-wise false-positive rate. Symptom: a long list of "significant" hits dominated by chance. Fix: use GAT's built-in FDR, FDR-adjust LOLA ranks, and read GREAT's binomial+hypergeometric q-values; require both GREAT tests.

GREAT term significant by only one test

Trigger: reporting a GO term significant by binomial OR hypergeometric alone. Mechanism: the two tests have opposite biases (domain size vs gene count). Symptom: a term driven by one large-domain gene, or by gene-counting alone. Fix: require significance by both; distrust single-test hits.

Quantitative Thresholds

ThresholdSourceRationale
Permutation N >= 1000empirical p resolution (Gel 2016; Phipson & Smyth 2010)stable empirical p down to ~0.001; the (hits+1)/(N+1) estimator avoids p=0
GAT --num-samples >= 10000GAT practiceneeded for stable small q across many annotations
50% reciprocal overlap (-f 0.5 -r) for CNV/SV concordancefield convention"same event"; one-sided fractions let a giant call swallow a tiny one
Universe = candidate/callable pool, not the genomeLOLA/GAT/regioneR designthe dominant lever; whole-genome universe inflates every enrichment
GREAT basal 5 kb up / 1 kb down, extend to 1 MbMcLean 2010 defaultregulatory-domain model; user-configurable, report the values used
ENCODE blacklist excluded before any testAmemiya 2019high-signal artifact regions otherwise manufacture overlap
FDR control across tracks/termsmultiple-testingmany-track / many-term tests inflate false positives

Common Errors

Error / symptomCauseSolution
Every annotation "significant," p absurdly smallwhole-genome universe for a biased queryset the universe to the callable/accessible pool
bedtools fisher errorinputs not sorted, or missing genome filesort -k1,1 -k2,2n; pass -g genome.txt
Empty/zero overlap in shuffle nullchrom naming mismatch (chr1 vs 1) across query/genome/blacklistharmonize chromosome naming across all files
permTest much more significant than expecteduniform randomizer on clustered regionsuse circularRandomizeRegions
GAT reports no GC effectno --isochores suppliedpass an isochore/GC segmentation of the workspace
GREAT term looks real but is fragilesignificant by only one of the two testsrequire binomial AND hypergeometric
regioneR toGRanges/getGenomeAndMask errorgenome name not recognized / mask chrom mismatchuse a supported genome id or supply explicit GRanges; match chrom naming

References

  • Quinlan AR, Hall IM. 2010. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26:841-842.
  • Heger A, Webber C, Goodson M, Ponting CP, Lunter G. 2013. GAT: a simulation framework for testing the association of genomic intervals. Bioinformatics 29:2046-2048.
  • Gel B, Diez-Villanueva A, Serra E, Buschbeck M, Peinado MA, Malinverni R. 2016. regioneR: an R/Bioconductor package for the association analysis of genomic regions based on permutation tests. Bioinformatics 32:289-291.
  • Sheffield NC, Bock C. 2016. LOLA: enrichment analysis for genomic region sets and regulatory elements in R and Bioconductor. Bioinformatics 32:587-589.
  • McLean CY, Bristor D, Hiller M, Clarke SL, Schaar BT, Lowe CB, Wenger AM, Bejerano G. 2010. GREAT improves functional interpretation of cis-regulatory regions. Nat Biotechnol 28:495-501.
  • Gu Z, Hubschmann D. 2023. rGREAT: an R/Bioconductor package for functional enrichment on genomic regions. Bioinformatics 39:btac745.
  • Amemiya HM, Kundaje A, Boyle AP. 2019. The ENCODE blacklist: identification of problematic regions of the genome. Sci Rep 9:9354.
  • Fulcher BD, Arnatkeviciute A, Fornito A. 2021. Overcoming false-positive gene-category enrichment in the analysis of spatially resolved transcriptomic brain atlas data. Nat Commun 12:2669.
  • Phipson B, Smyth GK. 2010. Permutation P-values should never be zero: calculating exact P-values when permutations are randomly drawn. Stat Appl Genet Mol Biol 9:Article 39.
  • interval-arithmetic - The intersect/shuffle/jaccard/fisher mechanics this skill turns into a test
  • bed-file-basics - BED format, coordinate systems, and sorting the inputs every test requires
  • proximity-operations - Nearest-feature assignment when the question is distance, not overlap enrichment
  • chip-seq/peak-calling - Source of the peak query sets tested for enrichment
  • chip-seq/peak-annotation - Assign enriched peaks to genes/features
  • atac-seq/atac-peak-calling - Source of ATAC peak sets and the accessible-region universe
  • pathway-analysis/go-enrichment - Gene-list ontology enrichment; GREAT is the region-based analog
  • causal-genomics/colocalization-analysis - GWAS/eQTL statistical colocalization (shared causal variant) - a distinct problem from interval overlap
  • data-visualization/genome-tracks - Render the query and annotation tracks behind an enrichment claim

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in genome-intervals/overlap-significance of GPTomics/bioSkills.

  • SKILL.md
  • examples/permutation_overlap.py
  • examples/regioner_permtest.R
  • examples/shuffle_permutation_test.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Intervals Overlap Significance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Intervals Overlap Significance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Intervals Overlap Significance this skillGPTomics/bioSkills1.2k1 repos~5.6kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Genome Intervals Overlap Significance

What does Bio Genome Intervals Overlap Significance do?

Tests whether two genomic interval sets overlap (colocalize) more than expected by chance using a permutation test against a structured-genome null model. Bio Genome Intervals Overlap Significance is an agent skill from GPTomics/bioSkills. Tests whether two genomic interval sets overlap (colocalize) more than expected by chance using a permutation test against a structured-genome null model.

When should I use Bio Genome Intervals Overlap Significance?

Bio Genome Intervals Overlap Significance fits situations like: asking whether peaks/regions are enriched at enhancers/TFBS/features; scoring region-set colocalization; region-set enrichment; comparing CNV/SV concordance.

How do I install Bio Genome Intervals Overlap Significance in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-intervals-overlap-significance -a claude-code`. Or copy the skill folder (genome-intervals/overlap-significance in GPTomics/bioSkills) into .claude/skills/bio-genome-intervals-overlap-significance in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Intervals Overlap Significance in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-intervals-overlap-significance -a codex`. Or copy the skill folder (genome-intervals/overlap-significance in GPTomics/bioSkills) into .agents/skills/bio-genome-intervals-overlap-significance in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Intervals Overlap Significance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-intervals-overlap-significance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-intervals-overlap-significance, .gemini/skills/bio-genome-intervals-overlap-significance, .github/skills/bio-genome-intervals-overlap-significance and .opencode/skills/bio-genome-intervals-overlap-significance in your project.

What does Bio Genome Intervals Overlap Significance need to run?

Going by SKILL.md and its folder, Bio Genome Intervals Overlap Significance needs Python, R and a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Genome Intervals Overlap Significance access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Genome Intervals Overlap Significance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Intervals Overlap Significance use?

Bio Genome Intervals Overlap Significance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Intervals Overlap Significance use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Intervals Overlap Significance?

Skills that share tags, products or a category with Bio Genome Intervals Overlap Significance: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Intervals Overlap Significance?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.