Gwas Database
davila7/claude-code-templates
Query NHGRI-EBI GWAS Catalog for SNP-trait associations. An agent skill from davila7/claude-code-templates.
Scans genomes for natural selection with SFS tests (Tajima's D, Fay & Wu H, Zeng E, SweepFinder2 CLR), haplotype tests (iHS, nSL, XP-EHH, Rsb, H12), and differentiation (FST, PBS) using…
$ npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-population-genetics-selection-statistics --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/population-genetics/selection-statistics .claude/skills/bio-population-genetics-selection-statistics && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-population-genetics-selection-statistics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/population-genetics/selection-statistics into .claude/skills/bio-population-genetics-selection-statistics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-population-genetics-selection-statistics", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/population-genetics/selection-statisticsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-population-genetics-selection-statistics --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/population-genetics/selection-statistics .agents/skills/bio-population-genetics-selection-statistics && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-population-genetics-selection-statistics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/population-genetics/selection-statistics into .agents/skills/bio-population-genetics-selection-statistics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-population-genetics-selection-statistics", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-population-genetics-selection-statistics --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/population-genetics/selection-statistics .cursor/skills/bio-population-genetics-selection-statistics && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-population-genetics-selection-statistics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/population-genetics/selection-statistics into .cursor/skills/bio-population-genetics-selection-statistics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-population-genetics-selection-statistics", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path population-genetics/selection-statistics--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-population-genetics-selection-statistics --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/population-genetics/selection-statistics .gemini/skills/bio-population-genetics-selection-statistics && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-population-genetics-selection-statistics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/population-genetics/selection-statistics into .gemini/skills/bio-population-genetics-selection-statistics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-population-genetics-selection-statistics", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-population-genetics-selection-statisticsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/population-genetics/selection-statistics .github/skills/bio-population-genetics-selection-statistics && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-population-genetics-selection-statistics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/population-genetics/selection-statistics into .github/skills/bio-population-genetics-selection-statistics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-population-genetics-selection-statistics", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-population-genetics-selection-statistics --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/population-genetics/selection-statistics .opencode/skills/bio-population-genetics-selection-statistics && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-population-genetics-selection-statistics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/population-genetics/selection-statistics into .opencode/skills/bio-population-genetics-selection-statistics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-population-genetics-selection-statistics", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-population-genetics-selection-statisticsScans genomes for natural selection with SFS tests (Tajima's D, Fay & Wu H, Zeng E, SweepFinder2 CLR), haplotype tests (iHS, nSL, XP-EHH, Rsb, H12), and differentiation (FST, PBS) using…
Bio Population Genetics Selection Statistics is an agent skill from GPTomics/bioSkills. Scans genomes for natural selection with SFS tests (Tajima's D, Fay & Wu H, Zeng E, SweepFinder2 CLR), haplotype tests (iHS, nSL, XP-EHH, Rsb, H12), and differentiation (FST, PBS) using scikit-allel, selscan, and SweepFinder2. No single statistic separates selection from demography at one locus, so the deliverable is empirical genome-wide outliers plus multiple orthogonal signals, not an absolute cutoff. iHS detects incomplete sweeps and collapses to zero at fixation while XP-EHH catches fixed sweeps; iHS/nSL…
Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/selection_scan.py` and `usage-guide.md`).
It sits in Data & Analytics, covering Statistics and Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Population Genetics Selection Statistics loads about 5.6k tokens when it runs. Until then it costs about 244 tokens; SKILL.md has 2,216 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,216 words, ~5,583 tokens.
.claude/skills/bio-population-genetics-selection-statistics/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: scikit-allel 1.3+, numpy 1.26+, selscan 2.0+, SweepFinder2 1.0+.
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signatures<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Version traps that change results, not just syntax: allel.standardize_by_allele_count(score, aac, ...) is for iHS and nSL (it bins by derived-allele count), while XP-EHH and iHH12 use plain genome-wide allel.standardize(score); conflating them is a real bug. allel.nsl(h) takes no pos/map_pos (nSL is map-free by construction) whereas allel.ihs(h, pos, map_pos=...) distorts without a genetic map. selscan 2.0 adds --unphased for multilocus genotypes; selscan 1.x requires phased haplotypes. The single source of truth for versions is this block, not headings.
"Scan my population for signatures of natural selection" -> Contrast each locus against the genome-wide neutral expectation that already carries the demographic history, using statistics that read complementary features of a sweep.
allel.ihs(), allel.xpehh(), allel.nsl(), allel.garud_h(), allel.windowed_tajima_d(), allel.hudson_fst() (scikit-allel)selscan --ihs|--xpehh|--nsl then norm for standardization; SweepFinder2 -lrb for CLR with a background-selection mapScope: selection-scan DESIGN - SFS neutrality tests, haplotype/EHH sweep statistics, differentiation outliers (FST, PBS), CLR scans, and polygenic Qx, with standardization and demography-aware outlier interpretation. PCA/ADMIXTURE population assignment routes to population-structure; LD/EHH mechanics to linkage-disequilibrium; the scikit-allel array-API mechanics (signatures, accessibility masks, data loading) to scikit-allel-analysis; phasing to phasing-imputation/haplotype-phasing; dN/dS and McDonald-Kreitman to comparative-genomics/positive-selection.
| Family | Statistic / tool | Citation | Reads / role | Phasing |
|---|---|---|---|---|
| SFS | Tajima's D (allel.tajima_d) | Tajima 1989 | theta_pi - theta_W; rare-vs-intermediate variant skew | no |
| SFS | Fay & Wu H, Zeng E | Fay & Wu 2000; Zeng 2006 | high-frequency-DERIVED excess; needs polarization | no |
| SFS-CLR | SweepFinder2 (-f/-l/-lr/-lrb) | DeGiorgio 2016 | local SFS vs empirical background; localizes target; -lrb deconfounds BGS | no |
| SFS-CLR | SweeD, OmegaPlus | Pavlidis 2013; Alachiotis 2012 | parallel CLR; LD-shoulder omega | no |
| Haplotype | iHS (allel.ihs, selscan --ihs) | Voight 2006 | INCOMPLETE sweeps; standardize by DAF bin | YES |
| Haplotype | nSL (allel.nsl, selscan --nsl) | Ferrer-Admetlla 2014 | incomplete hard+soft; map-free | YES |
| Haplotype | XP-EHH (allel.xpehh, selscan --xpehh) | Sabeti 2007 | FIXED/near-fixed sweeps; genome-wide z-score | YES |
| Haplotype | Rsb | Tang 2007 | XP-EHH niche; iES ratio | YES |
| Haplotype | H12 / H2H1 (allel.garud_h) | Garud 2015 | SOFT sweeps; hard-vs-soft tendency | YES |
| Differentiation | FST (allel.hudson_fst, allel.weir_cockerham_fst) | Weir & Cockerham 1984; Bhatia 2013 | local-adaptation divergence | no |
| Differentiation | PBS | Yi 2010 | which branch the divergence is on (3 pops) | no |
| Polygenic | Qx | Berg & Coop 2014 | over-dispersed polygenic scores across pops | no |
| Scenario | Use | Why |
|---|---|---|
| Localized recent sweep, ongoing/incomplete, one population, genetic map available | iHS | reads the frequency contrast between long derived and short ancestral haplotypes |
| Same but no reliable genetic map | nSL | integrates over segregating-site count; map-free, robust to recombination-rate variation |
| Sweep complete/fixed in one of two populations | XP-EHH or Rsb | maximum power exactly where iHS is blind (no within-population contrast remains) |
| Selection on standing variation (soft sweep) | H12 / H2H1 | collapses co-rising haplotypes; hard-sweep stats lose power |
| Localize the target with a demographic/BGS model | SweepFinder2 CLR (-lrb) or SweeD; OmegaPlus | empirical background SFS as null; B-value map deconfounds background selection |
| Differentiation outliers, unequal sample sizes | FST Hudson estimator + PBS for 3 pops | Hudson is robust to unequal n; PBS localizes the branch |
| Quick SFS reconnaissance | windowed Tajima's D + pi + Fay & Wu H (substitution-model polarization) | cheap, but never an absolute cutoff |
| Polygenic trait shift across populations | Qx with within-family GWAS effect sizes only | sweep scans are blind to coordinated tiny shifts; standard GWAS betas carry stratification bias |
| Ancient/recurrent coding selection across species | hand off to comparative-genomics/positive-selection | dN/dS and MK are a different timescale, not a within-population sweep |
Goal: Quantify allele-frequency divergence between populations and flag local-adaptation candidates without being fooled by rare variants or unequal samples.
Approach: Count alleles per subpopulation, compute the Hudson estimator per SNP, and report mean FST as a ratio-of-averages over windows after an MAF filter (rare variants deflate FST, Bhatia 2013).
import allel
import numpy as np
callset = allel.read_vcf('data.vcf.gz')
gt = allel.GenotypeArray(callset['calldata/GT'])
pos = callset['variants/POS']
subpops = {'pop1': [0, 1, 2, 3, 4], 'pop2': [5, 6, 7, 8, 9]}
ac_subpops = gt.count_alleles_subpops(subpops)
ac1, ac2 = ac_subpops['pop1'], ac_subpops['pop2']
# MAF filter first: rare variants systematically deflate FST (Bhatia 2013).
maf = np.minimum(ac1.to_frequencies()[:, 1], ac2.to_frequencies()[:, 1])
keep = (maf > 0.05) | (1 - maf > 0.05)
# Hudson estimator: robust to unequal sample sizes. Mean FST = ratio-of-averages, never mean of per-SNP ratios.
num, den = allel.hudson_fst(ac1[keep], ac2[keep])
fst_mean = np.nansum(num) / np.nansum(den)
# Windowed scan for outlier localization.
fst_win, windows, n_snps = allel.windowed_hudson_fst(pos[keep], ac1[keep], ac2[keep], size=100000, step=50000)
outlier = fst_win > np.nanpercentile(fst_win, 99)PBS (Population Branch Statistic, Yi 2010) needs three populations: convert three pairwise FST values to branch lengths via T = -log(1 - FST) and isolate the focal branch as PBS = (T_12 + T_13 - T_23) / 2. A long focal branch in a small/bottlenecked population can be drift, not selection.
Goal: Detect ongoing sweeps from extended haplotype homozygosity around derived core alleles.
Approach: Filter to segregating biallelic SNPs, compute the raw integrated-EHH score, then standardize WITHIN derived-allele-frequency bins (the raw score depends strongly on derived frequency and is uninterpretable unbinned).
import allel
import numpy as np
h = gt.to_haplotypes()
ac = h.count_alleles()
flt = (ac[:, 0] > 1) & (ac[:, 1] > 1)
h_flt, pos_flt, ac_flt = h.compress(flt, axis=0), pos[flt], ac.compress(flt, axis=0)
# iHS needs a genetic map (map_pos) where recombination rate varies; physical distance distorts iHH.
ihs_raw = allel.ihs(h_flt, pos_flt, min_maf=0.05, include_edges=True)
# Standardize WITHIN derived-allele-count bins (NOT genome-wide) - the most common iHS bug.
ihs_std, _ = allel.standardize_by_allele_count(ihs_raw, ac_flt[:, 1])
ihs_outlier = np.abs(ihs_std) > 2
# nSL: map-free (no pos/map_pos), robust to recombination-rate variation; still bin-standardized.
nsl_raw = allel.nsl(h_flt)
nsl_std, _ = allel.standardize_by_allele_count(nsl_raw, ac_flt[:, 1])A null iHS does NOT mean no selection: once the derived allele fixes, the ancestral haplotype is gone and iHS collapses to zero. Switch to XP-EHH/Rsb for completed sweeps. Report the windowed PROPORTION of |iHS|>2 SNPs, not single noisy hits.
Goal: Catch fixed or near-fixed sweeps in one of two populations, where within-population haplotype tests are blind.
Approach: Compute the cross-population integrated-EHH ratio on shared positions, then apply a plain GENOME-WIDE z-score (XP-EHH is not strongly correlated with derived frequency, so frequency-bin standardization is the wrong rule).
import allel
import numpy as np
h1 = h_flt.take(pop1_hap_idx, axis=1)
h2 = h_flt.take(pop2_hap_idx, axis=1)
xpehh_raw = allel.xpehh(h1, h2, pos_flt, include_edges=True)
# Genome-wide z-score for XP-EHH/iHH12 - NOT standardize_by_allele_count (that is for iHS/nSL).
xpehh_std = allel.standardize(xpehh_raw)
completed_sweep = np.abs(xpehh_std) > 2A shared ancestral sweep at the same locus in BOTH populations cancels in the contrast (false negative). Differential phasing quality between the two populations systematically biases the difference statistic.
Goal: Recover sweeps on standing variation that iHS and CLR miss because no single long haplotype dominates.
Approach: Compute Garud's H over haplotype-frequency spectra; H12 collapses the two most common haplotypes into one, and H2/H1 tends (not classifies) toward hard-vs-soft.
h1, h12, h123, h2_h1 = allel.garud_h(h_flt)
h12_win = allel.moving_garud_h(h_flt, size=100) # SNP-count windows, not bp# selscan haplotype scans (phased VCF + genetic map), then norm for standardization.
selscan --ihs --vcf phased.vcf --map genetic.map --out scan # within-pop incomplete sweeps
selscan --xpehh --vcf pop1.vcf --vcf-ref pop2.vcf --map genetic.map --out xp # completed sweeps
selscan --nsl --vcf phased.vcf --out nsl_scan # no map; map-free
selscan --ihs --unphased --vcf unphased.vcf --map genetic.map --out scan_unphased # selscan 2.0 multilocus-genotype
# norm bins iHS/nSL by derived-allele frequency (--bins default 100); XP-EHH/iHH12 get a genome-wide z-score.
norm --ihs --files scan.ihs.out --bins 100
norm --xpehh --files xp.xpehh.out
# SweepFinder2 CLR: build a genome-WIDE background SFS, then scan with recombination + B-value (BGS) map.
SweepFinder2 -f genome.freq genome.spect # empirical background spectrum
# -lrb takes N1 (current ingroup Ne), N2 (ancestral Ne), T (divergence time in generations) before OutFile.
SweepFinder2 -lrb 2000 region.freq genome.spect region.rec bvalue.map "$N1" "$N2" "$T" out.clr # G=2000 grid; -lrb deconfounds BGSDerived-allele tests (Fay & Wu H, Zeng E, unfolded SFS, the iHS sign) are POLARIZATION traps: one mispolarized CpG site masquerades as a high-frequency-derived variant and fakes a strongly negative H. Polarize with a probabilistic substitution model and two or more outgroups (Hernandez 2007), never a single chimp allele. Conservation scores (phyloP, GERP, phastCons) measure CONSTRAINT (purifying selection), the opposite sign of a recent positive sweep, and must not be read as "under selection." Polygenic Qx (Berg & Coop 2014) inherits every stratification bias in the GWAS effect sizes it uses: the celebrated European height-selection signal was shown to be a stratification artifact (Sohail 2019; Berg 2019), so use within-family/sibling effect estimates and treat standard-GWAS-beta Qx as provisional.
Trigger: flagging windows on |D|>2 as selection. Mechanism: growth/bottleneck/structure produce the same sign as a sweep/balancing selection (non-identifiable at one locus). Symptom: genome-wide false positives that are pure demography. Fix: use the genome-wide empirical distribution and intersect orthogonal statistics; report D relative to the genomic background.
Trigger: reporting raw iHS, or applying a genome-wide z-score to iHS. Mechanism: raw iHH ratio depends strongly on derived-allele frequency. Symptom: uninterpretable scores, false tails at particular frequencies. Fix: standardize_by_allele_count (DAF bins) for iHS/nSL; reserve standardize (genome-wide) for XP-EHH/iHH12.
Trigger: concluding neutrality from absent iHS signal. Mechanism: a completed sweep fixes the derived allele, erasing the within-population contrast iHS needs. Symptom: real fixed sweeps missed. Fix: add XP-EHH/Rsb against a second population for completed sweeps.
Trigger: deriving ancestral state from one chimp allele for H/E/unfolded SFS. Mechanism: recurrent mutation at CpG/hypermutable sites biases the single-allele ancestral estimator. Symptom: spurious strongly-negative Fay & Wu H (fake sweep). Fix: substitution-model polarization with >=2 outgroups (Hernandez 2007; est-sfs).
Trigger: reading an FST or CLR peak in a low-recombination region as positive selection. Mechanism: purifying selection against linked deleterious variants reduces local diversity. Symptom: FST/PBS/CLR peaks with no positive selection. Fix: supply SweepFinder2 a B-value map (-lrb); control for recombination rate; treat low-recombination outliers with suspicion.
Trigger: EHH statistics on poorly phased data. Mechanism: a switch error truncates long haplotypes and deflates EHH. Symptom: iHS/nSL/XP-EHH biased toward null; XP-EHH biased by differential phasing between populations. Fix: read-backed/high-quality phasing (see phasing-imputation/haplotype-phasing), or selscan 2.0 --unphased multilocus-genotype mode.
Trigger: running Qx on standard meta-analysis effect sizes. Mechanism: residual population stratification correlates effect sizes with ancestry axes. Symptom: spurious polygenic-adaptation clines (the height story). Fix: within-family/sibling effect sizes; treat PC-corrected meta-analysis betas as insufficient.
| Quantity | Typical value | Rationale |
|---|---|---|
| iHS / nSL flag | |standardized score| > 2 | ~top 2.5% two-sided of N(0,1); an empirical convention, NOT a calibrated p-value (Voight 2006) |
| Empirical outlier tail | top 1% (top 0.1% for stringency) | genome-wide bulk absorbs demography; percentile is a sensitivity/specificity tradeoff, not an alpha |
| iHS/nSL standardization | within derived-allele-frequency bins (selscan --bins 100) | raw score is frequency-dependent; binning makes scores comparable |
| XP-EHH/iHH12 standardization | genome-wide z-score | not frequency-correlated; bin-standardizing inverts the rule |
| Window size | 10-100 kb, 50% step (scale to LD decay) | wide enough for stable SFS estimates, narrow enough to localize |
| Haplotype-stat MAF | drop core alleles MAF < ~0.05 | EHH on singletons is undefined/uninformative (min_maf=0.05) |
| SweepFinder2 grid G | finer than expected sweep width (~1-2 kb in humans) | coarse grids step over narrow sweeps; significance from demographic-model simulation |
| FST mean | ratio-of-averages, MAF-filtered, Hudson estimator | rare variants deflate FST; arithmetic mean of per-SNP ratios is biased (Bhatia 2013) |
Thresholds are conventions, not laws - inspect distributions and calibrate against a demographic-model simulation or the genome-wide empirical tail before quoting any number.
| Error / symptom | Cause | Solution |
|---|---|---|
| "D < -2 = sweep" | absolute cutoff ignoring demography | use the genome-wide empirical tail; D sign never identifies cause |
| iHS scores incomparable across SNPs | unstandardized or genome-wide-standardized iHS | standardize_by_allele_count by DAF bin |
| XP-EHH never standardized | applying the iHS rule or no standardization | allel.standardize(xpehh) genome-wide z-score |
| "no iHS so no selection" | completed sweep has no within-pop contrast | add XP-EHH/Rsb against a second population |
| spurious negative Fay & Wu H | single-outgroup polarization at CpG sites | substitution-model polarization, >=2 outgroups (Hernandez 2007) |
| FST/CLR peak misread as positive selection | background selection in low-recombination region | B-value map (-lrb); recombination control |
allel.nsl(h, pos) TypeError | nSL is map-free; takes no pos/map_pos | call allel.nsl(h) |
| FST tail driven by rare variants | no MAF filter; mean of per-SNP ratios | MAF filter + Hudson ratio-of-averages |
| Qx height-style false signal | stratified GWAS effect sizes | within-family/sibling effect estimates |
| "high phyloP = under selection" | conflating constraint with positive selection | phyloP/GERP measure purifying constraint (opposite sign) |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in population-genetics/selection-statistics of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Population Genetics Selection Statistics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Population Genetics Selection Statistics this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.6k | Automated safety check: Pass | MIT | |
| Gwas Databasedavila7/claude-code-templates | 33k | 10 repos | ~5k | Automated safety check: Pass | MIT | |
| Bio Spatial Transcriptomics Spatial StatisticsFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | 1 repos | ~1.8k | Automated safety check: Pass | None | |
| Scikit Bioaipoch/medical-research-skills | 1.9k | — | ~1.4k | Automated safety check: Pass | MIT | |
| PyDESeq2 Differential Expressiondavila7/claude-code-templates | 33k | 11 repos | ~4k | Automated safety check: Pass | MIT | |
| Bio Sequence StatisticsFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | 1 repos | ~2.4k | Automated safety check: Pass | None |
davila7/claude-code-templates
Query NHGRI-EBI GWAS Catalog for SNP-trait associations. An agent skill from davila7/claude-code-templates.
FreedomIntelligence/OpenClaw-Medical-Skills
Compute spatial statistics for spatial transcriptomics data using Squidpy.
aipoch/medical-research-skills
A Python bioinformatics toolkit for sequence, phylogeny, and microbiome/community-ecology analysis; use it when you need to compute diversity/ordination/statistics from biological data and standard…
davila7/claude-code-templates
Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.
FreedomIntelligence/OpenClaw-Medical-Skills
Calculate sequence statistics (N50, length distribution, GC content, summary reports) using Biopython.
bioMate-AI/biomate-bioconductor-kb
This package implements a suite of methods to preprocess data from PTR-TOF-MS instruments (HDF5 format) and generates the 'sample by features' table of peak intensities in addition to the sample and…
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Scans genomes for natural selection with SFS tests (Tajima's D, Fay & Wu H, Zeng E, SweepFinder2 CLR), haplotype tests (iHS, nSL, XP-EHH, Rsb, H12), and differentiation (FST, PBS) using…. Bio Population Genetics Selection Statistics is an agent skill from GPTomics/bioSkills. Scans genomes for natural selection with SFS tests (Tajima's D, Fay & Wu H, Zeng E, SweepFinder2 CLR), haplotype tests (iHS, nSL, XP-EHH, Rsb, H12), and differentiation (FST, PBS) using scikit-allel, selscan, and SweepFinder2.
Bio Population Genetics Selection Statistics fits situations like: computing selection statistics like FST; scanning for selective sweeps.
Run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a claude-code`. Or copy the skill folder (population-genetics/selection-statistics in GPTomics/bioSkills) into .claude/skills/bio-population-genetics-selection-statistics in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a codex`. Or copy the skill folder (population-genetics/selection-statistics in GPTomics/bioSkills) into .agents/skills/bio-population-genetics-selection-statistics in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-selection-statistics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-population-genetics-selection-statistics, .gemini/skills/bio-population-genetics-selection-statistics, .github/skills/bio-population-genetics-selection-statistics and .opencode/skills/bio-population-genetics-selection-statistics in your project.
Going by SKILL.md and its folder, Bio Population Genetics Selection Statistics needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Population Genetics Selection Statistics is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Population Genetics Selection Statistics: Gwas Database (davila7/claude-code-templates, 33k stars), Bio Spatial Transcriptomics Spatial Statistics (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Scikit Bio (aipoch/medical-research-skills, 1.9k stars) and PyDESeq2 Differential Expression (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.