Agent skill

Bio Population Genetics Population Structure

by GPTomics in GPTomics/bioSkills

Infers and describes population structure with PCA (plink2 --pca, smartpca/EIGENSOFT, FlashPCA2), model-based clustering (ADMIXTURE, fastSTRUCTURE), FST estimators (Weir-Cockerham vs Hudson), and…

MITAuto-check passedData & Analytics

Install Bio Population Genetics Population Structure

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-population-genetics-population-structure -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-population-genetics-population-structure --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/population-genetics/population-structure .claude/skills/bio-population-genetics-population-structure && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-population-genetics-population-structure
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.6k tokens
SKILL.md length
2,332 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Infers and describes population structure with PCA (plink2 --pca, smartpca/EIGENSOFT, FlashPCA2), model-based clustering (ADMIXTURE, fastSTRUCTURE), FST estimators (Weir-Cockerham vs Hudson), and…

  • Works in 4 steps: The output is a deterministic function… → PCA does not find ancestry, it finds the… → ADMIXTURE Q-values are not ancestry… → …
  • F-statistics on QCd genotypes
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 10 more sections
  • Runs Shell scripts from its folder; calls pip

What it does

Bio Population Genetics Population Structure is an agent skill from GPTomics/bioSkills. Infers and describes population structure with PCA (plink2 --pca, smartpca/EIGENSOFT, FlashPCA2), model-based clustering (ADMIXTURE, fastSTRUCTURE), FST estimators (Weir-Cockerham vs Hudson), and f-statistics (f3/f4/D via AdmixTools/admixr), plus Python plotting of PCs and Q barplots. Every output is a model-conditioned description of variance, not truth: PCs conflate ancestry with LD/inversions/relatedness/batch, ADMIXTURE Q-values are panel- and K-dependent artifacts, and CV-minimum K is a guide not the true…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/structure_analysis.sh` and `usage-guide.md`).

It sits in Data & Analytics, covering Statistics and Data visualization. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • F-statistics on QCd genotypes
  • Tasks that involve Statistics
  • Tasks that involve Data visualization

Example prompts

  • “Use the bio-population-genetics-population-structure skill to infer and describes population structure with PCA (plink2 --pca, smartpca/EIGENSOFT…”
  • “/bio-population-genetics-population-structure”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. The output is a deterministic function of three silent choices: which samples are in the panel, which SNPs survive…
  2. PCA does not find ancestry, it finds the directions of greatest genotype covariance, which conflate ancestry with LD blocks, inversions…
  3. ADMIXTURE Q-values are not ancestry fractions but maximum-likelihood weights on K abstract allele-frequency vectors that exist only…
  4. Any interpreted statistic needs its own uncertainty: FST combines across SNPs as a ratio of averages (never an average of per-SNP FST)…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Population Genetics Population Structure loads about 5.6k tokens when it runs. Until then it costs about 258 tokens; SKILL.md has 2,332 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~258
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,332 words, ~5,567 tokens.

Download SKILL.mdSave it as .claude/skills/bio-population-genetics-population-structure/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-population-genetics-population-structure
description
Infers and describes population structure with PCA (plink2 --pca, smartpca/EIGENSOFT, FlashPCA2), model-based clustering (ADMIXTURE, fastSTRUCTURE), FST estimators (Weir-Cockerham vs Hudson), and f-statistics (f3/f4/D via AdmixTools/admixr), plus Python plotting of PCs and Q barplots. Every output is a model-conditioned description of variance, not truth: PCs conflate ancestry with LD/inversions/relatedness/batch, ADMIXTURE Q-values are panel- and K-dependent artifacts, and CV-minimum K is a guide not the true population count. FST must combine SNPs as a ratio of averages (sum numerators / sum denominators), never an average of per-SNP FST; negative per-SNP FST is normal and must not be clamped. f3/f4/D need a block jackknife or the significance is fake. Use when running PCA, ADMIXTURE, FST, or f-statistics on QC'd genotypes. For QC and KING relatedness see plink-basics; for LD pruning see linkage-disequilibrium; for array-based Python pipelines see scikit-allel-analysis.
tool_type
mixed
primary_tool
plink2

Version Compatibility

Reference examples tested with: PLINK 2.0 (alpha 6+), ADMIXTURE 1.3+, EIGENSOFT 7.2+, scikit-allel 1.3+, numpy 1.26+, pandas 2.2+, matplotlib 3.8+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Version traps that change results, not just syntax: ADMIXTURE is a standalone CLI (admixture --cv input.bed K), never an R or Python package, and --cv defaults to 5-fold. plink2 --pca builds on the variance-standardized relationship matrix from --make-rel/--make-grm and has no Tracy-Widom test; smartpca does. plink2 --make-king gives kinship (cutoff 0.0884 = second-degree), not the deprecated PI_HAT from PLINK 1.9 --genome. f-statistics live in AdmixTools (CLI) wrapped by admixr (R), not in plink. The single source of truth for versions is this block, not headings.

Population Structure

"Analyze the population structure in my genotypes" -> Project genotype covariance into continuous axes or discrete clusters, after pruning the artifacts the model would otherwise mistake for ancestry, and attach the uncertainty machinery any interpreted statistic requires.

  • CLI: plink2 --pca 20 approx (top eigenvectors of the relationship matrix; LD-pruned, relatives removed first)
  • CLI: admixture --cv data_pruned.bed 3 (maximum-likelihood mixing weights for K abstract clusters)
  • Python: allel.hudson_fst(ac1, ac2) then num.sum()/den.sum() (FST as a ratio of averages)

Scope: PCA, model-based clustering (ADMIXTURE/fastSTRUCTURE), FST estimators, and f-/D-statistics, with Python plotting. QC and KING relatedness route to plink-basics; LD pruning to linkage-disequilibrium; array-scale Python diversity/FST windows to scikit-allel-analysis; phased haplotype work to phasing-imputation/haplotype-phasing; introgression detection to comparative-genomics/introgression-detection.

The Single Most Important Insight -- every structure method returns a model-conditioned description of variance, not truth

  1. The output is a deterministic function of three silent choices: which samples are in the panel, which SNPs survive ascertainment/QC/LD-pruning, and which model is imposed (continuous PCs vs K discrete clusters vs a tree-with-admixture); change any one and the "answer" changes.
  2. PCA does not find ancestry, it finds the directions of greatest genotype covariance, which conflate ancestry with LD blocks, inversions, relatedness, batch, and differential missingness, so the mandatory work is pruning those out before reading axes.
  3. ADMIXTURE Q-values are not ancestry fractions but maximum-likelihood weights on K abstract allele-frequency vectors that exist only because K of them were requested, and the CV-minimum K is a prediction-accuracy guide, not the true number of populations.
  4. Any interpreted statistic needs its own uncertainty: FST combines across SNPs as a ratio of averages (never an average of per-SNP FST), negative per-SNP FST is kept not clamped, and every f3/f4/D needs a block-jackknife standard error or the significance is fabricated.

Tool Taxonomy

MethodCitationMechanism / roleWhen
plink2 --pcaPatterson 2006; Price 2006Eigenvectors of the variance-standardized relationship matrix; approx for large NStratification covariates, gross structure, QC outliers
smartpca (EIGENSOFT)Patterson 2006PCA plus Tracy-Widom significance, outlier removal, lsqproject projectionRigorous PCA, aDNA projection, per-PC p-values
FlashPCA2Abraham 2017Randomized PCA for biobank N (>100k)Very large cohorts
ADMIXTUREAlexander 2009Fast ML point estimate of Q (ancestry weights) and P (cluster frequencies)Genome-wide discrete ancestry proportions
fastSTRUCTURERaj 2014Variational Bayes clustering with chooseK.py K guidanceFast K exploration
Hudson FSTHudson 1992; Bhatia 2013Per-SNP heterozygosity estimator; ratio of averagesPairwise differentiation, unequal sample sizes, rare variants / SNP-array ascertainment
Weir-Cockerham FSTWeir & Cockerham 1984ANOVA estimator (a/(a+b+c)); ratio of averagesClassic variance-partition framing, balanced n
f3 / f4 / DPatterson 2012; Durand 2011Drift-distance tests of admixture and tree-ness; block jackknifeAdmixture detection, gene-flow tests
TreeMixPickrell 2012ML tree plus migration edges from frequency covarianceTree + migration hypotheses
admixr / AdmixToolsPetr 2019; Patterson 2012Reproducible R wrappers for qp3Pop/qpDstat/qpAdm/qpGraphf-statistics and admixture graphs

Decision Tree by Scenario

ScenarioRecommendedWhy
Stratification covariates for GWASplink2 --pca (approx >5000)assumption-light, fast, feeds --glm directly
Per-PC significance, outlier removal, aDNA projectionsmartpcaTracy-Widom test plus lsqproject for shrinkage-robust projection
Biobank-scale PCA (>100k)plink2 --pca approx or FlashPCA2randomized algorithms scale; exact PCA does not
Discrete ancestry proportionsADMIXTURE over a K span (--cv as guide)present a span, never the CV argmin alone
Fast K explorationfastSTRUCTURE + chooseK.pyvariational, very fast
Pairwise differentiation, unequal n or rare variantsHudson FST, ratio of averagesBhatia 2013; WC is sensitive to n, population count, and rare variants
"Is population C admixed?"f3(C; A,B) (qp3Pop / admixr)f3 < 0 proves admixture; block-jackknife Z
"Is there gene flow / introgression?"D / f4 (qpDstat, ABBA-BABA)`
Clinal / spatial structureEEMS, Mantelclines are isolation-by-distance, not discrete demes
Relatedness QC before everythingplink2 --king-cutoff (see plink-basics)a relative cluster grabs a spurious PC

Mandatory Preprocessing Before PCA

Goal: Compute PCs that track ancestry rather than inversions, relatedness, or batch.

Approach: LD-prune, exclude long-range-LD/inversion regions by coordinate, remove relatives before computing axes, drop very-low-MAF variants, then run PCA on the survivors.

bash
# LD-prune so a single dense block cannot dominate a PC (route detail to linkage-disequilibrium).
plink2 --bfile data --indep-pairwise 50 5 0.1 --out prune

# Exclude long-range-LD regions and inversions that survive pruning and create karyotype PCs.
# range_lrld.txt holds MHC chr6:25-35 Mb, 8p23.1, 17q21.31, and LCT/2q21 in plink --exclude range format.
plink2 --bfile data --extract prune.prune.in --exclude range range_lrld.txt --maf 0.01 \
    --make-bed --out data_for_pca

# PCA on the LD-pruned, inversion-stripped, relatedness-pruned set. approx is near-required above ~50k.
plink2 --bfile data_for_pca --pca 20 approx --out pca
# Outputs: pca.eigenvec (FID IID PC1..PCn), pca.eigenval (variance per PC).

Relatives must be removed BEFORE computing axes (a cluster of cousins forms its own high-covariance PC); compute the KING cutoff with plink2 --king-cutoff 0.0884 from plink-basics, build PCs on the unrelated set, then project relatives back. plink2 has no Tracy-Widom test: feed the eigenvalues to smartpca twstats or read a scree elbow to decide which PCs are real.

Rigorous PCA and Projection (smartpca)

Goal: Attach per-PC significance and project new/ancient samples without shrinkage artifacts.

Approach: Run smartpca with outlier iterations and Tracy-Widom output for the reference build, and use lsqproject: YES to place additional samples robustly to missingness.

bash
# smartpca parameter file (key params verified against the EIGENSOFT POPGEN README):
#   numoutevec: 20            # PCs to output (default 10)
#   numoutlieriter: 5         # outlier-removal iterations (default 5; 0 disables)
#   outliersigmathresh: 6.0   # SD threshold for outlier removal (default 6.0)
#   lsqproject: YES           # least-squares projection, robust to missing data (aDNA standard)
#   poplistname: ref_pops.txt # which populations build the axes (others are projected)
#   altnormstyle: NO          # NO = Price 2006 EIGENSTRAT normalization; YES = Patterson 2006
smartpca -p smartpca.par
# Tracy-Widom test on the eigenvalues: only PCs with p < ~0.05 plus a scree elbow are interpretable.
twstats -t twtable -i out.eval -o out.tw

Projected scores shrink toward the origin, worse with a small reference panel (<~5000) and more missing data, so an ancient sample plotting "between" two clusters may be shrunk, not admixed; lsqproject is the standard fix. plink2 projects via --pca allele-wts then --score on the .eigenvec.allele weights, but does not correct shrinkage.

ADMIXTURE Over a K Span

Goal: Estimate discrete ancestry proportions while treating K as a model-selection choice, not a discovery.

Approach: Run ADMIXTURE on LD-pruned data across a span of K with cross-validation, plot CV error as a guide, and check Q stability across seeds before interpreting any single K.

bash
# admixture is a standalone CLI; --cv defaults to 5-fold. -jN threads, -BN bootstrap SEs.
for K in $(seq 2 8); do
    admixture --cv -j4 data_pruned.bed "$K" 2>&1 | tee "log_K${K}.out"
done
# Outputs per K: data_pruned.K.Q (N x K ancestry weights) and data_pruned.K.P (cluster frequencies).
# CV error prints as: CV error (K=3): 0.512 -- a guide, never "the true number of populations".
grep -h "CV error" log_K*.out

Supervised mode (admixture --supervised data_pruned.bed K, reading data_pruned.pop with one label per individual, blank for unknowns) fixes labeled individuals to their population but assumes the reference populations are themselves unadmixed. Replicate Q at the same K can land in different local optima; align cluster labels across runs with CLUMPP or pong before averaging or plotting, and treat unstable Q as a sign of mis-specified K, not noise to smooth away.

FST as a Ratio of Averages

Goal: Estimate pairwise differentiation without the average-of-ratios bias.

Approach: Compute per-SNP Hudson numerators and denominators, keep negative numerators, then divide summed numerators by summed denominators across SNPs.

python
import allel
import numpy as np

# ac1, ac2 are AlleleCountsArrays for the two populations at the same SNPs (see scikit-allel-analysis).
num, den = allel.hudson_fst(ac1, ac2)   # per-SNP Hudson numerator and denominator (Bhatia 2013)
fst = num.sum() / den.sum()             # RATIO OF AVERAGES across SNPs; never np.mean(num/den)
# Negative per-SNP numerators are normal sampling behavior near FST=0 and stay in the sum.
# Use Weir-Cockerham only with balanced n: allel.weir_cockerham_fst returns a, b, c variance components.
a, b, c = allel.weir_cockerham_fst(genotype_array, subpops)
fst_wc = np.nansum(a) / np.nansum(a + b + c)   # still a ratio of averages

The Hudson estimator is preferred under sample-size asymmetry because Weir-Cockerham's finite-sample correction makes it sensitive to n and to the number of populations; the two can disagree enough to cross a Wright differentiation band. SNP-array ascertainment compresses and warps FST relative to whole-genome sequencing, so cross-study comparisons require matched ascertainment.

f-statistics with a Block Jackknife (admixr)

Goal: Test admixture and gene flow with honest standard errors.

Approach: Run f3/f4/D through AdmixTools (wrapped by admixr) so each statistic carries a block-jackknife SE, and read Z, never a raw point estimate.

r
library(admixr)
# admixr wraps AdmixTools (Petr 2019); each call returns the statistic with a block-jackknife SE and Z.
data <- eigenstrat('prefix')
res_f3 <- f3(A = 'PopA', B = 'PopB', C = 'PopC', data = data)   # f3(C; A,B) < 0 with Z < -3 proves C admixed
res_d  <- d(W = 'PopW', X = 'PopX', Y = 'PopY', Z = 'PopZ', data = data)  # |Z| > 3 indicates gene flow

A significantly negative f3 proves the target is admixed (no tree produces a negative f3), but a non-negative f3 is inconclusive, not proof of a clean tree: post-admixture drift in the target itself (a bottleneck after the admixture event), or heavily drifted sources, adds a positive term that can mask the negative cross-product even when admixture is real. A nonzero D or f4 is evidence of a tree violation, not specifically recent introgression: symmetric ancient structure mimics the same ABBA/BABA asymmetry (Durand 2011), so separating them needs admixture-LD decay or explicit modeling.

Per-Method Failure Modes

PCA tracks an inversion, not a deme

Trigger: PCA on LD-pruned data with MHC/8p23/17q21.31/LCT regions still in. Mechanism: megabase-long LD in inversions survives --indep-pairwise and dominates a PC. Symptom: a PC loads almost entirely on one chromosome arm and splits samples by karyotype. Fix: --exclude range the long-range-LD and inversion regions by coordinate before PCA.

Relatives grab a principal component

Trigger: computing PCs before relatedness pruning. Mechanism: a cluster of relatives forms a high-covariance bundle. Symptom: a tight outlier cluster on a top PC that is not a real population. Fix: --king-cutoff 0.0884 first, build PCs on the unrelated set, project relatives back.

Show full SKILL.md (925 more words)Show less
Projection shrinkage misread as admixture

Trigger: projecting new/ancient/low-coverage samples onto reference PCs naively. Mechanism: projected scores are biased toward the origin, worse with small panels and missing data. Symptom: a sample plots "between" two clusters and is narrated as admixed. Fix: smartpca lsqproject: YES; do not interpret shrunk scores as intermediacy.

CV-minimum K over-splits

Trigger: picking K at the CV argmin when the curve plateaus or keeps falling. Mechanism: CV error is prediction accuracy, not a population count. Symptom: uninterpretable extra clusters at high K. Fix: run K across a span with at least 10 seeds each (-s), align replicates with pong/CLUMPP, and report the K where CV error plateaus AND Q is seed-stable; if CV and stability disagree, present both and let sampling design plus orthogonal evidence pick the interpreted K.

Label switching corrupts averaged barplots

Trigger: averaging Q across runs or seeds without alignment. Mechanism: cluster labels permute arbitrarily between runs. Symptom: a smeared, meaningless mean barplot. Fix: align labels with CLUMPP or pong before averaging; treat unstable Q as a mis-specification warning.

Average-of-ratios FST

Trigger: combining per-SNP FST as mean(num/den). Mechanism: low-MAF SNPs have tiny denominators and dominate the mean. Symptom: badly biased genome-wide FST. Fix: ratio of averages, num.sum()/den.sum(); never clamp negative per-SNP values first.

f-statistic significance without a jackknife

Trigger: naive SNP-level standard errors for f3/f4/D. Mechanism: LD correlates neighboring SNPs, so per-SNP SEs are far too small. Symptom: everything looks significant. Fix: block-jackknife SE (drop ~5 cM blocks); report Z with |Z| > 3.

Clinal sample forced into discrete clusters

Trigger: running K-cluster ADMIXTURE on an isolation-by-distance continuum. Mechanism: smooth clines have no discrete demes. Symptom: phantom populations and "admixed" intermediates that are really IBD. Fix: describe clinal structure with EEMS / Mantel, not a STRUCTURE barplot.

Quantitative Thresholds

QuantityThresholdSource / rationale
LD pruning for PCA/ADMIXTURE--indep-pairwise 50 5 0.1 (range 0.05-0.2)near-independence so structure is not double-counted (see linkage-disequilibrium)
MAF floor before PCAdrop MAF < ~0.01 (often < 0.05)plink2 docs: very-low-MAF variants destabilize PCA
Relatedness removalKING --king-cutoff 0.0884 (2nd-degree)KING boundaries 0.354/0.177/0.0884/0.0442 = MZ/1st/2nd/3rd
PC significanceTracy-Widom p < 0.05 plus scree elbowPatterson 2006; TW null for the largest eigenvalue
smartpca outlier removaloutliersigmathresh 6.0, numoutlieriter 5EIGENSOFT defaults
ADMIXTURE CVdefault 5-fold; --cv=10 for stabilityAlexander/Shringarpure manual
f3/f4/D significance`Z
FST bands (Wright guideposts)0-0.05 little, 0.05-0.15 moderate, 0.15-0.25 great, >0.25 very greatheuristic only; estimator-, MAF-, ascertainment-dependent
Negative per-SNP FSTkeep (do not clamp)unbiased estimators yield negatives near FST=0 by sampling

Thresholds are conventions, not laws; the FST bands predate SNP arrays, and the estimator, MAF spectrum, ascertainment, and sample size each move FST by a whole band. Verify current best practice before applying numbers blindly.

Common Errors

Error / symptomCauseSolution
import admixture failsADMIXTURE is a CLI, not a packagerun admixture --cv data.bed K on the shell
CV argmin reported as the true Kreading CV error as a population countpresent a K span; CV is a guide, check Q stability
A PC tracks one chromosome arminversion/long-range-LD region left in--exclude range MHC/8p23/17q21.31/LCT before PCA
Outlier cluster is "a new population"relatives not removed before PCA--king-cutoff 0.0884 first, project relatives back
Genome-wide FST is biased highaverage-of-ratios and/or clamped negativesratio of averages; keep negative per-SNP values
WC and Hudson FST disagreeunequal sample sizes between populationsprefer Hudson under sample-size asymmetry (Bhatia 2013)
Everything is f3/D-significantnaive (non-jackknife) standard errorsuse block-jackknife SE; report Z, `
Non-negative f3 read as "no admixture"post-admixture drift in the target (or drifted sources) masks the negative termnon-negative f3 is inconclusive, not a clean tree
Smeared mean Q barplotlabel switching across replicatesalign with CLUMPP/pong before averaging

References

  1. Patterson N, Price AL, Reich D. Population structure and eigenanalysis. PLoS Genetics 2006; 2(12):e190. DOI:10.1371/journal.pgen.0020190.
  2. Price AL, Patterson NJ, Plenge RM, Weinblatt ME, Shadick NA, Reich D. Principal components analysis corrects for stratification in genome-wide association studies. Nature Genetics 2006; 38(8):904-909. DOI:10.1038/ng1847.
  3. Alexander DH, Novembre J, Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Research 2009; 19(9):1655-1664. DOI:10.1101/gr.094052.109.
  4. Raj A, Stephens M, Pritchard JK. fastSTRUCTURE: variational inference of population structure in large SNP data sets. Genetics 2014; 197(2):573-589. DOI:10.1534/genetics.114.164350.
  5. Weir BS, Cockerham CC. Estimating F-statistics for the analysis of population structure. Evolution 1984; 38(6):1358-1370. DOI:10.1111/j.1558-5646.1984.tb05657.x.
  6. Hudson RR, Slatkin M, Maddison WP. Estimation of levels of gene flow from DNA sequence data. Genetics 1992; 132(2):583-589. DOI:10.1093/genetics/132.2.583.
  7. Bhatia G, Patterson N, Sankararaman S, Price AL. Estimating and interpreting FST: the impact of rare variants. Genome Research 2013; 23(9):1514-1521. DOI:10.1101/gr.154831.113.
  8. Patterson N, Moorjani P, Luo Y, Mallick S, Rohland N, Zhan Y, Genschoreck T, Webster T, Reich D. Ancient admixture in human history. Genetics 2012; 192(3):1065-1093. DOI:10.1534/genetics.112.145037.
  9. Durand EY, Patterson N, Reich D, Slatkin M. Testing for ancient admixture between closely related populations. Molecular Biology and Evolution 2011; 28(8):2239-2252. DOI:10.1093/molbev/msr048.
  10. Pickrell JK, Pritchard JK. Inference of population splits and mixtures from genome-wide allele frequency data. PLoS Genetics 2012; 8(11):e1002967. DOI:10.1371/journal.pgen.1002967.
  11. Petr M, Vernot B, Kelso J. admixr - R package for reproducible analyses using ADMIXTOOLS. Bioinformatics 2019; 35(17):3194-3195. DOI:10.1093/bioinformatics/btz030.
  12. Abraham G, Qiu Y, Inouye M. FlashPCA2: principal component analysis of Biobank-scale genotype datasets. Bioinformatics 2017; 33(17):2776-2778. DOI:10.1093/bioinformatics/btx299.
  • plink-basics - QC, KING relatedness pruning, and fileset preparation before structure analysis
  • linkage-disequilibrium - LD pruning the SNP set that PCA and ADMIXTURE require
  • scikit-allel-analysis - array-scale FST and diversity windows in Python
  • phasing-imputation/haplotype-phasing - phased haplotypes for haplotype-based structure methods
  • comparative-genomics/introgression-detection - D/f4 introgression scans beyond pairwise tests

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in population-genetics/population-structure of GPTomics/bioSkills.

  • SKILL.md
  • examples/structure_analysis.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Population Genetics Population Structure next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Population Genetics Population Structure compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Population Genetics Population Structure this skillGPTomics/bioSkills1.2k1 repos~5.6kAutomated safety check: PassMIT
Scientific Toolkit SkillzLanqing/codex-claude-academic-skills4.7k—~1.2kAutomated safety check: PassMIT
Microsim Generatordmccreary/ibook-skills105—~11kAutomated safety check: PassNone
StatsmodelszLanqing/codex-claude-academic-skills4.7k15 repos~4.9kAutomated safety check: PassBSD-3-Clause
Scientific Figure MakingChenLiu-1996/figures4papers8.3k—~557Automated safety check: PassCustom licence
Plot From ImageTrae1ounG/paper-plot-skills8721 repos~868Automated safety check: PassNone

Similar skills

  • Scientific Toolkit Skill

    zLanqing/codex-claude-academic-skills

    Research computing toolkit for optoelectronic information science and engineering, MATLAB/Octave, Python scientific analysis, signal processing, image processing, statistics, simulation…

    4.7k GitHub stars~1.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Microsim Generator

    dmccreary/ibook-skills

    Creates interactive educational MicroSims, routing to the best-matched generator - p5.js, Chart.js, Plotly, Mermaid, vis-network, timelines, maps, Venn, causal-loop/feedback-loop diagrams (CLD)…

    105 GitHub stars~11k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 15 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Figure Making

    ChenLiu-1996/figures4papers

    Covers publication-ready matplotlib figures for academic papers, slides, and reports—bars, trends, scatter, heatmaps, and multi-panel layouts—with this…

    8.3k GitHub stars~557 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Plot From Image

    Trae1ounG/paper-plot-skills

    Reproduce any academic paper figure from an uploaded image using accumulated style experience.

    872 GitHub starsUsed in 1 repo~868 tokens
    Data & AnalyticsAuto-check passed
  • Analysis Graphing

    clshortfuse/renodx

    RenoDX workflow for creating readable analysis graphs and plots from shader math, CSVs, EXRs, LUTs, hue sweeps, tone curves, gamut comparisons, energy/scalar maps, and test-pattern statistics.

    4.5k GitHub stars~1.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Population Genetics Population Structure

What does Bio Population Genetics Population Structure do?

Infers and describes population structure with PCA (plink2 --pca, smartpca/EIGENSOFT, FlashPCA2), model-based clustering (ADMIXTURE, fastSTRUCTURE), FST estimators (Weir-Cockerham vs Hudson), and…. Bio Population Genetics Population Structure is an agent skill from GPTomics/bioSkills. Infers and describes population structure with PCA (plink2 --pca, smartpca/EIGENSOFT, FlashPCA2), model-based clustering (ADMIXTURE, fastSTRUCTURE), FST estimators (Weir-Cockerham vs Hudson), and f-statistics (f3/f4/D via AdmixTools/admixr), plus Python plotting of PCs and Q barplots.

When should I use Bio Population Genetics Population Structure?

Bio Population Genetics Population Structure fits situations like: F-statistics on QCd genotypes; tasks that involve Statistics; tasks that involve Data visualization.

How do I install Bio Population Genetics Population Structure in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-population-structure -a claude-code`. Or copy the skill folder (population-genetics/population-structure in GPTomics/bioSkills) into .claude/skills/bio-population-genetics-population-structure in your project. Claude Code loads it when a task matches its description.

How do I install Bio Population Genetics Population Structure in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-population-structure -a codex`. Or copy the skill folder (population-genetics/population-structure in GPTomics/bioSkills) into .agents/skills/bio-population-genetics-population-structure in your project. Codex loads it when a task matches its description.

Can I use Bio Population Genetics Population Structure in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-population-structure -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-population-genetics-population-structure, .gemini/skills/bio-population-genetics-population-structure, .github/skills/bio-population-genetics-population-structure and .opencode/skills/bio-population-genetics-population-structure in your project.

What does Bio Population Genetics Population Structure need to run?

Going by SKILL.md and its folder, Bio Population Genetics Population Structure needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Population Genetics Population Structure access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Population Genetics Population Structure safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Population Genetics Population Structure use?

Bio Population Genetics Population Structure is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Population Genetics Population Structure use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Population Genetics Population Structure?

Skills that share tags, products or a category with Bio Population Genetics Population Structure: Scientific Toolkit Skill (zLanqing/codex-claude-academic-skills, 4.7k stars), Microsim Generator (dmccreary/ibook-skills, 105 stars), Statsmodels (zLanqing/codex-claude-academic-skills, 4.7k stars) and Scientific Figure Making (ChenLiu-1996/figures4papers, 8.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Population Genetics Population Structure?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.