Agent skill

Bio Population Genetics Association Testing

by GPTomics in GPTomics/bioSkills

Single-variant common-variant GWAS with plink2 --glm (linear/logistic, Firth) and the linear mixed models GEMMA, BOLT-LMM, SAIGE, regenie (SPA).

MITAuto-check passedResearch & Science

Install Bio Population Genetics Association Testing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-population-genetics-association-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-population-genetics-association-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/population-genetics/association-testing .claude/skills/bio-population-genetics-association-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-population-genetics-association-testing
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.8k tokens
SKILL.md length
1,898 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Single-variant common-variant GWAS with plink2 --glm (linear/logistic, Firth) and the linear mixed models GEMMA, BOLT-LMM, SAIGE, regenie (SPA).

  • Works in 4 steps: Every classic GWAS pathology is one… → The corollary that organizes the… → The fix differs per distortion, and the… → …
  • Running single-variant GWAS
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 8 more sections
  • Runs Shell scripts from its folder; calls pip

What it does

Bio Population Genetics Association Testing is an agent skill from GPTomics/bioSkills. Single-variant common-variant GWAS with plink2 --glm (linear/logistic, Firth) and the linear mixed models GEMMA, BOLT-LMM, SAIGE, regenie (SPA). A GWAS statistic is valid only when genotype is independent of unmodeled phenotype drivers after the chosen covariates and random effects, so the engine follows sample structure and case:control imbalance, not taste: PC covariates absorb continuous ancestry but cannot remove relatedness (a covariance structure needing an LMM), genomic inflation above 1 is mostly true…

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/gwas_pipeline.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Statistics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Running single-variant GWAS
  • Choosing between a GLM and a mixed model
  • Controlling stratification
  • Case:control imbalance

Example prompts

  • “/bio-population-genetics-association-testing”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Every classic GWAS pathology is one violation of that assumption with a specific signature: structure or cryptic relatedness lifts the…
  2. The corollary that organizes the toolchain: under a real polygenic trait genomic inflation (lambda) is EXPECTED to exceed 1 and is mostly…
  3. The fix differs per distortion, and the wrong fix makes it worse: lambda-correcting a polygenic trait, or adding PCs to a family sample…
  4. Choosing the engine is choosing the assumption: a PC-adjusted GLM assumes structure is a low-rank mean shift, an LMM assumes a polygenic…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Population Genetics Association Testing loads about 4.8k tokens when it runs. Until then it costs about 267 tokens; SKILL.md has 1,898 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~267
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,898 words, ~4,827 tokens.

Download SKILL.mdSave it as .claude/skills/bio-population-genetics-association-testing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-population-genetics-association-testing
description
Single-variant common-variant GWAS with plink2 --glm (linear/logistic, Firth) and the linear mixed models GEMMA, BOLT-LMM, SAIGE, regenie (SPA). A GWAS statistic is valid only when genotype is independent of unmodeled phenotype drivers after the chosen covariates and random effects, so the engine follows sample structure and case:control imbalance, not taste: PC covariates absorb continuous ancestry but cannot remove relatedness (a covariance structure needing an LMM), genomic inflation above 1 is mostly true polygenic signal not confounding (the LDSC intercept is the diagnostic), LOCO prevents proximal contamination, and SPA/Firth keep the tail calibrated at extreme imbalance and low MAC. Use when running single-variant GWAS, choosing between a GLM and a mixed model, or controlling stratification, relatedness, and case:control imbalance. For rare-variant aggregation (burden, SKAT, SKAT-O, ACAT) see rare-variant-association; fine-mapping and MR see causal-genomics; PRS see clinical-databases/polygenic-risk.
tool_type
cli
primary_tool
plink2

Version Compatibility

Reference examples tested with: PLINK 2.0 (alpha 6+), SAIGE 1.3+, regenie 3.4+, BOLT-LMM 2.4+, GEMMA 0.98+, numpy 1.26+, pandas 2.2+, scipy 1.12+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Version traps that change results, not just syntax: PLINK 2.0 --glm firth-fallback is the DEFAULT for binary traits and writes .glm.logistic.hybrid (mixed logistic and Firth rows), not .glm.logistic. SAIGE --LOCO=TRUE is the recommended default and the step-1 and step-2 sample IDs plus variance-ratio file must match exactly. regenie applies Firth/SPA only when --firth/--spa are explicitly set in step 2 (not automatic for every variant). BOLT-LMM is calibrated only for quantitative traits at large N (case fraction >= 10%, MAF > 0.1% for binary coding). The single source of truth for versions is this block, not headings.

Association Testing

"Run a GWAS on my genotypes" -> Fit one regression per variant (additive dosage as predictor) under a confounder model chosen to make genotype independent of unmodeled phenotype drivers, then read effect sizes and p-values that are honest only insofar as that model holds.

  • CLI: plink2 --glm firth-fallback hide-covar --covar pcs.eigenvec for an unrelated sample whose structure is captured by PCs
  • CLI: SAIGE (SPA) or regenie --step 1/2 --firth for relatedness and/or case:control imbalance; GEMMA -lmm or BOLT-LMM --lmm for related quantitative traits

Scope: single-variant common-variant GWAS (linear/logistic GLM, linear mixed models, SPA, Firth) and the choice between them. Rare-variant GENE-BASED AGGREGATION (burden, SKAT, SKAT-O, ACAT, STAAR) routes to rare-variant-association. Fine-mapping, Mendelian randomization, and TWAS route to causal-genomics/fine-mapping and causal-genomics/mendelian-randomization. Polygenic risk scores route to clinical-databases/polygenic-risk. Imputed dosages enter from phasing-imputation/genotype-imputation.

The Single Most Important Insight -- a GWAS statistic is valid only if genotype is independent of unmodeled phenotype drivers AFTER the chosen covariates and random effects

  1. Every classic GWAS pathology is one violation of that assumption with a specific signature: structure or cryptic relatedness lifts the whole QQ plot; a heritable covariate adjusted as a collider manufactures direction-flipped hits; case:control imbalance with low MAC under a score test makes the significant tail anti-conservative.
  2. The corollary that organizes the toolchain: under a real polygenic trait genomic inflation (lambda) is EXPECTED to exceed 1 and is mostly true signal, so genomic control over-corrects and erases discoveries; the LDSC intercept (not lambda) separates confounding from polygenicity (Bulik-Sullivan 2015).
  3. The fix differs per distortion, and the wrong fix makes it worse: lambda-correcting a polygenic trait, or adding PCs to a family sample, both silently degrade the result.
  4. Choosing the engine is choosing the assumption: a PC-adjusted GLM assumes structure is a low-rank mean shift, an LMM assumes a polygenic covariance, SPA/Firth assume the score/Wald tail needs the cumulant-generating-function correction.

Tool Taxonomy

MethodCitationMechanismWhen
plink2 --glmChang 2015Per-variant linear/logistic GLM; Firth fallback for separationUnrelated sample, common variants, structure captured by PCs
GEMMA -lmmZhou & Stephens 2012Exact LMM; fits sigma_g^2 once under the null, then per-variant Wald/LRT/scoreRelatedness/structure, quantitative trait, small-medium N
BOLT-LMM --lmmLoh 2015Variational LMM with a non-infinitesimal Bayesian mixture prior; LD-Score calibratedQuantitative trait, very large N; gains power with large-effect loci
SAIGEZhou 2018Sparse-GRM LMM + saddlepoint approximation (SPA) on the score statisticBinary trait with relatedness AND case:control imbalance
regenie --step 1/2Mbatchou 2021Whole-genome ridge (LOCO) predictor in step 1, Firth/SPA test in step 2, no GRM eigendecompositionBiobank scale, mixed binary+quantitative; one pipeline

Decision Tree by Scenario

ScenarioUseWhy
Unrelated, common variants, PCs absorb structure (LDSC intercept ~ 1)plink2 --glm + PC covariatesFast and exact; an LMM is unnecessary when PCs suffice
Relatedness, family, or fine-scale structureany LMM (GEMMA/BOLT/SAIGE/regenie)PCs cannot remove a covariance structure; the GRM random effect can
Quantitative trait, related, small-medium NGEMMA -lmm or GCTA --mlma-locoExact LMM; LOCO avoids proximal contamination
Quantitative trait, biobank N (>5000)BOLT-LMM --lmmScales; non-infinitesimal model adds power; calibrated only at large N
Binary trait, related AND imbalanced case:controlSAIGE (SPA) or regenie (--firth/--spa)SPA/Firth keep the tail calibrated under imbalance and low MAC
Binary trait, unbalanced, NO relatednessplink2 --glm firth-fallback (default)Firth handles separation; no GRM needed
Imputed variantsregress on DOSAGES not hard callshard-calling discards imputation uncertainty and biases the SE
Rare-variant signal at low MACrare-variant-association (burden/SKAT/SKAT-O/ACAT)single-variant tests are powerless at low MAC; aggregate in a region/gene

plink2 GLM (unrelated sample, PC-adjusted)

bash
# Logistic for binary, linear for quantitative is auto-detected from the phenotype coding.
# firth-fallback is the binary-trait default: ordinary logistic, falling back to Firth only on
# non-convergence (separation). Output is .glm.logistic.hybrid (mixed logistic and Firth rows).
plink2 --bfile qc --pheno pheno.txt --glm firth-fallback hide-covar \
    --covar pcs.eigenvec --covar-name PC1-PC10 --out gwas

# Regress on imputed DOSAGES, not hard calls, so imputation uncertainty enters the SE. Reporting
# columns add A1 frequency and the imputation R2 (machr2 is meaningful only on dosage input) for
# downstream QC and effect-allele bookkeeping.
plink2 --pfile imputed --pheno pheno.txt --glm firth-fallback hide-covar cols=+a1freq,+machr2 \
    --covar covars.txt --covar-name PC1-PC10,age,sex --out gwas_dosage

PCs must be computed on LD-pruned, MAF-filtered genotypes with long-range-LD regions excluded (see population-structure); too few PCs leave residual stratification, too many absorb real signal. --glm sex adds sex as a covariate on chrX (no-x-sex suppresses it); split the pseudoautosomal region before any chrX test (see plink-basics).

Linear mixed model with LOCO (relatedness / structure)

bash
# GEMMA: -gk builds the GRM (1=centered, 2=standardized), then -lmm fits one variance component.
# The covariate file passed to -c MUST contain an explicit intercept column of 1s (GEMMA does not add one).
gemma -bfile qc -gk 1 -o grm
gemma -bfile qc -k output/grm.cXX.txt -c covars_with_intercept.txt -lmm 4 -o lmm   # -lmm 4 = Wald+LRT+score

# BOLT-LMM: --lmm decides by cross-validation whether the non-infinitesimal model adds power.
# --lmmInfOnly forces the standard infinitesimal model; --lmmForceNonInf forces the mixture.
bolt --bfile=qc --phenoFile=pheno.txt --phenoCol=trait \
    --covarFile=covars.txt --qCovarCol=PC{1:10} --lmm --LDscoresFile=LDSCORE.1000G_EUR.tab.gz \
    --statsFile=bolt.stats

# SAIGE: step 1 fits the null GLMM (sparse GRM, variance ratio); step 2 runs the SPA score test per variant.
# --LOCO=TRUE (default, recommended) excludes the tested chromosome from the polygenic predictor.
step1_fitNULLGLMM.R --plinkFile=qc --phenoFile=pheno.txt --phenoCol=trait --traitType=binary \
    --covarColList=PC1,PC2,age,sex --sampleIDColinphenoFile=IID --outputPrefix=step1 --LOCO=TRUE
step2_SPAtests.R --vcfFile=chr1.vcf.gz --GMMATmodelFile=step1.rda --varianceRatioFile=step1.varianceRatio.txt \
    --minMAC=20 --is_Firth_beta=TRUE --LOCO=TRUE --SAIGEOutputFile=chr1.saige

# regenie: step 1 builds the whole-genome LOCO ridge predictor; step 2 tests with Firth (--approx for speed) or SPA.
regenie --step 1 --bed qc --phenoFile pheno.txt --covarFile covars.txt --bt --bsize 1000 --out step1
regenie --step 2 --bed qc --phenoFile pheno.txt --covarFile covars.txt --bt \
    --firth --approx --pThresh 0.01 --pred step1_pred.list --bsize 400 --out step2

LOCO (leave-one-chromosome-out) is not optional: if the tested variant's chromosome is in the GRM/predictor, the random effect explains part of the variant's own signal and deflates power (proximal contamination). A hand-rolled "GRM from all SNPs" silently throws away power at every true locus.

Diagnose inflation: lambda vs the LDSC intercept

python
import numpy as np
from scipy import stats

def lambda_gc(pvalues):
    chisq = stats.chi2.ppf(1 - pvalues, 1)
    return np.median(chisq) / stats.chi2.ppf(0.5, 1)

# lambda > 1 under a polygenic trait is EXPECTED and mostly true signal. Do NOT divide statistics by it.
# Rescale to lambda_1000 (per 1000 cases/1000 controls) before comparing studies of different N.
def lambda_1000(lam, n_cases, n_controls):
    return 1 + (lam - 1) * (1 / n_cases + 1 / n_controls) / (1 / 1000 + 1 / 1000)

# The confounding diagnostic is the LDSC INTERCEPT (run ldsc on the sumstats), not lambda:
# intercept ~ 1 with high lambda = polygenicity (clean); intercept materially > 1 = confounding.
# Prefer the attenuation ratio = (intercept - 1) / (mean(chi2) - 1): the fraction of inflation NOT due
# to polygenicity (~0-0.2 acceptable). The intercept is also inflated by sample overlap in bivariate LDSC.

Per-Method Failure Modes

Genomic control on a polygenic trait

Trigger: dividing every chi-square by lambda_GC because lambda > 1.1. Mechanism: lambda rises with N and h2 under true polygenicity, so it is mostly signal. Symptom: discoveries vanish; power destroyed. Fix: never lambda-correct on lambda alone; use the LDSC intercept and only deflate if intercept-minus-1 is materially > 0.

Trigger: adding "10 PCs" to a family or cryptically-related cohort. Mechanism: relatedness is a pairwise covariance, not a low-rank mean shift, so no finite PC set removes it. Symptom: inflated, miscalibrated tail statistics despite the PCs. Fix: switch to an LMM (GEMMA/BOLT/SAIGE/regenie); more PCs is the wrong lever.

LMM without LOCO

Trigger: building the GRM from all chromosomes including the candidate. Mechanism: the variant partly explains itself as a random effect. Symptom: deflated statistics, lost power at true loci. Fix: use the LOCO variant (--mlma-loco, SAIGE --LOCO=TRUE, regenie step-1 predictor).

Wald collapse / score anti-conservatism at imbalance

Trigger: plain logistic for a rare or near-monomorphic-in-one-arm variant, or a score test at extreme case:control. Mechanism: the MLE diverges (Wald BETA/SE -> 0) or the score null is wrong in the tail. Symptom: real rare-variant signal looks non-significant, or the significant tail fills with artifacts. Fix: Firth (plink2 firth-fallback, regenie --firth) or SPA (SAIGE, regenie --spa).

HWE filtered in cases

Trigger: HWE filter on the case-only or combined sample as a discovery QC step. Mechanism: a true non-additive disease variant legitimately deviates from HWE in cases. Symptom: real associations removed before testing. Fix: compute HWE in CONTROLS (or founders) only (see plink-basics).

Show full SKILL.md (782 more words)Show less
Adjusting for a heritable covariate (collider)

Trigger: GWAS of a trait while adjusting for a genetically influenced covariate (BMI, smoking). Mechanism: opens a non-causal genotype-phenotype path. Symptom: spurious, often direction-flipped hits at loci affecting the covariate that replicate within the conditioned design. Fix: only adjust for covariates not affected by genotype, or interpret as a different (interventional) question (Aschard 2015).

BOLT-LMM on a binary trait

Trigger: treating BOLT as a logistic engine. Mechanism: it runs linear regression on case/control coding. Symptom: miscalibration and rare-variant false positives under imbalance. Fix: use it only for quantitative traits at case fraction >= 10%, MAF > 0.1%, large N; otherwise SAIGE/regenie.

Quantitative Thresholds

QuantityTypical valueRationale
Genome-wide significancep < 5e-8Bonferroni over ~1e6 independent common-variant tests in European HapMap LD (Pe'er 2008); ancestry/array-dependent
...for African ancestry~1e-8 to 3e-8less LD = more independent tests; 5e-8 is too lax
...for WGS / rare-variant scans~5e-9 or stricterfar larger effective test count; 5e-8 too permissive
Suggestivep < 1e-5follow-up convention, not a calibrated threshold
lambda_GC~1.0-1.05 fine; 1.05-1.10 inspect; >1.10 investigatescales with N and h2; pair with lambda_1000 and the LDSC intercept
MAC floor (single-variant)MAC >= 20 (>= 10 with SPA/Firth)below this even SPA/Firth are unstable; aggregate instead
SPA/Firth triggercase:control more extreme than ~1:10, or low MACthe score/Wald tail breaks exactly there (SAIGE motivated by ~1:600, Zhou 2018)
HWE filter (controls only)p < 1e-6catches genotyping artifacts; never case-only as a discovery filter

Thresholds are conventions, not laws; inspect distributions and verify current best practice before applying numbers blindly.

Common Errors

Error / symptomCauseSolution
.glm.logistic.hybrid expected .glm.logisticfirth-fallback is the binary default and mixes Firth rowsread the .hybrid file; the FIRTH? column flags Firth rows
0/1 phenotype gives a null GWASPLINK reads 1=control, 2=case by defaultpass --1 for 0/1 coding (see plink-basics)
Inflated tail despite 10 PCsrelatedness in the sampleswitch to an LMM; PCs cannot remove relatedness
Lost power at top loci in an LMMGRM includes the candidate chromosomeenable LOCO
Rare-variant hits look non-significantWald collapse under separationuse Firth or SPA
Flipped BETA cancels in meta-analysiseffect-allele/strand not harmonizedcarry CHR, POS, EA, OA, EAF; resolve A/T and C/G palindromes by frequency or drop
Discovery effect size too large downstreamwinner's curseuse out-of-sample or shrinkage-corrected effects for PRS/power/MR
chrX mis-codedmales hemizygous, PAR diploid--glm sex, split PAR first, handle X-inactivation coding explicitly
GEMMA model wrong, no error-c does not add an interceptthe covariate file must contain a column of 1s
Fixed-effect pooled OR with I^2 = 80%heterogeneous true effectscheck Cochran's Q / I^2; use random-effects or MR-MEGA for trans-ancestry

References

  1. Devlin B, Roeder K. Genomic control for association studies. Biometrics 1999; 55(4):997-1004. DOI:10.1111/j.0006-341X.1999.00997.x.
  2. Purcell S, Neale B, Todd-Brown K, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. American Journal of Human Genetics 2007; 81(3):559-575. DOI:10.1086/519795.
  3. Pe'er I, Yelensky R, Altshuler D, Daly MJ. Estimation of the multiple testing burden for genomewide association studies of nearly all common variants. Genetic Epidemiology 2008; 32(4):381-385. DOI:10.1002/gepi.20303.
  4. Zhou X, Stephens M. Genome-wide efficient mixed-model analysis for association studies. Nature Genetics 2012; 44(7):821-824. DOI:10.1038/ng.2310.
  5. Yang J, Lee SH, Goddard ME, Visscher PM. GCTA: a tool for genome-wide complex trait analysis. American Journal of Human Genetics 2011; 88(1):76-82. DOI:10.1016/j.ajhg.2010.11.011.
  6. Aschard H, Vilhjalmsson BJ, Joshi AD, Price AL, Kraft P. Adjusting for heritable covariates can bias effect estimates in genome-wide association studies. American Journal of Human Genetics 2015; 96(2):329-339. DOI:10.1016/j.ajhg.2014.12.021.
  7. Chang CC, Chow CC, Tellier LCAM, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience 2015; 4:7. DOI:10.1186/s13742-015-0047-8.
  8. Loh PR, Tucker G, Bulik-Sullivan BK, et al. Efficient Bayesian mixed-model analysis increases association power in large cohorts. Nature Genetics 2015; 47(3):284-290. DOI:10.1038/ng.3190.
  9. Bulik-Sullivan BK, Loh PR, Finucane HK, et al. LD Score regression distinguishes confounding from polygenicity in genome-wide association studies. Nature Genetics 2015; 47(3):291-295. DOI:10.1038/ng.3211.
  10. Zhou W, Nielsen JB, Fritsche LG, et al. Efficiently controlling for case-control imbalance and sample relatedness in large-scale genetic association studies. Nature Genetics 2018; 50(9):1335-1341. DOI:10.1038/s41588-018-0184-y.
  11. Mbatchou J, Barnard L, Backman J, et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nature Genetics 2021; 53(7):1097-1103. DOI:10.1038/s41588-021-00870-7.
  • plink-basics - QC, phenotype encoding, and the fileset that enters association
  • population-structure - PCA covariates for stratification control
  • linkage-disequilibrium - LD pruning before PCA and clumping of GWAS hits
  • rare-variant-association - gene-based aggregation (burden, SKAT, SKAT-O, ACAT) below the single-variant MAC floor
  • causal-genomics/fine-mapping - from an associated locus to a credible set of causal variants
  • causal-genomics/mendelian-randomization - GWAS variants as instruments for causal inference
  • clinical-databases/polygenic-risk - PRS built from association sumstats
  • phasing-imputation/genotype-imputation - imputed dosages that enter --glm

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in population-genetics/association-testing of GPTomics/bioSkills.

  • SKILL.md
  • examples/gwas_pipeline.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Population Genetics Association Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Population Genetics Association Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Population Genetics Association Testing this skillGPTomics/bioSkills1.2k1 repos~4.8kAutomated safety check: PassMIT
PyDESeq2 Differential Expressiondavila7/claude-code-templates33k11 repos~4kAutomated safety check: PassMIT
Ukb Ppp Region FetchClawBio/ClawBio1.2k—~4.6kAutomated safety check: PassMIT
TiledbvcfK-Dense-AI/scientific-agent-skills48k1 repos~3.5kAutomated safety check: PassMIT
Volcano Plot Scriptaipoch/medical-research-skills1.9k—~2.5kAutomated safety check: PassMIT
Tooluniverse Epigenomicswu-yc/LabClaw1.1k2 repos~14kAutomated safety check: PassNone

Similar skills

  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    33k GitHub starsUsed in 11 repos~4k tokens
    Research & ScienceAuto-check passed
  • Ukb Ppp Region Fetch

    ClawBio/ClawBio

    Fetch a regional slice of plasma pQTL summary statistics from the UK Biobank Pharma Proteomics Project (UKB-PPP; Sun 2023 Nature) for a specific (protein, ancestry) measurement.

    1.2k GitHub stars~4.6k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Tiledbvcf

    K-Dense-AI/scientific-agent-skills

    Stores and retrieves genomic variant calls with TileDB-VCF. An agent skill from K-Dense-AI/scientific-agent-skills.

    48k GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed
  • Volcano Plot Script

    aipoch/medical-research-skills

    Generate R/Python code for volcano plots from DEG (Differentially Expressed Genes) analysis results.

    1.9k GitHub stars~2.5k tokensUpdated 23 days ago
    Research & ScienceAuto-check passed
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • Analyze metabolomics data including metabolite identification, quantification, pathway analysis, and metabolic flux.

    1.1k GitHub starsUsed in 2 repos~5.9k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Population Genetics Association Testing

What does Bio Population Genetics Association Testing do?

Single-variant common-variant GWAS with plink2 --glm (linear/logistic, Firth) and the linear mixed models GEMMA, BOLT-LMM, SAIGE, regenie (SPA). Bio Population Genetics Association Testing is an agent skill from GPTomics/bioSkills. Single-variant common-variant GWAS with plink2 --glm (linear/logistic, Firth) and the linear mixed models GEMMA, BOLT-LMM, SAIGE, regenie (SPA).

When should I use Bio Population Genetics Association Testing?

Bio Population Genetics Association Testing fits situations like: running single-variant GWAS; choosing between a GLM and a mixed model; controlling stratification; case:control imbalance.

How do I install Bio Population Genetics Association Testing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-association-testing -a claude-code`. Or copy the skill folder (population-genetics/association-testing in GPTomics/bioSkills) into .claude/skills/bio-population-genetics-association-testing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Population Genetics Association Testing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-association-testing -a codex`. Or copy the skill folder (population-genetics/association-testing in GPTomics/bioSkills) into .agents/skills/bio-population-genetics-association-testing in your project. Codex loads it when a task matches its description.

Can I use Bio Population Genetics Association Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-population-genetics-association-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-population-genetics-association-testing, .gemini/skills/bio-population-genetics-association-testing, .github/skills/bio-population-genetics-association-testing and .opencode/skills/bio-population-genetics-association-testing in your project.

What does Bio Population Genetics Association Testing need to run?

Going by SKILL.md and its folder, Bio Population Genetics Association Testing needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Population Genetics Association Testing access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Population Genetics Association Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Population Genetics Association Testing use?

Bio Population Genetics Association Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Population Genetics Association Testing use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Population Genetics Association Testing?

Skills that share tags, products or a category with Bio Population Genetics Association Testing: PyDESeq2 Differential Expression (davila7/claude-code-templates, 33k stars), Ukb Ppp Region Fetch (ClawBio/ClawBio, 1.2k stars), Tiledbvcf (K-Dense-AI/scientific-agent-skills, 48k stars) and Volcano Plot Script (aipoch/medical-research-skills, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Population Genetics Association Testing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.