Agent skill

Bio Phasing Imputation Genotype Imputation

by GPTomics in GPTomics/bioSkills

Imputes untyped genotypes against a phased reference panel with Beagle, Minimac4, or IMPUTE5 (array data) or from genotype likelihoods with GLIMPSE2, QUILT2, or STITCH (low-coverage WGS), producing…

MITAuto-check passed

Install Bio Phasing Imputation Genotype Imputation

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-phasing-imputation-genotype-imputation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-phasing-imputation-genotype-imputation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/phasing-imputation/genotype-imputation .claude/skills/bio-phasing-imputation-genotype-imputation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-phasing-imputation-genotype-imputation
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.7k tokens
SKILL.md length
2,321 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Imputes untyped genotypes against a phased reference panel with Beagle, Minimac4, or IMPUTE5 (array data) or from genotype likelihoods with GLIMPSE2, QUILT2, or STITCH (low-coverage WGS), producing…

  • Works in 3 steps: Downstream analysis uses dosages, not… → The quality metric (Beagle DR2, Minimac… → Low-coverage WGS (0.5-4x) plus GLIMPSE2…
  • Increasing variant density for GWAS
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 10 more sections
  • Runs Shell scripts from its folder; calls java

What it does

Bio Phasing Imputation Genotype Imputation is an agent skill from GPTomics/bioSkills. Imputes untyped genotypes against a phased reference panel with Beagle, Minimac4, or IMPUTE5 (array data) or from genotype likelihoods with GLIMPSE2, QUILT2, or STITCH (low-coverage WGS), producing per-variant dosages (DS) with a self-estimated quality (Beagle DR2, Minimac R2, IMPUTE INFO). Covers why the honest output is a dosage posterior not a hard call, why GWAS regresses on DS, why the quality metric is an ESTIMATE of r2 from posterior spread (not validation against truth), the DS/GP/HDS fields, the phasing…

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/run_beagle_imputation.sh` and `usage-guide.md`).

The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Increasing variant density for GWAS
  • Harmonizing arrays
  • Inferring untyped variants
  • Imputing low-coverage sequence

Example prompts

  • “Use the bio-phasing-imputation-genotype-imputation skill to impute untyped genotypes against a phased reference panel with Beagle, Minimac4, or…”
  • “/bio-phasing-imputation-genotype-imputation”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Downstream analysis uses dosages, not hard genotypes. The dosage DS is the conditional expectation E[genotype | data, panel], the…
  2. The quality metric (Beagle DR2, Minimac R2, IMPUTE INFO) is an ESTIMATE of r2 from the posterior spread, computed without ever seeing the…
  3. Low-coverage WGS (0.5-4x) plus GLIMPSE2 has become a credible array replacement. Because it samples the whole genome rather than a fixed…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • java

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Phasing Imputation Genotype Imputation loads about 4.7k tokens when it runs. Until then it costs about 264 tokens; SKILL.md has 2,321 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~264
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,321 words, ~4,740 tokens.

Download SKILL.mdSave it as .claude/skills/bio-phasing-imputation-genotype-imputation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-phasing-imputation-genotype-imputation
description
Imputes untyped genotypes against a phased reference panel with Beagle, Minimac4, or IMPUTE5 (array data) or from genotype likelihoods with GLIMPSE2, QUILT2, or STITCH (low-coverage WGS), producing per-variant dosages (DS) with a self-estimated quality (Beagle DR2, Minimac R2, IMPUTE INFO). Covers why the honest output is a dosage posterior not a hard call, why GWAS regresses on DS, why the quality metric is an ESTIMATE of r2 from posterior spread (not validation against truth), the DS/GP/HDS fields, the phasing prerequisite, chunking, chrX ploidy, the Michigan/TOPMed servers (the only access to HRC/TOPMed), and low-coverage WGS as the modern array replacement. Use when increasing variant density for GWAS, harmonizing arrays, inferring untyped variants, or imputing low-coverage sequence. Phase first with haplotype-phasing; prepare the panel with reference-panels; filter with imputation-qc; the GWAS test is population-genetics/association-testing; end-to-end orchestration is workflows/gwas-pipeline.
tool_type
cli
primary_tool
Beagle

Version Compatibility

Reference examples tested with: Beagle 5.4 (22Jul22), Minimac4 4.1+, IMPUTE5 1.2, GLIMPSE2, bcftools 1.19+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Minimac4 (4.x) uses POSITIONAL arguments (minimac4 panel.msav target.vcf.gz); the old --refHaps/--haps/--prefix/--cpus style is Minimac3 and obsolete. Beagle 5.x emits DR2, AF, and IMP only (AR2 is a legacy 4.x field) and its default ne=100000 (not 1,000,000). The panel build (GRCh37 vs GRCh38) must match the data; record the panel name, version, and build with every result.

Genotype Imputation -- Inferring Untyped Genotypes as Dosages

"Fill in the variants I did not directly measure" -> Align the (phased or low-coverage) sample to a reference panel of phased haplotypes and infer the untyped alleles via the Li-Stephens HMM - because the output is a posterior over genotypes summarized as a dosage with a self-estimated quality, not a measured call, so the uncertainty must be carried downstream.

  • CLI: java -jar beagle.jar gt=phased.vcf.gz ref=panel.bref3 map=plink.chr20.map out=imputed (or minimac4 panel.msav phased.vcf.gz, or GLIMPSE2 for low-coverage WGS)

Scope: imputing untyped genotypes from a panel (array data) or from genotype likelihoods (low-coverage WGS), the dosage/quality output, chunking, chrX, and the servers. Phasing the input -> haplotype-phasing. Panel selection/preparation/strand -> reference-panels. Quality metrics and filtering thresholds -> imputation-qc. The GWAS test on the dosages -> population-genetics/association-testing. The genotype likelihoods that low-coverage imputation consumes -> variant-calling/vcf-basics. End-to-end orchestration -> workflows/gwas-pipeline.

The Single Most Important Modern Insight -- An Imputed Genotype Is a Posterior, and the Deliverable Is a Dosage Plus a Self-Estimated Quality, Not a Hard Call

Imputation aligns a sparsely-genotyped (or low-coverage-sequenced) sample to a densely-typed reference panel of phased haplotypes and infers, via a Li-Stephens HMM, the alleles at positions the sample never observed (Browning 2018 Am J Hum Genet 103:338). The output at each untyped variant is a distribution, summarized as an expected allelic dosage in [0,2]. Three facts define the field:

  1. Downstream analysis uses dosages, not hard genotypes. The dosage DS is the conditional expectation E[genotype | data, panel], the minimum-variance summary; hard-calling forces an uncertain 0.5 dosage to 0 or 1, injecting genotype error that attenuates effects and inflates standard errors. GWAS regresses the trait on DS -> population-genetics/association-testing.
  2. The quality metric (Beagle DR2, Minimac R2, IMPUTE INFO) is an ESTIMATE of r2 from the posterior spread, computed without ever seeing the truth. Poorly-imputed dosages shrink toward the allele-frequency mean 2p, so low posterior variance relative to the binomial expectation 2p(1-p) flags a low-confidence site. This is NOT a validation against held-out genotypes (that is empirical r2 / EmpRsq, a masked-site quantity). Say "DR2/R2/INFO is an estimate of imputation quality," never "the imputation accuracy was 0.9" as if measured. The metric also cannot detect panel-ancestry mismatch -> imputation-qc.
  3. Low-coverage WGS (0.5-4x) plus GLIMPSE2 has become a credible array replacement. Because it samples the whole genome rather than a fixed ascertained SNP set, it imputes rare variants and under-represented ancestries better than a dense array at comparable cost (Rubinacci 2023 Nat Genet 55:1088). The input is genotype likelihoods, not calls; the array-vs-low-coverage-WGS choice is an ascertainment decision (see Array vs Low-Coverage WGS below).

Tool Taxonomy

ToolCitationMechanism / roleWhen
Minimac4Das 2016 Nat Genet 48:1284array imputation; msav/m3vcf panel; the server engine; positional-arg CLIserver-style imputation; meta-imputation
Beagle 5.xBrowning 2018 Am J Hum Genet 103:338Java; phases unphased input AND imputes; bref3 panelone tool for phase + impute, no compile
IMPUTE5Rubinacci 2020 PLoS Genet 16:e1009049PBWT pre-selection then LS HMM; sub-linear in panel sizevery large reference panels; local speed
GLIMPSE2Rubinacci 2023 Nat Genet 55:1088low-coverage WGS imputation from genotype likelihoods; chunk/split/phase/ligate0.5-4x WGS with a panel
QUILT2Davies 2021 Nat Genet 53:1104low-coverage, panel-based, read-awarelong-read / haplotagged / ancient DNA / cfDNA
STITCHDavies 2016 Nat Genet 48:965low-coverage, REFERENCE-FREE; learns ancestral haplotypes by EMno panel exists (non-model organisms)
Michigan / TOPMed serversDas 2016 Nat Genet 48:1284Eagle2 phasing + Minimac4; the only access to HRC/TOPMedturnkey, access-controlled panels

Decision Tree by Scenario

ScenarioRecommendedWhy
Array data, want HRC/TOPMed and a turnkey pipelineTOPMed or Michigan Imputation Serverthe only sanctioned access to those panels; runs Eagle2 + Minimac4
Array data, local run, very large panel, want speedIMPUTE5 (PBWT) or Minimac4sub-linear scaling in panel size
Array data, local, one tool for phase + imputeBeagle 5.xphases unphased gt= input itself; bref3 panel
Low-coverage WGS (0.5-4x), have a panelGLIMPSE2 (chunk -> split-reference -> phase -> ligate)the standard; imputes from genotype likelihoods
Low-coverage, read-aware / long-read / ancient DNA / cfDNAQUILT2per-read, base-quality-aware
Low-coverage, NO reference panel (non-model organism)STITCHlearns ancestral haplotypes reference-free
Need the panel selected/prepared first-> reference-panelsthe panel is the prior
Need the input phased first (Minimac4, IMPUTE5)-> haplotype-phasingthose engines require a phased target
Filter the imputed output before analysis-> imputation-qcDR2/R2/INFO + MAF floor
The GWAS test on the dosages-> population-genetics/association-testingdownstream

Array vs Low-Coverage WGS: the Imputation-Input Fork

The upstream decision is how to generate the genotypes that will be imputed, and it is an ascertainment question, not just an accuracy one. An array assays a fixed, designed SNP set (biased to its design population); low-coverage WGS samples whatever is in the genome.

SNP array + pre-phase + imputeLow-coverage WGS (~0.5-4x) + impute from genotype likelihoods
Input to the HMMhard genotype calls (array error is tiny)genotype LIKELIHOODS (PL/GL); a hard call at 1x is mostly noise
AscertainmentFIXED - only the designed SNPs, biased to the design populationUNBIASED - whatever is in the genome is observed
Rare variantslimited by the array scaffold and panelmatches or beats dense arrays (Rubinacci 2021 Nat Genet 53:120)
Under-represented ancestrypoor (no good array, panel-mismatched)the main route around array/panel bias
ToolsBeagle / Minimac4 / IMPUTE5GLIMPSE2 (panel) / STITCH (no panel)

The judgment: common-variant GWAS in a well-paneled ancestry -> array plus imputation is cheap and adequate; rare variants, under-represented ancestry, or a need for unbiased genome-wide ascertainment -> low-coverage WGS plus genotype-likelihood imputation, the direction the field is moving as sequencing costs fall. Low-coverage WGS is only as good as its panel and its likelihoods (bad mapping, contamination, or damage produce garbage GLs that impute garbage).

Output Formats and Why Dosages

The central object is the posterior genotype distribution; everything else summarizes it. Request the fields up front (Minimac4 -f GT,DS,HDS,GP; Beagle gp=true ap=true).

FORMATMeaningShape
GPgenotype probabilities P(0/0),P(0/1),P(1/1); the full posterior3 values summing to 1
DSallelic dosage = P(0/1) + 2*P(1/1) = E[genotype]; the GWAS field1 value in [0,2]
HDShaploid (phased per-haplotype) dosage; DS = HDS1 + HDS2 (Minimac4/GLIMPSE)2 values, each [0,1]
AP1/AP2Beagle allele probabilities (P(ALT) per haplotype); DS = AP1 + AP2 (with ap=true)1 value each [0,1]
GThard best-guess genotype (argmax); lossy, discards uncertainty0/0, 0/1, 1/1

GP is the distribution; DS is its mean - two variants with different GP spreads can share a DS. Use DS for association (it propagates the uncertainty); use HDS/AP for phased/allele-specific analyses. Beagle computes GP from allele probabilities assuming Hardy-Weinberg and sets GT from the per-haplotype argmax, so its GT can occasionally disagree with the argmax of its own GP.

The Phasing Prerequisite

The reference panel is phased haplotypes; the target must align to that haplotype structure two ways:

  • Pre-phase then impute (Minimac4, IMPUTE5): phase the target FIRST (Eagle2 or SHAPEIT) into haplotypes, then impute. The server default (Eagle2 -> Minimac4) and the fast local pattern -> haplotype-phasing.
  • Phase-and-impute together (Beagle, GLIMPSE2): the tool phases internally. Low-coverage tools MUST do this, because there is no confident genotype to phase up front; GLIMPSE2 alternates haploid imputation and phasing, and gains accuracy by imputing all target samples jointly.

Low-Coverage WGS Workflow (GLIMPSE2)

The input is genotype likelihoods (PL/GL), not calls, because at 0.5-4x no genotype is certain. GLIMPSE2 can read BAM/CRAM directly (computing GLs internally) or a GL BCF made with bcftools mpileup ... -T panel_sites.vcf.gz | bcftools call -Aim -C alleles -T panel_sites.tsv.gz (the -C alleles constraint needs the panel sites supplied to call via -T; note the two -T files differ in format - a VCF for mpileup, a tab-delimited sites file for call). The pipeline:

  1. GLIMPSE2_chunk defines windows with buffers.
  2. GLIMPSE2_split_reference precomputes a binary panel per chunk (the speed innovation that made UK Biobank-scale imputation feasible).
  3. GLIMPSE2_phase imputes and phases per chunk (--bam-list or --input-gl; --ne default 100000).
  4. GLIMPSE2_ligate stitches chunks using the overlap buffers to keep phase. Output FORMAT: GT, DS, GP, HS plus a per-variant INFO score.

For chrX with GLIMPSE2, declare each sample's ploidy with --samples-file (sample and copy number) and run the PAR/nonPAR split as for the array tools (male nonPAR is haploid) -> reference-panels.

Show full SKILL.md (889 more words)Show less

Imputation Servers

The Michigan (now MIS2) and TOPMed servers run Eagle2 phasing + Minimac4 imputation server-side and are the ONLY sanctioned access to HRC and TOPMed (those panels are controlled-access, not downloadable). Upload a per-chromosome VCF, select the panel, build, and population; the server runs allele-frequency QC and strand-flip detection, phases, imputes in chunks, and returns per-chromosome VCFs in GT,DS,GP plus a Minimac info file with R2 and a QC report. Results are encrypted with a one-time password and auto-deleted after a few days. The reproducibility cost: the panel version (HRC r1.1 vs TOPMed r2 vs r3), tool version, and build can change between runs, so record exactly which server/panel/version produced a result.

Per-Method Failure Modes

Obsolete Minimac4 syntax

Trigger: minimac4 --refHaps panel.m3vcf --haps study.vcf --prefix out. Mechanism: that is Minimac3; Minimac4 4.x takes positional args. Symptom: the command errors or is not recognized. Fix: minimac4 panel.msav target.phased.vcf.gz -o imputed.vcf.gz -f GT,DS,HDS,GP -t 8; build the panel with minimac4 --compress-reference.

Imputing unphased input to a pre-phase engine

Trigger: feeding unphased genotypes to Minimac4 or IMPUTE5. Mechanism: those engines assume a phased target aligned to the panel haplotypes. Symptom: garbage or refused input. Fix: phase first (Eagle2/SHAPEIT) -> haplotype-phasing, or use Beagle/GLIMPSE2 which phase internally.

Imputing cases and controls separately

Trigger: running imputation per batch (cases, then controls, or per cohort). Mechanism: batch-differential imputation quality at a variant creates artifactual genotype structure correlated with phenotype. Symptom: genome-wide-significant hits that fail to replicate; every single-batch QC metric passes. Fix: impute all samples together (or harmonize panels/versions and check that quality does not differ by batch) -> imputation-qc.

Hard-calling the dosage

Trigger: thresholding DS to 0/1/2 for association. Mechanism: discards the posterior uncertainty, worst at low-R2 rare variants. Symptom: lost power read as a true null. Fix: regress on DS (PLINK2 dosage=DS, SNPTEST, REGENIE, BOLT-LMM all accept dosages).

Missing DS field

Trigger: a downstream tool cannot find dosages. Mechanism: the FORMAT fields were not requested. Symptom: only GT or GP present. Fix: request -f GT,DS,HDS,GP (Minimac4) or gp=true ap=true (Beagle) at run time.

Genome build or strand not aligned to the panel

Trigger: GRCh37 data against a GRCh38 panel, or unflipped palindromic SNPs. Mechanism: positions/alleles disagree with the panel; the HMM copies wrong templates. Symptom: near-zero accuracy across regions, no error. Fix: align build and strand before imputing -> reference-panels.

Quantitative Thresholds

ThresholdSourceRationale
Regress on DS (dosage), not hard GTBrowning 2018 Am J Hum Genet 103:338DS = E[genotype
Beagle ne=100000 (default)Beagle 5.x defaulteffective population size for the HMM; not 1,000,000
Beagle window=40.0 / overlap=2.0 cMBeagle 5.x defaultswindow must be >= 1.1x overlap; rarely tuned
Impute all samples togetherBrowning 2018 Am J Hum Genet 103:338 (framing)separate case/control imputation manufactures false associations -> imputation-qc
Low-coverage sweet spot ~0.5-4xRubinacci 2023 Nat Genet 55:1088GLIMPSE2 accuracy range; ~1x is array-competitive
Request DS explicitly (Minimac4 default is GT,DS)Minimac4 docsHDS/GP for phased/probabilistic uses must be named
Post-imputation R2/DR2/INFO filter (a QC decision, not a default)-> imputation-qcthe imputer's number is the INPUT to filtering, not a tool default

Common Errors

Error / symptomCauseSolution
minimac4 --refHaps not recognizedMinimac3 syntaxuse positional args: minimac4 panel.msav target.vcf.gz -o out
Beagle OutOfMemoryErrorJVM heap too small / whole genome one jobraise -Xmx; impute per chromosome
No DS in outputfields not requested-f GT,DS,HDS,GP (Minimac4) / gp=true ap=true (Beagle)
Imputation accuracy near zero across a regionbuild/strand mismatch to the panelalign build and strand first -> reference-panels
Hits do not replicatecases/controls imputed separately, or hard-calledimpute together; regress on dosages -> imputation-qc
Engine errors on multiallelic sitesnon-biallelic inputbcftools norm -m -any first -> variant-calling/variant-normalization
Cannot download HRC/TOPMedcontrolled-access panelsuse the imputation server

References

  • Das S, Forer L, Schonherr S, et al. 2016. Next-generation genotype imputation service and methods. Nat Genet 48:1284-1287.
  • Browning BL, Zhou Y, Browning SR. 2018. A one-penny imputed genome from next-generation reference panels. Am J Hum Genet 103:338-348.
  • Browning BL, Tian X, Zhou Y, Browning SR. 2021. Fast two-stage phasing of large-scale sequence data. Am J Hum Genet 108:1880-1890.
  • Rubinacci S, Delaneau O, Marchini J. 2020. Genotype imputation using the Positional Burrows-Wheeler Transform. PLoS Genet 16:e1009049.
  • Rubinacci S, Ribeiro DM, Hofmeister RJ, Delaneau O. 2021. Efficient phasing and imputation of low-coverage sequencing data using large reference panels. Nat Genet 53:120-126.
  • Rubinacci S, Hofmeister RJ, Sousa da Mota B, Delaneau O. 2023. Imputation of low-coverage sequencing data from 150,119 UK Biobank genomes. Nat Genet 55:1088-1090.
  • Davies RW, Kucka M, Su D, et al. 2021. Rapid genotype imputation from sequence with reference panels. Nat Genet 53:1104-1111.
  • Davies RW, Flint J, Myers S, Mott R. 2016. Rapid genotype imputation from sequence without reference panels. Nat Genet 48:965-969.
  • McCarthy S, Das S, Kretzschmar W, et al. 2016. A reference panel of 64,976 haplotypes for genotype imputation. Nat Genet 48:1279-1283.
  • Taliun D, Harris DN, Kessler MD, et al. 2021. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature 590:290-299.
  • haplotype-phasing - Pre-phasing the target (required by Minimac4 and IMPUTE5)
  • reference-panels - Select and prepare the panel (the prior) and align build/strand
  • imputation-qc - Filter by DR2/R2/INFO and MAF; the metric is an estimate, not truth
  • variant-calling/vcf-basics - Genotype likelihoods (PL/GL) for low-coverage imputation
  • variant-calling/variant-normalization - Split multiallelics before imputation
  • population-genetics/association-testing - GWAS test on the imputed dosages
  • clinical-databases/polygenic-risk - Polygenic scores from imputed dosages
  • workflows/gwas-pipeline - End-to-end QC -> phase -> impute -> associate

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in phasing-imputation/genotype-imputation of GPTomics/bioSkills.

  • SKILL.md
  • examples/run_beagle_imputation.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Phasing Imputation Genotype Imputation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Phasing Imputation Genotype Imputation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Phasing Imputation Genotype Imputation this skillGPTomics/bioSkills1.2k1 repos~4.7kAutomated safety check: PassMIT
Gsd Phaseopen-gsd/gsd-core10k1 repos~603Automated safety check: NotesMIT
Gsd Execute Phaseopen-gsd/gsd-core10k1 repos~801Automated safety check: NotesMIT
Gsd Spec Phaseopen-gsd/gsd-core10k1 repos~658Automated safety check: NotesMIT
Gsd Mvp Phaseopen-gsd/gsd-core10k1 repos~505Automated safety check: NotesMIT
Gsd Secure Phaseopen-gsd/gsd-core10k2 repos~250Automated safety check: NotesMIT

Similar skills

  • Gsd Phase

    open-gsd/gsd-core

    Multi-phase management — add, insert, remove, or edit phases in ROADMAP.md (roadmap phase CRUD)

    10k GitHub starsUsed in 1 repo~603 tokens
    Product & Project ManagementAuto-check: notes
  • Gsd Execute Phase

    open-gsd/gsd-core

    SDD phase execution — execute all plans in a phase with dependency-aware wave parallelization

    10k GitHub starsUsed in 1 repo~801 tokens
    Agent WorkflowsAuto-check: notes
  • Gsd Spec Phase

    open-gsd/gsd-core

    Clarify WHAT a phase delivers with ambiguity scoring; produces a SPEC.md before discuss-phase.

    10k GitHub starsUsed in 1 repo~658 tokens
    Auto-check: notes
  • Gsd Mvp Phase

    open-gsd/gsd-core

    Plan a phase as a vertical MVP slice — user story, SPIDR splitting, then plan-phase

    10k GitHub starsUsed in 1 repo~505 tokens
    Product & Project ManagementAuto-check: notes
  • Gsd Secure Phase

    open-gsd/gsd-core

    Retroactively verify threat mitigations for a completed phase

    10k GitHub starsUsed in 2 repos~250 tokens
    Auto-check: notes
  • Gsd Ultraplan Phase

    open-gsd/gsd-core

    [BETA] Offload plan phase to Claude Code's ultraplan cloud; review in browser and import back.

    10k GitHub starsUsed in 1 repo~286 tokens
    Auto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Phasing Imputation Genotype Imputation

What does Bio Phasing Imputation Genotype Imputation do?

Imputes untyped genotypes against a phased reference panel with Beagle, Minimac4, or IMPUTE5 (array data) or from genotype likelihoods with GLIMPSE2, QUILT2, or STITCH (low-coverage WGS), producing…. Bio Phasing Imputation Genotype Imputation is an agent skill from GPTomics/bioSkills. Imputes untyped genotypes against a phased reference panel with Beagle, Minimac4, or IMPUTE5 (array data) or from genotype likelihoods with GLIMPSE2, QUILT2, or STITCH (low-coverage WGS), producing per-variant dosages (DS) with a self-estimated quality (Beagle DR2, Minimac R2, IMPUTE INFO).

When should I use Bio Phasing Imputation Genotype Imputation?

Bio Phasing Imputation Genotype Imputation fits situations like: increasing variant density for GWAS; harmonizing arrays; inferring untyped variants; imputing low-coverage sequence.

How do I install Bio Phasing Imputation Genotype Imputation in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-phasing-imputation-genotype-imputation -a claude-code`. Or copy the skill folder (phasing-imputation/genotype-imputation in GPTomics/bioSkills) into .claude/skills/bio-phasing-imputation-genotype-imputation in your project. Claude Code loads it when a task matches its description.

How do I install Bio Phasing Imputation Genotype Imputation in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-phasing-imputation-genotype-imputation -a codex`. Or copy the skill folder (phasing-imputation/genotype-imputation in GPTomics/bioSkills) into .agents/skills/bio-phasing-imputation-genotype-imputation in your project. Codex loads it when a task matches its description.

Can I use Bio Phasing Imputation Genotype Imputation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-phasing-imputation-genotype-imputation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-phasing-imputation-genotype-imputation, .gemini/skills/bio-phasing-imputation-genotype-imputation, .github/skills/bio-phasing-imputation-genotype-imputation and .opencode/skills/bio-phasing-imputation-genotype-imputation in your project.

What does Bio Phasing Imputation Genotype Imputation need to run?

Going by SKILL.md and its folder, Bio Phasing Imputation Genotype Imputation needs a shell for the scripts in its folder and the command-line tools its instructions call (java). Our summary lists: A Bash shell.

Does Bio Phasing Imputation Genotype Imputation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Phasing Imputation Genotype Imputation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Phasing Imputation Genotype Imputation use?

Bio Phasing Imputation Genotype Imputation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Phasing Imputation Genotype Imputation use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Phasing Imputation Genotype Imputation?

Skills that share tags, products or a category with Bio Phasing Imputation Genotype Imputation: Gsd Phase (open-gsd/gsd-core, 10k stars), Gsd Execute Phase (open-gsd/gsd-core, 10k stars), Gsd Spec Phase (open-gsd/gsd-core, 10k stars) and Gsd Mvp Phase (open-gsd/gsd-core, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Phasing Imputation Genotype Imputation?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.