Agent skill

Bio Phasing Imputation Haplotype Phasing

by GPTomics in GPTomics/bioSkills

Estimates haplotype phase from population linkage disequilibrium with SHAPEIT5, SHAPEIT4, Eagle2, or Beagle - turning unphased genotypes (0/1) into phased haplotypes (0|1) for imputation input…

MITAuto-check passedResearch & Science

Install Bio Phasing Imputation Haplotype Phasing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-phasing-imputation-haplotype-phasing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-phasing-imputation-haplotype-phasing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/phasing-imputation/haplotype-phasing .claude/skills/bio-phasing-imputation-haplotype-phasing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-phasing-imputation-haplotype-phasing
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.2k tokens
SKILL.md length
2,045 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Estimates haplotype phase from population linkage disequilibrium with SHAPEIT5, SHAPEIT4, Eagle2, or Beagle - turning unphased genotypes (0/1) into phased haplotypes (0|1) for imputation input…

  • Works in 3 steps: The genome-wide switch-error rate lies,… → The deliverable is a switch-error rate… → The modern arc is the scaffold design,…
  • Phasing genotypes before imputation
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 8 more sections
  • Runs Shell scripts from its folder

What it does

Bio Phasing Imputation Haplotype Phasing is an agent skill from GPTomics/bioSkills. Estimates haplotype phase from population linkage disequilibrium with SHAPEIT5, SHAPEIT4, Eagle2, or Beagle - turning unphased genotypes (0/1) into phased haplotypes (0|1) for imputation input, compound-heterozygote calls, HLA typing, or population genetics. Covers why statistical phase is an INFERENCE (not a measurement) whose error concentrates at rare variants, why a genome-wide switch-error rate hides catastrophic rare-variant error and must be reported MAC-stratified, the SHAPEIT5 common-scaffold-then-rare…

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/run_shapeit.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Phasing genotypes before imputation
  • For compound-het/ASE/HLA
  • Benchmarking against trios

Example prompts

  • “Use the bio-phasing-imputation-haplotype-phasing skill to estimate haplotype phase from population linkage disequilibrium with SHAPEIT5, SHAPEIT4…”
  • “/bio-phasing-imputation-haplotype-phasing”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The genome-wide switch-error rate lies, because it is dominated by easy common sites. A headline "switch error rate 0.3%" is averaged over…
  2. The deliverable is a switch-error rate against an independent truth set, not the tool name. "We used SHAPEIT" is not a switch-error rate…
  3. The modern arc is the scaffold design, and rare-variant phasing needs biobank scale to work at all. SHAPEIT5 phases common variants into a…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Phasing Imputation Haplotype Phasing loads about 4.2k tokens when it runs. Until then it costs about 256 tokens; SKILL.md has 2,045 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~256
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,045 words, ~4,164 tokens.

Download SKILL.mdSave it as .claude/skills/bio-phasing-imputation-haplotype-phasing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-phasing-imputation-haplotype-phasing
description
Estimates haplotype phase from population linkage disequilibrium with SHAPEIT5, SHAPEIT4, Eagle2, or Beagle - turning unphased genotypes (0/1) into phased haplotypes (0|1) for imputation input, compound-heterozygote calls, HLA typing, or population genetics. Covers why statistical phase is an INFERENCE (not a measurement) whose error concentrates at rare variants, why a genome-wide switch-error rate hides catastrophic rare-variant error and must be reported MAC-stratified, the SHAPEIT5 common-scaffold-then-rare design (phase_common, ligate, phase_rare, switch), reference-based vs within-cohort phasing, the build-matched genetic map, chrX male-haploid handling, and the switch-vs-flip-vs-Hamming distinction. Use when phasing genotypes before imputation, for compound-het/ASE/HLA, or benchmarking against trios. Read-backed / molecular phasing (long reads, Hi-C) is long-read-sequencing/haplotype-phasing; panel choice is reference-panels; imputation is genotype-imputation.
tool_type
cli
primary_tool
SHAPEIT5

Version Compatibility

Reference examples tested with: SHAPEIT5 5.1.1, Eagle 2.4.1, Beagle 5.4 (22Jul22), bcftools 1.19+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

SHAPEIT4 to SHAPEIT5 changed the CLI substantially: SHAPEIT5 is a SUITE of binaries (phase_common, phase_rare, ligate, switch), not a single shapeit command, and phase_common is the engine formerly known as SHAPEIT4. The genetic map and the reference panel must match the data's genome build (GRCh37 vs GRCh38); a build-mismatched map silently degrades phasing. PBWT and Ne defaults have drifted between betas; confirm against the installed --help.

Statistical Haplotype Phasing -- Inferring Phase From Population LD

"Resolve which alleles sit together on each chromosome" -> Estimate haplotype phase from population linkage disequilibrium via the Li-Stephens HMM - because phase is INFERRED statistically from how haplotypes are shared across a population, not read off the genotype, so a switch error is a model uncertainty (the rate, not zero, is the deliverable), not a typo.

  • CLI: phase_common --input target.bcf --filter-maf 0.001 --map chr20.b38.gmap.gz --region chr20 --output scaffold.bcf then ligate then phase_rare (SHAPEIT5), or Eagle2/Beagle for common-variant phasing

Scope: population/statistical phasing of array or sequence genotypes for imputation input, compound-het/ASE/HLA, and population genetics. Read-backed / molecular single-sample phasing (long reads, Hi-C, 10x linked reads) is a PHYSICALLY DIFFERENT signal -> long-read-sequencing/haplotype-phasing (the two are easily conflated; do not run SHAPEIT on long-read evidence or trust statistical phase for a private clinical variant). Panel choice -> reference-panels. Imputation against a panel -> genotype-imputation. The input VCF and biallelic normalization -> variant-calling/variant-normalization. End-to-end orchestration -> workflows/gwas-pipeline.

The Single Most Important Modern Insight -- A Phased Haplotype Is a Statistical Estimate, and Its Error Concentrates Exactly Where the Biology of Interest Lives

Statistical phasing reconstructs which alleles are on the same chromosome by borrowing LD across many individuals or a reference panel (Delaneau 2019 Nat Commun 10:5436). That works beautifully for common variants in LD with their neighbors and fails, by construction, for rare variants - which are young, carried by few people, and in LD with almost nothing. Three facts drive every decision:

  1. The genome-wide switch-error rate lies, because it is dominated by easy common sites. A headline "switch error rate 0.3%" is averaged over millions of common heterozygous sites and says nothing about the singleton or doubleton that is most likely to be the compound-het, the de-novo, or the pathogenic allele of interest - those are phased at MAC-dependent accuracy an order of magnitude worse, and a true singleton is essentially a coin flip without special machinery (Hofmeister 2023 Nat Genet 55:1243). Report accuracy stratified by minor allele count, never as one number.
  2. The deliverable is a switch-error rate against an independent truth set, not the tool name. "We used SHAPEIT" is not a switch-error rate. A switch error changes which haplotype an allele sits on without changing any genotype, so it is invisible to every per-site genotype QC; for any phase-dependent claim, measure the rate against a trio (Mendelian truth via switch --pedigree) or read-backed truth.
  3. The modern arc is the scaffold design, and rare-variant phasing needs biobank scale to work at all. SHAPEIT5 phases common variants into a fixed, near-perfect scaffold, then places each rare allele onto it by PBWT/IBD haplotype matching - which depends on finding a long shared haplotype, itself a function of cohort size. This is why rare-variant phasing in a small cohort cannot be trusted for a cis/trans call without orthogonal (trio or read-backed) evidence.

Tool Taxonomy

ToolCitationMechanism / roleWhen
SHAPEIT5Hofmeister 2023 Nat Genet 55:1243suite (phase_common/phase_rare/ligate/switch); scaffold design for rare/singleton phasing; PBWTbiobank-scale WGS/WES; rare-variant phasing
SHAPEIT4 (= phase_common engine)Delaneau 2019 Nat Commun 10:5436sub-linear common-variant phasing; integrates panels, scaffolds, read-backed phasecommon-variant phasing / pre-phasing; legacy
Eagle2Loh 2016 Nat Genet 48:1443HMM + PBWT-derived HapHedge; reference-based (--vcfRef) and within-cohortarray data; the classic imputation-server phaser
Beagle 5.xBrowning 2021 Am J Hum Genet 108:1880Java; does BOTH phasing (gt=, no ref=) and imputation; two-stage for sequenceone tool for phase and impute; no compile
Trio / pedigree phasing(Mendelian transmission)deterministic phase where the trio is informativegold standard; validating other phasers via switch
WhatsHap (boundary)Patterson 2015 J Comput Biol 22:498read-backed phasing (weighted MEC) from aligned reads-> long-read-sequencing/haplotype-phasing; can seed SHAPEIT as a scaffold

Decision Tree by Scenario

ScenarioRecommendedWhy
Array data, small-to-modest cohort, have a panelEagle2 --vcfRef or phase_common --referencea panel models LD better than a few thousand samples
Array data, large cohort, no panelEagle2 or phase_common within-cohortLD is modeled from the cohort; accuracy rises with N
WGS/WES, biobank scale, need rare variants phasedSHAPEIT5: phase_common -> ligate -> phase_rarethe scaffold design is the only route to accurate rare-variant phase
Pre-phasing as imputation inputEagle2 or Beagle 5small switch errors largely wash out in imputation -> genotype-imputation
One tool for phase and impute, no compileBeagle 5.x (gt= to phase, add ref= to impute)pragmatic single tool
Trio/pedigree availabletrio/pedigree phasing; use switch to benchmarkdeterministic where informative; the truth ruler
Long reads on the same sample-> long-read-sequencing/haplotype-phasing (then seed SHAPEIT as a scaffold)read-backed phase is local and deterministic; combine, do not replace
Common-variant phasing only, modest dataSHAPEIT4 or Beaglerare-variant machinery is unnecessary overhead

The Common-Scaffold-Then-Rare Design (SHAPEIT5)

Rare variants carry too little LD to phase in a joint model, and a joint HMM over millions of rare sites does not scale, so SHAPEIT5 splits the problem. Use the full pipeline when N > ~2,000; below that, phase_common alone suffices (too few rare-allele carriers for the rare step to add value).

  1. phase_common phases the common variants (e.g. --filter-maf 0.001) into accurate haplotypes - the scaffold. Run per chunk for large chromosomes, with OVERLAPPING regions.
  2. ligate stitches the per-chunk common scaffolds into one chromosome; chunks must overlap so ligate can resolve phase across the seam (a non-overlapping seam is a guaranteed switch).
  3. phase_rare takes the FULL genotypes plus the fixed scaffold and places each rare allele onto the already-phased common haplotypes by IBD matching. Do not filter rare variants out of the phase_rare input - placing them is the whole point.

Switch Error vs Flip vs Hamming -- the Metrics

A single rate hides the failure mode. Report more than one, and look at the distribution of switch positions.

MetricWhat it countsInflates on
Switch error rate (SER)fraction of consecutive het-site pairs whose phase relationship is wrongmany small local errors; the standard headline
Flip erroran isolated het phased wrong then immediately corrected (two switches one site apart)noisy single sites; double-counts in raw SER
Hamming errorfraction of het sites on the wrong haplotype under the best global alignmenta few LARGE block swaps - high Hamming, low switch count
Long switch / block flipa sustained segment on the wrong haplotypepoor long-range LD; ruinous for cis/trans yet only 2 switches

SER and Hamming measure different sins: many tiny flips give high SER but modest Hamming; one half-chromosome block swap gives catastrophic Hamming but only two switches. Het density matters too - SER is per-het-pair, so sparse het sites mean the same SER spans more bp. Typical magnitudes (order-of-magnitude, dataset-specific): Eagle2 + HRC reference, European array 1.36%; Eagle2 within-cohort N5,000 1.5%; within-cohort N150,000 (UK Biobank) ~0.27-0.35%; SHAPEIT5 for a variant in ~1 of 100,000 < ~5%. The pattern: common-variant phasing in a big cohort is sub-1%; rare-variant phasing is single-digit-percent at best and worsens steeply as MAC approaches 1.

Show full SKILL.md (796 more words)Show less

Reference-Based vs Within-Cohort

Reference-based phasing wins when the cohort is small (a few thousand samples cannot model LD as well as a 32k-100k+ haplotype panel); phase against the biggest ancestry-matched panel available (Eagle2 --vcfRef). Within-cohort phasing wins when the cohort is large and ancestry-matched to itself, because accuracy rises monotonically with N; by UK-Biobank scale within-cohort is more accurate than any external panel. The crossover is in the tens of thousands. Ancestry match dominates either way - a mismatched panel phases worse than a smaller matched one or within-cohort -> reference-panels.

Per-Method Failure Modes

Genome-wide SER trusted for a rare-variant call

Trigger: quoting one switch-error rate and treating all haplotypes as equally trustworthy. Mechanism: SER is dominated by easy common sites; rare-variant phase is far worse and MAC-dependent. Symptom: a confident compound-het (cis/trans) call from a small-cohort statistical phase that is actually near chance. Fix: stratify accuracy by MAC; confirm rare-variant cis/trans with a trio or read-backed phase.

Wrong-build or flat genetic map

Trigger: a GRCh37 map on GRCh38 data, or a uniform map "for simplicity". Mechanism: the map sets the HMM's recombination (transition) rates; wrong coordinates or a flat rate mis-place where haplotype breaks are expected. Symptom: degraded phasing, more long switches, no error message. Fix: use the build-matched per-chromosome map shipped with the tool; the default population map is right.

Non-overlapping ligate seam

Trigger: chunking a chromosome with abutting (non-overlapping) regions. Mechanism: ligate needs overlap to resolve the phase relationship across the seam. Symptom: a guaranteed switch at every chunk boundary. Fix: make --region / --input-region / --scaffold-region overlap between adjacent chunks.

chrX male coded diploid

Trigger: phasing male chrX non-PAR as diploid heterozygous. Mechanism: males are haploid outside the PARs; a het call there is biologically impossible. Symptom: corrupted male chrX phase. Fix: pass the male sample list (SHAPEIT5 --haploids; Eagle handles mixed ploidy); keep PAR1/PAR2 as separate diploid regions with build-correct coordinates.

Multiallelic records fed to a phaser

Trigger: phasing raw multiallelic sites. Mechanism: phasers expect biallelic records; a multiallelic record is undefined behavior. Symptom: tool errors or mis-phased sites. Fix: bcftools norm -m -any to split and left-align first -> variant-calling/variant-normalization.

Quantitative Thresholds

ThresholdSourceRationale
--filter-maf 0.001 defines the common/rare scaffold splitSHAPEIT5 docscommon variants build the accurate scaffold; rarer variants are phased onto it
Use phase_common -> ligate -> phase_rare when N > ~2,000SHAPEIT5 docsbelow that, too few rare-allele carriers for the rare step to help
Report SER stratified by MAC, not genome-wideHofmeister 2023 Nat Genet 55:1243phasing quality is a steep function of MAC; a single number hides rare-variant failure
Eagle2 --Kpbwt default 10000 (raise at large N)Loh 2016 Nat Genet 48:1443more conditioning haplotypes raise accuracy at biobank scale
phase_rare --effective-size ~15000 (verify)SHAPEIT5 docsNe sets expected recombination; often tuned per dataset, confirm with --help
Genetic map must match the data buildDelaneau 2019 Nat Commun 10:5436a build-mismatched map mis-assigns recombination rates silently

Common Errors

Error / symptomCauseSolution
Switch at every chunk boundarynon-overlapping ligate seamsoverlap adjacent chunk regions
Corrupted male chrX phasemale non-PAR coded diploidpass --haploids; split PAR/nonPAR
Phaser errors on some sitesmultiallelic recordsbcftools norm -m -any first
Rare-variant cis/trans call does not replicatesmall-cohort statistical phase of rare variantsuse SHAPEIT5 at scale; confirm with trio/read-backed
Phasing mysteriously bad in one regionwrong-build or flat genetic mapbuild-match the map
SHAPEIT4 syntax fails under SHAPEIT5SHAPEIT5 split into phase_common/phase_rare/ligateuse the suite binaries, not a single shapeit
Beagle OutOfMemoryErrorJVM heap too small / whole genome in one jobraise -Xmx; phase per chromosome

References

  • Hofmeister RJ, Ribeiro DM, Rubinacci S, Delaneau O. 2023. Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank. Nat Genet 55:1243-1249.
  • Delaneau O, Zagury JF, Robinson MR, Marchini JL, Dermitzakis ET. 2019. Accurate, scalable and integrative haplotype estimation. Nat Commun 10:5436.
  • Loh PR, Danecek P, Palamara PF, et al. 2016. Reference-based phasing using the Haplotype Reference Consortium panel. Nat Genet 48:1443-1448.
  • Browning BL, Tian X, Zhou Y, Browning SR. 2021. Fast two-stage phasing of large-scale sequence data. Am J Hum Genet 108:1880-1890.
  • Durbin R. 2014. Efficient haplotype matching and storage using the positional Burrows-Wheeler transform (PBWT). Bioinformatics 30:1266-1272.
  • Patterson M, Marschall T, Pisanti N, et al. 2015. WhatsHap: weighted haplotype assembly for future-generation sequencing reads. J Comput Biol 22:498-509.
  • Li N, Stephens M. 2003. Modeling linkage disequilibrium and identifying recombination hotspots using single-nucleotide polymorphism data. Genetics 165:2213-2233.
  • reference-panels - Select the ancestry-matched panel that reference-based phasing copies from
  • genotype-imputation - Imputation consumes the phased haplotypes (pre-phasing)
  • imputation-qc - Switch-error benchmarking sits alongside imputation quality QC
  • long-read-sequencing/haplotype-phasing - Read-backed / molecular single-sample phasing (a different signal)
  • variant-calling/variant-normalization - Split multiallelics and left-align before phasing
  • causal-genomics/fine-mapping - Phased haplotypes feed haplotype-level fine-mapping
  • clinical-databases/hla-typing - HLA typing is a high-stakes consumer of long-range phase
  • workflows/gwas-pipeline - End-to-end QC -> phase -> impute -> associate

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in phasing-imputation/haplotype-phasing of GPTomics/bioSkills.

  • SKILL.md
  • examples/run_shapeit.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Phasing Imputation Haplotype Phasing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Phasing Imputation Haplotype Phasing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Phasing Imputation Haplotype Phasing this skillGPTomics/bioSkills1.2k1 repos~4.2kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Phasing Imputation Haplotype Phasing

What does Bio Phasing Imputation Haplotype Phasing do?

Estimates haplotype phase from population linkage disequilibrium with SHAPEIT5, SHAPEIT4, Eagle2, or Beagle - turning unphased genotypes (0/1) into phased haplotypes (0|1) for imputation input…. Bio Phasing Imputation Haplotype Phasing is an agent skill from GPTomics/bioSkills. Estimates haplotype phase from population linkage disequilibrium with SHAPEIT5, SHAPEIT4, Eagle2, or Beagle - turning unphased genotypes (0/1) into phased haplotypes (0|1) for imputation input, compound-heterozygote calls, HLA typing, or population genetics.

When should I use Bio Phasing Imputation Haplotype Phasing?

Bio Phasing Imputation Haplotype Phasing fits situations like: phasing genotypes before imputation; for compound-het/ASE/HLA; benchmarking against trios.

How do I install Bio Phasing Imputation Haplotype Phasing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-phasing-imputation-haplotype-phasing -a claude-code`. Or copy the skill folder (phasing-imputation/haplotype-phasing in GPTomics/bioSkills) into .claude/skills/bio-phasing-imputation-haplotype-phasing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Phasing Imputation Haplotype Phasing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-phasing-imputation-haplotype-phasing -a codex`. Or copy the skill folder (phasing-imputation/haplotype-phasing in GPTomics/bioSkills) into .agents/skills/bio-phasing-imputation-haplotype-phasing in your project. Codex loads it when a task matches its description.

Can I use Bio Phasing Imputation Haplotype Phasing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-phasing-imputation-haplotype-phasing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-phasing-imputation-haplotype-phasing, .gemini/skills/bio-phasing-imputation-haplotype-phasing, .github/skills/bio-phasing-imputation-haplotype-phasing and .opencode/skills/bio-phasing-imputation-haplotype-phasing in your project.

What does Bio Phasing Imputation Haplotype Phasing need to run?

Going by SKILL.md and its folder, Bio Phasing Imputation Haplotype Phasing needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Bio Phasing Imputation Haplotype Phasing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Phasing Imputation Haplotype Phasing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Phasing Imputation Haplotype Phasing use?

Bio Phasing Imputation Haplotype Phasing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Phasing Imputation Haplotype Phasing use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Phasing Imputation Haplotype Phasing?

Skills that share tags, products or a category with Bio Phasing Imputation Haplotype Phasing: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Phasing Imputation Haplotype Phasing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.