Agent skill

Bio Variant Calling

by GPTomics in GPTomics/bioSkills

Call germline SNPs and indels from a BAM/CRAM with bcftools mpileup and call, and select the right calling engine for the job.

MITAuto-check passedResearch & Science

Install Bio Variant Calling

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-variant-calling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-variant-calling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/variant-calling/variant-calling .claude/skills/bio-variant-calling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-variant-calling
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.1k tokens
SKILL.md length
1,697 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Call germline SNPs and indels from a BAM/CRAM with bcftools mpileup and call, and select the right calling engine for the job.

  • Works in 3 steps: Normalize - left-align and split… → Filter - bcftools produces no VQSR/DL… → Inspect / query - counts, Ti/Tv,…
  • Generating a VCF from aligned reads
  • SKILL.md covers Version Compatibility, The Governing Principle, Engine Selection (the decision… and bcftools mpileup + call, plus 9 more sections
  • Runs Shell scripts from its folder

What it does

Bio Variant Calling is an agent skill from GPTomics/bioSkills. Call germline SNPs and indels from a BAM/CRAM with bcftools mpileup and call, and select the right calling engine for the job. Use when generating a VCF from aligned reads, choosing between bcftools, GATK HaplotypeCaller, DeepVariant, and DRAGEN, setting ploidy for haploid/organelle/polyploid/sex-chromosome calling, or deciding whether pileup-based calling is good enough versus a local-reassembly caller for indels and difficult regions. Not for cohort joint genotyping (see variant-calling/joint-calling)…

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/call_variants.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Generating a VCF from aligned reads
  • Choosing between bcftools
  • GATK HaplotypeCaller
  • Setting ploidy for haploid/organelle/polyploid/sex-chromosome calling

Example prompts

  • “/bio-variant-calling”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Normalize - left-align and split multiallelics so identical variants have identical records: bcftools norm -f reference.fa -m -any…
  2. Filter - bcftools produces no VQSR/DL score, so apply quality/depth/strand hard filters (e.g. QUAL, FORMAT/DP, SP) suited to the depth and…
  3. Inspect / query - counts, Ti/Tv, per-sample stats. See variant-calling/vcf-basics and variant-calling/vcf-statistics.

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Variant Calling loads about 4.1k tokens when it runs. Until then it costs about 171 tokens; SKILL.md has 1,697 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~171
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,697 words, ~4,125 tokens.

Download SKILL.mdSave it as .claude/skills/bio-variant-calling/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-variant-calling
description
Call germline SNPs and indels from a BAM/CRAM with bcftools mpileup and call, and select the right calling engine for the job. Use when generating a VCF from aligned reads, choosing between bcftools, GATK HaplotypeCaller, DeepVariant, and DRAGEN, setting ploidy for haploid/organelle/polyploid/sex-chromosome calling, or deciding whether pileup-based calling is good enough versus a local-reassembly caller for indels and difficult regions. Not for cohort joint genotyping (see variant-calling/joint-calling), GATK-specific workflows (see variant-calling/gatk-variant-calling), deep-learning calling (see variant-calling/deepvariant), or somatic/low-VAF detection.
tool_type
cli
primary_tool
bcftools

Version Compatibility

Reference examples tested with: bcftools 1.19+

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Note: bcftools mpileup applies BAQ (per-Base Alignment Quality) by default; this is a real behavior that changes calls, not a nuisance flag (see The Governing Principle).

Variant Calling from a BAM

"Call SNPs and indels from my aligned reads" -> Compute per-position genotype likelihoods from a BAM/CRAM against the reference, then call variant sites under a Bayesian model at the assumed ploidy.

  • CLI (fast, position-based): bcftools mpileup -f ref.fa in.bam | bcftools call -mv
  • CLI (reassembly, higher indel accuracy): GATK HaplotypeCaller (variant-calling/gatk-variant-calling)
  • CLI (deep learning): DeepVariant (variant-calling/deepvariant)

This skill does the bcftools calling and is the engine-selection hub: it tells the agent when pileup calling is the right tool and when to hand off to a reassembly or deep-learning caller.

The Governing Principle

There are two families of short-variant caller, and the choice between them is the single most consequential decision here.

  • Position-based genotype-likelihood callers (bcftools mpileup|call, the old samtools/UnifiedGenotyper lineage) trust the aligner's per-read placement. At each reference position they tally the pileup, compute P(reads | genotype) per site, and call under a Bayesian model. Fast, transparent, no training data. But the mapper places each read greedily and independently, so an indel near a read end or inside a repeat is placed inconsistently across reads, and the per-position model cannot repair that. bcftools mitigates it with BAQ (per-Base Alignment Quality: base qualities near a likely misalignment are downweighted so a shaky column does not produce a confident false SNP), but BAQ suppresses false positives rather than reconstructing the true indel.
  • Local-reassembly / haplotype callers (GATK HaplotypeCaller, DeepVariant, DRAGEN) discard the local alignment in an active region and re-derive it: assemble candidate haplotypes, realign every read to them (PairHMM or a learned model), then genotype. The indel is represented once, on the assembled haplotype, instead of (mis)placed per read. This is precisely why they beat pileup callers on indels, clustered variants, and difficult regions.

Consequence: bcftools is fine-to-excellent for simple germline SNPs and quick genome-wide scans, materially weaker on indels and in low-complexity / segmental-duplication / MHC regions, and not built for somatic low-VAF detection or scalable cohort joint calling. Pick the engine from the analysis, not from habit.

Engine Selection (the decision that comes before any command)

Guidance, not dogma; on a production human pipeline, validate against current GIAB/GA4GH benchmarks (hap.py + vcfeval) before committing.

EngineBest whenFails / weak whenHand off to
bcftools mpileup|callSimple germline SNPs; non-model/organelle/microbial genomes (no training data, any ploidy); quick exploratory scans; low compute; small multi-sample setsIndels in homopolymers/STRs; segdups, MHC, low-mappability; low-VAF somatic/mosaic; cohorts beyond ~100 samplesthis skill
GATK HaplotypeCallerAuditable open-source human WGS/WES; every parameter inspectable; the joint-calling/best-practices orthodoxy (GVCF -> GenomicsDBImport -> GenotypeGVCFs)Lower indel/difficult-region accuracy than DeepVariant/DRAGEN; local assembly can abort in pathological high-depth/repeat regionsvariant-calling/gatk-variant-calling; variant-calling/joint-calling
DeepVariantBest open-source accuracy on indels and difficult regions; PacBio HiFi / ONT (platform-specific trained models); generalizes off one training sampleNeeds the correct platform model (wrong model degrades accuracy); GPU helps; cohort merge needs GLnexus, not GenotypeGVCFsvariant-calling/deepvariant
DRAGENMaximum throughput on Illumina (FPGA, ~20-25 min/genome); leads difficult-to-map benchmarks (alt-aware mapping)Proprietary/hardware- or license-gated; ML recalibrator trained on GIAB truth (benchmark-overfitting caveat)vendor pipeline; HaplotypeCaller --dragen-mode for an open-source approximation

Honest state of the field: DeepVariant and DRAGEN lead on indels and difficult regions; GATK is the joint-calling and best-practices reference everyone else is measured against; bcftools wins on speed, simplicity, non-model organisms, and organelle/haploid calling. On easy SNPs every modern caller exceeds F1 0.999, so a caller's headline SNP number is rarely the deciding factor - indels and hard regions are.

bcftools mpileup + call

Goal: Detect germline SNPs and indels from aligned reads with the pileup-and-call pipeline.

Approach: Generate per-position genotype likelihoods with mpileup (BAQ on by default), pipe as uncompressed BCF into the multiallelic caller.

Basic calling
bash
bcftools mpileup -f reference.fa input.bam | bcftools call -mv -Oz -o variants.vcf.gz
bcftools index variants.vcf.gz
bash
# -Ou between steps avoids VCF (de)serialization; -q/-Q drop poorly-supported reads/bases;
# -a requests the FORMAT tags downstream filtering needs (DP, allelic depths, strand-bias p)
bcftools mpileup -Ou -f reference.fa \
    -q 20 -Q 20 \
    -a FORMAT/DP,FORMAT/AD,FORMAT/SP \
    input.bam | \
bcftools call -mv -Oz -o variants.vcf.gz
bcftools index variants.vcf.gz
Region-restricted and multi-sample calling
bash
# Single region / BED targets
bcftools mpileup -f reference.fa -r chr1:1000000-2000000 input.bam | bcftools call -mv -Oz -o region.vcf.gz
bcftools mpileup -f reference.fa -R targets.bed input.bam | bcftools call -mv -Oz -o targets.vcf.gz

# Multiple BAMs (small cohorts only; see The Governing Principle for the scaling limit)
bcftools mpileup -f reference.fa sample1.bam sample2.bam sample3.bam | bcftools call -mv -Oz -o cohort.vcf.gz

# BAM list file: one path per line
bcftools mpileup -f reference.fa -b bams.txt | bcftools call -mv -Oz -o cohort.vcf.gz

The mpileup / call flags that change results

StageFlagEffect
mpileup-f ref.faReference FASTA (required); must be the exact one used for alignment
mpileup-q INTMin mapping quality; -q 20 drops ambiguously placed reads (paralog mismapping)
mpileup-Q INTMin base quality; -Q 20 drops low-confidence base calls
mpileup-a LISTExtra FORMAT/INFO tags: FORMAT/AD (allelic depths), FORMAT/DP, FORMAT/SP (Phred strand-bias p), FORMAT/ADF/ADR (per-strand), INFO/AD
mpileup-d INTMax per-file depth (default 250); set to 3-4x expected mean coverage to avoid truncating high-coverage sites
mpileup-B / -E-B disables BAQ (more raw indel signal, more false SNPs near indels); -E recomputes BAQ on the fly (more sensitive, slower)
call-mMultiallelic caller - default, recommended for all new work
call-cConsensus caller - legacy; only for reproducing old pipelines
call-vEmit variant sites only (omit to emit all sites, e.g. for hom-ref confidence)
call-O z|b|u|vOutput: z bgzipped VCF, b BCF, u uncompressed BCF (piping), v VCF
call--ploidy / --ploidy-fileSample/region ploidy (below)
call-P FLOATMutation-rate prior (default 1.1e-3, human); lower for inbred lines, raise for diverse/outbred populations

The multiallelic caller (-m) handles sites with several ALT alleles natively and is statistically superior; the consensus caller (-c) exists only for backward reproducibility.

Ploidy: sample, organelle, and sex-chromosome calling

Goal: Match the caller's ploidy to the biology so genotypes are representable.

Approach: Set a scalar ploidy for uniform samples, or a ploidy file (or built-in preset) to vary ploidy by region and sex.

Wrong ploidy silently corrupts calls: calling a diploid as haploid halves heterozygous sensitivity; calling a haploid/hemizygous region as diploid manufactures false heterozygous calls from every error and paralog mismap.

bash
# Haploid: bacteria, mitochondria (nuclear germline heteroplasmy caveat below), non-PAR chrX/chrY in a male
bcftools mpileup -f reference.fa input.bam | bcftools call -m --ploidy 1 -Oz -o haploid.vcf.gz

# Built-in human preset applies karyotype-aware sex-chromosome ploidy
bcftools call -m --ploidy GRCh38 ...

# Ploidy file: CHROM  FROM  TO  SEX  PLOIDY  (chrY absent in females -> 0)
#   chrX  1  -1  M  1
#   chrX  1  -1  F  2
#   chrY  1  -1  M  1
#   chrY  1  -1  F  0
#   *     1  -1  *  2
bcftools mpileup -f reference.fa input.bam | bcftools call -m --ploidy-file ploidy.txt -Oz -o sexaware.vcf.gz

Scope notes: true mitochondrial heteroplasmy is continuous-VAF (not 0/0.5/1) and is a somatic-shaped signal - a diploid or haploid genotype model cannot express it; use a somatic caller (GATK Mutect2 --mitochondria-mode) for real heteroplasmy work. Polyploid/pooled samples need --ploidy N set to the true copy number so dosage/allele-count is preserved rather than collapsed to het.

Show full SKILL.md (690 more words)Show less

After calling: the pipeline map

A raw caller VCF is not a finished callset. The standard downstream order:

  1. Normalize - left-align and split multiallelics so identical variants have identical records: bcftools norm -f reference.fa -m -any variants.vcf.gz -Oz -o norm.vcf.gz. Do this before ANY comparison, annotation, or merge. See variant-calling/variant-normalization.
  2. Filter - bcftools produces no VQSR/DL score, so apply quality/depth/strand hard filters (e.g. QUAL, FORMAT/DP, SP) suited to the depth and platform. See variant-calling/filtering-best-practices.
  3. Inspect / query - counts, Ti/Tv, per-sample stats. See variant-calling/vcf-basics and variant-calling/vcf-statistics.

Comparing callers honestly

If the point of choosing bcftools vs a reassembly caller is accuracy, compare them correctly - this is where naive analyses go wrong:

  • Normalize both callsets first (bcftools norm -f ref.fa -m -any). Two VCFs can encode the identical haplotype with different records (indel placement in repeats, MNP vs split SNVs); un-normalized records mismatch spuriously.
  • Use haplotype-aware benchmarking, not bcftools isec. A line-diff / isec on raw records overcounts both false positives and false negatives from representation alone. Score against a GIAB truth set with hap.py + vcfeval inside the confident-region BED (Krusche 2019), reporting SNVs and indels separately.
  • Stratify. A genome-wide F1 hides the differences that matter - they live in indels-in-repeats, segdups, and MHC. Report per-region, not one headline number.

Performance

Goal: Speed up calling on large inputs.

Approach: Pipe uncompressed BCF between stages, thread both tools, and shard by chromosome.

bash
# Threaded, uncompressed-BCF pipe
bcftools mpileup -Ou -f reference.fa --threads 4 input.bam | \
    bcftools call -mv --threads 4 -Oz -o variants.vcf.gz

# Parallel by chromosome, then concatenate
for chr in chr1 chr2 chr3; do
    bcftools mpileup -Ou -f reference.fa -r "$chr" input.bam | \
        bcftools call -mv -Oz -o "${chr}.vcf.gz" &
done
wait
bcftools concat -Oz -o all.vcf.gz chr*.vcf.gz
bcftools index all.vcf.gz

Difficult regions (know where pileup calling breaks)

  • Homopolymers / STRs - the dominant indel false-positive source; slippage + mapping ambiguity + representation ambiguity all concentrate here. Validate indels in homopolymers >6 bp with a reassembly caller or manual review, or switch engines.
  • Segmental duplications / low mappability - paralog reads pile up and manufacture false SNPs; sites with mean MQ <40 signal ambiguous mapping. -q 20 helps; a reassembly/alt-aware caller helps more.
  • MHC and other hyper-polymorphic loci - extreme divergence from the reference; expect reduced recall from any linear-reference caller.
  • High-depth regions - set -d to 3-4x expected mean coverage; the default 250 truncates deep targeted panels and can bias likelihoods.

Common Errors

SymptomCauseFix
no FASTA reference-f omittedAdd -f reference.fa
[E::faidx] ... different number of sequences / reference mismatchmpileup reference != alignment referenceUse the exact FASTA the BAM was aligned to; compare @SQ in samtools view -H against grep '^>' ref.fa
No variants calledCoverage too low, -q/-Q too strict, empty/wrong BAMCheck samtools depth; relax -q/-Q; confirm reference build
False heterozygous calls everywhere on chrX/chrY (male)Non-PAR sex chromosome called as diploidSet --ploidy 1 for non-PAR, or use a --ploidy-file / --ploidy GRCh38
Excess indel false positives in repeatsPosition-based limitation, not a bugNormalize + hard-filter; validate or recall indels with a reassembly caller
Downstream tools disagree on the same variantRecords not normalizedbcftools norm -f ref.fa -m -any before comparing/merging/annotating
  • variant-calling/vcf-basics - View and query the resulting VCF
  • variant-calling/variant-normalization - Left-align and split multiallelics before comparison
  • variant-calling/filtering-best-practices - Hard-filter a bcftools callset (no VQSR/DL score)
  • variant-calling/vcf-statistics - Ti/Tv, counts, and callset QC
  • variant-calling/gatk-variant-calling - Local-reassembly calling with HaplotypeCaller and DRAGEN-GATK mode
  • variant-calling/deepvariant - Deep-learning caller; best indel/difficult-region accuracy, long-read models
  • variant-calling/joint-calling - Scalable cohort genotyping (GVCF workflow, GLnexus)
  • alignment-files/pileup-generation - Alternative pileup generation
  • read-alignment/bwa-alignment - Upstream mapping that determines calling quality

References

  • Li H. A statistical framework for SNP calling, mutation discovery, association mapping and population genetical parameter estimation from sequencing data. Bioinformatics 27(21):2987-2993 (2011). DOI 10.1093/bioinformatics/btr509. (The mpileup genotype-likelihood model.)
  • Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, McCarthy SA, Davies RM, Li H. Twelve years of SAMtools and BCFtools. GigaScience 10(2):giab008 (2021). DOI 10.1093/gigascience/giab008. (bcftools mpileup/call/norm implementation.)
  • DePristo MA, Banks E, Poplin R, Garimella KV, Maguire JR, Hartl C, et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nature Genetics 43(5):491-498 (2011). DOI 10.1038/ng.806. (Local-reassembly genotyping framework - the reassembly contrast.)
  • Poplin R, Chang P-C, Alexander D, Schwartz S, Colthurst T, Ku A, et al. A universal SNP and small-indel variant caller using deep neural networks. Nature Biotechnology 36(10):983-987 (2018). DOI 10.1038/nbt.4235. (DeepVariant.)
  • Krusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, et al. Best practices for benchmarking germline small-variant calls in human genomes. Nature Biotechnology 37:555-560 (2019). DOI 10.1038/s41587-019-0054-x. (hap.py/vcfeval, confident-region model, normalize-before-compare.)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in variant-calling/variant-calling of GPTomics/bioSkills.

  • SKILL.md
  • examples/call_variants.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Variant Calling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Variant Calling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Variant Calling this skillGPTomics/bioSkills1.2k1 repos~4.1kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Variant Calling

What does Bio Variant Calling do?

Call germline SNPs and indels from a BAM/CRAM with bcftools mpileup and call, and select the right calling engine for the job. Bio Variant Calling is an agent skill from GPTomics/bioSkills. Call germline SNPs and indels from a BAM/CRAM with bcftools mpileup and call, and select the right calling engine for the job.

When should I use Bio Variant Calling?

Bio Variant Calling fits situations like: generating a VCF from aligned reads; choosing between bcftools; GATK HaplotypeCaller; setting ploidy for haploid/organelle/polyploid/sex-chromosome calling.

How do I install Bio Variant Calling in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-variant-calling -a claude-code`. Or copy the skill folder (variant-calling/variant-calling in GPTomics/bioSkills) into .claude/skills/bio-variant-calling in your project. Claude Code loads it when a task matches its description.

How do I install Bio Variant Calling in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-variant-calling -a codex`. Or copy the skill folder (variant-calling/variant-calling in GPTomics/bioSkills) into .agents/skills/bio-variant-calling in your project. Codex loads it when a task matches its description.

Can I use Bio Variant Calling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-variant-calling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-variant-calling, .gemini/skills/bio-variant-calling, .github/skills/bio-variant-calling and .opencode/skills/bio-variant-calling in your project.

What does Bio Variant Calling need to run?

Going by SKILL.md and its folder, Bio Variant Calling needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Bio Variant Calling access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Variant Calling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Variant Calling use?

Bio Variant Calling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Variant Calling use?

About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Variant Calling?

Skills that share tags, products or a category with Bio Variant Calling: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Variant Calling?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.