Agent skill

Gatk Variant Calling

by jaechang-hits in jaechang-hits/SciAgent-Skills

GATK Best Practices for germline SNP/indel calling from WGS/WES BAMs.

BSD-3-ClauseAuto-check passedResearch & Science

Install Gatk Variant Calling

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill gatk-variant-calling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills gatk-variant-calling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/gatk-variant-calling .claude/skills/gatk-variant-calling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gatk-variant-calling
GitHub stars
374
Used in
1 other repo
Token cost
~3.7k tokens
SKILL.md length
782 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

GATK Best Practices for germline SNP/indel calling from WGS/WES BAMs.

  • Works in 6 steps: Base Quality Score Recalibration (BQSR) → Call Variants with HaplotypeCaller (GVCF… → Consolidate GVCFs with GenomicsDBImport → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 7 more sections
  • Calls wget and conda; reaches github.com

What it does

Gatk Variant Calling is an agent skill from jaechang-hits/SciAgent-Skills. GATK Best Practices for germline SNP/indel calling from WGS/WES BAMs. Per-sample HaplotypeCaller GVCFs, GenomicsDBImport, GenotypeGVCFs joint calling, VQSR or hard filters. Requires BWA-MEM2-aligned, markdup, BQSR BAMs. Use DeepVariant for a faster DL alternative; GATK is the NIH/ENCODE standard.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/gatk-variant-calling”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Base Quality Score Recalibration (BQSR)
  2. Call Variants with HaplotypeCaller (GVCF Mode)
  3. Consolidate GVCFs with GenomicsDBImport
  4. Joint Genotyping with GenotypeGVCFs
  5. Variant Filtration (Hard Filters)
  6. Parse VCF Results with Python

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • wget
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • gatk.broadinstitute.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gatk Variant Calling loads about 3.7k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 782 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 782 words, ~3,699 tokens.

Download SKILL.mdSave it as .claude/skills/gatk-variant-calling/SKILL.md (or your agent's skills folder).
name
gatk-variant-calling
description
GATK Best Practices for germline SNP/indel calling from WGS/WES BAMs. Per-sample HaplotypeCaller GVCFs, GenomicsDBImport, GenotypeGVCFs joint calling, VQSR or hard filters. Requires BWA-MEM2-aligned, markdup, BQSR BAMs. Use DeepVariant for a faster DL alternative; GATK is the NIH/ENCODE standard.
license
BSD-3-Clause

GATK — Germline Variant Calling Pipeline

Overview

GATK (Genome Analysis Toolkit) implements the GATK Best Practices workflow for calling SNPs and indels from Illumina WGS and WES data. The pipeline runs HaplotypeCaller per sample (producing GVCF files), consolidates GVCFs with GenomicsDBImport, performs joint genotyping with GenotypeGVCFs, and filters variants with VQSR (Variant Quality Score Recalibration) or hard filters. GATK requires BWA-MEM2-aligned, duplicate-marked, and base quality score recalibrated (BQSR) BAM files as input. It integrates with Picard tools, samtools, and bcftools for pre- and post-processing. The GATK4 workflow is the NIH/ENCODE standard for germline variant calling in research and clinical genomics.

When to Use

  • Calling germline SNPs and indels from WGS or WES samples for population genetics or clinical variant analysis
  • Running joint genotyping across multiple samples for cohort-scale studies (families, case-control)
  • Applying base quality score recalibration (BQSR) to improve variant calling accuracy before HaplotypeCaller
  • Generating GVCF files for scalable cohort expansion: add new samples without reprocessing existing ones
  • Producing variant call sets for downstream annotation with Ensembl VEP, ANNOVAR, or SnpEff
  • Use DeepVariant (Google) instead for a faster deep-learning approach with comparable accuracy
  • Use bcftools call instead for rapid variant calling without assembly-based local realignment

Prerequisites

  • Software: GATK4, Java 17+, samtools, BWA-MEM2
  • Reference files: genome FASTA + known variants VCF (dbSNP, 1000G, Mills indels)
  • Input: duplicate-marked, sorted BAM with @RG read group headers (from BWA-MEM2)

Check before installing: The tool may already be available in the current environment (e.g., inside a pixi / conda env). Run command -v gatk first and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool via pixi run gatk rather than bare gatk.

bash
# Install GATK4
wget https://github.com/broadinstitute/gatk/releases/download/4.6.0.0/gatk-4.6.0.0.zip
unzip gatk-4.6.0.0.zip
export GATK="$PWD/gatk-4.6.0.0/gatk"

# Or with conda
conda install -c bioconda gatk4

# Verify
gatk --version
# GATK v4.6.0.0

# Download GATK resource bundle files (GRCh38)
# From gs://gcp-public-data--broad-references/hg38/v0/ (requires gsutil or Broad FTP)

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: referenceBundle
    kind: derived
    source: upstream
    ask: "Which reference FASTA were the BAMs aligned to?"
    default: "carried from the alignment stage; must match exactly"

  - id: D2
    param: callingIntervals
    kind: required
    source: user
    ask: "Call across the whole genome, or restrict to capture targets or regions of interest?"
    default: null

  - id: D3
    param: cohortMode
    kind: required
    source: user
    ask: "Call each sample on its own, or emit GVCFs now and genotype the cohort jointly afterwards?"
    default: "joint genotyping when more than one sample is being analysed together"

  - id: D4
    param: filteringStrategy
    kind: required
    source: user
    depends_on: [D3]
    ask: "Filter calls with fixed hard thresholds, or train a variant-recalibration model on the cohort?"
    default: "hard filters; recalibration needs a cohort large enough to train on"

  - id: D5
    param: knownSites
    kind: optional
    source: user
    ask: "Annotate calls against a known-variant resource such as dbSNP?"
    default: "none"

  - id: D6
    param: callConfidence
    kind: optional
    source: user
    ask: "How confident must a genotype be before the variant is emitted at all?"
    default: 30

  - id: D7
    param: threadsAndHeap
    kind: never_ask
    source: data
    reason: "Pair-HMM threads and JVM heap size affect runtime and memory, not the calls"
    default: "4 threads, heap sized to the genome"

D4 hangs on D3 because variant recalibration is a cohort method - it has nothing to train on for a single sample, so the filtering choice is only meaningful once the cohort question is answered.

Quick Start

bash
GENOME="GRCh38.fa"
DBSNP="dbsnp_146.hg38.vcf.gz"

# Run HaplotypeCaller in GVCF mode
gatk HaplotypeCaller \
    -R $GENOME \
    -I sample1.markdup.bam \
    -O sample1.g.vcf.gz \
    -ERC GVCF \
    --dbsnp $DBSNP \
    --native-pair-hmm-threads 4

echo "GVCF: sample1.g.vcf.gz"

Workflow

Step 1: Base Quality Score Recalibration (BQSR)

Correct systematic errors in base quality scores before variant calling.

bash
GENOME="GRCh38.fa"
KNOWN_SITES="dbsnp_146.hg38.vcf.gz Mills_and_1000G_gold_standard.indels.hg38.vcf.gz"
KNOWN_FLAGS=$(printf -- '--known-sites %s ' $KNOWN_SITES)

# Step 1a: Build recalibration table
gatk BaseRecalibrator \
    -R $GENOME \
    -I sample1.markdup.bam \
    $KNOWN_FLAGS \
    -O sample1.recal.table

# Step 1b: Apply recalibration
gatk ApplyBQSR \
    -R $GENOME \
    -I sample1.markdup.bam \
    --bqsr-recal-file sample1.recal.table \
    -O sample1.bqsr.bam

echo "BQSR BAM: sample1.bqsr.bam"
samtools flagstat sample1.bqsr.bam | head -3
Step 2: Call Variants with HaplotypeCaller (GVCF Mode)

Run per-sample variant calling, producing an intermediate GVCF for joint genotyping.

bash
# HaplotypeCaller in GVCF mode (recommended for cohort analysis)
gatk HaplotypeCaller \
    -R GRCh38.fa \
    -I sample1.bqsr.bam \
    -O gvcfs/sample1.g.vcf.gz \
    -ERC GVCF \
    --dbsnp dbsnp_146.hg38.vcf.gz \
    --native-pair-hmm-threads 4

# For WES: specify target intervals
# gatk HaplotypeCaller ... -L exome_targets.interval_list --interval-padding 100

echo "GVCF: gvcfs/sample1.g.vcf.gz"
zcat gvcfs/sample1.g.vcf.gz | grep -v "^#" | wc -l
Step 3: Consolidate GVCFs with GenomicsDBImport

Merge per-sample GVCFs for efficient joint genotyping.

bash
# Create sample map file: sample_name\tpath_to_gvcf
printf "ctrl_1\tgvcfs/ctrl_1.g.vcf.gz\n" > sample_map.txt
printf "ctrl_2\tgvcfs/ctrl_2.g.vcf.gz\n" >> sample_map.txt
printf "treat_1\tgvcfs/treat_1.g.vcf.gz\n" >> sample_map.txt
printf "treat_2\tgvcfs/treat_2.g.vcf.gz\n" >> sample_map.txt

# Import GVCFs into GenomicsDB for each chromosome
for CHR in chr1 chr2 chr3; do
    gatk GenomicsDBImport \
        --sample-name-map sample_map.txt \
        --genomicsdb-workspace-path genomicsdb/${CHR} \
        -L $CHR \
        --reader-threads 4
done

echo "GenomicsDB created for $(ls genomicsdb/ | wc -l) chromosomes"
Step 4: Joint Genotyping with GenotypeGVCFs

Genotype all samples simultaneously across the GenomicsDB.

bash
# Joint genotype all samples
mkdir -p vcfs
for CHR in chr1 chr2 chr3; do
    gatk GenotypeGVCFs \
        -R GRCh38.fa \
        -V gendb://genomicsdb/${CHR} \
        --dbsnp dbsnp_146.hg38.vcf.gz \
        -O vcfs/cohort_${CHR}.vcf.gz
done

# Merge per-chromosome VCFs
gatk MergeVcfs \
    $(ls vcfs/cohort_chr*.vcf.gz | sed 's/^/-I /') \
    -O vcfs/cohort_all.vcf.gz

echo "Joint genotyping complete: vcfs/cohort_all.vcf.gz"
gatk CountVariants -V vcfs/cohort_all.vcf.gz
Step 5: Variant Filtration (Hard Filters)

Apply hard filters for small cohorts where VQSR is underpowered.

bash
# Separate SNPs and indels
gatk SelectVariants -V vcfs/cohort_all.vcf.gz --select-type-to-include SNP -O vcfs/snps.vcf.gz
gatk SelectVariants -V vcfs/cohort_all.vcf.gz --select-type-to-include INDEL -O vcfs/indels.vcf.gz

# Apply hard filters: SNPs
gatk VariantFiltration \
    -V vcfs/snps.vcf.gz \
    --filter-expression "QD < 2.0" --filter-name "QD2" \
    --filter-expression "FS > 60.0" --filter-name "FS60" \
    --filter-expression "MQ < 40.0" --filter-name "MQ40" \
    --filter-expression "MQRankSum < -12.5" --filter-name "MQRankSum-12.5" \
    -O vcfs/snps_filtered.vcf.gz

# Apply hard filters: Indels
gatk VariantFiltration \
    -V vcfs/indels.vcf.gz \
    --filter-expression "QD < 2.0" --filter-name "QD2" \
    --filter-expression "FS > 200.0" --filter-name "FS200" \
    --filter-expression "ReadPosRankSum < -20.0" --filter-name "ReadPosRankSum-20" \
    -O vcfs/indels_filtered.vcf.gz

# Merge filtered SNPs + indels
gatk MergeVcfs -I vcfs/snps_filtered.vcf.gz -I vcfs/indels_filtered.vcf.gz \
    -O vcfs/cohort_filtered.vcf.gz

echo "PASS variants: $(bcftools view -f PASS vcfs/cohort_filtered.vcf.gz | grep -v '^#' | wc -l)"
Step 6: Parse VCF Results with Python

Extract variants, annotate with gene info, and prepare a DataFrame.

python
import subprocess
import pandas as pd
import io

# Use bcftools query to extract fields from filtered VCF
result = subprocess.run(
    ["bcftools", "query",
     "-f", "%CHROM\t%POS\t%ID\t%REF\t%ALT\t%QUAL\t%FILTER\t%INFO/QD\t%INFO/FS\n",
     "-i", "FILTER='PASS'",
     "vcfs/cohort_filtered.vcf.gz"],
    capture_output=True, text=True
)

cols = ["CHR", "POS", "ID", "REF", "ALT", "QUAL", "FILTER", "QD", "FS"]
df = pd.read_csv(io.StringIO(result.stdout), sep="\t", names=cols)
df["QUAL"] = pd.to_numeric(df["QUAL"], errors="coerce")
df["QD"] = pd.to_numeric(df["QD"], errors="coerce")

print(f"PASS variants: {len(df)}")
print(f"SNPs:  {(df['REF'].str.len() == 1) & (df['ALT'].str.len() == 1)).sum()}")
print(f"Indels: {((df['REF'].str.len() > 1) | (df['ALT'].str.len() > 1)).sum()}")
print(df.head())
df.to_csv("pass_variants.tsv", sep="\t", index=False)
Show full SKILL.md (363 more words)Show less

Key Parameters

ParameterDefaultRange/OptionsEffect
-ERCNONEGVCF, BP_RESOLUTIONEmit reference confidence mode; GVCF for cohort workflows
--native-pair-hmm-threads41–32Threads for pair-HMM in HaplotypeCaller (most CPU-intensive step)
-Lwhole genomeinterval file or chrRestrict calling to intervals (exome targets, BED regions)
--dbsnp—VCF pathdbSNP VCF for rsID annotation in output
--stand-call-conf300–100Min genotype quality score to emit a variant call
-GStandardAnnotationannotation groupAnnotation modules to apply (StandardHCAnnotation for HaplotypeCaller)
--sample-name-map—TSV fileSample-to-GVCF mapping for GenomicsDBImport
--reader-threads11–16Threads for GenomicsDBImport reading
--filter-expression—JEXL expressionHard filter expression (e.g., "QD < 2.0")
--java-options—-Xmx4gJava heap size; use -Xmx16g for large genomes

Common Recipes

Recipe 1: Single-Sample Variant Calling (No Cohort)
bash
# For a single sample, skip GenomicsDBImport and call directly
gatk HaplotypeCaller \
    -R GRCh38.fa \
    -I sample.bqsr.bam \
    -O sample.vcf.gz \
    --dbsnp dbsnp_146.hg38.vcf.gz \
    --native-pair-hmm-threads 8

# Hard filter the single-sample VCF
gatk VariantFiltration \
    -V sample.vcf.gz \
    --filter-expression "QD < 2.0" --filter-name "QD2" \
    --filter-expression "FS > 60.0" --filter-name "FS60" \
    -O sample_filtered.vcf.gz

echo "PASS variants: $(bcftools view -f PASS sample_filtered.vcf.gz | grep -v '^#' | wc -l)"
Recipe 2: Run Full Pipeline with Snakemake
python
# Snakefile — GATK BQSR + HaplotypeCaller
configfile: "config.yaml"
SAMPLES = config["samples"]
GENOME  = config["genome"]
DBSNP   = config["dbsnp"]
KNOWN   = config["known_sites"]

rule all:
    input: expand("gvcfs/{sample}.g.vcf.gz", sample=SAMPLES)

rule bqsr:
    input: bam="markdup/{sample}.markdup.bam"
    output: bam="bqsr/{sample}.bqsr.bam"
    shell:
        """
        gatk BaseRecalibrator -R {GENOME} -I {input.bam} \
             --known-sites {KNOWN} -O bqsr/{wildcards.sample}.recal.table &&
        gatk ApplyBQSR -R {GENOME} -I {input.bam} \
             --bqsr-recal-file bqsr/{wildcards.sample}.recal.table \
             -O {output.bam}
        """

rule haplotypecaller:
    input: bam="bqsr/{sample}.bqsr.bam"
    output: gvcf="gvcfs/{sample}.g.vcf.gz"
    threads: 4
    shell:
        "gatk HaplotypeCaller -R {GENOME} -I {input.bam} -O {output.gvcf} "
        "-ERC GVCF --dbsnp {DBSNP} --native-pair-hmm-threads {threads}"

Expected Outputs

OutputFormatDescription
*.g.vcf.gzGVCFPer-sample GVCF with reference confidence blocks; input to GenomicsDBImport
cohort_all.vcf.gzVCFJoint-genotyped multi-sample VCF; unfiltered
cohort_filtered.vcf.gzVCFFiltered VCF; FILTER=PASS for passing variants
*.recal.tableBQSRBase quality recalibration table from BaseRecalibrator
*.bqsr.bamBAMRecalibrated BAM; use as HaplotypeCaller input
genomicsdb/DirectoryGenomicsDB workspace per chromosome for joint genotyping

Troubleshooting

ProblemCauseSolution
SAM/BAM file has no @RG headerMissing read group from BWA-MEM2Re-align with -R "@RG\tID:...\tSM:...\tPL:ILLUMINA"
Java OutOfMemoryErrorInsufficient heap sizeAdd --java-options "-Xmx16g" or more
HaplotypeCaller very slowSingle-threaded HMMAdd --native-pair-hmm-threads 8; use interval lists to parallelize by chr
Empty GVCF outputWrong interval or no reads in regionCheck samtools view -c sample.bam chrN for read counts
VQSR fails (< 10k variants)Too few variants for trainingUse hard filters instead of VQSR for small cohorts or exomes
GenomicsDB import fails on existing pathGenomicsDB workspace already existsDelete existing workspace: rm -rf genomicsdb/chr1 before re-running
IndexOutOfBoundsExceptionChromosome name mismatch between BAM and referenceEnsure genome FASTA and BAM use same chr naming (chr1 vs 1)
BCFtools/tabix index missingTabix index (.tbi) not createdRun gatk IndexFeatureFile -I file.vcf.gz or tabix -p vcf file.vcf.gz

References

© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/variant/gatk-variant-calling of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Gatk Variant Calling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gatk Variant Calling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gatk Variant Calling this skilljaechang-hits/SciAgent-Skills3741 repos~3.7kAutomated safety check: PassBSD-3-Clause
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Gatk Variant Calling

What does Gatk Variant Calling do?

GATK Best Practices for germline SNP/indel calling from WGS/WES BAMs. Gatk Variant Calling is an agent skill from jaechang-hits/SciAgent-Skills. GATK Best Practices for germline SNP/indel calling from WGS/WES BAMs.

When should I use Gatk Variant Calling?

Gatk Variant Calling fits situations like: tasks that involve Bioinformatics.

How do I install Gatk Variant Calling in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill gatk-variant-calling -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/variant/gatk-variant-calling in jaechang-hits/SciAgent-Skills) into .claude/skills/gatk-variant-calling in your project. Claude Code loads it when a task matches its description.

How do I install Gatk Variant Calling in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill gatk-variant-calling -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/variant/gatk-variant-calling in jaechang-hits/SciAgent-Skills) into .agents/skills/gatk-variant-calling in your project. Codex loads it when a task matches its description.

Can I use Gatk Variant Calling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill gatk-variant-calling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gatk-variant-calling, .gemini/skills/gatk-variant-calling, .github/skills/gatk-variant-calling and .opencode/skills/gatk-variant-calling in your project.

What does Gatk Variant Calling need to run?

Going by SKILL.md and its folder, Gatk Variant Calling needs the command-line tools its instructions call (wget and conda). Our summary lists: Python 3.

Does Gatk Variant Calling access the network?

SKILL.md names 2 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: gatk.broadinstitute.org. This is read from the text; nothing was executed.

Is Gatk Variant Calling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gatk Variant Calling use?

Gatk Variant Calling is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gatk Variant Calling use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gatk Variant Calling?

Skills that share tags, products or a category with Gatk Variant Calling: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gatk Variant Calling?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.