Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
CLI for VCF/BCF: filter, merge, annotate, query, normalize, compute stats.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bcftools-variant-manipulation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation .claude/skills/bcftools-variant-manipulation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bcftools-variant-manipulation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation into .claude/skills/bcftools-variant-manipulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bcftools-variant-manipulation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/bcftools-variant-manipulationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bcftools-variant-manipulation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation .agents/skills/bcftools-variant-manipulation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bcftools-variant-manipulation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation into .agents/skills/bcftools-variant-manipulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bcftools-variant-manipulation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bcftools-variant-manipulation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation .cursor/skills/bcftools-variant-manipulation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bcftools-variant-manipulation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation into .cursor/skills/bcftools-variant-manipulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bcftools-variant-manipulation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/variant/bcftools-variant-manipulation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bcftools-variant-manipulation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation .gemini/skills/bcftools-variant-manipulation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bcftools-variant-manipulation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation into .gemini/skills/bcftools-variant-manipulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bcftools-variant-manipulation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills bcftools-variant-manipulationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation .github/skills/bcftools-variant-manipulation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bcftools-variant-manipulation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation into .github/skills/bcftools-variant-manipulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bcftools-variant-manipulation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills bcftools-variant-manipulation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation .opencode/skills/bcftools-variant-manipulation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bcftools-variant-manipulation" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/bcftools-variant-manipulation into .opencode/skills/bcftools-variant-manipulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bcftools-variant-manipulation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bcftools-variant-manipulationCLI for VCF/BCF: filter, merge, annotate, query, normalize, compute stats.
Bcftools Variant Manipulation is an agent skill from jaechang-hits/SciAgent-Skills. CLI for VCF/BCF: filter, merge, annotate, query, normalize, compute stats. Core post-variant-calling: quality filtering, multi-sample merging, rsID annotation, genotype extraction. Samtools companion in HTSlib. Use GATK for complex indel realignment during calling; use VCFtools for population genetics stats.
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
condabrewFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
samtools.github.iogithub.comdoi.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bcftools Variant Manipulation loads about 4.8k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 990 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its MIT licence (© jaechang-hits). 990 words, ~4,796 tokens.
.claude/skills/bcftools-variant-manipulation/SKILL.md (or your agent's skills folder).bcftools is the standard command-line toolkit for processing VCF (Variant Call Format) and BCF (Binary Call Format) files in the HTSlib ecosystem. It covers the complete post-variant-calling workflow: format conversion, quality filtering, variant normalization, multi-sample merging, annotation with external databases, genotype extraction, and QC statistics. bcftools uses streaming by design — most commands read from stdin and write to stdout, making it ideal for memory-efficient pipelines on large cohorts.
GATK HaplotypeCaller instead when calling variants with local realignment in human samplesVCFtools instead for population genetics statistics (Fst, LD, Hardy-Weinberg)bcftools in the HTSlib pipeline; use picard for duplicate-marking and library metrics.vcf.gz + .vcf.gz.tbi) for region queriessamtools for BAM processing; tabix for VCF indexingCheck before installing: The tool may already be available in the current environment (e.g., inside a
pixi/condaenv). Runcommand -v bcftoolsfirst and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool viapixi run bcftoolsrather than barebcftools.
# Bioconda (recommended — installs HTSlib suite)
conda install -c bioconda bcftools
# Homebrew (macOS)
brew install bcftools
# Verify
bcftools --version | head -1
# bcftools 1.20
# Index a VCF for region queries
bcftools index -t variants.vcf.gz # creates .tbi
bcftools index -c variants.vcf.gz # creates .csi (for chromosomes > 512 Mb)Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: filterExpression
kind: required
source: user
ask: "Which variants should survive - by depth, genotype quality, allele balance, or population frequency?"
default: null
- id: D2
param: multiallelicHandling
kind: required
source: user
ask: "Should multi-allelic records be split into one row per alternate allele before filtering and annotation?"
default: "split - most annotation and comparison tools assume one alternate per row"
- id: D3
param: leftAlignment
kind: required
source: upstream
depends_on: [D2]
ask: "Should indels be left-aligned against the reference so the same variant is written identically across callers?"
default: "left-aligned when a reference FASTA is available"
- id: D4
param: sampleSubset
kind: optional
source: user
ask: "Should the output be restricted to particular samples?"
default: "all samples"
- id: D5
param: regionSubset
kind: optional
source: user
ask: "Should records be restricted to particular regions?"
default: "whole file"
- id: D6
param: outputFormat
kind: optional
source: upstream
ask: "Which container should the result be written in?"
default: "bgzip-compressed VCF with an index"
- id: D7
param: threads
kind: never_ask
source: data
reason: "Compression threads affect runtime only, not the records"
default: "min(4, available_cores)"D2 and D3 together decide whether two VCFs of the same sample can be compared at all. Unsplit multi-allelic rows and right-aligned indels produce apparent private variants that are the same variant written differently - and every downstream intersection then reports a difference that does not exist.
# Typical post-calling workflow: normalize → filter → annotate → extract
bcftools norm -d any -f reference.fa variants.vcf.gz \
| bcftools filter -i 'QUAL>20 && DP>10' \
| bcftools annotate -a dbSNP.vcf.gz -c ID \
| bcftools view -O z -o final.vcf.gz
# Index the output
bcftools index -t final.vcf.gz
# Count variants at each stage
bcftools stats final.vcf.gz | grep "^SN"Convert between text VCF and binary BCF; compress and index for random access.
# VCF → compressed BCF (fastest format for piping)
bcftools view -O b -o variants.bcf variants.vcf
# BCF → VCF (for human-readable output)
bcftools view -O v -o variants.vcf variants.bcf
# VCF → bgzipped + indexed (standard archive format)
bcftools view -O z -W -o variants.vcf.gz variants.vcf
# -W automatically creates .tbi index after writing# Extract specific samples
bcftools view -s sample1,sample2 -O z -o subset.vcf.gz variants.vcf.gz
# Exclude samples (prefix with ^)
bcftools view -s ^outlier_sample -O z -o cleaned.vcf.gz variants.vcf.gz
# Extract by region (fast; requires index)
bcftools view -r chr1:1000000-2000000 variants.vcf.gz -O v -o chr1_region.vcf
# Streaming pipeline: no intermediate files
samtools mpileup -Ou input.bam | bcftools call -m -Oz -o calls.vcf.gzApply quality thresholds and FLAG-based filters to retain high-confidence calls.
# Expression-based filter (include)
bcftools filter -i 'QUAL>20 && DP>10' variants.vcf.gz -O z -o filtered.vcf.gz
# Expression-based filter (exclude)
bcftools filter -e 'QUAL<10 || DP<5' variants.vcf.gz -O v -o filtered.vcf
# Soft filter: mark but keep (sets FILTER field to label)
bcftools filter -s LowQual -e 'QUAL<20' variants.vcf.gz -O z -o soft_filtered.vcf.gz
# Variants with QUAL<20 get FILTER="LowQual"; others get FILTER=PASS# Keep only PASS variants
bcftools view -f PASS variants.vcf.gz -O z -o pass_only.vcf.gz
# SNP-only output
bcftools view --type snps variants.vcf.gz -O z -o snps.vcf.gz
# Indel-only output
bcftools view --type indels variants.vcf.gz -O z -o indels.vcf.gz
# Filter by allele frequency and depth
bcftools filter -i 'AF>0.1 && DP>20 && MQ>40' variants.vcf.gz -O z -o confident.vcf.gz
# Remove SNPs within 3 bp of indels
bcftools filter --SnpGap 3 variants.vcf.gz -O z -o gapfiltered.vcf.gzTransform VCF content into tabular text for downstream analysis.
# Extract chrom, position, ref, alt, quality
bcftools query -f '%CHROM\t%POS\t%REF\t%ALT\t%QUAL\n' variants.vcf.gz > variants.txt
# With header row (-H adds #-prefixed column names)
bcftools query -H -f '%CHROM\t%POS\t%REF\t%ALT\t%QUAL\n' variants.vcf.gz > variants.tsv
# Per-sample genotypes and allele depths
bcftools query -f '[%SAMPLE\t%GT\t%AD\n]' variants.vcf.gz > genotypes.txt
# Output: sample1 0/1 25,18 (ref_depth,alt_depth)# Rare variants (AF < 1%)
bcftools query -i 'AF<0.01' -f '%CHROM\t%POS\t%REF\t%ALT\t%AF\n' \
variants.vcf.gz > rare_variants.txt
# Count variants per chromosome
bcftools query -f '%CHROM\n' variants.vcf.gz | sort | uniq -c | sort -rn
# Extract genotype matrix across all samples
bcftools query -f '%CHROM:%POS\t[%GT\t]\n' -H variants.vcf.gz > genotype_matrix.tsvCombine VCF files from multiple samples (merge) or chromosomes (concat).
# Merge: join VCFs from DIFFERENT sample sets (same variants)
bcftools merge sample1.vcf.gz sample2.vcf.gz sample3.vcf.gz \
-O z -o cohort.vcf.gz
# Merge with auto-indexing and threading
bcftools merge -O b -W --threads 4 sample*.vcf.gz > cohort.bcf
# Concat: join VCFs from SAME sample set (different chromosomes or batches)
bcftools concat chr1.vcf.gz chr2.vcf.gz chr3.vcf.gz -O z -o full.vcf.gz
# Concat with overlap handling (from batched calling)
bcftools concat -a --threads 4 batch*.vcf.gz -O z -o concat.vcf.gz# Pipeline: merge → filter → normalize
bcftools merge sample1.vcf.gz sample2.vcf.gz \
| bcftools filter -i 'QUAL>20' \
| bcftools norm -d any -f genome.fa \
| bcftools view -O z -o merged_clean.vcf.gz
bcftools index -t merged_clean.vcf.gz
# Extract genotype matrix from merged cohort
bcftools merge cohort*.vcf.gz | bcftools query -f '[%GT\t]\n' > gt_matrix.tsvAdd identifiers, gene annotations, or external data to VCF records.
# Add rsIDs from dbSNP
bcftools annotate -a dbSNP.vcf.gz -c ID variants.vcf.gz -O z -o rsid_annotated.vcf.gz
# Annotate with BED file (adds gene names)
bcftools annotate -a genes.bed.gz \
-h <(echo '##INFO=<ID=GENE,Number=1,Type=String,Description="Gene name">') \
-c CHROM,FROM,TO,GENE \
variants.vcf.gz -O z -o gene_annotated.vcf.gz
# Remove unwanted INFO fields
bcftools annotate -x INFO/AC,INFO/AN,INFO/MQ variants.vcf.gz -O v -o stripped.vcf# Normalize: left-align indels, split multi-allelic records
bcftools norm -f reference.fa variants.vcf.gz -O v -o normalized.vcf
# Split multi-allelic sites into separate records
bcftools norm -m -any variants.vcf.gz -O v -o split.vcf
# Deduplicate overlapping records
bcftools norm -d any variants.vcf.gz -O z -o deduped.vcf.gz
# Full normalize pipeline
bcftools norm -m -any variants.vcf.gz | \
bcftools norm -d any -f reference.fa | \
bcftools view -O z -o normalized_split.vcf.gzGenerate summary metrics and per-sample variant counts.
# Full VCF statistics report
bcftools stats variants.vcf.gz > qc.stats.txt
# Extract Summary Numbers section only
grep "^SN" qc.stats.txt | cut -f3,4
# number of records: 45231
# number of SNPs: 38941
# number of indels: 6290
# ...
# Per-sample stats (PSC = per-sample counts)
bcftools stats -s - variants.vcf.gz | grep "^PSC" > per_sample.txt
# cols: id sample hom_RR het hom_AA ts tv indel missing singleton# Transition/transversion ratio (genome-wide QC)
bcftools stats variants.vcf.gz | grep "Ts/Tv"
# Ts/Tv ratio: 2.06 (healthy WGS; <1.8 or >2.2 suggests quality issues)
# Check for sample contamination (F-statistic per sample)
bcftools stats -s - variants.vcf.gz | grep "^PSC" | awk '{print $2, $9}'
# sample F_missing (high = poor sample quality)
# Variant calling (mpileup → call pipeline)
samtools mpileup -Ou -f genome.fa *.bam | bcftools call -m -v -Oz -o calls.vcf.gz
bcftools stats calls.vcf.gz | grep "^SN"-O v → VCF text (uncompressed) default for human inspection
-O z → bgzipped VCF (.vcf.gz) standard for archiving
-O b → binary BCF (compressed) fastest for piping
-O u → binary BCF (uncompressed) fastest output (no compression)Rule: Use -O b or -O u for intermediate pipeline steps (no I/O overhead). Use -O z for files you will store or index with tabix.
Expressions use INFO and FORMAT fields with comparison operators:
# INFO fields (one value per variant)
QUAL>20 # quality score
DP>10 # total depth
AF<0.05 # allele frequency
MQ>40 # mapping quality
# FORMAT fields (per-sample; use [] to iterate)
GT=="1/1" # homozygous alternate
AD[1]>5 # alt allele depth > 5
GQ>20 # genotype quality
# Combined
QUAL>20 && DP>10 && AF>0.01
(GT=="0/0" || GT=="1/1") && GQ>30Goal: Normalize, filter, and annotate a raw variant call set for downstream analysis.
#!/bin/bash
VCF="raw_calls.vcf.gz"
REF="reference.fa"
DBSNP="dbSNP_hg38.vcf.gz"
FINAL="variants_filtered_annotated.vcf.gz"
# 1. Normalize: left-align indels, split multi-allelic, deduplicate
bcftools norm -m -any $VCF \
| bcftools norm -d any -f $REF \
| bcftools view -O z -o normalized.vcf.gz
bcftools index -t normalized.vcf.gz
# 2. Filter by quality and depth
bcftools filter -i 'QUAL>20 && DP>10' normalized.vcf.gz \
| bcftools filter --SnpGap 3 \
| bcftools view -f PASS -O z -o filtered.vcf.gz
bcftools index -t filtered.vcf.gz
# 3. Annotate with rsIDs
bcftools annotate -a $DBSNP -c ID filtered.vcf.gz -O z -o $FINAL
bcftools index -t $FINAL
# Report variant counts
echo "Final variant count:"
bcftools stats $FINAL | grep "number of records"Goal: Merge per-sample VCFs into a cohort VCF; extract a genotype matrix for GWAS.
#!/bin/bash
SAMPLES=(sample1 sample2 sample3 sample4 sample5)
# 1. Ensure all VCFs are indexed
for s in "${SAMPLES[@]}"; do
bcftools index -t ${s}.vcf.gz
done
# 2. Merge into cohort VCF (only sites present in ALL samples: -m none)
bcftools merge -m none "${SAMPLES[@]/%/.vcf.gz}" \
-O z --threads 8 -o cohort.vcf.gz
bcftools index -t cohort.vcf.gz
# 3. Filter: PASS, SNPs only, MAF > 1%
bcftools view -f PASS --type snps cohort.vcf.gz \
| bcftools filter -i 'AF>0.01 && AF<0.99' \
| bcftools view -O z -o cohort_snps_filtered.vcf.gz
# 4. Extract numeric genotype matrix (for plink/R/Python)
bcftools query -H \
-f '%CHROM\t%POS\t%REF\t%ALT\t[%GT\t]\n' \
cohort_snps_filtered.vcf.gz > genotype_matrix.tsv
echo "Genotype matrix: $(wc -l < genotype_matrix.tsv) variants x ${#SAMPLES[@]} samples"| Parameter | Command | Default | Range/Options | Effect |
|---|---|---|---|---|
-O | Most | v | v,z,b,u | Output format: VCF, bgzip-VCF, BCF, uncompressed BCF |
-r | Most | — | chr:pos-end | Region filter (requires tabix index) |
-s | Most | All | sample names | Include specific samples (prefix ^ to exclude) |
-i | filter | — | expression | Include variants matching expression |
-e | filter | — | expression | Exclude variants matching expression |
-f | query | — | format string | Custom output format string |
--threads | Most | 1 | 1–N | Compression/decompression threads |
-a | annotate | — | file path | Annotation source (BED, VCF, TSV) |
-m | norm | none | -any, +any | Split (−) or join (+) multi-allelic records |
-d | norm | — | all,any,snps | Deduplication strategy |
-W | view | — | flag | Auto-create index after writing |
-v | call | — | flag | Output variant sites only (skip reference sites) |
Index before region queries: bcftools view -r chr1:... requires a .tbi or .csi index. Always run bcftools index -t output.vcf.gz after creating any bgzipped VCF.
Normalize before merging or annotation: Different callers represent the same indel differently. Run bcftools norm -m -any | bcftools norm -d any -f ref.fa before merging to prevent duplicate records.
Use -O u for pipeline intermediates: Uncompressed BCF output (-O u) eliminates compression/decompression overhead in multi-step pipes — typically 2-3× faster than -O z.
Verify Ts/Tv ratio after calling: bcftools stats variants.vcf.gz | grep Ts/Tv. For human WGS, expect 2.0–2.1; exome 2.5–3.0. Values outside these ranges indicate quality problems.
Filter before merge for large cohorts: Filtering per-sample VCFs before merging reduces memory and I/O. Apply site-level QC (QUAL>20 && DP>5) to each sample before bcftools merge.
Check chromosome naming consistency: bcftools fails silently if merging chr1-style with 1-style VCFs. Verify with bcftools view -h file.vcf.gz | grep "^##contig".
echo "Before:" $(bcftools view -c 1 raw.vcf.gz | wc -l)
echo "PASS only:" $(bcftools view -f PASS raw.vcf.gz | bcftools view -c 1 | wc -l)
echo "SNPs PASS:" $(bcftools view -f PASS --type snps raw.vcf.gz | bcftools view -c 1 | wc -l)bcftools view -s SAMPLE_A variants.vcf.gz \
| bcftools filter -i 'GT="0/1"' \
| bcftools query -f '%CHROM\t%POS\t%REF\t%ALT\n' > het_sites.txt
echo "$(wc -l < het_sites.txt) heterozygous sites"# Find variants unique to each file and shared
bcftools isec -p isec_dir file1.vcf.gz file2.vcf.gz
ls isec_dir/
# 0000.vcf: private to file1
# 0001.vcf: private to file2
# 0002.vcf: shared (from file1 perspective)
# 0003.vcf: shared (from file2 perspective)
echo "Shared: $(wc -l < isec_dir/0002.vcf) variants"| Problem | Cause | Solution |
|---|---|---|
Missing index error | VCF not indexed (needed for -r region queries) | Run bcftools index -t file.vcf.gz |
[E::vcf_parse_format] parse error | Malformed VCF FORMAT or INFO field | Validate: bcftools view file.vcf.gz 2>&1 | head -5; check source tool version |
| Empty filter output | Expression too strict or field missing | Test expression: bcftools view -h file.vcf.gz | grep "##INFO=<ID=QUAL" |
| Merge: sample duplication | Duplicate sample names across input VCFs | Rename with bcftools reheader -s new_names.txt sample.vcf.gz before merging |
| Wrong Ts/Tv ratio (<1.8) | Low-quality calls or poor coverage | Apply stricter quality filter (QUAL>30 && DP>15); check alignment quality |
concat fails with overlap | Overlapping regions in input VCFs | Use -a flag: bcftools concat -a file1.vcf.gz file2.vcf.gz |
| Annotation mismatch | chr naming conflict (chr1 vs 1) | Check: bcftools view -h file.vcf.gz | grep contig; rename with bcftools annotate --rename-chrs |
query returns empty fields | FORMAT field not populated for sample | Check VCF header: bcftools view -h file.vcf.gz | grep FORMAT |
© jaechang-hits, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/variant/bcftools-variant-manipulation of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Bcftools Variant Manipulation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bcftools Variant Manipulation this skilljaechang-hits/SciAgent-Skills | 371 | 1 repos | ~4.8k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
CLI for VCF/BCF: filter, merge, annotate, query, normalize, compute stats. Bcftools Variant Manipulation is an agent skill from jaechang-hits/SciAgent-Skills. CLI for VCF/BCF: filter, merge, annotate, query, normalize, compute stats.
Bcftools Variant Manipulation fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/variant/bcftools-variant-manipulation in jaechang-hits/SciAgent-Skills) into .claude/skills/bcftools-variant-manipulation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/variant/bcftools-variant-manipulation in jaechang-hits/SciAgent-Skills) into .agents/skills/bcftools-variant-manipulation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill bcftools-variant-manipulation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bcftools-variant-manipulation, .gemini/skills/bcftools-variant-manipulation, .github/skills/bcftools-variant-manipulation and .opencode/skills/bcftools-variant-manipulation in your project.
Going by SKILL.md and its folder, Bcftools Variant Manipulation needs the command-line tools its instructions call (conda and brew).
SKILL.md names 3 domains. As links in the text: samtools.github.io, github.com and doi.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bcftools Variant Manipulation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bcftools Variant Manipulation: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 371 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.