Agent skill

Snpeff Variant Annotation

by jaechang-hits in jaechang-hits/SciAgent-Skills

Annotate and filter VCF variants with SnpEff and SnpSift. An agent skill from jaechang-hits/SciAgent-Skills.

MITAuto-check passedResearch & Science

Install Snpeff Variant Annotation

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill snpeff-variant-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills snpeff-variant-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/snpeff-variant-annotation .claude/skills/snpeff-variant-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
snpeff-variant-annotation
GitHub stars
374
Used in
1 other repo
Token cost
~5.4k tokens
SKILL.md length
1,357 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
MIT

At a glance

Annotate and filter VCF variants with SnpEff and SnpSift. An agent skill from jaechang-hits/SciAgent-Skills.

  • Works in 7 steps: Install SnpEff and Download Genome… → Annotate VCF with Functional Effects → Filter HIGH-Impact Variants with SnpSift → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 8 more sections
  • Calls java, wget and conda; reaches ftp.ncbi.nlm.nih.gov and snpeff.blob.core.windows.net

What it does

Snpeff Variant Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate and filter VCF variants with SnpEff and SnpSift. SnpEff predicts functional effects (HIGH/MODERATE/LOW/MODIFIER), genes, transcripts, AA changes, HGVS; SnpSift filters and adds ClinVar/dbSNP. Java CLI with Python subprocess integration. Use ANNOVAR for multi-database annotation; Ensembl VEP for REST API; SnpEff for fast CLI with pre-built genomes.

Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics and REST APIs. It works with Java, Python and Ensembl. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve REST APIs

Example prompts

  • “/snpeff-variant-annotation”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Install SnpEff and Download Genome Database
  2. Annotate VCF with Functional Effects
  3. Filter HIGH-Impact Variants with SnpSift
  4. Add ClinVar and dbSNP Annotations
  5. Extract Fields to Tab-Delimited Output
  6. Parse Annotated VCF in Python with cyvcf2
  7. Summary Statistics and Consequence Visualization

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • java
    • wget
    • conda
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ftp.ncbi.nlm.nih.gov
    • snpeff.blob.core.windows.net

    Also links to:

    • pcingola.github.io
    • github.com
    • doi.org
    • brentp.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Snpeff Variant Annotation loads about 5.4k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 1,357 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its MIT licence (© jaechang-hits). 1,357 words, ~5,436 tokens.

Download SKILL.mdSave it as .claude/skills/snpeff-variant-annotation/SKILL.md (or your agent's skills folder).
name
snpeff-variant-annotation
description
Annotate and filter VCF variants with SnpEff and SnpSift. SnpEff predicts functional effects (HIGH/MODERATE/LOW/MODIFIER), genes, transcripts, AA changes, HGVS; SnpSift filters and adds ClinVar/dbSNP. Java CLI with Python subprocess integration. Use ANNOVAR for multi-database annotation; Ensembl VEP for REST API; SnpEff for fast CLI with pre-built genomes.
license
MIT

SnpEff + SnpSift — Variant Annotation and Filtering

Overview

SnpEff annotates variants in VCF files by predicting their functional consequences: impact level (HIGH, MODERATE, LOW, MODIFIER), affected gene and transcript, amino acid change, and HGVS notation. SnpSift is the companion tool for filtering, sorting, and enriching annotated VCFs with external databases such as ClinVar and dbSNP. Together they form a fast, self-contained pipeline for going from raw variant calls to biologically interpretable, filtered variant sets. Both tools are Java-based and are invoked from the command line or Python subprocess; pre-built genome databases (hg38, GRCh37, mm10, and 100+ others) are downloaded with a single command.

When to Use

  • Annotating VCF files from GATK, DeepVariant, bcftools, or other callers with predicted gene-level functional consequences before manual review or downstream filtering
  • Prioritizing clinically relevant variants by filtering to HIGH-impact stop-gain, frameshift, and splice-site variants for rare disease or cancer gene panel analysis
  • Adding ClinVar pathogenicity classifications and dbSNP rsIDs to a variant set for cross-study comparison or clinical reporting
  • Extracting structured, tab-delimited fields (gene, protein change, AF, ClinSig) from annotated VCFs into pandas DataFrames for statistical analysis
  • Identifying candidate de novo variants in trio analysis by combining allele frequency thresholds, impact filters, and parent VCF exclusion
  • Use omics-plotting SKILL to render consequence/impact bar charts from the annotated variant tables
  • Use ANNOVAR instead when comprehensive annotation from multiple databases (gnomAD, CADD, SpliceAI) in a single run is required
  • Use Ensembl VEP instead when REST API access or VEP-specific plugins (CADD, LOFTEE, SpliceRegion) are needed

Prerequisites

  • Java: Java 11+ (required for SnpEff 5.x)
  • SnpEff JAR: downloaded from the SnpEff releases page or via conda/bioconda
  • Python packages (optional): cyvcf2, pandas, matplotlib, seaborn for Python-side parsing and visualization
  • Reference genome database: downloaded once per assembly (e.g., hg38, GRCh37, mm10)

Check before installing: The tool may already be available in the current environment (e.g., inside a pixi / conda env). Run command -v snpEff first and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool via pixi run snpEff rather than bare snpEff.

bash
# Download SnpEff JAR
wget https://snpeff.blob.core.windows.net/versions/snpEff_latest_core.zip
unzip snpEff_latest_core.zip
# JAR is at snpEff/snpEff.jar and snpEff/SnpSift.jar

# Or via conda (recommended for reproducibility)
conda install -c bioconda snpeff

# Verify
java -jar snpEff/snpEff.jar -version
# SnpEff 5.2a (build 2024-02-06)

# Install Python packages for downstream parsing
pip install cyvcf2 pandas matplotlib seaborn

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: annotationDatabase
    kind: required
    source: upstream
    ask: "Which genome database should annotate these variants, and does its assembly match the VCF coordinates?"
    default: "the assembly the variants were called on"

  - id: D2
    param: transcriptScope
    kind: required
    source: user
    ask: "Report the consequence on every transcript of a gene, or only on the canonical one?"
    default: "all transcripts - the worst consequence across them is what most filters use"

  - id: D3
    param: filterExpression
    kind: required
    source: user
    ask: "Which annotated variants should survive into the final table - by predicted impact, population frequency, or clinical significance?"
    default: null

  - id: D4
    param: frequencySkip
    kind: optional_conditional
    source: user
    ask: "Skip annotating variants already common in the population?"
    default: "annotate everything"

  - id: D5
    param: summaryReport
    kind: optional
    source: user
    ask: "Generate the HTML summary alongside the annotated VCF?"
    default: "generated"

  - id: D6
    param: javaHeap
    kind: never_ask
    source: data
    reason: "Prevents out-of-memory on large VCFs; does not change annotations"
    default: "4-8 GB"

D3 has no safe default because the filter is the deliverable - a rare-disease question keeps high-impact rare variants, a pharmacogenomic one keeps known clinical annotations regardless of impact. Applying somebody else's filter silently answers their question instead.

Quick Start

bash
# 1. Download the hg38 genome database (one-time setup, ~1 GB)
java -jar snpEff/snpEff.jar download hg38

# 2. Annotate variants
java -jar snpEff/snpEff.jar \
    -v hg38 \
    input.vcf.gz \
    > annotated.vcf

# 3. Filter to HIGH-impact variants
java -jar snpEff/SnpSift.jar filter \
    "ANN[*].IMPACT = 'HIGH'" \
    annotated.vcf \
    > high_impact.vcf

echo "Done. Check snpEff_summary.html and snpEff_genes.txt for QC stats."

Workflow

Step 1: Install SnpEff and Download Genome Database

Download and verify the pre-built genome database for your target assembly. Databases include transcript models from Ensembl or UCSC and are built into SnpEff's local cache directory (~/.snpEff/data/ or snpEff/data/).

bash
# List all available databases (grep for your assembly)
java -jar snpEff/snpEff.jar databases | grep -i "GRCh38\|hg38"

# Download hg38 (Homo sapiens, GRCh38)
java -jar snpEff/snpEff.jar download hg38

# Download GRCh37 (older assemblies still widely used in clinical pipelines)
java -jar snpEff/snpEff.jar download GRCh37.75

# Download mouse mm10
java -jar snpEff/snpEff.jar download mm10

# List installed databases
ls ~/.snpEff/data/
Step 2: Annotate VCF with Functional Effects

Run SnpEff annotation to add ANN INFO fields to each variant. The ANN field encodes pipe-separated annotations per transcript: Allele|Effect|Impact|Gene|GeneID|Feature|FeatureID|BioType|Rank|HGVS.c|HGVS.p|cDNA_pos|CDS_pos|Protein_pos|Distance|Errors.

bash
# Annotate a gzipped VCF (outputs to stdout, pipe or redirect)
java -Xmx8g -jar snpEff/snpEff.jar \
    -v \
    -stats snpeff_summary.html \
    hg38 \
    input.vcf.gz \
    > annotated.vcf

# Compress and index the output for downstream tools
bgzip annotated.vcf
tabix -p vcf annotated.vcf.gz

echo "Annotated variants in annotated.vcf.gz"
echo "QC report: snpeff_summary.html"
echo "Gene table: snpEff_genes.txt"
Step 3: Filter HIGH-Impact Variants with SnpSift

SnpSift filter evaluates boolean expressions over INFO and FORMAT fields. The ANN[*] syntax iterates over all transcript annotations for each variant — any matching transcript qualifies the variant.

bash
# Filter for HIGH-impact variants only (stop-gain, frameshift, splice-site)
java -jar snpEff/SnpSift.jar filter \
    "ANN[*].IMPACT = 'HIGH'" \
    annotated.vcf.gz \
    > high_impact.vcf

# Filter HIGH or MODERATE impact and allele frequency < 1%
# (requires AF in INFO field, e.g., from GATK or gnomAD annotation)
java -jar snpEff/SnpSift.jar filter \
    "(ANN[*].IMPACT = 'HIGH' | ANN[*].IMPACT = 'MODERATE') & (AF < 0.01)" \
    annotated.vcf.gz \
    > rare_functional.vcf

# Count filtered variants
grep -v "^#" high_impact.vcf | wc -l
Step 4: Add ClinVar and dbSNP Annotations

SnpSift annotate transfers INFO fields from a reference VCF (ClinVar, dbSNP) to the target VCF by matching on chromosome and position. Download ClinVar and dbSNP VCFs from NCBI FTP before running.

bash
# Download reference databases (one-time setup)
# ClinVar (GRCh38)
wget https://ftp.ncbi.nlm.nih.gov/pub/clinvar/vcf_GRCh38/clinvar.vcf.gz
wget https://ftp.ncbi.nlm.nih.gov/pub/clinvar/vcf_GRCh38/clinvar.vcf.gz.tbi

# dbSNP (GRCh38, b156 or later)
wget https://ftp.ncbi.nlm.nih.gov/snp/latest_release/VCF/GCF_000001405.40.gz
wget https://ftp.ncbi.nlm.nih.gov/snp/latest_release/VCF/GCF_000001405.40.gz.tbi

# Annotate with ClinVar CLNSIG and CLNDN fields
java -jar snpEff/SnpSift.jar annotate \
    -info CLNSIG,CLNDN,CLNREVSTAT \
    clinvar.vcf.gz \
    annotated.vcf.gz \
    > annotated_clinvar.vcf

# Annotate with dbSNP rsIDs (adds RS field to INFO)
java -jar snpEff/SnpSift.jar annotate \
    GCF_000001405.40.gz \
    annotated_clinvar.vcf \
    > annotated_full.vcf

echo "ClinVar + dbSNP annotations added to annotated_full.vcf"
Step 5: Extract Fields to Tab-Delimited Output

SnpSift extractFields converts VCF annotations into a flat table. Use ANN[0] to take the first (most severe) transcript annotation, or ANN[*] to expand all transcripts.

bash
# Extract key fields: CHROM, POS, REF, ALT, impact, gene, protein change, AF, ClinVar sig
java -jar snpEff/SnpSift.jar extractFields \
    -s "," \
    -e "." \
    annotated_full.vcf \
    CHROM POS REF ALT \
    "ANN[0].GENE" \
    "ANN[0].EFFECT" \
    "ANN[0].IMPACT" \
    "ANN[0].HGVS_P" \
    "ANN[0].HGVS_C" \
    "ANN[0].FEATUREID" \
    AF \
    CLNSIG \
    CLNDN \
    > variants_table.tsv

# Preview
head -3 variants_table.tsv
# CHROM  POS   REF  ALT  ANN[0].GENE  ANN[0].EFFECT            ANN[0].IMPACT  ...
# chr1   69511 A    G    OR4F5        synonymous_variant        LOW            ...
# chr1   925952 G   A    SAMD11       missense_variant          MODERATE       ...
Step 6: Parse Annotated VCF in Python with cyvcf2

Use cyvcf2 (or PyVCF2) for programmatic VCF parsing in Python. cyvcf2 wraps htslib and is significantly faster than pure-Python parsers.

python
import re
import pandas as pd
from cyvcf2 import VCF

records = []

for variant in VCF("annotated_full.vcf.gz"):
    ann_field = variant.INFO.get("ANN")
    if ann_field is None:
        continue

    # Take the first (most severe) annotation transcript
    first_ann = ann_field.split(",")[0].split("|")
    # ANN field columns (0-indexed): allele, effect, impact, gene, gene_id,
    # feature_type, feature_id, biotype, rank, hgvs_c, hgvs_p, ...
    effect   = first_ann[1] if len(first_ann) > 1 else "."
    impact   = first_ann[2] if len(first_ann) > 2 else "."
    gene     = first_ann[3] if len(first_ann) > 3 else "."
    hgvs_p   = first_ann[10] if len(first_ann) > 10 else "."

    records.append({
        "chrom":   variant.CHROM,
        "pos":     variant.POS,
        "ref":     variant.REF,
        "alt":     ",".join(variant.ALT),
        "gene":    gene,
        "effect":  effect,
        "impact":  impact,
        "hgvs_p":  hgvs_p,
        "af":      variant.INFO.get("AF", None),
        "clnsig":  variant.INFO.get("CLNSIG", "."),
    })

df = pd.DataFrame(records)
print(f"Total variants: {len(df)}")
print(df["impact"].value_counts())
# HIGH        342
# MODERATE   4218
# LOW        8903
# MODIFIER  51204
Step 7: Summary Statistics and Consequence Visualization

Summarize variant consequences by impact tier, then read skills/data-visualization/omics-plotting/SKILL.md and follow its "Box / Violin / Bar" recipe on the exported tables (→ figures/variant_consequence_summary.png).

python
# Consequence counts by impact (df from Step 6)
impact_order = ["HIGH", "MODERATE", "LOW", "MODIFIER"]
impact_counts = df["impact"].value_counts().reindex(impact_order, fill_value=0)

# Top 15 effects within HIGH and MODERATE
top_effects = (
    df[df["impact"].isin(["HIGH", "MODERATE"])]
    ["effect"]
    .str.replace("_", " ")
    .value_counts()
    .head(15)
)

# Save summary tables; render bar charts with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Box / Violin / Bar" recipe.
impact_counts.rename("count").to_csv("variant_impact_counts.csv")
top_effects.rename("count").to_csv("variant_top_effects.csv")
print(impact_counts.to_string())

# HIGH-impact gene table
high_genes = (
    df[df["impact"] == "HIGH"]["gene"]
    .value_counts()
    .head(20)
    .rename("n_high_variants")
)
print("\nTop genes with HIGH-impact variants:")
print(high_genes.to_string())

Key Parameters

ParameterDefaultRange / OptionsEffect
-Xmx (JVM heap)JVM default (~256 MB)4g–32gPrevents OutOfMemoryError for large WGS VCFs; set to ~4–8 GB per run
-v (verbose)offflagPrints per-chromosome progress and summary counts to stderr
-statssnpEff_summary.htmlany .html pathWrites interactive QC summary with consequence pie charts and gene tables
-noStatsoffflagDisables HTML report generation (faster for batch runs)
-canonoffflagAnnotates only the canonical transcript per gene (reduces ANN field size)
-maxAFdisabled0.0–1.0Skip annotation for variants with population AF above threshold (requires gnomAD in database)
filter expression (SnpSift)—boolean expressionANN[*].IMPACT = 'HIGH', AF < 0.001 & DP > 10, CLNSIG has 'Pathogenic'
-s separator (extractFields)tabany stringField separator in output table; use "," for CSV
-e empty value (extractFields)emptyany stringPlaceholder for missing fields; use "." for consistency with VCF convention

Key Concepts

ANN Field Structure

Each variant's ANN INFO field contains one annotation per predicted transcript consequence, separated by commas. Each annotation is a pipe-delimited string with 16 fields:

ANN=T|missense_variant|MODERATE|BRCA1|ENSG00000012048|transcript|ENST00000357654|protein_coding|12/23|c.1234A>T|p.Lys412Met|1234|1234|412|.|

The first annotation (index 0) is the most severe consequence. Multiple annotations arise when a variant affects multiple transcripts or genes.

Show full SKILL.md (546 more words)Show less
Impact Levels
ImpactVariant TypesTypical Priority
HIGHStop-gained, frameshift, splice-donor/acceptor, start-lostDisease candidate; always review
MODERATEMissense, in-frame indel, splice-regionPathogenicity depends on conservation and domain context
LOWSynonymous, stop-retained, start-retainedRarely causal; useful as controls
MODIFIERIntronic, intergenic, UTR, downstreamBackground; filter out for most analyses
HGVS Notation in ANN

SnpEff provides both cDNA (HGVS.c, e.g., c.1234A>T) and protein (HGVS.p, e.g., p.Lys412Met) notation. HGVS fields are empty (.) for non-coding variants. The FEATUREID field contains the Ensembl or RefSeq transcript ID enabling cross-database lookup.

SnpSift Filter Syntax

SnpSift filter expressions support comparison operators (=, !=, <, >), logical operators (&, |, !), and the has operator for substring matching. Use ANN[*] to match any annotation transcript, ANN[0] for first only:

# Compound filter: rare + damaging + not in ClinVar benign
(AF < 0.001) & (ANN[*].IMPACT = 'HIGH' | ANN[*].IMPACT = 'MODERATE') & !(CLNSIG has 'Benign')

Common Recipes

Recipe: Filter De Novo Candidates (Trio Analysis)

Identify rare, functionally damaging variants in a proband that are absent from both parents. Combine allele frequency filtering, impact filtering, and VCF subtraction using SnpSift.

bash
# Step 1: Annotate proband VCF
java -jar snpEff/snpEff.jar -v hg38 proband.vcf.gz > proband_ann.vcf

# Step 2: Filter rare HIGH/MODERATE variants not in either parent
# Requires proband, mother, and father VCFs to be jointly genotyped or
# use SnpSift filter on a multi-sample VCF
java -jar snpEff/SnpSift.jar filter \
    "(ANN[*].IMPACT = 'HIGH' | ANN[*].IMPACT = 'MODERATE') \
     & (AF < 0.001 | AF = '.') \
     & (GEN[proband].GT != './.') \
     & (GEN[mother].GT = '0/0' | GEN[mother].GT = '0|0') \
     & (GEN[father].GT = '0/0' | GEN[father].GT = '0|0')" \
    proband_ann.vcf \
    > denovo_candidates.vcf

# Step 3: Extract to table for review
java -jar snpEff/SnpSift.jar extractFields \
    denovo_candidates.vcf \
    CHROM POS REF ALT "ANN[0].GENE" "ANN[0].EFFECT" "ANN[0].IMPACT" \
    "ANN[0].HGVS_P" AF CLNSIG \
    > denovo_candidates.tsv

echo "De novo candidates: $(grep -v '^CHROM' denovo_candidates.tsv | wc -l)"
Recipe: Extract Protein-Changing Variants to pandas DataFrame

Load an annotated VCF directly into pandas using SnpSift extractFields output for downstream statistical or ML analysis.

python
import subprocess
import io
import pandas as pd

VCF_IN   = "annotated_full.vcf.gz"
SNPSIFT  = "snpEff/SnpSift.jar"

# Run SnpSift extractFields via subprocess and capture stdout
cmd = [
    "java", "-jar", SNPSIFT, "extractFields",
    "-s", ",", "-e", ".",
    VCF_IN,
    "CHROM", "POS", "REF", "ALT",
    "ANN[0].GENE", "ANN[0].EFFECT", "ANN[0].IMPACT",
    "ANN[0].HGVS_P", "ANN[0].HGVS_C", "ANN[0].FEATUREID",
    "AF", "DP", "CLNSIG", "CLNDN",
]
result = subprocess.run(cmd, capture_output=True, text=True, check=True)
df = pd.read_csv(io.StringIO(result.stdout), sep="\t")
df.columns = [c.replace("ANN[0].", "") for c in df.columns]  # clean column names

# Keep protein-changing variants only
protein_changing = df[df["IMPACT"].isin(["HIGH", "MODERATE"])].copy()
protein_changing["AF"] = pd.to_numeric(protein_changing["AF"], errors="coerce")
protein_changing["DP"] = pd.to_numeric(protein_changing["DP"], errors="coerce")

print(f"Protein-changing variants: {len(protein_changing)}")
print(protein_changing[["GENE", "EFFECT", "IMPACT", "HGVS_P", "AF"]].head(10).to_string(index=False))
Recipe: Build a Custom Genome Database

When working with a non-standard organism or custom assembly, build a SnpEff database from a FASTA + GTF file.

bash
# 1. Set up directory structure under SnpEff's data directory
GENOME_NAME="my_organism"
mkdir -p snpEff/data/${GENOME_NAME}
cp my_genome.fa snpEff/data/${GENOME_NAME}/sequences.fa
cp my_annotation.gtf snpEff/data/${GENOME_NAME}/genes.gtf

# 2. Add genome entry to snpEff.config
echo "${GENOME_NAME}.genome : My Organism" >> snpEff/snpEff.config

# 3. Build the database
java -jar snpEff/snpEff.jar build \
    -gtf22 \
    -v \
    ${GENOME_NAME}

# 4. Verify
java -jar snpEff/snpEff.jar dump ${GENOME_NAME} | head -20
echo "Custom database built for ${GENOME_NAME}"

Expected Outputs

FileFormatContents
annotated.vcf.gzbgzipped VCFOriginal variants with ANN INFO field; one record per variant
snpEff_summary.htmlHTMLInteractive QC report: consequence pie charts, top genes, transition/transversion ratio
snpEff_genes.txtTSVPer-gene counts of variants by consequence type
high_impact.vcfVCFSubset of variants with ANN[*].IMPACT = 'HIGH'
annotated_clinvar.vcfVCFAnnotated VCF enriched with CLNSIG, CLNDN, CLNREVSTAT from ClinVar
variants_table.tsvTSVFlat table of extracted fields; one row per variant (first transcript)
variant_consequence_summary.pngPNGBar charts of impact tiers and top consequence types

Troubleshooting

ProblemCauseSolution
OutOfMemoryError during annotationDefault JVM heap too small for WGS VCFsAdd -Xmx8g (or -Xmx16g) before -jar: java -Xmx8g -jar snpEff.jar ...
ERROR_CHROMOSOME_NOT_FOUND for every variantChromosome naming mismatch (e.g., chr1 vs 1)Pass -noCheckChr flag; or rename contigs in VCF with bcftools annotate --rename-chrs
Database not found errorGenome database not downloadedRun java -jar snpEff.jar download hg38; check snpEff/data/ directory exists
Empty ANN field on all variantsWrong genome database version relative to VCF referenceConfirm VCF uses the same assembly as the database (e.g., hg38 vs GRCh38 — use GRCh38.86 for Ensembl builds)
SnpSift filter returns zero variantsFilter expression syntax error or wrong field nameTest expression on small VCF; check ANN[*].IMPACT vs ANN[0].IMPACT; wrap expression in quotes
CLNSIG field missing after snpSift annotateClinVar VCF not indexed, or contig mismatchRun tabix -p vcf clinvar.vcf.gz before annotating; ensure chromosome names match
extractFields output missing HGVS_P columnVariant is non-coding or affects UTR/intronUse -e "." flag to fill empty fields; filter to coding variants first
Very large ANN field (thousands of transcripts)Variant in a region with many overlapping transcriptsAdd -canon flag to annotate only canonical transcripts, or use ANN[0] in downstream filters

References

  • SnpEff Documentation — comprehensive usage guide, ANN field specification, filter syntax
  • SnpEff GitHub Repository — source code, issue tracker, release notes
  • Cingolani P, et al. (2012) "A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118." Fly 6(2):80–92. doi:10.4161/fly.19695
  • ClinVar VCF FTP — ClinVar VCF files by assembly
  • cyvcf2 Documentation — Python VCF parsing library used in Step 6

© jaechang-hits, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/variant/snpeff-variant-annotation of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Snpeff Variant Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Snpeff Variant Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Snpeff Variant Annotation this skilljaechang-hits/SciAgent-Skills3741 repos~5.4kAutomated safety check: PassMIT
Bio Ensembl RESTGPTomics/bioSkills1.2k2 repos~3.6kAutomated safety check: PassMIT
Tooluniverse Epigenomicswu-yc/LabClaw1.1k2 repos~14kAutomated safety check: PassNone
Ensembl Databaseaipoch/medical-research-skills1.9k—~1.5kAutomated safety check: PassMIT
UniProt Database Accessdavila7/claude-code-templates33k14 repos~1.7kAutomated safety check: PassMIT
Ggetdavila7/claude-code-templates33k10 repos~6.3kAutomated safety check: PassMIT

Similar skills

  • Bio Ensembl REST

    GPTomics/bioSkills

    Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • Ensembl Database

    aipoch/medical-research-skills

    Access Ensembl REST API for vertebrate genomic data; use when you need gene/ID lookups, sequence retrieval, variant effect prediction (VEP), or homology/assembly coordinate mapping.

    1.9k GitHub stars~1.5k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    33k GitHub starsUsed in 14 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 10 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs.

    276k GitHub starsUsed in 1 repo~1.3k tokens
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Snpeff Variant Annotation

What does Snpeff Variant Annotation do?

Annotate and filter VCF variants with SnpEff and SnpSift. An agent skill from jaechang-hits/SciAgent-Skills. Snpeff Variant Annotation is an agent skill from jaechang-hits/SciAgent-Skills. Annotate and filter VCF variants with SnpEff and SnpSift.

When should I use Snpeff Variant Annotation?

Snpeff Variant Annotation fits situations like: tasks that involve Bioinformatics; tasks that involve REST APIs.

How do I install Snpeff Variant Annotation in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill snpeff-variant-annotation -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/variant/snpeff-variant-annotation in jaechang-hits/SciAgent-Skills) into .claude/skills/snpeff-variant-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Snpeff Variant Annotation in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill snpeff-variant-annotation -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/variant/snpeff-variant-annotation in jaechang-hits/SciAgent-Skills) into .agents/skills/snpeff-variant-annotation in your project. Codex loads it when a task matches its description.

Can I use Snpeff Variant Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill snpeff-variant-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/snpeff-variant-annotation, .gemini/skills/snpeff-variant-annotation, .github/skills/snpeff-variant-annotation and .opencode/skills/snpeff-variant-annotation in your project.

What does Snpeff Variant Annotation need to run?

Going by SKILL.md and its folder, Snpeff Variant Annotation needs the command-line tools its instructions call (java, wget, conda and pip). Our summary lists: Python 3.

Does Snpeff Variant Annotation access the network?

SKILL.md names 6 domains. In commands or code: ftp.ncbi.nlm.nih.gov and snpeff.blob.core.windows.net; the agent is likely to contact these when it follows the instructions. As links in the text: pcingola.github.io, github.com, doi.org and brentp.github.io. This is read from the text; nothing was executed.

Is Snpeff Variant Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Snpeff Variant Annotation use?

Snpeff Variant Annotation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Snpeff Variant Annotation use?

About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Snpeff Variant Annotation?

Skills that share tags, products or a category with Snpeff Variant Annotation: Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars), Tooluniverse Epigenomics (wu-yc/LabClaw, 1.1k stars), Ensembl Database (aipoch/medical-research-skills, 1.9k stars) and UniProt Database Access (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Snpeff Variant Annotation?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.