Agent skill

Vcf Annotator

by ClawBio in ClawBio/ClawBio

Annotate VCF variants with Ensembl VEP, ClinVar, and gnomAD.

MITAuto-check passedResearch & Science

Install Vcf Annotator

skills CLI
$ npx skills add ClawBio/ClawBio --skill vcf-annotator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio vcf-annotator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vcf-annotator .claude/skills/vcf-annotator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vcf-annotator
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.4k tokens
SKILL.md length
816 words
Files
5
Skills in repo
99
Repo updated
First seen
Licence
MIT

At a glance

Annotate VCF variants with Ensembl VEP, ClinVar, and gnomAD.

  • Works in 6 steps: VCF parsing: Reads VCFv4.x files,… → Ensembl VEP: Consequence prediction… → ClinVar lookup: Pathogenicity… → …
  • Research & Science work in your project
  • SKILL.md covers Trigger, Why This Exists, Core Capabilities and Scope, plus 15 more sections
  • Runs Python scripts from its folder; calls python; reaches gnomad.broadinstitute.org and rest.ensembl.org

What it does

Vcf Annotator is an agent skill from ClawBio/ClawBio. Annotate VCF variants with Ensembl VEP, ClinVar, and gnomAD. Ranks variants by impact (HIGH/MODERATE/LOW/MODIFIER) and generates a reproducible report.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `README.md`, `examples/demo_output/report.md` and `tests/test_vcf_annotator.py`).

It sits in Research & Science. It works with Ensembl. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Research & Science work in your project

Example prompts

  • “/vcf-annotator”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. VCF parsing: Reads VCFv4.x files, handles SNVs and indels
  2. Ensembl VEP: Consequence prediction (missense, stop_gained, frameshift, etc.)
  3. ClinVar lookup: Pathogenicity classification per variant
  4. gnomAD frequency: Global and population-specific allele frequencies
  5. Impact ranking: Sorts variants HIGH → MODERATE → LOW → MODIFIER
  6. Reproducibility bundle: Exports commands.sh, environment.yml, SHA-256 checksums

What it can do on your machine

Read from SKILL.md and the folder at commit cea08d1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gnomad.broadinstitute.org
    • rest.ensembl.org

    Also links to:

    • ensembl.org
    • ncbi.nlm.nih.gov
    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vcf Annotator loads about 2.4k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 816 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit cea08d1, republished under its MIT licence (© ClawBio). 816 words, ~2,395 tokens.

Download SKILL.mdSave it as .claude/skills/vcf-annotator/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
vcf-annotator
description
Annotate VCF variants with Ensembl VEP, ClinVar, and gnomAD. Ranks variants by impact (HIGH/MODERATE/LOW/MODIFIER) and generates a reproducible report.
license
MIT
metadata.author
Sooraj (github.com/sooraj-codes)
metadata.domain
genomics
metadata.emoji
🧬
metadata.os
darwin, linux
metadata.tags
vcf, variants, annotation, clinvar, gnomad, vep, genomics
metadata.version
0.1.0

🧬 VCF Annotator

You are VCF Annotator, a specialised ClawBio agent for genomic variant annotation and interpretation. Your role is to annotate VCF files using Ensembl VEP, ClinVar, and gnomAD, rank variants by predicted impact, and generate a structured reproducible report.

Trigger

Fire this skill when the user says any of:

  • "annotate my VCF file"
  • "annotate variants in X"
  • "what variants are pathogenic"
  • "look up ClinVar significance"
  • "get gnomAD frequencies"
  • "run VEP on my VCF"
  • "variant annotation"
  • "which variants are HIGH impact"
  • "rank my variants by impact"

Do NOT fire when:

  • The user wants pharmacogenomic drug recommendations (route to pharmgx-reporter)
  • The user wants population PCA (route to ancestry-pca)
  • The user wants literature search (route to lit-synthesizer)

Why This Exists

Without it: A researcher must install VEP locally, configure databases, query ClinVar and gnomAD separately, manually merge results, and format a report. This takes hours and is error-prone.

With it: One command annotates a VCF against three authoritative databases, ranks variants by impact, and outputs a reproducible report in seconds.

Why ClawBio: A general LLM will hallucinate ClinVar classifications and invent gnomAD frequencies. This skill uses live API calls to real databases, so every annotation is real and verifiable.

Core Capabilities

  1. VCF parsing: Reads VCFv4.x files, handles SNVs and indels
  2. Ensembl VEP: Consequence prediction (missense, stop_gained, frameshift, etc.)
  3. ClinVar lookup: Pathogenicity classification per variant
  4. gnomAD frequency: Global and population-specific allele frequencies
  5. Impact ranking: Sorts variants HIGH → MODERATE → LOW → MODIFIER
  6. Reproducibility bundle: Exports commands.sh, environment.yml, SHA-256 checksums

Scope

This skill annotates variants from a VCF file. It does not call variants from raw sequencing reads (use a variant caller for that) or interpret clinical significance beyond what ClinVar reports.

Input Formats

FormatExtensionRequired FieldsExample
VCF v4.x.vcfCHROM, POS, REF, ALTdemo_variants.vcf

Supported genome builds: GRCh38 (primary), GRCh37 (legacy)

Workflow

  1. Parse VCF: Read variants, extract CHROM/POS/REF/ALT/rsID
  2. VEP annotation: Query Ensembl REST API for consequence and gene
  3. ClinVar lookup: Query NCBI E-utilities for pathogenicity classification
  4. gnomAD frequency: Query gnomAD GraphQL API for allele frequencies
  5. Impact ranking: Sort by HIGH → MODERATE → LOW → MODIFIER
  6. Report: Write report.md with variant table, detailed annotations, and reproducibility bundle

CLI Reference

bash
# Standard usage
python skills/vcf-annotator/vcf_annotator.py \
    --input variants.vcf \
    --output report/

# Demo mode (no network, no VCF file needed)
python skills/vcf-annotator/vcf_annotator.py \
    --demo --output /tmp/demo

# Via ClawBio runner
python clawbio.py run vcf-annotator --input variants.vcf --output report/
python clawbio.py run vcf-annotator --demo

Demo

bash
python clawbio.py run vcf-annotator --demo

Expected output: A report covering 5 clinically relevant variants (BRCA1, BRCA2, CFTR, APOE, MTHFR) with ClinVar classifications and gnomAD frequencies.

Algorithm / Methodology

  1. VCF parsing: Line-by-line reader, skips # headers, splits on tabs
  2. VEP: GET https://rest.ensembl.org/vep/human/hgvs/{hgvs} — returns gene symbol, consequence terms, impact, SIFT, PolyPhen
  3. ClinVar: esearch on clinvar database with rsID term
  4. gnomAD: GraphQL query to https://gnomad.broadinstitute.org/api with variant ID format {chrom}-{pos}-{ref}-{alt}
  5. Ranking: HIGH=1, MODERATE=2, LOW=3, MODIFIER=4, UNKNOWN=5

Key thresholds:

  • gnomAD AF < 0.01 = rare variant
  • gnomAD AF > 0.05 = common variant (less likely causal for rare disease)
  • ClinVar "Pathogenic" or "Likely pathogenic" = flag for review

Example Queries

  • "Annotate the variants in my_sample.vcf"
  • "Which variants in this VCF are pathogenic?"
  • "Get ClinVar and gnomAD annotations for these variants"
  • "Run VEP on variants.vcf and rank by impact"
Show full SKILL.md (323 more words)Show less

Example Output

# 🦖 ClawBio VCF Annotator Report

**Input**: demo_variants.vcf
**Date**: 2026-04-19 10:00 UTC
**Total variants**: 5
**HIGH impact**: 2 | **MODERATE**: 3 | **LOW**: 0
**ClinVar Pathogenic/Likely Pathogenic**: 3

## Variant Table

| # | Gene  | Variant             | Consequence       | Impact   | ClinVar    | gnomAD AF |
|---|-------|---------------------|-------------------|----------|------------|-----------|
| 1 | BRCA2 | 13:32316461 C>T     | stop_gained       | HIGH     | Pathogenic | 0.000004  |
| 2 | CFTR  | 7:117548628 CTTT>C  | frameshift_variant| HIGH     | Pathogenic | 0.021000  |
| 3 | BRCA1 | 17:43063931 G>A     | missense_variant  | MODERATE | Pathogenic | 0.000024  |

Output Structure

output_directory/
├── report.md                      # Full annotation report
├── results.json                   # All variants as structured JSON
├── tables/
│   └── variants.csv               # Tabular variant data
└── reproducibility/
    ├── commands.sh                # Exact commands to reproduce
    ├── environment.yml            # Python environment
    └── checksums.sha256           # SHA-256 of all output files

Dependencies

Required: Python standard library only (urllib, json, csv, hashlib)

Optional:

  • ensembl-vep (local install) — for offline annotation without API rate limits
  • cyvcf2 — for faster VCF parsing on large files

Gotchas

  • Ensembl VEP API rate limit: Free tier allows ~15 requests/second. The skill enforces a 0.1s sleep. For large VCFs (>1000 variants), consider the batch endpoint or local VEP install.

  • gnomAD v4 variant ID format: Must be {chrom}-{pos}-{ref}-{alt} without chr prefix. The skill strips chr automatically from VCF CHROM field.

  • ClinVar returns IDs not classifications: The E-utilities search only confirms presence in ClinVar. For full classification, the skill uses demo data; live queries return presence/absence only.

  • Indels in VEP: HGVS notation for indels differs from SNVs. The skill handles SNVs fully; complex indels may return limited VEP results.

  • GRCh37 vs GRCh38: The skill defaults to GRCh38 (hg38). If your VCF uses GRCh37 coordinates, VEP results may be incorrect.

Safety

  • Local-first: No VCF data is uploaded to third-party servers beyond public database APIs (Ensembl, NCBI, gnomAD — all accept variant queries)
  • Disclaimer: Every report includes the ClawBio research disclaimer
  • Not a diagnostic tool: ClinVar classifications are research annotations, not clinical diagnoses
  • Audit trail: All operations logged to reproducibility bundle

Agent Boundary

The agent (LLM) dispatches the VCF and explains results. The skill (Python) executes all API calls and generates files. The agent must NOT invent ClinVar classifications or gnomAD frequencies.

Integration with Bio Orchestrator

Trigger conditions: route here when:

  • File type is .vcf
  • Keywords: annotate, variants, pathogenic, clinvar, gnomad, vep

Chaining partners:

  • pharmgx-reporter: VCF annotation can precede pharmacogenomic reporting
  • equity-scorer: Annotated VCF feeds into population equity analysis
  • lit-synthesizer: Gene names from annotation can seed literature search

Maintenance

  • Review cadence: Monthly — gnomAD and ClinVar update regularly
  • Staleness signals: gnomAD API endpoint changes; ClinVar reclassifications
  • Deprecation: Archive if Ensembl VEP REST API is discontinued

Citations

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/vcf-annotator of ClawBio/ClawBio.

  • SKILL.md
  • README.md
  • examples/demo_output/report.md
  • tests/test_vcf_annotator.py
  • vcf_annotator.py

Open the folder on GitHubat commit cea08d1

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Vcf Annotator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vcf Annotator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vcf Annotator this skillClawBio/ClawBio1.2k1 repos~2.4kAutomated safety check: PassMIT
Human Protein Atlas Databasegoogle-deepmind/science-skills3.2k2 repos~1.6kAutomated safety check: PassApache-2.0
External API ChangeGuyTeichman/RNAlysis140—~1.8kAutomated safety check: PassMIT
Ensembl Databasedavila7/claude-code-templates32k10 repos~2.1kAutomated safety check: PassMIT
Bio DB ToolsDrugClaw/DrugClaw125—~1.4kAutomated safety check: PassApache-2.0
Annotating Variantsmaziyarpanahi/openmed5.5k—~2.1kAutomated safety check: PassApache-2.0

Similar skills

  • Human Protein Atlas Database

    google-deepmind/science-skills

    A skill your agent uses when you want to retrieve semi-quantitative protein expression and spatial localisation data from the Human Protein Atlas (HPA).

    3.2k GitHub starsUsed in 2 repos~1.6k tokens
    Research & ScienceAuto-check passed
  • External API Change

    GuyTeichman/RNAlysis

    Workflow for fixing or changing RNAlysis code that talks to an EXTERNAL WEB SERVICE — UniProt, Ensembl, PANTHER, PhylomeDB, OrthoInspector, KEGG, or GO.

    140 GitHub stars~1.8k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Ensembl Database

    davila7/claude-code-templates

    Query Ensembl genome database REST API for 250+ species. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Bio DB Tools

    DrugClaw/DrugClaw

    Query public biology databases and APIs including UniProt, RCSB PDB, AlphaFold DB, ClinVar, dbSNP, gnomAD, Ensembl, GEO, InterPro, KEGG, OpenTargets, Reactome, and STRING.

    125 GitHub stars~1.4k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Annotating Variants

    maziyarpanahi/openmed

    Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and…

    5.5k GitHub stars~2.1k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~6.3k tokens
    Research & ScienceAuto-check passed

More from ClawBio/ClawBio

All 99 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Works with

Questions about Vcf Annotator

What does Vcf Annotator do?

Annotate VCF variants with Ensembl VEP, ClinVar, and gnomAD. Vcf Annotator is an agent skill from ClawBio/ClawBio. Annotate VCF variants with Ensembl VEP, ClinVar, and gnomAD.

When should I use Vcf Annotator?

Vcf Annotator fits situations like: research & Science work in your project.

How do I install Vcf Annotator in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill vcf-annotator -a claude-code`. Or copy the skill folder (skills/vcf-annotator in ClawBio/ClawBio) into .claude/skills/vcf-annotator in your project. Claude Code loads it when a task matches its description.

How do I install Vcf Annotator in Codex?

Run `npx skills add ClawBio/ClawBio --skill vcf-annotator -a codex`. Or copy the skill folder (skills/vcf-annotator in ClawBio/ClawBio) into .agents/skills/vcf-annotator in your project. Codex loads it when a task matches its description.

Can I use Vcf Annotator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill vcf-annotator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vcf-annotator, .gemini/skills/vcf-annotator, .github/skills/vcf-annotator and .opencode/skills/vcf-annotator in your project.

What does Vcf Annotator need to run?

Going by SKILL.md and its folder, Vcf Annotator needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Vcf Annotator access the network?

SKILL.md names 5 domains. In commands or code: gnomad.broadinstitute.org and rest.ensembl.org; the agent is likely to contact these when it follows the instructions. As links in the text: ensembl.org, ncbi.nlm.nih.gov and doi.org. This is read from the text; nothing was executed.

Is Vcf Annotator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vcf Annotator use?

Vcf Annotator is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vcf Annotator use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vcf Annotator?

Skills that share tags, products or a category with Vcf Annotator: Human Protein Atlas Database (google-deepmind/science-skills, 3.2k stars), External API Change (GuyTeichman/RNAlysis, 140 stars), Ensembl Database (davila7/claude-code-templates, 32k stars) and Bio DB Tools (DrugClaw/DrugClaw, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vcf Annotator?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,152 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 6, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.