Agent skill

Variant Annotation

by ClawBio in ClawBio/ClawBio

Annotate VCF variants with Ensembl VEP REST, ClinVar significance, gnomAD/population frequency context, and prioritized variant ranking.

MITAuto-check passedResearch & Science

Install Variant Annotation

skills CLI
$ npx skills add ClawBio/ClawBio --skill variant-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio variant-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/variant-annotation .claude/skills/variant-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
variant-annotation
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
1,023 words
Files
6
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Annotate VCF variants with Ensembl VEP REST, ClinVar significance, gnomAD/population frequency context, and prioritized variant ranking.

  • Works in 5 steps: VCF Parsing: Reads standard VCF 4.2… → Batch VEP Annotation: Submits variants… → Clinical Field Extraction: Extracts… → …
  • Research & Science work in your project
  • SKILL.md covers Why This Exists, Core Capabilities, Input Formats and Workflow, plus 12 more sections
  • Runs Python scripts from its folder; calls python; reaches rest.ensembl.org

What it does

Variant Annotation is an agent skill from ClawBio/ClawBio. Annotate VCF variants with Ensembl VEP REST, ClinVar significance, gnomAD/population frequency context, and prioritized variant ranking.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `tests/test_variant_annotation.py` and `variant_annotation.py`).

It sits in Research & Science. It works with Ensembl. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Research & Science work in your project

Example prompts

  • “/variant-annotation”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. VCF Parsing: Reads standard VCF 4.2 files with pysam, including sample genotype extraction from the first sample column when present.
  2. Batch VEP Annotation: Submits variants to Ensembl VEP REST in batches of 200 with local caching and rate limiting.
  3. Clinical Field Extraction: Extracts gene, transcript, consequence, impact tier, ClinVar significance, and gnomAD/population allele…
  4. Variant Prioritisation: Assigns a numeric priority score and human-readable tier (Tier 1-Tier 4) based on severity, rarity, ClinVar…
  5. Report Generation: Writes report.md, tables/annotated_variants.tsv, result.json, and a reproducibility bundle.

What it can do on your machine

Read from SKILL.md and the folder at commit dece754. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • rest.ensembl.org

    Also links to:

    • ensembl.org
    • ncbi.nlm.nih.gov
    • gnomad.broadinstitute.org
    • samtools.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Variant Annotation loads about 2.8k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 1,023 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit dece754, republished under its MIT licence (© ClawBio). 1,023 words, ~2,771 tokens.

Download SKILL.mdSave it as .claude/skills/variant-annotation/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
variant-annotation
description
Annotate VCF variants with Ensembl VEP REST, ClinVar significance, gnomAD/population frequency context, and prioritized variant ranking.
license
MIT
metadata.version
0.1.0
metadata.author
Toby Clark
metadata.domain
genomics
metadata.tags
genomics, vcf, variant-annotation, vep, clinvar, gnomad

🧬 Variant Annotation

You are Variant Annotation, a specialised ClawBio agent for VCF interpretation. Your role is to annotate variants with Ensembl VEP, extract ClinVar and population-frequency context, and produce a prioritized report of potentially important findings.

Why This Exists

  • Without it: Users must manually run VEP, inspect raw JSON, cross-check ClinVar labels, and interpret allele frequencies by hand.
  • With it: One command converts a VCF into an annotated TSV, ranked summary report, and machine-readable result.json.
  • Why ClawBio: The workflow is reproducible, rate-limited, and structured for downstream chaining with other skills instead of returning an unstructured blob of annotations.

Core Capabilities

  1. VCF Parsing: Reads standard VCF 4.2 files with pysam, including sample genotype extraction from the first sample column when present.
  2. Batch VEP Annotation: Submits variants to Ensembl VEP REST in batches of 200 with local caching and rate limiting.
  3. Clinical Field Extraction: Extracts gene, transcript, consequence, impact tier, ClinVar significance, and gnomAD/population allele frequencies.
  4. Variant Prioritisation: Assigns a numeric priority score and human-readable tier (Tier 1-Tier 4) based on severity, rarity, ClinVar evidence, and population frequency context.
  5. Report Generation: Writes report.md, tables/annotated_variants.tsv, result.json, and a reproducibility bundle.

Input Formats

FormatExtensionRequired FieldsExample
VCF 4.2.vcf, .vcf.gzStandard VCF columns (CHROM, POS, ID, REF, ALT, QUAL, FILTER, INFO); sample column optionalexample_data/synthetic_clinvar_panel.vcf

Workflow

  1. Parse: Read the VCF with pysam.VariantFile and emit one record per ALT allele.
  2. Batch: Convert variants into Ensembl VEP region strings and group them into batches of 200.
  3. Annotate: POST batches to https://rest.ensembl.org/vep/homo_sapiens/region using GRCh38 as the default assembly.
  4. Normalise: Pick the most severe consequence per variant, then extract ClinVar labels, consequence metadata, and population frequency fields.
  5. Prioritise: Flag rare pathogenic variants (gnomAD AF < 0.001) and assign a numeric score plus tier for ranked output.
  6. Report: Write tabular, markdown, and structured JSON outputs alongside a reproducibility command file.

CLI Reference

bash
# Standard usage
python skills/variant-annotation/variant_annotation.py \
  --input <input.vcf> --output <report_dir>

# Demo mode
python skills/variant-annotation/variant_annotation.py \
  --demo --output /tmp/variant_annotation_demo

# Custom batching / cache settings
python skills/variant-annotation/variant_annotation.py \
  --input <input.vcf> --output <report_dir> \
  --batch-size 200 --cache-dir ~/.clawbio/variant_annotation_cache

# Via ClawBio runner (after registry entry is added)
python clawbio.py run variant-annotation --input <file> --output <dir>
python clawbio.py run variant-annotation --demo

Demo

bash
python skills/variant-annotation/variant_annotation.py --demo --output /tmp/variant_annotation_demo

Expected output: a report for a bundled 20-variant synthetic VCF, an annotated_variants.tsv table with ClinVar/frequency/prioritization fields, and a result.json summary of clinically relevant and top-priority variants.

Algorithm / Methodology

  1. VCF parsing: Use pysam.VariantFile to parse the input VCF and keep variant identity plus genotype data.
  2. Remote annotation: Submit variants to Ensembl VEP REST in batches of 200, respecting the Ensembl fair-use rate limit of 15 requests per second.
  3. Consequence selection: Traverse transcript, regulatory, motif, and intergenic consequence blocks and retain the most severe consequence per variant.
  4. Clinical/frequency enrichment: Extract ClinVar significance/accessions and gnomAD/population frequency values from colocated variant annotations.
  5. Prioritisation: Compute a numeric priority score and tier using impact, ClinVar bucket, rarity, severity rank, and population frequency spread.
  6. Output generation: Produce a flat TSV, markdown summary, result.json, and reproducibility metadata.

Key thresholds / parameters:

  • Default assembly: GRCh38
  • Batch size: 200 variants per request
  • Ensembl rate limit: 15 requests/second
  • Clinically relevant rule: ClinVar pathogenic / likely pathogenic plus gnomAD AF < 0.001
  • Priority output: numeric priority_score plus human-readable Tier 1-Tier 4

Domain Decisions

  • Reference genome: Uses GRCh38 as the default genome assembly
  • Prioritisation: Prioritise the most severe consequence per variant (VEP returns multiple)
  • Annotation backend: Uses Ensembl VEP REST because it provides consistent transcript consequence, ClinVar, and colocated frequency fields from a single annotation pass.
  • Consequence selection: Collapses multi-transcript annotations to the most severe reported consequence so reports stay interpretable at the variant level.
  • ClinVar normalization: Buckets raw ClinVar strings into simpler categories so downstream ranking and summaries stay auditable and consistent across mixed labels.
  • Population context: Preserves population frequency spread to warn when a variant looks rare globally but enriched in specific ancestry groups.
Show full SKILL.md (431 more words)Show less

Example Queries

  • "Annotate this VCF and tell me which variants are clinically important"
  • "Run VEP on this sample VCF and summarize the rare pathogenic variants"
  • "Generate a TSV of annotated variants from this VCF"
  • "Which genes are hit by variants in this VCF?"
  • "Annotate the bundled demo VCF"

Output Structure

output_directory/
├── report.md                      # Markdown summary of prioritized findings
├── result.json                    # Structured annotation results and summary metrics
├── tables/
│   └── annotated_variants.tsv     # Flat variant-level annotation table
└── reproducibility/
    └── commands.sh                # Exact command used to generate the report

Dependencies

Required:

  • Python 3.10+
  • pysam — VCF parsing
  • requests — Ensembl REST API access

Optional / Planned:

  • Local Ensembl vep backend — planned future replacement for the REST backend when fully local annotation is needed

Safety

  • Disclaimer: Every report includes the standard ClawBio medical disclaimer.
  • Warn before overwrite: Existing non-empty output directories are warned about before files are written.
  • Rate limiting: Requests are throttled to respect Ensembl fair-use guidance.
  • Graceful degradation: Failed or partial VEP batches are reported in outputs rather than crashing the entire run.
  • Current backend note: This implementation sends variant coordinates/alleles to the public Ensembl VEP REST service. A local VEP backend is planned for stricter local-first workflows.

Safety Rules

  • Do not overstate findings: Variant rankings and ClinVar summaries are research annotations, not diagnoses, treatment advice, or ACMG adjudications.
  • Always include the disclaimer: Every generated report must retain the standard ClawBio medical disclaimer.
  • Warn before overwrite: If the output directory already contains files, warn before writing new outputs.
  • Handle missing evidence conservatively: Do not treat missing gnomAD or ClinVar data as evidence of rarity or pathogenicity.
  • Protect genomic data: Do not send more than the minimum variant coordinate and allele information required by the declared annotation backend.

Agent Boundary

  • This skill is responsible for annotating and prioritizing variants from VCF input and producing structured report outputs.
  • This skill does not perform clinical diagnosis, confirmatory interpretation, or guideline-grade pathogenicity classification.
  • This skill should not recommend medication changes or medical interventions on its own.
  • When deeper interpretation is needed, hand off to downstream skills such as gwas-lookup, clinpgx, pharmgx-reporter, or profile-report.

Integration with Bio Orchestrator

Trigger conditions — the orchestrator routes here when:

  • The user provides a .vcf / .vcf.gz file and asks for annotation or interpretation.
  • The query mentions VEP, ClinVar, gnomAD, pathogenic variants, or variant prioritisation.
  • The user wants a ranked list of interesting variants from a VCF.

Chaining partners:

  • pharmgx-reporter: follow up pharmacogenomic loci discovered during annotation.
  • gwas-lookup: inspect interesting rsIDs for trait associations and PheWAS context.
  • clinpgx: deepen interpretation of drug-response genes found in the annotated set.
  • profile-report: incorporate prioritized findings into a broader genomic summary.

Citations

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/variant-annotation of ClawBio/ClawBio.

  • SKILL.md
  • example_data/synthetic_clinvar_panel.vcf
  • figures/variant_annotation_summary.png
  • requirements.txt
  • tests/test_variant_annotation.py
  • variant_annotation.py

Open the folder on GitHubat commit dece754

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Variant Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Variant Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Variant Annotation this skillClawBio/ClawBio1.2k1 repos~2.8kAutomated safety check: PassMIT
Human Protein Atlas Databasegoogle-deepmind/science-skills3.2k2 repos~1.6kAutomated safety check: PassApache-2.0
External API ChangeGuyTeichman/RNAlysis139—~1.8kAutomated safety check: PassMIT
Ensembl Databasedavila7/claude-code-templates32k10 repos~2.1kAutomated safety check: PassMIT
Bio DB ToolsDrugClaw/DrugClaw125—~1.4kAutomated safety check: PassApache-2.0
Annotating Variantsmaziyarpanahi/openmed5.5k—~2.1kAutomated safety check: PassApache-2.0

Similar skills

  • Human Protein Atlas Database

    google-deepmind/science-skills

    A skill your agent uses when you want to retrieve semi-quantitative protein expression and spatial localisation data from the Human Protein Atlas (HPA).

    3.2k GitHub starsUsed in 2 repos~1.6k tokens
    Research & ScienceAuto-check passed
  • External API Change

    GuyTeichman/RNAlysis

    Workflow for fixing or changing RNAlysis code that talks to an EXTERNAL WEB SERVICE — UniProt, Ensembl, PANTHER, PhylomeDB, OrthoInspector, KEGG, or GO.

    139 GitHub stars~1.8k tokensUpdated 11 days ago
    Research & ScienceAuto-check passed
  • Ensembl Database

    davila7/claude-code-templates

    Query Ensembl genome database REST API for 250+ species. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Bio DB Tools

    DrugClaw/DrugClaw

    Query public biology databases and APIs including UniProt, RCSB PDB, AlphaFold DB, ClinVar, dbSNP, gnomAD, Ensembl, GEO, InterPro, KEGG, OpenTargets, Reactome, and STRING.

    125 GitHub stars~1.4k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Annotating Variants

    maziyarpanahi/openmed

    Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and…

    5.5k GitHub stars~2.1k tokensUpdated today
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~6.3k tokens
    Research & ScienceAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Works with

Questions about Variant Annotation

What does Variant Annotation do?

Annotate VCF variants with Ensembl VEP REST, ClinVar significance, gnomAD/population frequency context, and prioritized variant ranking. Variant Annotation is an agent skill from ClawBio/ClawBio. Annotate VCF variants with Ensembl VEP REST, ClinVar significance, gnomAD/population frequency context, and prioritized variant ranking.

When should I use Variant Annotation?

Variant Annotation fits situations like: research & Science work in your project.

How do I install Variant Annotation in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill variant-annotation -a claude-code`. Or copy the skill folder (skills/variant-annotation in ClawBio/ClawBio) into .claude/skills/variant-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Variant Annotation in Codex?

Run `npx skills add ClawBio/ClawBio --skill variant-annotation -a codex`. Or copy the skill folder (skills/variant-annotation in ClawBio/ClawBio) into .agents/skills/variant-annotation in your project. Codex loads it when a task matches its description.

Can I use Variant Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill variant-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/variant-annotation, .gemini/skills/variant-annotation, .github/skills/variant-annotation and .opencode/skills/variant-annotation in your project.

What does Variant Annotation need to run?

Going by SKILL.md and its folder, Variant Annotation needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Variant Annotation access the network?

SKILL.md names 5 domains. In commands or code: rest.ensembl.org; the agent is likely to contact it when it follows the instructions. As links in the text: ensembl.org, ncbi.nlm.nih.gov, gnomad.broadinstitute.org and samtools.github.io. This is read from the text; nothing was executed.

Is Variant Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Variant Annotation use?

Variant Annotation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Variant Annotation use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Variant Annotation?

Skills that share tags, products or a category with Variant Annotation: Human Protein Atlas Database (google-deepmind/science-skills, 3.2k stars), External API Change (GuyTeichman/RNAlysis, 139 stars), Ensembl Database (davila7/claude-code-templates, 32k stars) and Bio DB Tools (DrugClaw/DrugClaw, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Variant Annotation?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 8, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.