Agent skill

Annotating Variants

by maziyarpanahi in maziyarpanahi/openmed

Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and…

Apache-2.0Auto-check passedResearch & Science

Install Annotating Variants

skills CLI
$ npx skills add maziyarpanahi/openmed --skill annotating-variants -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed annotating-variants --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/annotating-variants .claude/skills/annotating-variants && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
annotating-variants
GitHub stars
5.5k
Token cost
~2.1k tokens
SKILL.md length
671 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and…

  • Works in 5 steps: Normalize input to a canonical form:… → Annotate — REST (/vep/human/hgvs or… → Attach frequencies from gnomAD; flag… → …
  • The user wants to predict variant consequences
  • SKILL.md covers When to use, Quick start (real Ensembl VEP…, Population frequencies via… and Offline annotation at scale, plus 4 more sections
  • Calls curl; reaches rest.ensembl.org and gnomad.broadinstitute.org

What it does

Annotating Variants is an agent skill from maziyarpanahi/openmed. Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and the clinical context OpenMed extracts. Use when the user wants to predict variant consequences, map HGVS to genomic coordinates, annotate a VCF, attach allele frequencies, or pair variants with phenotype/oncology context. Trigger keywords: VCF, HGVS, variant annotation, VEP, SnpEff, ANNOVAR, consequence, missense…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. It works with Ensembl. The repository describes itself as: Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data…. The licence is Apache-2.0.

When your agent uses it

  • The user wants to predict variant consequences
  • Map HGVS to genomic coordinates
  • Attach allele frequencies
  • Pair variants with phenotype/oncology context

Example prompts

  • “Use the annotating-variants skill to annotate VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST…”
  • “/annotating-variants”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Normalize input to a canonical form: left-align/trim VCF alleles; for HGVS,
  2. Annotate — REST (/vep/human/hgvs or /vep/human/region) for a handful,
  3. Attach frequencies from gnomAD; flag common variants (e.g. AF > 1%).
  4. Filter/prioritize by most_severe_consequence, impact, and rarity.
  5. Join to clinical context from OpenMed (gene/variant mentions, oncology,

What it can do on your machine

Read from SKILL.md and the folder at commit ea920f3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • rest.ensembl.org
    • gnomad.broadinstitute.org
    • grch37.rest.ensembl.org

    Also links to:

    • samtools.github.io
    • hgvs-nomenclature.org
    • ensembl.org
    • pcingola.github.io
    • ncbi.nlm.nih.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Annotating Variants loads about 2.1k tokens when it runs. Until then it costs about 195 tokens; SKILL.md has 671 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~195
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit ea920f3, republished under its Apache-2.0 licence (© maziyarpanahi). 671 words, ~2,118 tokens.

Download SKILL.mdSave it as .claude/skills/annotating-variants/SKILL.md (or your agent's skills folder).
name
annotating-variants
description
Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and the clinical context OpenMed extracts. Use when the user wants to predict variant consequences, map HGVS to genomic coordinates, annotate a VCF, attach allele frequencies, or pair variants with phenotype/oncology context. Trigger keywords: VCF, HGVS, variant annotation, VEP, SnpEff, ANNOVAR, consequence, missense, gnomAD, allele frequency, GRCh38, rsID, transcript. Pairs adjacent to OpenMed: combine annotated variants with Genomics/Oncology entities and phenotype from openmed.analyze_text. Tools used are free; restricted clinical databases are user-supplied.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
research-genomics
metadata.pairs
adjacent
metadata.version
1.0

Annotating variants & normalizing HGVS

Turn raw genomic variants — VCF rows, rsIDs, or HGVS strings — into annotated, consequence-predicted records, and link them to the clinical context OpenMed extracts from text (genes, variants, oncology findings, phenotype). The workhorse for a quick, no-install annotation is the Ensembl VEP REST API; for scale, run VEP, SnpEff, or ANNOVAR offline.

These annotators are free and license-permissive. Restricted clinical interpretation databases (e.g. licensed HGMD) are user-supplied — this skill sticks to open resources (Ensembl, gnomAD, ClinVar).

When to use

  • You have a VCF / HGVS / rsID and need consequence predictions (missense, stop-gain, splice), affected transcripts, and protein change.
  • You need to normalize HGVS to genomic coordinates (and back) on a known build (GRCh38 by default; GRCh37 via the dedicated endpoint).
  • You want gnomAD population allele frequencies to flag common vs rare.
  • You are pairing molecular findings with the phenotype/oncology context that OpenMed pulls from notes or literature.

Quick start (real Ensembl VEP REST call)

Base URL: https://rest.ensembl.org (GRCh38). For GRCh37 use https://grch37.rest.ensembl.org. Default species is human/homo_sapiens.

python
import requests

REST = "https://rest.ensembl.org"
HEADERS = {"Content-Type": "application/json", "Accept": "application/json"}

def vep_hgvs(hgvs: str) -> list[dict]:
    """Annotate a single HGVS variant (GET)."""
    r = requests.get(f"{REST}/vep/human/hgvs/{hgvs}", headers=HEADERS, timeout=30)
    r.raise_for_status()
    return r.json()

# Transcript-level HGVS (coding) — note the build-aware default transcript set
ann = vep_hgvs("ENST00000269305.9:c.215C>G")   # TP53 example
v = ann[0]
print(v["most_severe_consequence"])            # e.g. "missense_variant"
for tc in v.get("transcript_consequences", []):
    print(tc["gene_symbol"], tc.get("hgvsp"), tc.get("sift_prediction"),
          tc.get("polyphen_prediction"))

Batch many variants with the POST endpoint (region "CHROM POS ID REF ALT . . ." format, up to 200 per request):

python
def vep_region_batch(variants: list[str]) -> list[dict]:
    body = {"variants": variants}   # ["17 7676154 . C G . . .", ...] 1-based
    r = requests.post(f"{REST}/vep/human/region", headers=HEADERS,
                      json=body, timeout=60)
    r.raise_for_status()
    return r.json()

Equivalent cURL:

bash
curl 'https://rest.ensembl.org/vep/human/hgvs/ENST00000269305.9:c.215C>G' \
  -H 'Content-Type:application/json'

Response highlights per variant: most_severe_consequence, transcript_consequences[] (gene_symbol, hgvsc, hgvsp, sift_prediction, polyphen_prediction, impact), and colocated_variants[] (rsIDs and population frequencies). Request gnomAD frequencies and ClinVar via VEP options / plugins.

Population frequencies via gnomAD (GraphQL)

For authoritative allele frequencies, query the gnomAD GraphQL API at https://gnomad.broadinstitute.org/api. Use variant IDs in chrom-pos-ref-alt form. Frequencies are derived from ac/an (allele count / number) — request those, not a non-existent af on subpopulations.

python
GNOMAD = "https://gnomad.broadinstitute.org/api"

QUERY = """
query Variant($id: String!, $ds: DatasetId!) {
  variant(variantId: $id, dataset: $ds) {
    variant_id rsids
    genome { ac an af homozygote_count }
    exome  { ac an af homozygote_count }
  }
}"""

def gnomad_freq(variant_id: str, dataset: str = "gnomad_r4") -> dict:
    r = requests.post(GNOMAD, json={"query": QUERY,
        "variables": {"id": variant_id, "ds": dataset}}, timeout=30)
    r.raise_for_status()
    return r.json()["data"]["variant"]

# gnomad_freq("17-7676154-C-G")  -> ac/an/af for exome and genome

Offline annotation at scale

For whole-VCF jobs, run a local annotator instead of per-variant REST calls:

ToolStrengthsNotes
Ensembl VEP (offline)richest, plugin ecosystem (gnomAD, CADD, SpliceAI), HGVSneeds cache download per build
SnpEfffast, self-contained genome databasesgreat for bulk consequence calling
ANNOVARmany annotation databasesregistration required; license terms apply

All emit per-variant gene, consequence, and (with the right database) frequency and clinical fields. Keep the reference build (GRCh38) consistent end to end.

Workflow

  1. Normalize input to a canonical form: left-align/trim VCF alleles; for HGVS, confirm the reference transcript and build.
  2. Annotate — REST (/vep/human/hgvs or /vep/human/region) for a handful, offline VEP/SnpEff for a VCF.
  3. Attach frequencies from gnomAD; flag common variants (e.g. AF > 1%).
  4. Filter/prioritize by most_severe_consequence, impact, and rarity.
  5. Join to clinical context from OpenMed (gene/variant mentions, oncology, phenotype) to assemble an interpretable record.
Show full SKILL.md (274 more words)Show less

Hand-off to / from OpenMed

  • OpenMed → variant context. openmed.analyze_text(report, model_name=<a Genomics or Oncology model>) extracts gene symbols, variant mentions (e.g. "EGFR L858R"), and tumor/oncology findings from pathology or molecular reports. Use those to (a) select which VCF variants matter and (b) attach phenotype context to each annotation.
  • Variant → OpenMed. Free-text variant descriptions in reports can be normalized to HGVS here, then the surrounding clinical narrative is structured by OpenMed — linking genotype to extracted phenotype/diagnosis.
  • Keep genomic + clinical data local. The REST/GraphQL calls carry only the variant coordinates (public allele data), never patient identifiers — and any narrative is de-identified with openmed.deidentify first.

Edge cases & gotchas

  • Build mismatch is the #1 error. GRCh38 coordinates against a GRCh37 endpoint (or cache) give wrong genes. Use grch37.rest.ensembl.org only for GRCh37 data; default REST is GRCh38.
  • Transcript choice changes the HGVS. c./p. notation depends on the reference transcript (MANE Select vs others). Pin the transcript explicitly.
  • Normalize before annotating. Un-left-aligned indels and multi-allelic VCF rows produce inconsistent annotations — decompose and normalize first (e.g. bcftools norm).
  • REST is rate-limited. ~15 req/s and 200 variants/POST on the Ensembl REST server; switch to offline VEP for large VCFs. Honor Retry-After on 429.
  • gnomAD subpopulation fields. Query ac/an (and compute AF) for subpopulations; some schema paths reject af directly — track the current schema version, which changes between gnomAD releases.
  • No clinical interpretation here. Consequence ≠ pathogenicity. Pathogenicity classification (ACMG/AMP) uses curated evidence and licensed databases the user supplies; this skill produces annotations, not diagnoses.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/annotating-variants of maziyarpanahi/openmed.

Open the folder on GitHubat commit ea920f3

Compare with similar skills

Annotating Variants next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Annotating Variants compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Annotating Variants this skillmaziyarpanahi/openmed5.5k—~2.1kAutomated safety check: PassApache-2.0
External API ChangeGuyTeichman/RNAlysis140—~1.8kAutomated safety check: PassMIT
Ensembl Databasedavila7/claude-code-templates32k10 repos~2.1kAutomated safety check: PassMIT
Ggetdavila7/claude-code-templates32k11 repos~6.3kAutomated safety check: PassMIT
Ensembl Databasegoogle-deepmind/science-skills3.2k1 repos~2.2kAutomated safety check: PassApache-2.0
Scientific Pkg Ggetaffaan-m/ECC275k1 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • External API Change

    GuyTeichman/RNAlysis

    Workflow for fixing or changing RNAlysis code that talks to an EXTERNAL WEB SERVICE — UniProt, Ensembl, PANTHER, PhylomeDB, OrthoInspector, KEGG, or GO.

    140 GitHub stars~1.8k tokensUpdated 10 days ago
    Research & ScienceAuto-check passed
  • Ensembl Database

    davila7/claude-code-templates

    Query Ensembl genome database REST API for 250+ species. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Ensembl Database

    google-deepmind/science-skills

    Query the Ensembl database to resolve gene, transcript, and protein IDs, fetch genomic or protein sequences, retrieve gene structures (exons), and get variant consequence and effect predictions (VEP).

    3.2k GitHub starsUsed in 1 repo~2.2k tokens
    Research & ScienceAuto-check passed
  • gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs.

    275k GitHub starsUsed in 1 repo~1.3k tokens
    Research & ScienceAuto-check passed
  • Production-ready phylogenetics and sequence analysis skill for alignment processing, tree analysis, and evolutionary metrics.

    1.1k GitHub starsUsed in 2 repos~4.2k tokens
    Research & ScienceAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Annotating Variants

What does Annotating Variants do?

Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and…. Annotating Variants is an agent skill from maziyarpanahi/openmed. Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and the clinical context OpenMed extracts.

When should I use Annotating Variants?

Annotating Variants fits situations like: the user wants to predict variant consequences; map HGVS to genomic coordinates; attach allele frequencies; pair variants with phenotype/oncology context.

How do I install Annotating Variants in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill annotating-variants -a claude-code`. Or copy the skill folder (skills/annotating-variants in maziyarpanahi/openmed) into .claude/skills/annotating-variants in your project. Claude Code loads it when a task matches its description.

How do I install Annotating Variants in Codex?

Run `npx skills add maziyarpanahi/openmed --skill annotating-variants -a codex`. Or copy the skill folder (skills/annotating-variants in maziyarpanahi/openmed) into .agents/skills/annotating-variants in your project. Codex loads it when a task matches its description.

Can I use Annotating Variants in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill annotating-variants -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/annotating-variants, .gemini/skills/annotating-variants, .github/skills/annotating-variants and .opencode/skills/annotating-variants in your project.

What does Annotating Variants need to run?

Going by SKILL.md and its folder, Annotating Variants needs the command-line tools its instructions call (curl). Our summary lists: Python 3.

Does Annotating Variants access the network?

SKILL.md names 8 domains. In commands or code: rest.ensembl.org, gnomad.broadinstitute.org and grch37.rest.ensembl.org; the agent is likely to contact these when it follows the instructions. As links in the text: samtools.github.io, hgvs-nomenclature.org, ensembl.org, pcingola.github.io and ncbi.nlm.nih.gov. This is read from the text; nothing was executed.

Is Annotating Variants safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Annotating Variants use?

Annotating Variants is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Annotating Variants use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Annotating Variants?

Skills that share tags, products or a category with Annotating Variants: External API Change (GuyTeichman/RNAlysis, 140 stars), Ensembl Database (davila7/claude-code-templates, 32k stars), Gget (davila7/claude-code-templates, 32k stars) and Ensembl Database (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Annotating Variants?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,457 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 7, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.