Agent skill

Ensembl Database

by google-deepmind in google-deepmind/science-skills

Query the Ensembl database to resolve gene, transcript, and protein IDs, fetch genomic or protein sequences, retrieve gene structures (exons), and get variant consequence and effect predictions (VEP).

Apache-2.0Auto-check passedResearch & Science

Install Ensembl Database

skills CLI
$ npx skills add google-deepmind/science-skills --skill ensembl-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-deepmind/science-skills ensembl-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ensembl_database .claude/skills/ensembl-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ensembl-database
GitHub stars
3.2k
Used in
1 other repo
Token cost
~2.2k tokens
SKILL.md length
824 words
Files
4 (incl. scripts, references)
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Query the Ensembl database to resolve gene, transcript, and protein IDs, fetch genomic or protein sequences, retrieve gene structures (exons), and get variant consequence and effect predictions (VEP).

  • Works in 2 steps: uv: Read the uv skill and follow its… → User Notification: If…
  • Tasks that involve Bioinformatics
  • SKILL.md covers Prerequisites, Overview, Core Rules and Parsing Outputs, plus 1 more section
  • Runs Python scripts from its folder; calls uv

What it does

Ensembl Database is an agent skill from google-deepmind/science-skills. Query the Ensembl database to resolve gene, transcript, and protein IDs, fetch genomic or protein sequences, retrieve gene structures (exons), and get variant consequence and effect predictions (VEP). Use this skill as a primary ID translator, genomic sequence database and variant effect prediction tool.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/ensembl_rest_api_reference.md` and `scripts/ensembl_api.py`).

It sits in Research & Science, covering Bioinformatics and Translation. It works with Ensembl. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve Translation

Example prompts

  • “/ensembl-database”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. uv: Read the uv skill and follow its Setup instructions to ensure
  2. User Notification: If .licenses/ensembl_database_LICENSE.txt does not

What it can do on your machine

Read from SKILL.md and the folder at commit 6883275. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • useast.ensembl.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ensembl Database loads about 2.2k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 81 tokens; SKILL.md has 824 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google-deepmind/science-skills at commit 6883275, republished under its Apache-2.0 licence (© google-deepmind). 824 words, ~2,202 tokens.

Download SKILL.mdSave it as .claude/skills/ensembl-database/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ensembl-database
description
Query the Ensembl database to resolve gene, transcript, and protein IDs, fetch genomic or protein sequences, retrieve gene structures (exons), and get variant consequence and effect predictions (VEP). Use this skill as a primary ID translator, genomic sequence database and variant effect prediction tool.

Ensembl Database: ID Mapping and Genomic Features

Prerequisites

  1. uv: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH.
  2. User Notification: If .licenses/ensembl_database_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://useast.ensembl.org/index.html and https://github.com/Ensembl/ensembl-rest/wiki, then (2) create the file recording the notification text and timestamp.

Overview

The Ensembl database is a resource for genome annotation. This skill allows you to interact with the Ensembl REST API to resolve ambiguous symbols, cross-reference IDs (RefSeq, HGNC, UniProt, ENSG), fetch raw sequences, and retrieve detailed transcript structures.

Key Concepts:

  • ENSG (Gene): Stable identifier for a human gene. Other species will have different three-letter species codes.
  • ENST (Transcript): Stable identifier for a transcript (splicing isoform).
  • ENSP (Protein): Stable identifier for a translated protein.
  • MANE Select: The consensus primary transcript agreed upon by Ensembl and NCBI.
  • Canonical: Ensembl's representative transcript (used if MANE is not available or non-human).

Core Rules

  • Use the Wrapper: ALWAYS execute the provided helper scripts to query the database rather than accessing the database directly. The scripts automatically enforce the required rate limit gracefully.
  • Default Species: If the species is absent or ambiguous in the prompt, default to "human". You MUST explicitly flag this default to the user to ensure they are aware.
  • Primary Transcripts: When listing transcripts for a gene, only return the MANE Select transcript (for human) or the Canonical transcript (for others) unless the user explicitly asks for all alternative isoforms. You MUST flag to the user when multiple transcripts are available and you are defaulting to the primary one.
  • Assembly Handling: The default assembly is GRCh38. For GRCh37 requests, you MUST use the --assembly GRCh37 flag. You MUST explicitly flag to the user when a non-default assembly is being used.
  • Output Location: The script writes full JSON/FASTA output to temporary files in /tmp by default, or to a user-specified file using the --output flag. It also prints a concise summary to stdout.
  • Notification: If this skill is used, ensure this is mentioned in the output.
Available Commands

1. Resolve Gene ID — Resolve a symbol, alias, or RefSeq ID to ENSG ID(s). Automatically falls back to resolving synonyms if primary symbol is not found.

bash
uv run scripts/ensembl_api.py resolve-gene TP53 --species human --output tp53.json
uv run scripts/ensembl_api.py resolve-gene PCL2 --output pcl2.json # Falls back to synonym resolution

2. Map ID to External Database — Cross-reference an Ensembl ID to UniProt, HGNC, RefSeq, etc.

bash
uv run scripts/ensembl_api.py map-id ENSG00000141510 --external-db UniProt --output uniprot_map.json
uv run scripts/ensembl_api.py map-id ENST00000269305 --external-db RefSeq_mRNA --output refseq_map.json

3. Get Genomic Sequence — Fetch raw DNA for a coordinate window. Supports GRCh37 via --assembly GRCh37.

bash
uv run scripts/ensembl_api.py get-sequence 17:7661779-7687550 --species human --output seq.txt
uv run scripts/ensembl_api.py get-sequence chr9:21971100-21971200 --assembly GRCh37 --output seq_grch37.txt

4. Gene Summary — High-level metadata: symbol, biotype, description, chromosomal location.

bash
uv run scripts/ensembl_api.py gene-summary ENSG00000141510 --output gene_summary.json

5. List Transcripts — All transcripts for a gene, with optional --only-mane or --only-canonical filters. Output includes Transcript Support Level (TSL).

bash
uv run scripts/ensembl_api.py transcripts ENSG00000141510 --only-mane --output transcripts_mane.json
uv run scripts/ensembl_api.py transcripts ENSG00000141510 --only-canonical --output transcripts_canonical.json
uv run scripts/ensembl_api.py transcripts ENSG00000141510 --output transcripts_all.json

5b. Canonical TSS — Get the single coordinate of the Transcription Start Site (TSS) for the canonical transcript of a gene.

[!NOTE] Unlike the standard transcripts command, canonical-tss accepts both symbols (e.g., TP53) and Ensembl IDs, and automatically resolves them. It also does the math for strand orientation (TSS is Start for + strand and End for - strand), outputting the single integer coordinate directly.

Show full SKILL.md (324 more words)Show less
bash
uv run scripts/ensembl_api.py canonical-tss TP53 --output tp53_tss.json
uv run scripts/ensembl_api.py canonical-tss ENSG00000141510 --output tss.json

6. Transcript Structure — Exon coordinates, CDS boundaries, and computed 5'/3' UTR regions for a transcript.

bash
uv run scripts/ensembl_api.py transcript-structure ENST00000269305 --output structure.json

7. Protein Info — ENSP ID and sequence length for a transcript.

bash
uv run scripts/ensembl_api.py protein-info ENST00000269305 --output protein_info.json

8. Protein Sequence — Amino acid FASTA for a transcript (ENST) or protein (ENSP) ID.

bash
uv run scripts/ensembl_api.py protein-sequence ENST00000269305 --output protein.fasta
uv run scripts/ensembl_api.py protein-sequence ENSP00000269305 --output protein_ensp.fasta

9. Variant Consequence (VEP) — Predict molecular consequences for a genomic variant. Includes open-licensed plugins: AlphaMissense, Conservation, DosageSensitivity, IntAct, MaveDB, OpenTargets, LoF (Loftee), NMD, UTRAnnotator, mutfunc, LOEUF.

bash
uv run scripts/ensembl_api.py vep 9:21971147:T:C --species human --output vep.json
uv run scripts/ensembl_api.py vep rs699 --species human --output vep_rs699.json

Example VEP stdout output:

[*] Variant: 9:21971147:T>C
[*] Most severe consequence: missense_variant
[*] Found 15 transcript consequences.

[*] VEP Predictions:

  - ENST00000304494 (CDKN2A): Consequence = missense_variant
  - ENST00000304494 (CDKN2A): Amino Acids = N/S
  - ENST00000304494 (CDKN2A): SIFT = deleterious (0.01)
  - ENST00000304494 (CDKN2A): AlphaMissense Class = likely_benign
  - ENST00000304494 (CDKN2A): AlphaMissense Pathogenicity = 0.2129
  - ENST00000304494 (CDKN2A): Conservation = 2.05
  - ENST00000304494 (CDKN2A): Dosage Sensitivity (Haplo) = 0.889228328567991
  - ENST00000304494 (CDKN2A): Dosage Sensitivity (Triplo) = 0.135514349094646
  - ENST00000304494 (CDKN2A): Loss of Function (LOEUF) = 0.791

Presenting VEP Results: After running the VEP command, you MUST present the full VEP Predictions list from stdout to the user. This list contains both standard VEP predictions (Consequence, Amino Acids, SIFT, PolyPhen) and open-license plugin results (AlphaMissense, Conservation, Dosage Sensitivity, LOEUF, Loftee LoF, NMD, UTRAnnotator, Mutfunc). Do NOT just summarize — show the complete list so the user can see all predictions. If the list is very long (many transcripts), show the MANE Select / canonical transcript rows in full and note that the complete data is in the JSON output.

Parsing Outputs

If the user needs detailed, nested structural data (like the precise integer coordinates of Exon 2 of a transcript) that isn't summarized in stdout:

  1. Locate the JSON file (either specified via --output or the temporary file path printed by the script).
  2. Use terminal tools like jq or write a quick, disposable python snippet to extract the specific data point requested. Do not attempt to read the entire JSON file into your context if it is very large.

Custom Queries

If you need to make an API call that the script does not support (e.g., fetching protein domain annotations, coordinate mapping between assemblies, homology searches, linkage disequilibrium, or phenotype lookups), read references/ensembl_rest_api_reference.md for a complete reference of available endpoints, parameters, and response fields.

CRITICAL: When writing custom scripts or using alternatives to the provided scripts, you MUST respect the Ensembl REST API rate limits (maximum 15 requests per second) and handle 429 Too Many Requests errors gracefully (e.g., with exponential backoff).

© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/ensembl_database of google-deepmind/science-skills.

  • SKILL.md
  • references/citation.bib
  • references/ensembl_rest_api_reference.md
  • scripts/ensembl_api.py

Open the folder on GitHubat commit 6883275

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in google-deepmind/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Ensembl Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ensembl Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ensembl Database this skillgoogle-deepmind/science-skills3.2k1 repos~2.2kAutomated safety check: PassApache-2.0
External API ChangeGuyTeichman/RNAlysis140—~1.8kAutomated safety check: PassMIT
Ensembl Databasedavila7/claude-code-templates32k10 repos~2.1kAutomated safety check: PassMIT
Annotating Variantsmaziyarpanahi/openmed5.5k—~2.1kAutomated safety check: PassApache-2.0
Ggetdavila7/claude-code-templates32k11 repos~6.3kAutomated safety check: PassMIT
Scientific Pkg Ggetaffaan-m/ECC274k1 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • External API Change

    GuyTeichman/RNAlysis

    Workflow for fixing or changing RNAlysis code that talks to an EXTERNAL WEB SERVICE — UniProt, Ensembl, PANTHER, PhylomeDB, OrthoInspector, KEGG, or GO.

    140 GitHub stars~1.8k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Ensembl Database

    davila7/claude-code-templates

    Query Ensembl genome database REST API for 250+ species. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Annotating Variants

    maziyarpanahi/openmed

    Annotates VCF variants and normalizes HGVS nomenclature with public, license-free annotators (Ensembl VEP REST, VEP/SnpEff/ANNOVAR offline) and links variants to gnomAD population frequencies and…

    5.5k GitHub stars~2.1k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs.

    274k GitHub starsUsed in 1 repo~1.3k tokens
    Research & ScienceAuto-check passed
  • Bulkrna Geneid Mapping

    TianGzlab/OmicsClaw

    Load when converting gene identifiers between Ensembl, Entrez, and HGNC symbol in a bulk RNA-seq count matrix.

    161 GitHub stars~1k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed

More from google-deepmind/science-skills

All 40 skills in this repo
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Auto-check: notes
  • Chembl Database

    google-deepmind/science-skills

    Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.

    3.2k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Auto-check: notes
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Auto-check: notes

Works with

Questions about Ensembl Database

What does Ensembl Database do?

Query the Ensembl database to resolve gene, transcript, and protein IDs, fetch genomic or protein sequences, retrieve gene structures (exons), and get variant consequence and effect predictions (VEP). Ensembl Database is an agent skill from google-deepmind/science-skills. Query the Ensembl database to resolve gene, transcript, and protein IDs, fetch genomic or protein sequences, retrieve gene structures (exons), and get variant consequence and effect predictions (VEP).

When should I use Ensembl Database?

Ensembl Database fits situations like: tasks that involve Bioinformatics; tasks that involve Translation.

How do I install Ensembl Database in Claude Code?

Run `npx skills add google-deepmind/science-skills --skill ensembl-database -a claude-code`. Or copy the skill folder (skills/ensembl_database in google-deepmind/science-skills) into .claude/skills/ensembl-database in your project. Claude Code loads it when a task matches its description.

How do I install Ensembl Database in Codex?

Run `npx skills add google-deepmind/science-skills --skill ensembl-database -a codex`. Or copy the skill folder (skills/ensembl_database in google-deepmind/science-skills) into .agents/skills/ensembl-database in your project. Codex loads it when a task matches its description.

Can I use Ensembl Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill ensembl-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ensembl-database, .gemini/skills/ensembl-database, .github/skills/ensembl-database and .opencode/skills/ensembl-database in your project.

What does Ensembl Database need to run?

Going by SKILL.md and its folder, Ensembl Database needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Ensembl Database access the network?

SKILL.md names 2 domains. As links in the text: useast.ensembl.org and github.com. This is read from the text; nothing was executed.

Is Ensembl Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ensembl Database use?

Ensembl Database is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ensembl Database use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Ensembl Database?

Skills that share tags, products or a category with Ensembl Database: External API Change (GuyTeichman/RNAlysis, 140 stars), Ensembl Database (davila7/claude-code-templates, 32k stars), Annotating Variants (maziyarpanahi/openmed, 5.5k stars) and Gget (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ensembl Database?

google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,216 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on September 15, 2026.

Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.