Agent skill

Ena Database

by aipoch in aipoch/medical-research-skills

Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need…

MITAuto-check passedBackend & APIs

Install Ena Database

skills CLI
$ npx skills add aipoch/medical-research-skills --skill ena-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills ena-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Evidence Insight/ena-database' .claude/skills/ena-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ena-database
GitHub stars
2k
Token cost
~1.9k tokens
SKILL.md length
538 words
Files
3 (incl. references)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need…

  • Works in 5 steps: Download raw sequencing reads (FASTQ)… → Find samples, runs, experiments, or… → Retrieve record metadata (XML/JSON/TSV)… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 1 more section
  • Calls python; reaches ebi.ac.uk

What it does

Ena Database is an agent skill from aipoch/medical-research-skills. Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need metadata-driven discovery for genomics pipelines.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `ena-database_audit_result_v1.json` and `references/api_reference.md`).

It sits in Backend & APIs, covering Bioinformatics and REST APIs. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve REST APIs

Example prompts

  • “/ena-database”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Download raw sequencing reads (FASTQ) for a run/experiment/study using ENA accessions (e.g., ERR..., SRR..., PRJ...).
  2. Find samples, runs, experiments, or assemblies by metadata filters (organism, platform, collection date, geography, etc.).
  3. Retrieve record metadata (XML/JSON/TSV) for reproducible reporting and pipeline inputs.
  4. Query taxonomic lineage/rank for organisms to drive filtering or grouping in analyses.
  5. Perform bulk discovery + bulk download workflows (search first, then fetch many files via FTP/Aspera/tools).

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ebi.ac.uk

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ena Database loads about 1.9k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 538 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 538 words, ~1,914 tokens.

Download SKILL.mdSave it as .claude/skills/ena-database/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
ena-database
description
Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need metadata-driven discovery for genomics pipelines.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

Use this skill when you need to:

  1. Download raw sequencing reads (FASTQ) for a run/experiment/study using ENA accessions (e.g., ERR..., SRR..., PRJ...).
  2. Find samples, runs, experiments, or assemblies by metadata filters (organism, platform, collection date, geography, etc.).
  3. Retrieve record metadata (XML/JSON/TSV) for reproducible reporting and pipeline inputs.
  4. Query taxonomic lineage/rank for organisms to drive filtering or grouping in analyses.
  5. Perform bulk discovery + bulk download workflows (search first, then fetch many files via FTP/Aspera/tools).

Key Features

  • Multi-object ENA coverage: studies/projects, samples, experiments, runs, assemblies, sequences, analyses, taxonomy records.
  • Two primary API styles:
    • Portal API for advanced search and metadata export (JSON/TSV/CSV).
    • Browser API for direct record retrieval by accession (XML).
  • Multiple data formats: FASTQ, FASTA, BAM/CRAM, EMBL flat file, plus metadata in XML/JSON/TSV.
  • Bulk transfer options: FTP/Aspera and command-line tooling patterns for large datasets.
  • Cross-references and reference retrieval: ENA xref service and CRAM reference registry endpoints.
  • Operational guidance: rate limiting awareness (HTTP 429) and best practices for robust pipelines.

For detailed endpoint and parameter documentation, see references/api_reference.md.

Dependencies

  • Python >=3.9
  • requests >=2.31.0

Optional (recommended for XML parsing when using the Browser API):

  • lxml >=4.9.0

Example Usage

The following script is a complete, runnable example that:

  1. searches ENA for runs in a study via the Portal API (JSON), then
  2. fetches one run’s record via the Browser API (XML), and
  3. retrieves taxonomy lineage via the Taxonomy REST API.
python
#!/usr/bin/env python3
import sys
import time
import requests

PORTAL_SEARCH = "https://www.ebi.ac.uk/ena/portal/api/search"
BROWSER_XML = "https://www.ebi.ac.uk/ena/browser/api/xml"
TAXONOMY = "https://www.ebi.ac.uk/ena/taxonomy/rest"

SESSION = requests.Session()
SESSION.headers.update({"User-Agent": "ena-database-skill/1.0"})

def get_with_backoff(url, params=None, max_retries=6, timeout=30):
    delay = 1.0
    for attempt in range(max_retries):
        r = SESSION.get(url, params=params, timeout=timeout)
        if r.status_code != 429:
            r.raise_for_status()
            return r
        time.sleep(delay)
        delay *= 2
    r.raise_for_status()

def search_runs_by_study(study_accession, limit=5):
    params = {
        "result": "read_run",
        "query": f"study_accession={study_accession}",
        "format": "json",
        "limit": limit,
        # Ask for a few useful fields; adjust as needed for your pipeline.
        "fields": "run_accession,study_accession,sample_accession,experiment_accession,tax_id,scientific_name,fastq_ftp"
    }
    r = get_with_backoff(PORTAL_SEARCH, params=params)
    return r.json()

def fetch_run_xml(run_accession):
    url = f"{BROWSER_XML}/{run_accession}"
    r = get_with_backoff(url)
    return r.text  # XML string

def fetch_taxonomy_lineage(tax_id):
    url = f"{TAXONOMY}/tax-id/{tax_id}"
    r = get_with_backoff(url)
    return r.json()

def main():
    if len(sys.argv) < 2:
        print("Usage: python ena_example.py <STUDY_ACCESSION>  (e.g., PRJEB1234)", file=sys.stderr)
        sys.exit(2)

    study = sys.argv[1]
    runs = search_runs_by_study(study_accession=study, limit=5)

    if not runs:
        print(f"No runs found for study {study}")
        return

    print(f"Found {len(runs)} runs for study {study}")
    first = runs[0]
    run_acc = first.get("run_accession")
    tax_id = first.get("tax_id")

    print("\nFirst run summary (Portal API JSON):")
    for k in ["run_accession", "sample_accession", "experiment_accession", "scientific_name", "tax_id", "fastq_ftp"]:
        print(f"  {k}: {first.get(k)}")

    if run_acc:
        xml = fetch_run_xml(run_acc)
        print("\nBrowser API XML (first 600 chars):")
        print(xml[:600])

    if tax_id:
        tax = fetch_taxonomy_lineage(tax_id)
        print("\nTaxonomy lineage (ENA Taxonomy REST API):")
        # Response is typically a list with one record
        rec = tax[0] if isinstance(tax, list) and tax else tax
        print(f"  scientificName: {rec.get('scientificName')}")
        print(f"  rank: {rec.get('rank')}")
        print(f"  lineage: {rec.get('lineage')}")

if __name__ == "__main__":
    main()

Run:

bash
python ena_example.py PRJEB1234

Implementation Details

ENA data model (what you query and retrieve)

ENA organizes records into common object types used in pipelines:

  • Study/Project: umbrella entity for a dataset; primary unit for citation.
  • Sample: biological material metadata.
  • Experiment: library prep + instrument metadata.
  • Run: the actual sequencing output files (often FASTQ) for one run.
  • Assembly: genome/transcriptome/metagenome assemblies.
  • Sequence/Record: annotated sequences (e.g., EMBL records).
  • Analysis: computational results derived from sequence data.
  • Taxonomy: lineage and rank information.
Show full SKILL.md (228 more words)Show less
API selection guidance
  • Portal API (/ena/portal/api/search): use for searching and exporting metadata at scale.
    • Typical outputs: json, tsv, csv.
    • Supports complex query expressions (see references/api_reference.md).
  • Browser API (/ena/browser/api/xml/{accession}): use for direct retrieval by accession.
    • Output: XML (parse with an XML parser, not regex).
  • Taxonomy REST API (/ena/taxonomy/rest/...): use for lineage/rank lookups.
  • Cross-reference service: https://www.ebi.ac.uk/ena/xref/rest/ for related records in external databases.
  • CRAM reference registry: https://www.ebi.ac.uk/ena/cram/ for reference sequence retrieval by checksum.
Query parameters and outputs (practical notes)
  • Portal API core parameters (commonly used):
    • result: record type (e.g., sample, read_run, assembly)
    • query: filter expression (e.g., study_accession=PRJEB1234, tax_tree(Escherichia coli))
    • fields: comma-separated fields to return (improves performance vs returning everything)
    • format: json/tsv/csv
    • limit (and pagination where applicable)
  • File retrieval:
    • For raw reads, prefer extracting file locations (e.g., fastq_ftp) from Portal results, then download via FTP/Aspera for scale.
Rate limiting and robustness
  • ENA APIs are rate-limited (commonly documented as 50 requests/second). Exceeding limits returns HTTP 429.
  • Implement:
    • exponential backoff on 429,
    • request consolidation (fetch multiple fields in one query),
    • bulk download mechanisms for large datasets instead of per-accession loops.
  1. Search with Portal API to obtain accessions and file URLs.
  2. Resolve any needed details (optional) via Browser API XML for specific accessions.
  3. Download large files via FTP/Aspera or tooling (rather than API streaming).
  4. Cache taxonomy lookups when processing many records to reduce repeated calls.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in scientific-skills/Evidence Insight/ena-database of aipoch/medical-research-skills.

  • SKILL.md
  • ena-database_audit_result_v1.json
  • references/api_reference.md

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Ena Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ena Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ena Database this skillaipoch/medical-research-skills2k—~1.9kAutomated safety check: PassMIT
Pride Databasemajiayu000/claude-skill-registry6662 repos~8.2kAutomated safety check: PassApache-2.0
Snpeff Variant Annotationjaechang-hits/SciAgent-Skills3701 repos~5.4kAutomated safety check: PassMIT
Encode Ccres Databasegoogle-deepmind/science-skills3.2k1 repos~1.7kAutomated safety check: PassApache-2.0
Bio Ensembl RESTGPTomics/bioSkills1.2k2 repos~3.6kAutomated safety check: PassMIT
Pride FetchClawBio/ClawBio1.2k—~4.2kAutomated safety check: PassMIT

Similar skills

  • Pride Database

    majiayu000/claude-skill-registry

    Search the PRIDE Archive v3 REST API for proteomics datasets: discover projects by keyword + faceted filters (organism, instrument, disease, software), fetch project metadata, list and download…

    666 GitHub starsUsed in 2 repos~8.2k tokens
    Backend & APIsAuto-check passed
  • Snpeff Variant Annotation

    jaechang-hits/SciAgent-Skills

    Annotate and filter VCF variants with SnpEff and SnpSift. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~5.4k tokens
    Research & ScienceAuto-check passed
  • Encode Ccres Database

    google-deepmind/science-skills

    Query the ENCODE Registry of cis-Regulatory Elements (cCREs) via the SCREEN GraphQL API, or make custom queries to the ENCODE Portal REST API for experiments and files (ChIP-seq peaks, etc.).

    3.2k GitHub starsUsed in 1 repo~1.7k tokens
    Backend & APIsAuto-check passed
  • Bio Ensembl REST

    GPTomics/bioSkills

    Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Pride Fetch

    ClawBio/ClawBio

    Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.

    1.2k GitHub stars~4.2k tokensUpdated today
    Research & ScienceAuto-check passed
  • Remap Database

    jaechang-hits/SciAgent-Skills

    Query ReMap 2022 TF ChIP-seq peak database via REST API and BED downloads.

    370 GitHub starsUsed in 2 repos~7.2k tokens
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 21 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 21 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 21 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 21 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed

Questions about Ena Database

What does Ena Database do?

Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need…. Ena Database is an agent skill from aipoch/medical-research-skills. Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need metadata-driven discovery for genomics pipelines.

When should I use Ena Database?

Ena Database fits situations like: tasks that involve Bioinformatics; tasks that involve REST APIs.

How do I install Ena Database in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill ena-database -a claude-code`. Or copy the skill folder (scientific-skills/Evidence Insight/ena-database in aipoch/medical-research-skills) into .claude/skills/ena-database in your project. Claude Code loads it when a task matches its description.

How do I install Ena Database in Codex?

Run `npx skills add aipoch/medical-research-skills --skill ena-database -a codex`. Or copy the skill folder (scientific-skills/Evidence Insight/ena-database in aipoch/medical-research-skills) into .agents/skills/ena-database in your project. Codex loads it when a task matches its description.

Can I use Ena Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill ena-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ena-database, .gemini/skills/ena-database, .github/skills/ena-database and .opencode/skills/ena-database in your project.

What does Ena Database need to run?

Going by SKILL.md and its folder, Ena Database needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Ena Database access the network?

SKILL.md names 1 domain. In commands or code: ebi.ac.uk; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Ena Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ena Database use?

Ena Database is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ena Database use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to Ena Database?

Skills that share tags, products or a category with Ena Database: Pride Database (majiayu000/claude-skill-registry, 666 stars), Snpeff Variant Annotation (jaechang-hits/SciAgent-Skills, 370 stars), Encode Ccres Database (google-deepmind/science-skills, 3.2k stars) and Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ena Database?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,974 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.