Agent skill

Gwas Database

by jaechang-hits in jaechang-hits/SciAgent-Skills

NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS.

Apache-2.0Auto-check passedResearch & Science

Install Gwas Database

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills gwas-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gwas-database .claude/skills/gwas-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gwas-database
GitHub stars
371
Used in
1 other repo
Token cost
~6.6k tokens
SKILL.md length
1,242 words
Files
2 (incl. references)
Skills in repo
169
Repo updated
First seen
Licence
Apache-2.0

At a glance

NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS.

  • Works in 7 steps: Paginate large result sets: Default page… → Use EFO IDs for precise trait queries:… → Always check p-values: The catalog… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 8 more sections
  • Calls pip; reaches ebi.ac.uk and ftp.ebi.ac.uk

What it does

Gwas Database is an agent skill from jaechang-hits/SciAgent-Skills. NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS. Query studies, associations, variants, traits, genes, summary stats. Build PRS candidates, analyze pleiotropy, fetch stats for Manhattan plots. No auth.

Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/api_endpoints.md`).

It sits in Research & Science, covering Bioinformatics and REST APIs. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve REST APIs

Example prompts

  • “/gwas-database”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Paginate large result sets: Default page size is 20; set size=500 and loop over pages for complete data retrieval. Check…
  2. Use EFO IDs for precise trait queries: Free-text search may return related but different traits. Look up the exact EFO ID first, then…
  3. Always check p-values: The catalog contains associations at various significance levels. Filter to p < 5e-8 for genome-wide significant…
  4. Be ancestry-aware: Effect sizes and allele frequencies vary across populations. Check the initialSampleSize and replicationSampleSize…
  5. Add rate-limiting delays: Although no official limit exists, time.sleep(0.2) between requests prevents server overload and avoids…
  6. Cache frequently accessed data: Study and trait metadata rarely change. Cache results locally when running batch analyses to reduce…
  7. Anti-pattern -- Don't rely solely on reported genes: Author-reported genes may not be the causal gene. Cross-reference with functional…

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ebi.ac.uk
    • ftp.ebi.ac.uk
    • pgscatalog.org

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gwas Database loads about 6.6k tokens when it runs, and up to ~8.7k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 1,242 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~6.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 1,242 words, ~6,632 tokens.

Download SKILL.mdSave it as .claude/skills/gwas-database/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
gwas-database
description
NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS. Query studies, associations, variants, traits, genes, summary stats. Build PRS candidates, analyze pleiotropy, fetch stats for Manhattan plots. No auth.
license
Apache-2.0

GWAS Catalog Database — SNP-Trait Association Queries

Overview

The NHGRI-EBI GWAS Catalog is a curated collection of published genome-wide association studies, mapping SNP-trait associations with genomic context. The REST API provides programmatic access to studies, associations, variants, traits, genes, and summary statistics. All responses are HAL+JSON with embedded _links for pagination.

When to Use

  • Finding genetic variants associated with a disease or trait (e.g., "which SNPs are linked to type 2 diabetes?")
  • Retrieving genome-wide significant associations for a specific variant (rs ID)
  • Exploring the genetic architecture of complex traits (number of loci, effect sizes)
  • Checking variant pleiotropy (how many traits a single SNP affects)
  • Downloading summary statistics for meta-analysis or polygenic risk score construction
  • Identifying published GWAS studies by disease, gene, or PubMed ID
  • Cross-referencing EFO trait ontology terms with GWAS evidence
  • Building candidate gene lists from GWAS association regions
  • Use omics-plotting SKILL to render Manhattan, QQ, and forest plots from the summary statistics / association results
  • For drug target validation from GWAS hits, use opentargets-database instead
  • For variant functional annotation (consequence prediction, regulatory impact), use Ensembl VEP via gget

Prerequisites

bash
pip install requests matplotlib numpy

API access:

  • No authentication required -- fully open access
  • Rate limits: no official limit, but add time.sleep(0.2) between requests to be courteous
  • Base URL: https://www.ebi.ac.uk/gwas/rest/api
  • Response format: HAL+JSON with _embedded data and _links for pagination
  • Pagination: default 20 results per page; max 500 via size parameter

Quick Start

python
import requests
import time

BASE = "https://www.ebi.ac.uk/gwas/rest/api"

def gwas_get(endpoint, params=None):
    """GWAS Catalog REST API helper with rate limiting and pagination support."""
    url = f"{BASE}/{endpoint}"
    resp = requests.get(url, params=params or {})
    resp.raise_for_status()
    time.sleep(0.2)
    return resp.json()

# Find studies for a trait keyword. Study records have no top-level `title`
# — the publication title lives at `publicationInfo.title`; the trait label
# lives at `diseaseTrait.trait`.
data = gwas_get("studies/search/findByDiseaseTrait", {"diseaseTrait": "diabetes"})
studies = data["_embedded"]["studies"]
print(f"Found {len(studies)} studies for 'diabetes'")
for s in studies[:3]:
    title = (s.get("publicationInfo") or {}).get("title", "N/A")
    trait = (s.get("diseaseTrait") or {}).get("trait", "N/A")
    print(f"  {s['accessionId']} | {trait[:40]:<40} | {title[:60]}")

Core API

Search GWAS studies by disease trait keyword or PubMed ID.

python
# Search studies by disease trait
data = gwas_get("studies/search/findByDiseaseTrait", {"diseaseTrait": "breast cancer"})
studies = data["_embedded"]["studies"]
for s in studies[:5]:
    pi = s.get("publicationInfo") or {}
    print(f"  {s['accessionId']} | PMID:{pi.get('pubmedId','N/A')} | {pi.get('title','')[:60]}")

time.sleep(0.2)

# Search by PubMed ID. NOTE: the older `findByPubmedId` 404s on /studies/;
# the working endpoint is `findByPublicationIdPubmedId`.
data = gwas_get("studies/search/findByPublicationIdPubmedId", {"pubmedId": "25673413"})
studies = data["_embedded"]["studies"]
print(f"Studies from PMID 25673413: {len(studies)}")
for s in studies:
    trait = (s.get("diseaseTrait") or {}).get("trait", "N/A")
    print(f"  {s['accessionId']}: {trait}")
Module 2: Association Queries

Retrieve SNP-trait associations filtered by trait (EFO term), variant, or p-value.

python
# Associations by EFO trait. The old path `efoTraits/{shortForm}/associations`
# also works *if* you have the current shortForm — but trait shortForms have
# been re-mapped to MONDO (e.g. EFO_0000249 → MONDO_0004975). The most reliable
# path is `associations/search/findByEfoTrait?efoTrait=<canonical trait name>`.
data = gwas_get("associations/search/findByEfoTrait",
                {"efoTrait": "type 2 diabetes mellitus", "size": 50})
assocs = data["_embedded"]["associations"]
print(f"Associations for 'type 2 diabetes mellitus': {len(assocs)}")

for a in assocs[:5]:
    pval = a.get("pvalue", None)
    genes = []
    for locus in a.get("loci", []) or []:
        for gene in locus.get("authorReportedGenes", []) or []:
            genes.append(gene.get("geneName", ""))
    loci = a.get("loci") or [{}]
    snps = [r.get("snps", [{}])[0].get("rsId", "N/A")
            for r in (loci[0].get("strongestRiskAlleles") or [])]
    print(f"  rs={snps} | p={pval} | genes={genes}")
python
# Associations for a specific variant. NOTE: association records do not embed
# `efoTraits` inline — they expose them via the `_links.efoTraits.href`
# HAL link. Follow the link (cached if needed) to resolve trait names.
data = gwas_get("singleNucleotidePolymorphisms/rs7903146/associations", {"size": 5})
assocs = data["_embedded"]["associations"]
print(f"Associations for rs7903146 (first page): {len(assocs)}")

def association_traits(assoc):
    """Resolve efoTraits via the HAL link on an association record."""
    href = (assoc.get("_links") or {}).get("efoTraits", {}).get("href")
    if not href:
        return []
    r = requests.get(href, timeout=15)
    if not r.ok:
        return []
    return [t.get("trait") for t in r.json().get("_embedded", {}).get("efoTraits", [])]

for a in assocs[:5]:
    traits = association_traits(a)
    print(f"  p={a.get('pvalue')} | OR={a.get('orPerCopyNum', 'N/A')} | traits={traits}")
    time.sleep(0.1)
Module 3: Variant Lookup

Query variant details by rsID, chromosomal region, or cytogenetic band.

python
# Lookup single variant
data = gwas_get("singleNucleotidePolymorphisms/rs7903146")
loc = data.get("locations", [{}])[0]
print(f"rs7903146: chr{loc.get('chromosomeName', '?')}:{loc.get('chromosomePosition', '?')}")
print(f"  Functional class: {data.get('functionalClass', 'N/A')}")
print(f"  Merged into: {data.get('merged', 'N/A')}")

time.sleep(0.2)

# Search variants by chromosomal region
data = gwas_get("singleNucleotidePolymorphisms/search/findByChromBpLocationRange",
                {"chrom": "10", "bpStart": "114750000", "bpEnd": "114800000", "size": 50})
snps = data["_embedded"]["singleNucleotidePolymorphisms"]
print(f"Variants in chr10:114750000-114800000: {len(snps)}")
for v in snps[:5]:
    print(f"  {v['rsId']}: {v.get('functionalClass', 'N/A')}")
python
# Search variants by gene name (the cytogenetic-band endpoint
# `findByCytogeneticBand` was removed — use gene or chromosome-range instead).
data = gwas_get("singleNucleotidePolymorphisms/search/findByGene",
                {"geneName": "TCF7L2", "size": 5})
snps = data["_embedded"]["singleNucleotidePolymorphisms"]
print(f"Variants in TCF7L2: {len(snps)}")

Browse and search EFO-mapped traits in the GWAS Catalog.

python
# Search traits by exact name (the older `findByDescription` endpoint was
# removed — search/efoTrait now expects the canonical trait label).
data = gwas_get("efoTraits/search/findByEfoTrait", {"trait": "Alzheimer disease"})
traits = data["_embedded"]["efoTraits"]
print(f"Traits matching 'Alzheimer disease': {len(traits)}")
for t in traits[:5]:
    print(f"  {t['shortForm']}: {t['trait']}  (uri={t['uri']})")

time.sleep(0.2)

# Get specific trait by shortForm. NOTE: many legacy EFO IDs have been
# re-mapped to MONDO (e.g. old `EFO_0000249` for Alzheimer is now
# `MONDO_0004975` — `efoTraits/EFO_0000249` returns 404). Resolve via search
# above first, then use the current shortForm:
short_form = traits[0]["shortForm"]   # e.g. 'MONDO_0004975'
data = gwas_get(f"efoTraits/{short_form}")
print(f"Trait: {data['trait']}")
print(f"  URI       : {data['uri']}")
print(f"  shortForm : {data['shortForm']}")
Module 5: Summary Statistics

Access study-level summary statistics for downstream analysis (meta-analysis, PRS).

python
# List studies with available summary statistics
data = gwas_get("studies/search/findByFullPvalueSet", {"fullPvalueSet": True, "size": 10})
studies = data["_embedded"]["studies"]
print(f"Studies with summary stats (first page): {len(studies)}")
for s in studies[:5]:
    trait = (s.get("diseaseTrait") or {}).get("trait", "N/A")
    print(f"  {s['accessionId']}: {trait[:50]}")

time.sleep(0.2)

# Summary statistics metadata is NOT exposed via the REST API
# (`studies/{acc}/summaryStatistics` returns 404). Use the GWAS Catalog FTP
# directly — paths are predictable by study accession:
ftp_base = "http://ftp.ebi.ac.uk/pub/databases/gwas/summary_statistics"
acc = studies[0]["accessionId"]
ftp_url = f"{ftp_base}/{acc[:-3]}001-{acc[:-3]}999/{acc}/"
print(f"Summary stats FTP directory: {ftp_url}")
python
# Download summary statistics file (FTP)
# Summary statistics are hosted on the GWAS Catalog FTP, not the REST API
import urllib.request

study_id = "GCST006867"  # Example study
ftp_base = "http://ftp.ebi.ac.uk/pub/databases/gwas/summary_statistics"
# Actual paths vary by study; check the study page for the download link
# Example: ftp_base/{study_id}/{study_id}.tsv.gz
url = f"{ftp_base}/{study_id}"
print(f"Summary stats FTP directory: {url}")
# Use requests.get() or urllib to download the .tsv.gz file

Find GWAS associations by gene name or retrieve publication metadata.

python
# Search associations by gene
data = gwas_get("singleNucleotidePolymorphisms/search/findByGene",
                {"geneName": "BRCA1", "size": 50})
snps = data["_embedded"]["singleNucleotidePolymorphisms"]
print(f"Variants near BRCA1: {len(snps)}")
for v in snps[:5]:
    locs = v.get("locations", [{}])
    pos = locs[0].get("chromosomePosition", "?") if locs else "?"
    print(f"  {v['rsId']}: chr{locs[0].get('chromosomeName', '?')}:{pos}")

time.sleep(0.2)

# Get study publication details
data = gwas_get("studies/GCST000392")
pub = data.get("publicationInfo", {})
print(f"Study: {data['accessionId']}")
print(f"  Author: {pub.get('author', {}).get('fullname', 'N/A')}")
print(f"  Journal: {pub.get('publication', 'N/A')}")
print(f"  PMID: {pub.get('pubmedId', 'N/A')}")
print(f"  Date: {pub.get('publicationDate', 'N/A')}")

Key Concepts

Data Entities and Relationships

The GWAS Catalog organizes data as interconnected entities:

EntityDescriptionKey IdentifierExample
StudyA published GWAS experimentGCST accession (e.g., GCST000392)Wellcome Trust Case Control Consortium study
AssociationA SNP-trait association with p-value and effect sizeInternal IDrs7903146 associated with T2D at p=1e-40
Variant (SNP)A single nucleotide polymorphismrs number (e.g., rs7903146)TCF7L2 variant
TraitA disease/phenotype mapped to EFO ontologyEFO ID (e.g., EFO_0001360)Type 2 diabetes mellitus
GeneA gene near or harboring GWAS variantsGene symbol (e.g., TCF7L2)Transcription factor 7-like 2

Relationships: Study --(reports)--> Association --(involves)--> Variant + Trait. Variants map to genomic positions and nearby genes.

HAL+JSON Response Structure

All API responses follow HAL (Hypertext Application Language) format:

Top-level SectionContentsAccess Pattern
_embeddedPrimary data objects (studies, associations, etc.)response["_embedded"]["studies"]
_linksNavigation links (self, next, prev, first, last)response["_links"]["next"]["href"]
pagePagination metadata (size, totalElements, totalPages, number)response["page"]["totalElements"]
Genome-wide Significance

The standard genome-wide significance threshold is p <= 5 x 10^-8, correcting for approximately 1 million independent tests across the human genome. Associations below this threshold are considered suggestive. The GWAS Catalog includes associations at various significance levels -- always check p-values when filtering results.

Key Identifiers
  • GCST IDs: GWAS Catalog study accessions (e.g., GCST000392)
  • rs numbers: dbSNP reference SNP identifiers (e.g., rs7903146)
  • EFO terms: Experimental Factor Ontology for trait standardization (e.g., EFO_0001360 for type 2 diabetes)
  • Cytogenetic bands: Chromosomal location notation (e.g., 10q25.2)

Common Workflows

Workflow 1: Disease Genetic Architecture

Goal: Map the genetic landscape of a disease by collecting all genome-wide significant loci.

python
import requests, time

BASE = "https://www.ebi.ac.uk/gwas/rest/api"

def gwas_get(endpoint, params=None):
    url = f"{BASE}/{endpoint}"
    resp = requests.get(url, params=params or {})
    resp.raise_for_status()
    time.sleep(0.2)
    return resp.json()

# Step 1: Resolve trait → current shortForm via findByEfoTrait
# (the older `findByDescription` endpoint was removed).
traits = gwas_get("efoTraits/search/findByEfoTrait", {"trait": "schizophrenia"})
efo_id = traits["_embedded"]["efoTraits"][0]["shortForm"]
print(f"Using EFO: {efo_id}")

# Step 2: Get all associations for this trait — nested path
# `efoTraits/{shortForm}/associations` works once you have the canonical shortForm.
all_assocs = []
page = 0
while True:
    data = gwas_get(f"efoTraits/{efo_id}/associations",
                    {"size": 500, "page": page})
    assocs = data["_embedded"]["associations"]
    all_assocs.extend(assocs)
    if page >= data["page"]["totalPages"] - 1:
        break
    page += 1

# Step 3: Filter genome-wide significant
significant = [a for a in all_assocs if a.get("pvalue") and a["pvalue"] < 5e-8]
print(f"Total associations: {len(all_assocs)}, genome-wide significant: {len(significant)}")

# Step 4: Extract variant and effect details
for a in significant[:10]:
    risk_alleles = a.get("loci", [{}])[0].get("strongestRiskAlleles", [])
    snp = risk_alleles[0].get("snps", [{}])[0].get("rsId", "N/A") if risk_alleles else "N/A"
    or_val = a.get("orPerCopyNum", "N/A")
    beta = a.get("betaNum", "N/A")
    print(f"  {snp} | p={a['pvalue']:.2e} | OR={or_val} | beta={beta}")
Workflow 2: Variant Pleiotropy Analysis

Goal: Determine how many distinct traits a single variant is associated with.

python
import requests, time

BASE = "https://www.ebi.ac.uk/gwas/rest/api"

def gwas_get(endpoint, params=None):
    url = f"{BASE}/{endpoint}"
    resp = requests.get(url, params=params or {})
    resp.raise_for_status()
    time.sleep(0.2)
    return resp.json()

rs_id = "rs7903146"  # Well-known pleiotropic variant

# Get all associations for this variant
data = gwas_get(f"singleNucleotidePolymorphisms/{rs_id}/associations", {"size": 500})
assocs = data["_embedded"]["associations"]

# Association records don't embed efoTraits inline — follow the HAL
# `_links.efoTraits.href` link per association. Cache per-href to avoid
# duplicate fetches.
trait_set = {}
href_cache = {}
def fetch_traits(href):
    if href in href_cache:
        return href_cache[href]
    r = requests.get(href, timeout=15)
    href_cache[href] = (r.json().get("_embedded", {}).get("efoTraits", [])) if r.ok else []
    return href_cache[href]

for a in assocs:
    href = (a.get("_links") or {}).get("efoTraits", {}).get("href")
    if not href:
        continue
    for t in fetch_traits(href):
        tid = t.get("shortForm", "unknown")
        if tid not in trait_set:
            trait_set[tid] = {"trait": t.get("trait", "N/A"),
                              "best_pval": a.get("pvalue", 1), "count": 0}
        trait_set[tid]["count"] += 1
        if a.get("pvalue") and a["pvalue"] < trait_set[tid]["best_pval"]:
            trait_set[tid]["best_pval"] = a["pvalue"]
    time.sleep(0.1)

print(f"{rs_id} is associated with {len(trait_set)} distinct traits:")
for tid, info in sorted(trait_set.items(), key=lambda x: x[1]["best_pval"]):
    print(f"  {tid}: {info['trait']} (best p={info['best_pval']:.2e}, n={info['count']})")
Workflow 3: Summary Statistics Manhattan Plot

Goal: Download summary statistics for a study and create a Manhattan plot.

python
import numpy as np
import pandas as pd

# Summary statistics as a CHR/BP/P table (simulated here for demonstration).
# In practice: download from GWAS Catalog FTP, e.g.
# url = "http://ftp.ebi.ac.uk/pub/databases/gwas/summary_statistics/GCSTXXXXXX/..."
# import gzip, requests; data = gzip.decompress(requests.get(url).content)
np.random.seed(42)
n_snps = 5000
pvalues = np.random.uniform(0, 1, size=n_snps)
pvalues[:20] = 10 ** np.random.uniform(-15, -8, size=20)   # add some "hits"
sumstats = pd.DataFrame({
    "CHR": np.random.choice(range(1, 23), size=n_snps),
    "BP": np.random.randint(1, 250_000_000, size=n_snps),
    "P": pvalues,
})
sumstats.to_csv("gwas.csv", index=False)
print(f"Summary stats: {len(sumstats)} SNPs, {(sumstats['P'] < 5e-8).sum()} genome-wide sig -> gwas.csv")
# Render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Manhattan" recipe -> figures/manhattan_plot.png

Key Parameters

ParameterFunction/EndpointDefaultRange / OptionsEffect
sizeAll paginated endpoints201-500Results per page
pageAll paginated endpoints00-totalPages-1Page number (0-indexed)
diseaseTraitstudies/search/findByDiseaseTrait--Any stringTrait keyword search
pubmedIdstudies/search/findByPublicationIdPubmedId--Valid PMIDStudy lookup by publication (findByPubmedId 404s on /studies/)
geneNamesnps/search/findByGene--Gene symbolVariants near a gene
chrom, bpStart, bpEndsnps/search/findByChromBpLocationRange--chr:start-endRegional variant query
fullPvalueSetstudies/search/findByFullPvalueSet--True/FalseFilter studies with summary stats
traitefoTraits/search/findByEfoTrait--Canonical trait nameTrait lookup by exact name (findByDescription was removed)
efoTraitassociations/search/findByEfoTrait--Canonical trait nameAll associations for a trait
Show full SKILL.md (546 more words)Show less

Best Practices

  1. Paginate large result sets: Default page size is 20; set size=500 and loop over pages for complete data retrieval. Check page.totalElements to know the full count before iterating.

  2. Use EFO IDs for precise trait queries: Free-text search may return related but different traits. Look up the exact EFO ID first, then query associations by EFO ID for precision.

  3. Always check p-values: The catalog contains associations at various significance levels. Filter to p < 5e-8 for genome-wide significant results unless you specifically need suggestive associations.

  4. Be ancestry-aware: Effect sizes and allele frequencies vary across populations. Check the initialSampleSize and replicationSampleSize fields to understand the ancestry composition of each study.

  5. Add rate-limiting delays: Although no official limit exists, time.sleep(0.2) between requests prevents server overload and avoids temporary blocks.

  6. Cache frequently accessed data: Study and trait metadata rarely change. Cache results locally when running batch analyses to reduce redundant API calls.

  7. Anti-pattern -- Don't rely solely on reported genes: Author-reported genes may not be the causal gene. Cross-reference with functional annotation tools and eQTL databases for biological interpretation.

Common Recipes

Recipe 1: Cross-Reference with PGS Catalog

Identify polygenic scores available for a GWAS trait.

python
import requests, time

# Step 1: Get EFO/MONDO shortForm from GWAS Catalog
# (the old `findByDescription` endpoint was removed; use `findByEfoTrait`)
BASE = "https://www.ebi.ac.uk/gwas/rest/api"
traits = requests.get(f"{BASE}/efoTraits/search/findByEfoTrait",
                      params={"trait": "coronary artery disease"}).json()
efo_id = traits["_embedded"]["efoTraits"][0]["shortForm"]
time.sleep(0.2)

# Step 2: Query PGS Catalog for scores using same EFO
pgs_url = f"https://www.pgscatalog.org/rest/score/search?trait_id={efo_id}"
pgs_data = requests.get(pgs_url).json()
print(f"PGS scores for {efo_id}: {pgs_data.get('count', 0)}")
for score in pgs_data.get("results", [])[:5]:
    print(f"  {score['id']}: {score['name']} (variants: {score.get('variants_number', '?')})")
Recipe 2: Regional Association Summary

Collect all GWAS associations in a genomic region.

python
import requests, time

BASE = "https://www.ebi.ac.uk/gwas/rest/api"

# Query region: chr9:21,900,000-22,200,000 (CDKN2A/2B locus)
data = requests.get(f"{BASE}/singleNucleotidePolymorphisms/search/findByChromBpLocationRange",
                    params={"chrom": "9", "bpStart": "21900000",
                            "bpEnd": "22200000", "size": 200}).json()
time.sleep(0.2)
snps = data["_embedded"]["singleNucleotidePolymorphisms"]
print(f"Variants in CDKN2A/2B locus: {len(snps)}")

# Get associations for each variant
region_assocs = []
for v in snps[:10]:  # Limit for demo
    rs = v["rsId"]
    try:
        a_data = requests.get(f"{BASE}/singleNucleotidePolymorphisms/{rs}/associations",
                              params={"size": 50}).json()
        time.sleep(0.2)
        for a in a_data["_embedded"]["associations"]:
            for t in a.get("efoTraits", []):
                region_assocs.append({"rsId": rs, "trait": t["trait"],
                                      "pvalue": a.get("pvalue")})
    except Exception:
        continue

print(f"Total associations in region: {len(region_assocs)}")
Recipe 3: Effect Size Forest Plot

Visualize effect sizes across studies for a single variant.

python
import re
import pandas as pd
import requests, time

BASE = "https://www.ebi.ac.uk/gwas/rest/api"
data = requests.get(f"{BASE}/singleNucleotidePolymorphisms/rs1801282/associations",
                    params={"size": 100}).json()
time.sleep(0.2)
assocs = data["_embedded"]["associations"]

def parse_ci(text):
    """Pull [low-high] confidence bounds out of the GWAS Catalog 'range' string."""
    nums = re.findall(r"[0-9]*\.?[0-9]+", text or "")
    return (float(nums[0]), float(nums[1])) if len(nums) >= 2 else (float("nan"), float("nan"))

# Build a forest-plot table: label, estimate (OR), ci_low, ci_high
rows = []
for a in assocs:
    or_val = a.get("orPerCopyNum")
    trait = a.get("efoTraits", [{}])[0].get("trait", "N/A") if a.get("efoTraits") else "N/A"
    if or_val and or_val > 0:
        lo, hi = parse_ci(a.get("range", ""))
        rows.append({"label": trait[:30], "estimate": or_val, "ci_low": lo, "ci_high": hi})

effects = pd.DataFrame(rows)
effects.to_csv("effects.csv", index=False)
print(f"Effect sizes: {len(effects)} entries -> effects.csv")
# Render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Forest" recipe (null line at OR=1) -> figures/forest_plot.png

Troubleshooting

ProblemCauseSolution
404 Not FoundInvalid GCST, rs ID, or EFO IDVerify identifier exists via search endpoint first
Empty _embeddedNo results for queryBroaden search terms or check spelling; try partial match
ConnectionError / timeoutEBI server temporarily downRetry with exponential backoff; check https://www.ebi.ac.uk/gwas/status
Truncated resultsDefault pagination (20 per page)Set size=500 and iterate over page parameter
Missing OR/beta valuesNot all associations report effect sizesCheck both orPerCopyNum and betaNum fields; some studies report only p-values
Stale dataCatalog updated quarterlyCheck lastUpdateDate on studies; re-query if needed
KeyError: '_embedded'Endpoint returns single object, not listDirect lookups (by ID) return a single object; _embedded only appears in search/list results
Summary stats unavailableNot all studies deposit summary statsFilter with findByFullPvalueSet=True to find studies with available data

Bundled Resources

references/api_endpoints.md

Covers: Complete endpoint catalog for all 6 REST API groups (studies, associations, variants, traits, genes, summary statistics) with query parameters, plus response field tables for the main entity types (association fields, study fields, variant fields), HTTP error codes, and pagination patterns.

Relocated inline: Core query patterns and the most common endpoints are demonstrated in Core API modules with full code examples. HAL+JSON structure and key identifier formats are in Key Concepts.

Omitted from original api_reference.md (794 lines): Advanced query composition patterns (multi-parameter chaining beyond what Core API shows), child/parent trait traversal details, detailed changelog/versioning notes. These are specialized and covered by EBI documentation.

  • gget-genomic-databases -- gene lookups, variant annotation, BLAST (upstream ID resolution)
  • ensembl-database (planned) -- variant effect prediction, regulatory annotations
  • opentargets-database (planned) -- drug target validation from GWAS hits
  • clinvar-database (planned) -- clinical variant interpretation
  • bioservices-multi-database -- cross-database queries integrating GWAS with pathway data

References

© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/genomics-bioinformatics/databases/gwas-database of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/api_endpoints.md

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Gwas Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gwas Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gwas Database this skilljaechang-hits/SciAgent-Skills3711 repos~6.6kAutomated safety check: PassApache-2.0
Bio Ensembl RESTGPTomics/bioSkills1.2k2 repos~3.6kAutomated safety check: PassMIT
Pride FetchClawBio/ClawBio1.2k—~4.2kAutomated safety check: PassMIT
Ensembl Databaseaipoch/medical-research-skills2k—~1.5kAutomated safety check: PassMIT
UniProt Database Accessdavila7/claude-code-templates32k14 repos~1.7kAutomated safety check: PassMIT
Singlecell Portalaipoch/medical-research-skills2k—~1.2kAutomated safety check: PassMIT

Similar skills

  • Bio Ensembl REST

    GPTomics/bioSkills

    Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Pride Fetch

    ClawBio/ClawBio

    Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.

    1.2k GitHub stars~4.2k tokensUpdated today
    Research & ScienceAuto-check passed
  • Ensembl Database

    aipoch/medical-research-skills

    Access Ensembl REST API for vertebrate genomic data; use when you need gene/ID lookups, sequence retrieval, variant effect prediction (VEP), or homology/assembly coordinate mapping.

    2k GitHub stars~1.5k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    32k GitHub starsUsed in 14 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • Singlecell Portal

    aipoch/medical-research-skills

    Programmatically query public single-cell study metadata from the Broad Institute Single Cell Portal REST API when you need to search and filter datasets by organism, tissue, disease, or cell type…

    2k GitHub stars~1.2k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • Ena Database

    aipoch/medical-research-skills

    Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need…

    2k GitHub stars~1.9k tokensUpdated 22 days ago
    Backend & APIsAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    371 GitHub stars~4k tokensUpdated 10 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    371 GitHub stars~3.2k tokensUpdated 10 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    371 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    371 GitHub stars~6.9k tokensUpdated 10 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub stars~2.3k tokensUpdated 10 days ago
    Auto-check passed

Questions about Gwas Database

What does Gwas Database do?

NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS. Gwas Database is an agent skill from jaechang-hits/SciAgent-Skills. NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS.

When should I use Gwas Database?

Gwas Database fits situations like: tasks that involve Bioinformatics; tasks that involve REST APIs.

How do I install Gwas Database in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gwas-database in jaechang-hits/SciAgent-Skills) into .claude/skills/gwas-database in your project. Claude Code loads it when a task matches its description.

How do I install Gwas Database in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gwas-database in jaechang-hits/SciAgent-Skills) into .agents/skills/gwas-database in your project. Codex loads it when a task matches its description.

Can I use Gwas Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gwas-database, .gemini/skills/gwas-database, .github/skills/gwas-database and .opencode/skills/gwas-database in your project.

What does Gwas Database need to run?

Going by SKILL.md and its folder, Gwas Database needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Gwas Database access the network?

SKILL.md names 4 domains. In commands or code: ebi.ac.uk, ftp.ebi.ac.uk and pgscatalog.org; the agent is likely to contact these when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Gwas Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gwas Database use?

Gwas Database is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gwas Database use?

About 6.6k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Gwas Database?

Skills that share tags, products or a category with Gwas Database: Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars), Pride Fetch (ClawBio/ClawBio, 1.2k stars), Ensembl Database (aipoch/medical-research-skills, 2k stars) and UniProt Database Access (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gwas Database?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 371 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.