Bio Ensembl REST
GPTomics/bioSkills
Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…
NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gwas-database --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gwas-database .claude/skills/gwas-database && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gwas-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gwas-database into .claude/skills/gwas-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gwas-database", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gwas-databaseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gwas-database --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gwas-database .agents/skills/gwas-database && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gwas-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gwas-database into .agents/skills/gwas-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gwas-database", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gwas-database --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gwas-database .cursor/skills/gwas-database && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gwas-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gwas-database into .cursor/skills/gwas-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gwas-database", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/databases/gwas-database--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gwas-database --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gwas-database .gemini/skills/gwas-database && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gwas-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gwas-database into .gemini/skills/gwas-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gwas-database", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills gwas-databaseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gwas-database .github/skills/gwas-database && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gwas-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gwas-database into .github/skills/gwas-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gwas-database", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gwas-database --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gwas-database .opencode/skills/gwas-database && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gwas-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gwas-database into .opencode/skills/gwas-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gwas-database", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gwas-databaseNHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS.
Gwas Database is an agent skill from jaechang-hits/SciAgent-Skills. NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS. Query studies, associations, variants, traits, genes, summary stats. Build PRS candidates, analyze pleiotropy, fetch stats for Manhattan plots. No auth.
Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/api_endpoints.md`).
It sits in Research & Science, covering Bioinformatics and REST APIs. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
ebi.ac.ukftp.ebi.ac.ukpgscatalog.orgAlso links to:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gwas Database loads about 6.6k tokens when it runs, and up to ~8.7k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 1,242 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 1,242 words, ~6,632 tokens.
.claude/skills/gwas-database/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.The NHGRI-EBI GWAS Catalog is a curated collection of published genome-wide association studies, mapping SNP-trait associations with genomic context. The REST API provides programmatic access to studies, associations, variants, traits, genes, and summary statistics. All responses are HAL+JSON with embedded _links for pagination.
opentargets-database insteadggetpip install requests matplotlib numpyAPI access:
time.sleep(0.2) between requests to be courteoushttps://www.ebi.ac.uk/gwas/rest/api_embedded data and _links for paginationsize parameterimport requests
import time
BASE = "https://www.ebi.ac.uk/gwas/rest/api"
def gwas_get(endpoint, params=None):
"""GWAS Catalog REST API helper with rate limiting and pagination support."""
url = f"{BASE}/{endpoint}"
resp = requests.get(url, params=params or {})
resp.raise_for_status()
time.sleep(0.2)
return resp.json()
# Find studies for a trait keyword. Study records have no top-level `title`
# — the publication title lives at `publicationInfo.title`; the trait label
# lives at `diseaseTrait.trait`.
data = gwas_get("studies/search/findByDiseaseTrait", {"diseaseTrait": "diabetes"})
studies = data["_embedded"]["studies"]
print(f"Found {len(studies)} studies for 'diabetes'")
for s in studies[:3]:
title = (s.get("publicationInfo") or {}).get("title", "N/A")
trait = (s.get("diseaseTrait") or {}).get("trait", "N/A")
print(f" {s['accessionId']} | {trait[:40]:<40} | {title[:60]}")Search GWAS studies by disease trait keyword or PubMed ID.
# Search studies by disease trait
data = gwas_get("studies/search/findByDiseaseTrait", {"diseaseTrait": "breast cancer"})
studies = data["_embedded"]["studies"]
for s in studies[:5]:
pi = s.get("publicationInfo") or {}
print(f" {s['accessionId']} | PMID:{pi.get('pubmedId','N/A')} | {pi.get('title','')[:60]}")
time.sleep(0.2)
# Search by PubMed ID. NOTE: the older `findByPubmedId` 404s on /studies/;
# the working endpoint is `findByPublicationIdPubmedId`.
data = gwas_get("studies/search/findByPublicationIdPubmedId", {"pubmedId": "25673413"})
studies = data["_embedded"]["studies"]
print(f"Studies from PMID 25673413: {len(studies)}")
for s in studies:
trait = (s.get("diseaseTrait") or {}).get("trait", "N/A")
print(f" {s['accessionId']}: {trait}")Retrieve SNP-trait associations filtered by trait (EFO term), variant, or p-value.
# Associations by EFO trait. The old path `efoTraits/{shortForm}/associations`
# also works *if* you have the current shortForm — but trait shortForms have
# been re-mapped to MONDO (e.g. EFO_0000249 → MONDO_0004975). The most reliable
# path is `associations/search/findByEfoTrait?efoTrait=<canonical trait name>`.
data = gwas_get("associations/search/findByEfoTrait",
{"efoTrait": "type 2 diabetes mellitus", "size": 50})
assocs = data["_embedded"]["associations"]
print(f"Associations for 'type 2 diabetes mellitus': {len(assocs)}")
for a in assocs[:5]:
pval = a.get("pvalue", None)
genes = []
for locus in a.get("loci", []) or []:
for gene in locus.get("authorReportedGenes", []) or []:
genes.append(gene.get("geneName", ""))
loci = a.get("loci") or [{}]
snps = [r.get("snps", [{}])[0].get("rsId", "N/A")
for r in (loci[0].get("strongestRiskAlleles") or [])]
print(f" rs={snps} | p={pval} | genes={genes}")# Associations for a specific variant. NOTE: association records do not embed
# `efoTraits` inline — they expose them via the `_links.efoTraits.href`
# HAL link. Follow the link (cached if needed) to resolve trait names.
data = gwas_get("singleNucleotidePolymorphisms/rs7903146/associations", {"size": 5})
assocs = data["_embedded"]["associations"]
print(f"Associations for rs7903146 (first page): {len(assocs)}")
def association_traits(assoc):
"""Resolve efoTraits via the HAL link on an association record."""
href = (assoc.get("_links") or {}).get("efoTraits", {}).get("href")
if not href:
return []
r = requests.get(href, timeout=15)
if not r.ok:
return []
return [t.get("trait") for t in r.json().get("_embedded", {}).get("efoTraits", [])]
for a in assocs[:5]:
traits = association_traits(a)
print(f" p={a.get('pvalue')} | OR={a.get('orPerCopyNum', 'N/A')} | traits={traits}")
time.sleep(0.1)Query variant details by rsID, chromosomal region, or cytogenetic band.
# Lookup single variant
data = gwas_get("singleNucleotidePolymorphisms/rs7903146")
loc = data.get("locations", [{}])[0]
print(f"rs7903146: chr{loc.get('chromosomeName', '?')}:{loc.get('chromosomePosition', '?')}")
print(f" Functional class: {data.get('functionalClass', 'N/A')}")
print(f" Merged into: {data.get('merged', 'N/A')}")
time.sleep(0.2)
# Search variants by chromosomal region
data = gwas_get("singleNucleotidePolymorphisms/search/findByChromBpLocationRange",
{"chrom": "10", "bpStart": "114750000", "bpEnd": "114800000", "size": 50})
snps = data["_embedded"]["singleNucleotidePolymorphisms"]
print(f"Variants in chr10:114750000-114800000: {len(snps)}")
for v in snps[:5]:
print(f" {v['rsId']}: {v.get('functionalClass', 'N/A')}")# Search variants by gene name (the cytogenetic-band endpoint
# `findByCytogeneticBand` was removed — use gene or chromosome-range instead).
data = gwas_get("singleNucleotidePolymorphisms/search/findByGene",
{"geneName": "TCF7L2", "size": 5})
snps = data["_embedded"]["singleNucleotidePolymorphisms"]
print(f"Variants in TCF7L2: {len(snps)}")Browse and search EFO-mapped traits in the GWAS Catalog.
# Search traits by exact name (the older `findByDescription` endpoint was
# removed — search/efoTrait now expects the canonical trait label).
data = gwas_get("efoTraits/search/findByEfoTrait", {"trait": "Alzheimer disease"})
traits = data["_embedded"]["efoTraits"]
print(f"Traits matching 'Alzheimer disease': {len(traits)}")
for t in traits[:5]:
print(f" {t['shortForm']}: {t['trait']} (uri={t['uri']})")
time.sleep(0.2)
# Get specific trait by shortForm. NOTE: many legacy EFO IDs have been
# re-mapped to MONDO (e.g. old `EFO_0000249` for Alzheimer is now
# `MONDO_0004975` — `efoTraits/EFO_0000249` returns 404). Resolve via search
# above first, then use the current shortForm:
short_form = traits[0]["shortForm"] # e.g. 'MONDO_0004975'
data = gwas_get(f"efoTraits/{short_form}")
print(f"Trait: {data['trait']}")
print(f" URI : {data['uri']}")
print(f" shortForm : {data['shortForm']}")Access study-level summary statistics for downstream analysis (meta-analysis, PRS).
# List studies with available summary statistics
data = gwas_get("studies/search/findByFullPvalueSet", {"fullPvalueSet": True, "size": 10})
studies = data["_embedded"]["studies"]
print(f"Studies with summary stats (first page): {len(studies)}")
for s in studies[:5]:
trait = (s.get("diseaseTrait") or {}).get("trait", "N/A")
print(f" {s['accessionId']}: {trait[:50]}")
time.sleep(0.2)
# Summary statistics metadata is NOT exposed via the REST API
# (`studies/{acc}/summaryStatistics` returns 404). Use the GWAS Catalog FTP
# directly — paths are predictable by study accession:
ftp_base = "http://ftp.ebi.ac.uk/pub/databases/gwas/summary_statistics"
acc = studies[0]["accessionId"]
ftp_url = f"{ftp_base}/{acc[:-3]}001-{acc[:-3]}999/{acc}/"
print(f"Summary stats FTP directory: {ftp_url}")# Download summary statistics file (FTP)
# Summary statistics are hosted on the GWAS Catalog FTP, not the REST API
import urllib.request
study_id = "GCST006867" # Example study
ftp_base = "http://ftp.ebi.ac.uk/pub/databases/gwas/summary_statistics"
# Actual paths vary by study; check the study page for the download link
# Example: ftp_base/{study_id}/{study_id}.tsv.gz
url = f"{ftp_base}/{study_id}"
print(f"Summary stats FTP directory: {url}")
# Use requests.get() or urllib to download the .tsv.gz fileFind GWAS associations by gene name or retrieve publication metadata.
# Search associations by gene
data = gwas_get("singleNucleotidePolymorphisms/search/findByGene",
{"geneName": "BRCA1", "size": 50})
snps = data["_embedded"]["singleNucleotidePolymorphisms"]
print(f"Variants near BRCA1: {len(snps)}")
for v in snps[:5]:
locs = v.get("locations", [{}])
pos = locs[0].get("chromosomePosition", "?") if locs else "?"
print(f" {v['rsId']}: chr{locs[0].get('chromosomeName', '?')}:{pos}")
time.sleep(0.2)
# Get study publication details
data = gwas_get("studies/GCST000392")
pub = data.get("publicationInfo", {})
print(f"Study: {data['accessionId']}")
print(f" Author: {pub.get('author', {}).get('fullname', 'N/A')}")
print(f" Journal: {pub.get('publication', 'N/A')}")
print(f" PMID: {pub.get('pubmedId', 'N/A')}")
print(f" Date: {pub.get('publicationDate', 'N/A')}")The GWAS Catalog organizes data as interconnected entities:
| Entity | Description | Key Identifier | Example |
|---|---|---|---|
| Study | A published GWAS experiment | GCST accession (e.g., GCST000392) | Wellcome Trust Case Control Consortium study |
| Association | A SNP-trait association with p-value and effect size | Internal ID | rs7903146 associated with T2D at p=1e-40 |
| Variant (SNP) | A single nucleotide polymorphism | rs number (e.g., rs7903146) | TCF7L2 variant |
| Trait | A disease/phenotype mapped to EFO ontology | EFO ID (e.g., EFO_0001360) | Type 2 diabetes mellitus |
| Gene | A gene near or harboring GWAS variants | Gene symbol (e.g., TCF7L2) | Transcription factor 7-like 2 |
Relationships: Study --(reports)--> Association --(involves)--> Variant + Trait. Variants map to genomic positions and nearby genes.
All API responses follow HAL (Hypertext Application Language) format:
| Top-level Section | Contents | Access Pattern |
|---|---|---|
_embedded | Primary data objects (studies, associations, etc.) | response["_embedded"]["studies"] |
_links | Navigation links (self, next, prev, first, last) | response["_links"]["next"]["href"] |
page | Pagination metadata (size, totalElements, totalPages, number) | response["page"]["totalElements"] |
The standard genome-wide significance threshold is p <= 5 x 10^-8, correcting for approximately 1 million independent tests across the human genome. Associations below this threshold are considered suggestive. The GWAS Catalog includes associations at various significance levels -- always check p-values when filtering results.
GCST000392)rs7903146)EFO_0001360 for type 2 diabetes)10q25.2)Goal: Map the genetic landscape of a disease by collecting all genome-wide significant loci.
import requests, time
BASE = "https://www.ebi.ac.uk/gwas/rest/api"
def gwas_get(endpoint, params=None):
url = f"{BASE}/{endpoint}"
resp = requests.get(url, params=params or {})
resp.raise_for_status()
time.sleep(0.2)
return resp.json()
# Step 1: Resolve trait → current shortForm via findByEfoTrait
# (the older `findByDescription` endpoint was removed).
traits = gwas_get("efoTraits/search/findByEfoTrait", {"trait": "schizophrenia"})
efo_id = traits["_embedded"]["efoTraits"][0]["shortForm"]
print(f"Using EFO: {efo_id}")
# Step 2: Get all associations for this trait — nested path
# `efoTraits/{shortForm}/associations` works once you have the canonical shortForm.
all_assocs = []
page = 0
while True:
data = gwas_get(f"efoTraits/{efo_id}/associations",
{"size": 500, "page": page})
assocs = data["_embedded"]["associations"]
all_assocs.extend(assocs)
if page >= data["page"]["totalPages"] - 1:
break
page += 1
# Step 3: Filter genome-wide significant
significant = [a for a in all_assocs if a.get("pvalue") and a["pvalue"] < 5e-8]
print(f"Total associations: {len(all_assocs)}, genome-wide significant: {len(significant)}")
# Step 4: Extract variant and effect details
for a in significant[:10]:
risk_alleles = a.get("loci", [{}])[0].get("strongestRiskAlleles", [])
snp = risk_alleles[0].get("snps", [{}])[0].get("rsId", "N/A") if risk_alleles else "N/A"
or_val = a.get("orPerCopyNum", "N/A")
beta = a.get("betaNum", "N/A")
print(f" {snp} | p={a['pvalue']:.2e} | OR={or_val} | beta={beta}")Goal: Determine how many distinct traits a single variant is associated with.
import requests, time
BASE = "https://www.ebi.ac.uk/gwas/rest/api"
def gwas_get(endpoint, params=None):
url = f"{BASE}/{endpoint}"
resp = requests.get(url, params=params or {})
resp.raise_for_status()
time.sleep(0.2)
return resp.json()
rs_id = "rs7903146" # Well-known pleiotropic variant
# Get all associations for this variant
data = gwas_get(f"singleNucleotidePolymorphisms/{rs_id}/associations", {"size": 500})
assocs = data["_embedded"]["associations"]
# Association records don't embed efoTraits inline — follow the HAL
# `_links.efoTraits.href` link per association. Cache per-href to avoid
# duplicate fetches.
trait_set = {}
href_cache = {}
def fetch_traits(href):
if href in href_cache:
return href_cache[href]
r = requests.get(href, timeout=15)
href_cache[href] = (r.json().get("_embedded", {}).get("efoTraits", [])) if r.ok else []
return href_cache[href]
for a in assocs:
href = (a.get("_links") or {}).get("efoTraits", {}).get("href")
if not href:
continue
for t in fetch_traits(href):
tid = t.get("shortForm", "unknown")
if tid not in trait_set:
trait_set[tid] = {"trait": t.get("trait", "N/A"),
"best_pval": a.get("pvalue", 1), "count": 0}
trait_set[tid]["count"] += 1
if a.get("pvalue") and a["pvalue"] < trait_set[tid]["best_pval"]:
trait_set[tid]["best_pval"] = a["pvalue"]
time.sleep(0.1)
print(f"{rs_id} is associated with {len(trait_set)} distinct traits:")
for tid, info in sorted(trait_set.items(), key=lambda x: x[1]["best_pval"]):
print(f" {tid}: {info['trait']} (best p={info['best_pval']:.2e}, n={info['count']})")Goal: Download summary statistics for a study and create a Manhattan plot.
import numpy as np
import pandas as pd
# Summary statistics as a CHR/BP/P table (simulated here for demonstration).
# In practice: download from GWAS Catalog FTP, e.g.
# url = "http://ftp.ebi.ac.uk/pub/databases/gwas/summary_statistics/GCSTXXXXXX/..."
# import gzip, requests; data = gzip.decompress(requests.get(url).content)
np.random.seed(42)
n_snps = 5000
pvalues = np.random.uniform(0, 1, size=n_snps)
pvalues[:20] = 10 ** np.random.uniform(-15, -8, size=20) # add some "hits"
sumstats = pd.DataFrame({
"CHR": np.random.choice(range(1, 23), size=n_snps),
"BP": np.random.randint(1, 250_000_000, size=n_snps),
"P": pvalues,
})
sumstats.to_csv("gwas.csv", index=False)
print(f"Summary stats: {len(sumstats)} SNPs, {(sumstats['P'] < 5e-8).sum()} genome-wide sig -> gwas.csv")
# Render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Manhattan" recipe -> figures/manhattan_plot.png| Parameter | Function/Endpoint | Default | Range / Options | Effect |
|---|---|---|---|---|
size | All paginated endpoints | 20 | 1-500 | Results per page |
page | All paginated endpoints | 0 | 0-totalPages-1 | Page number (0-indexed) |
diseaseTrait | studies/search/findByDiseaseTrait | -- | Any string | Trait keyword search |
pubmedId | studies/search/findByPublicationIdPubmedId | -- | Valid PMID | Study lookup by publication (findByPubmedId 404s on /studies/) |
geneName | snps/search/findByGene | -- | Gene symbol | Variants near a gene |
chrom, bpStart, bpEnd | snps/search/findByChromBpLocationRange | -- | chr:start-end | Regional variant query |
fullPvalueSet | studies/search/findByFullPvalueSet | -- | True/False | Filter studies with summary stats |
trait | efoTraits/search/findByEfoTrait | -- | Canonical trait name | Trait lookup by exact name (findByDescription was removed) |
efoTrait | associations/search/findByEfoTrait | -- | Canonical trait name | All associations for a trait |
Paginate large result sets: Default page size is 20; set size=500 and loop over pages for complete data retrieval. Check page.totalElements to know the full count before iterating.
Use EFO IDs for precise trait queries: Free-text search may return related but different traits. Look up the exact EFO ID first, then query associations by EFO ID for precision.
Always check p-values: The catalog contains associations at various significance levels. Filter to p < 5e-8 for genome-wide significant results unless you specifically need suggestive associations.
Be ancestry-aware: Effect sizes and allele frequencies vary across populations. Check the initialSampleSize and replicationSampleSize fields to understand the ancestry composition of each study.
Add rate-limiting delays: Although no official limit exists, time.sleep(0.2) between requests prevents server overload and avoids temporary blocks.
Cache frequently accessed data: Study and trait metadata rarely change. Cache results locally when running batch analyses to reduce redundant API calls.
Anti-pattern -- Don't rely solely on reported genes: Author-reported genes may not be the causal gene. Cross-reference with functional annotation tools and eQTL databases for biological interpretation.
Identify polygenic scores available for a GWAS trait.
import requests, time
# Step 1: Get EFO/MONDO shortForm from GWAS Catalog
# (the old `findByDescription` endpoint was removed; use `findByEfoTrait`)
BASE = "https://www.ebi.ac.uk/gwas/rest/api"
traits = requests.get(f"{BASE}/efoTraits/search/findByEfoTrait",
params={"trait": "coronary artery disease"}).json()
efo_id = traits["_embedded"]["efoTraits"][0]["shortForm"]
time.sleep(0.2)
# Step 2: Query PGS Catalog for scores using same EFO
pgs_url = f"https://www.pgscatalog.org/rest/score/search?trait_id={efo_id}"
pgs_data = requests.get(pgs_url).json()
print(f"PGS scores for {efo_id}: {pgs_data.get('count', 0)}")
for score in pgs_data.get("results", [])[:5]:
print(f" {score['id']}: {score['name']} (variants: {score.get('variants_number', '?')})")Collect all GWAS associations in a genomic region.
import requests, time
BASE = "https://www.ebi.ac.uk/gwas/rest/api"
# Query region: chr9:21,900,000-22,200,000 (CDKN2A/2B locus)
data = requests.get(f"{BASE}/singleNucleotidePolymorphisms/search/findByChromBpLocationRange",
params={"chrom": "9", "bpStart": "21900000",
"bpEnd": "22200000", "size": 200}).json()
time.sleep(0.2)
snps = data["_embedded"]["singleNucleotidePolymorphisms"]
print(f"Variants in CDKN2A/2B locus: {len(snps)}")
# Get associations for each variant
region_assocs = []
for v in snps[:10]: # Limit for demo
rs = v["rsId"]
try:
a_data = requests.get(f"{BASE}/singleNucleotidePolymorphisms/{rs}/associations",
params={"size": 50}).json()
time.sleep(0.2)
for a in a_data["_embedded"]["associations"]:
for t in a.get("efoTraits", []):
region_assocs.append({"rsId": rs, "trait": t["trait"],
"pvalue": a.get("pvalue")})
except Exception:
continue
print(f"Total associations in region: {len(region_assocs)}")Visualize effect sizes across studies for a single variant.
import re
import pandas as pd
import requests, time
BASE = "https://www.ebi.ac.uk/gwas/rest/api"
data = requests.get(f"{BASE}/singleNucleotidePolymorphisms/rs1801282/associations",
params={"size": 100}).json()
time.sleep(0.2)
assocs = data["_embedded"]["associations"]
def parse_ci(text):
"""Pull [low-high] confidence bounds out of the GWAS Catalog 'range' string."""
nums = re.findall(r"[0-9]*\.?[0-9]+", text or "")
return (float(nums[0]), float(nums[1])) if len(nums) >= 2 else (float("nan"), float("nan"))
# Build a forest-plot table: label, estimate (OR), ci_low, ci_high
rows = []
for a in assocs:
or_val = a.get("orPerCopyNum")
trait = a.get("efoTraits", [{}])[0].get("trait", "N/A") if a.get("efoTraits") else "N/A"
if or_val and or_val > 0:
lo, hi = parse_ci(a.get("range", ""))
rows.append({"label": trait[:30], "estimate": or_val, "ci_low": lo, "ci_high": hi})
effects = pd.DataFrame(rows)
effects.to_csv("effects.csv", index=False)
print(f"Effect sizes: {len(effects)} entries -> effects.csv")
# Render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Forest" recipe (null line at OR=1) -> figures/forest_plot.png| Problem | Cause | Solution |
|---|---|---|
404 Not Found | Invalid GCST, rs ID, or EFO ID | Verify identifier exists via search endpoint first |
Empty _embedded | No results for query | Broaden search terms or check spelling; try partial match |
ConnectionError / timeout | EBI server temporarily down | Retry with exponential backoff; check https://www.ebi.ac.uk/gwas/status |
| Truncated results | Default pagination (20 per page) | Set size=500 and iterate over page parameter |
| Missing OR/beta values | Not all associations report effect sizes | Check both orPerCopyNum and betaNum fields; some studies report only p-values |
| Stale data | Catalog updated quarterly | Check lastUpdateDate on studies; re-query if needed |
KeyError: '_embedded' | Endpoint returns single object, not list | Direct lookups (by ID) return a single object; _embedded only appears in search/list results |
| Summary stats unavailable | Not all studies deposit summary stats | Filter with findByFullPvalueSet=True to find studies with available data |
references/api_endpoints.mdCovers: Complete endpoint catalog for all 6 REST API groups (studies, associations, variants, traits, genes, summary statistics) with query parameters, plus response field tables for the main entity types (association fields, study fields, variant fields), HTTP error codes, and pagination patterns.
Relocated inline: Core query patterns and the most common endpoints are demonstrated in Core API modules with full code examples. HAL+JSON structure and key identifier formats are in Key Concepts.
Omitted from original api_reference.md (794 lines): Advanced query composition patterns (multi-parameter chaining beyond what Core API shows), child/parent trait traversal details, detailed changelog/versioning notes. These are specialized and covered by EBI documentation.
© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/genomics-bioinformatics/databases/gwas-database of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Gwas Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gwas Database this skilljaechang-hits/SciAgent-Skills | 371 | 1 repos | ~6.6k | Automated safety check: Pass | Apache-2.0 | |
| Bio Ensembl RESTGPTomics/bioSkills | 1.2k | 2 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Pride FetchClawBio/ClawBio | 1.2k | — | ~4.2k | Automated safety check: Pass | MIT | |
| Ensembl Databaseaipoch/medical-research-skills | 2k | — | ~1.5k | Automated safety check: Pass | MIT | |
| UniProt Database Accessdavila7/claude-code-templates | 32k | 14 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Singlecell Portalaipoch/medical-research-skills | 2k | — | ~1.2k | Automated safety check: Pass | MIT |
GPTomics/bioSkills
Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…
ClawBio/ClawBio
Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.
aipoch/medical-research-skills
Access Ensembl REST API for vertebrate genomic data; use when you need gene/ID lookups, sequence retrieval, variant effect prediction (VEP), or homology/assembly coordinate mapping.
davila7/claude-code-templates
Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.
aipoch/medical-research-skills
Programmatically query public single-cell study metadata from the Broad Institute Single Cell Portal REST API when you need to search and filter datasets by organism, tissue, disease, or cell type…
aipoch/medical-research-skills
Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need…
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS. Gwas Database is an agent skill from jaechang-hits/SciAgent-Skills. NHGRI-EBI GWAS Catalog REST API for SNP-trait associations from published GWAS.
Gwas Database fits situations like: tasks that involve Bioinformatics; tasks that involve REST APIs.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gwas-database in jaechang-hits/SciAgent-Skills) into .claude/skills/gwas-database in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gwas-database in jaechang-hits/SciAgent-Skills) into .agents/skills/gwas-database in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill gwas-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gwas-database, .gemini/skills/gwas-database, .github/skills/gwas-database and .opencode/skills/gwas-database in your project.
Going by SKILL.md and its folder, Gwas Database needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 4 domains. In commands or code: ebi.ac.uk, ftp.ebi.ac.uk and pgscatalog.org; the agent is likely to contact these when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gwas Database is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.6k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gwas Database: Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars), Pride Fetch (ClawBio/ClawBio, 1.2k stars), Ensembl Database (aipoch/medical-research-skills, 2k stars) and UniProt Database Access (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 371 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.