Bio Ensembl REST
GPTomics/bioSkills
Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…
Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping.
$ npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills hmdb-database --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/hmdb-database .claude/skills/hmdb-database && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hmdb-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/hmdb-database into .claude/skills/hmdb-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hmdb-database", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/hmdb-databaseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills hmdb-database --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/proteomics-protein-engineering/hmdb-database .agents/skills/hmdb-database && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hmdb-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/hmdb-database into .agents/skills/hmdb-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hmdb-database", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills hmdb-database --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/proteomics-protein-engineering/hmdb-database .cursor/skills/hmdb-database && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hmdb-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/hmdb-database into .cursor/skills/hmdb-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hmdb-database", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/proteomics-protein-engineering/hmdb-database--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills hmdb-database --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/proteomics-protein-engineering/hmdb-database .gemini/skills/hmdb-database && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hmdb-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/hmdb-database into .gemini/skills/hmdb-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hmdb-database", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills hmdb-databaseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/proteomics-protein-engineering/hmdb-database .github/skills/hmdb-database && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hmdb-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/hmdb-database into .github/skills/hmdb-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hmdb-database", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills hmdb-database --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/proteomics-protein-engineering/hmdb-database .opencode/skills/hmdb-database && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hmdb-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/hmdb-database into .opencode/skills/hmdb-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hmdb-database", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hmdb-databaseParse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping.
Hmdb Database is an agent skill from jaechang-hits/SciAgent-Skills. Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping. No REST API — uses ~6 GB XML download. Use drugbank-database-access for drugs; pubchem-compound-search for live lookups.
Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics and REST APIs. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
hmdb.caAlso links to:
doi.orgmetaboanalyst.caFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Hmdb Database loads about 6.1k tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 1,078 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 1,078 words, ~6,126 tokens.
.claude/skills/hmdb-database/SKILL.md (or your agent's skills folder).Query the Human Metabolome Database (HMDB, 220,000+ metabolite entries) by parsing locally downloaded XML with Python's ElementTree. Covers metabolite lookup, chemical properties, biological context (pathways, enzymes, biofluids), disease/biomarker associations, spectral data for metabolite identification, and cross-database ID mapping to KEGG, PubChem, ChEBI, and DrugBank.
drugbank-database-access insteadpubchem-compound-search insteadhmdb_metabolites.xml.zip (~6 GB uncompressed)lxml (faster XPath) or standard xml.etree.ElementTree, pandashmdbQuery on CRAN provides some query wrappers but is limited and outdatedpip install lxml pandasimport xml.etree.ElementTree as ET
NS = {'hmdb': 'http://www.hmdb.ca'}
tree = ET.parse('hmdb_metabolites.xml') # 60-120s for full XML
root = tree.getroot()
# Build lookup index (HMDB ID + lowercase name -> element)
metabolite_index = {}
for met in root.findall('hmdb:metabolite', NS):
accession = met.find('hmdb:accession', NS)
name = met.find('hmdb:name', NS)
if accession is not None and accession.text:
metabolite_index[accession.text] = met
if name is not None and name.text:
metabolite_index[name.text.lower()] = met
def find_metabolite(query):
"""Find metabolite by HMDB ID or name (case-insensitive)."""
return metabolite_index.get(query) or metabolite_index.get(query.lower())
met = find_metabolite('HMDB0000122') # Glucose
name = met.find('hmdb:name', NS).text
formula = met.find('hmdb:chemical_formula', NS).text
print(f"{name}: {formula}")
# Glucose: C6H12O6import xml.etree.ElementTree as ET
NS = {'hmdb': 'http://www.hmdb.ca'}
tree = ET.parse('hmdb_metabolites.xml')
root = tree.getroot()
print(f"Total metabolite entries: {len(root.findall('hmdb:metabolite', NS))}")
# Total metabolite entries: ~220000+For memory-constrained environments, use iterparse:
met_names = {}
for event, elem in ET.iterparse('hmdb_metabolites.xml', events=('end',)):
if elem.tag == '{http://www.hmdb.ca}metabolite':
acc = elem.find('{http://www.hmdb.ca}accession')
name = elem.find('{http://www.hmdb.ca}name')
if acc is not None and name is not None and acc.text and name.text:
met_names[acc.text] = name.text
elem.clear()
print(f"Parsed {len(met_names)} metabolites via iterparse")def get_chemical_properties(met_element):
"""Extract chemical properties from a metabolite entry."""
def txt(path):
el = met_element.find(path, NS)
return el.text if el is not None and el.text else None
# Fields: accession, name, chemical_formula, average_molecular_weight,
# monisotopic_molecular_weight, smiles, inchi, inchikey, state, iupac_name
return {tag: txt(f'hmdb:{tag}') for tag in [
'accession', 'name', 'chemical_formula', 'average_molecular_weight',
'monisotopic_molecular_weight', 'smiles', 'inchi', 'inchikey', 'state']}
props = get_chemical_properties(find_metabolite('HMDB0000158')) # L-Tyrosine
print(f"{props['name']}: MW={props['average_molecular_weight']}, SMILES={props['smiles']}")# Extract taxonomy / chemical classification (ClassyFire ontology)
def get_classification(met_element):
tax = met_element.find('hmdb:taxonomy', NS)
if tax is None:
return {}
def txt(tag):
el = tax.find(f'hmdb:{tag}', NS)
return el.text if el is not None and el.text else None
return {'kingdom': txt('kingdom'), 'super_class': txt('super_class'),
'class': txt('class'), 'sub_class': txt('sub_class'),
'direct_parent': txt('direct_parent')}
print(get_classification(find_metabolite('HMDB0000122')))
# {'kingdom': 'Organic compounds', 'super_class': 'Organooxygen compounds', ...}def get_pathways(met_element):
"""Extract metabolic pathway associations."""
pathways = []
for pw in met_element.findall('hmdb:pathways/hmdb:pathway', NS):
name = pw.find('hmdb:name', NS)
smpdb_id = pw.find('hmdb:smpdb_id', NS)
kegg_id = pw.find('hmdb:kegg_map_id', NS)
pathways.append({
'name': name.text if name is not None else None,
'smpdb_id': smpdb_id.text if smpdb_id is not None else None,
'kegg_map_id': kegg_id.text if kegg_id is not None else None,
})
return pathways
for pw in get_pathways(find_metabolite('HMDB0000122'))[:5]:
print(f" {pw['smpdb_id']}: {pw['name']}")def get_biolocations(met_element):
"""Extract biofluid, tissue, and cellular locations."""
bp = 'hmdb:biological_properties/hmdb:'
def texts(path):
return [el.text for el in met_element.findall(path, NS) if el.text]
return {'biofluids': texts(f'{bp}biospecimen_locations/hmdb:biospecimen'),
'tissues': texts(f'{bp}tissue_locations/hmdb:tissue'),
'cellular': texts(f'{bp}cellular_locations/hmdb:cellular')}
locs = get_biolocations(find_metabolite('HMDB0000122'))
print(f"Glucose biofluids: {locs['biofluids']}")def get_enzymes(met_element):
"""Extract associated enzymes/proteins with UniProt IDs."""
return [{'name': (p.find('hmdb:name', NS).text
if p.find('hmdb:name', NS) is not None else None),
'uniprot_id': (p.find('hmdb:uniprot_id', NS).text
if p.find('hmdb:uniprot_id', NS) is not None else None),
'gene_name': (p.find('hmdb:gene_name', NS).text
if p.find('hmdb:gene_name', NS) is not None else None)}
for p in met_element.findall('hmdb:protein_associations/hmdb:protein', NS)]
enz = get_enzymes(find_metabolite('HMDB0000122'))
print(f"Glucose-associated proteins: {len(enz)}")
for e in enz[:3]:
print(f" {e['gene_name']} ({e['uniprot_id']}): {e['name']}")def get_diseases(met_element):
"""Extract disease associations with OMIM IDs and PubMed references."""
diseases = []
for d in met_element.findall('hmdb:diseases/hmdb:disease', NS):
name = d.find('hmdb:name', NS)
omim = d.find('hmdb:omim_id', NS)
pmids = [r.find('hmdb:pubmed_id', NS).text
for r in d.findall('hmdb:references/hmdb:reference', NS)
if r.find('hmdb:pubmed_id', NS) is not None and r.find('hmdb:pubmed_id', NS).text]
diseases.append({'name': name.text if name is not None else None,
'omim_id': omim.text if omim is not None else None,
'pubmed_ids': pmids})
return diseases
diseases = get_diseases(find_metabolite('HMDB0000122'))
print(f"Glucose disease associations: {len(diseases)}")
for d in diseases[:3]:
print(f" {d['name']} (OMIM: {d['omim_id']})")def get_concentrations(met_element, biospecimen='Blood'):
"""Extract normal/abnormal concentration data filtered by biospecimen."""
result = {'normal': [], 'abnormal': []}
for ctype, key in [('normal_concentrations', 'normal'),
('abnormal_concentrations', 'abnormal')]:
for c in met_element.findall(f'hmdb:{ctype}/hmdb:concentration', NS):
bio = c.find('hmdb:biospecimen', NS)
if bio is None or bio.text != biospecimen:
continue
def txt(tag):
el = c.find(f'hmdb:{tag}', NS)
return el.text if el is not None else None
result[key].append({'value': txt('concentration_value'),
'units': txt('concentration_units'),
'condition': txt('subject_condition')})
return result
conc = get_concentrations(find_metabolite('HMDB0000122'), 'Blood')
print(f"Glucose blood: {len(conc['normal'])} normal, {len(conc['abnormal'])} abnormal")def get_ms_spectra(met_element):
"""Extract MS/MS spectral peak lists (m/z + intensity)."""
spectra = []
for spec in met_element.findall('hmdb:spectra/hmdb:spectrum', NS):
stype = spec.find('hmdb:type', NS)
if stype is None or 'MS' not in (stype.text or ''):
continue
peaks = [{'mz': float(p.find('hmdb:mass_charge', NS).text),
'intensity': float(p.find('hmdb:intensity', NS).text or 0)}
for p in spec.findall('hmdb:ms_ms_peaks/hmdb:ms_ms_peak', NS)
if p.find('hmdb:mass_charge', NS) is not None
and p.find('hmdb:mass_charge', NS).text]
spectra.append({'type': stype.text, 'num_peaks': len(peaks), 'peaks': peaks})
return spectra
spectra = get_ms_spectra(find_metabolite('HMDB0000122'))
print(f"Glucose MS spectra: {len(spectra)}")def get_nmr_spectra(met_element):
"""Extract NMR spectral peak lists (1H, 13C). Same pattern as MS above."""
spectra = []
for spec in met_element.findall('hmdb:spectra/hmdb:spectrum', NS):
stype = spec.find('hmdb:type', NS)
if stype is None or 'NMR' not in (stype.text or ''):
continue
nucleus = spec.find('hmdb:nucleus', NS)
shifts = [float(p.find('hmdb:chemical_shift', NS).text)
for p in spec.findall('hmdb:nmr_one_d_peaks/hmdb:nmr_one_d_peak', NS)
if p.find('hmdb:chemical_shift', NS) is not None
and p.find('hmdb:chemical_shift', NS).text]
spectra.append({'type': stype.text,
'nucleus': nucleus.text if nucleus is not None else None,
'num_peaks': len(shifts), 'chemical_shifts': shifts})
return spectra
nmr = get_nmr_spectra(find_metabolite('HMDB0000122'))
print(f"Glucose NMR spectra: {len(nmr)}")def get_external_ids(met_element):
"""Extract cross-database identifiers (KEGG, PubChem, ChEBI, DrugBank, CAS, etc.)."""
fields = {'kegg_id': 'KEGG', 'pubchem_compound_id': 'PubChem', 'chebi_id': 'ChEBI',
'drugbank_id': 'DrugBank', 'chemspider_id': 'ChemSpider',
'cas_registry_number': 'CAS', 'biocyc_id': 'BioCyc', 'pdb_id': 'PDB',
'foodb_id': 'FooDB', 'metlin_id': 'METLIN'}
ids = {}
for tag, label in fields.items():
el = met_element.find(f'hmdb:{tag}', NS)
if el is not None and el.text:
ids[label] = el.text
return ids
ids = get_external_ids(find_metabolite('HMDB0000122'))
print(f"Glucose cross-refs: {ids}")
# {'KEGG': 'C00031', 'PubChem': '5793', 'ChEBI': '17234', ...}# Build cross-reference table for a metabolite list
import pandas as pd
queries = ['HMDB0000122', 'HMDB0000158', 'HMDB0000167', 'HMDB0000148']
rows = []
for q in queries:
m = find_metabolite(q)
if m is None: continue
ids = get_external_ids(m)
rows.append({'name': m.find('hmdb:name', NS).text, 'hmdb_id': q,
'kegg': ids.get('KEGG'), 'pubchem': ids.get('PubChem'),
'chebi': ids.get('ChEBI')})
print(pd.DataFrame(rows).to_string(index=False))| Section | XPath | Content |
|---|---|---|
| Identity | hmdb:accession, hmdb:name, hmdb:iupac_name | Primary identifiers |
| Chemical | hmdb:chemical_formula, hmdb:smiles, hmdb:inchi | Structure descriptors |
| Properties | hmdb:average_molecular_weight, hmdb:state | Physical properties |
| Taxonomy | hmdb:taxonomy/hmdb:kingdom, class, etc. | Chemical classification |
| Biological | hmdb:biological_properties | Biofluids, tissues, cellular locations |
| Pathways | hmdb:pathways/hmdb:pathway | SMPDB + KEGG pathway links |
| Enzymes | hmdb:protein_associations/hmdb:protein | Associated proteins/enzymes |
| Diseases | hmdb:diseases/hmdb:disease | Disease associations + OMIM IDs |
| Concentrations | hmdb:normal_concentrations, hmdb:abnormal_concentrations | Biomarker reference ranges |
| Spectra | hmdb:spectra/hmdb:spectrum | NMR, MS/MS peak lists |
| External IDs | hmdb:kegg_id, hmdb:pubchem_compound_id, etc. | Cross-database identifiers |
| Ontology | hmdb:ontology | Physiological/disposition/process roles |
Not all entries have all fields populated. Coverage varies by metabolite class:
| Field Category | Approximate Coverage | Notes |
|---|---|---|
| Accession, name, formula | ~100% | Always present |
| SMILES, InChI, MW | ~90% | Missing for some lipids and complex metabolites |
| Taxonomy/classification | ~85% | Chemical ontology from ClassyFire |
| Biofluid locations | ~60% | Best for common human metabolites |
| Pathways | ~40% | Curated SMPDB + KEGG links |
| Protein associations | ~35% | Enzyme-metabolite relationships |
| Disease associations | ~25% | Primarily for biomarker metabolites |
| Normal concentrations | ~20% | Reference ranges for clinical metabolites |
| MS/MS spectra | ~15% | Experimental spectral libraries |
| NMR spectra | ~10% | 1H and 13C chemical shift data |
| Format | File | Size | Best For |
|---|---|---|---|
| Full XML | hmdb_metabolites.xml | ~6 GB | Complete data access (all fields) |
| SDF | structures.sdf | ~200 MB | Chemical structures + basic properties |
| CSV | Various exports | ~50-500 MB | Tabular data (properties, concentrations) |
| FASTA | hmdb_proteins.fasta | ~50 MB | Protein sequence lookups |
Goal: Match an observed m/z value to candidate metabolites using molecular weight. Assumes root, NS from Quick Start.
import pandas as pd
observed_mz = 180.063 # [M+H]+ for glucose
adduct_mass = 1.00728 # H+ adduct
target_mw = observed_mz - adduct_mass
tolerance_da = 0.01
candidates = []
for met in root.findall('hmdb:metabolite', NS):
mw_el = met.find('hmdb:monisotopic_molecular_weight', NS)
if mw_el is None or not mw_el.text:
continue
mw = float(mw_el.text)
if abs(mw - target_mw) <= tolerance_da:
candidates.append({
'hmdb_id': met.find('hmdb:accession', NS).text,
'name': met.find('hmdb:name', NS).text,
'mw': mw, 'delta_da': abs(mw - target_mw),
'formula': (met.find('hmdb:chemical_formula', NS).text
if met.find('hmdb:chemical_formula', NS) is not None else None),
})
df = pd.DataFrame(candidates).sort_values('delta_da')
print(f"Candidates within {tolerance_da} Da of {target_mw:.3f}: {len(df)}")
print(df.head(10).to_string(index=False))Goal: Find all metabolites associated with a disease and their concentration changes. Assumes root, NS from Quick Start.
import pandas as pd
disease_query = 'diabetes'
biomarkers = []
for met in root.findall('hmdb:metabolite', NS):
for d in met.findall('hmdb:diseases/hmdb:disease', NS):
dname = d.find('hmdb:name', NS)
if dname is None or not dname.text:
continue
if disease_query.lower() not in dname.text.lower():
continue
abnormal = [c for c in met.findall(
'hmdb:abnormal_concentrations/hmdb:concentration', NS)
if c.find('hmdb:subject_condition', NS) is not None
and disease_query.lower() in
(c.find('hmdb:subject_condition', NS).text or '').lower()]
biomarkers.append({
'hmdb_id': met.find('hmdb:accession', NS).text,
'metabolite': met.find('hmdb:name', NS).text,
'disease': dname.text,
'abnormal_measurements': len(abnormal),
})
df = pd.DataFrame(biomarkers).drop_duplicates(['hmdb_id', 'disease'])
print(f"Metabolites linked to '{disease_query}': {len(df)}")
print(df.sort_values('abnormal_measurements', ascending=False).head(15).to_string(index=False))Goal: Given a list of identified metabolites, find over-represented pathways. Assumes metabolite_index from Quick Start.
from collections import Counter
import pandas as pd
hit_ids = ['HMDB0000122', 'HMDB0000158', 'HMDB0000167',
'HMDB0000148', 'HMDB0000064', 'HMDB0000161']
hit_pathways = Counter()
for hid in hit_ids:
met = metabolite_index.get(hid)
if met is None:
continue
for pw in met.findall('hmdb:pathways/hmdb:pathway', NS):
pw_name = pw.find('hmdb:name', NS)
if pw_name is not None and pw_name.text:
hit_pathways[pw_name.text] += 1
enriched = [(pw, count) for pw, count in hit_pathways.most_common() if count >= 2]
df = pd.DataFrame(enriched, columns=['Pathway', 'Hit_Count'])
print(f"Pathways with 2+ hits from {len(hit_ids)} metabolites:")
print(df.to_string(index=False))| Parameter | Function/Endpoint | Default | Description |
|---|---|---|---|
NS (namespace dict) | All XPath queries | {'hmdb': 'http://www.hmdb.ca'} | Required for all find/findall calls |
tolerance_da | MW matching | 0.01 | Mass tolerance in Daltons for metabolite ID |
biospecimen | get_concentrations() | 'Blood' | Filter: Blood, Urine, Cerebrospinal Fluid (CSF), Saliva, etc. |
events | ET.iterparse() | ('end',) | Parse events; use ('end',) to fire on closing tags |
| adduct mass | MS identification | varies | 1.00728 [M+H]+, 22.9892 [M+Na]+, -1.00728 [M-H]- |
target_type | Spectral queries | 'MS' or 'NMR' | Filter spectra by type string |
find()/findall() needs NS = {'hmdb': 'http://www.hmdb.ca'}. Omitting it returns empty resultselem.clear() in iterparse to process incrementallyel is not None and el.text before accessingmonisotopic_molecular_weight is correct for mass spec; average_molecular_weight for other calculationsWhen to use: Create a flat table of all metabolites with key properties.
import pandas as pd
records = []
for met in root.findall('hmdb:metabolite', NS):
def txt(p):
el = met.find(p, NS)
return el.text if el is not None and el.text else None
records.append({'hmdb_id': txt('hmdb:accession'), 'name': txt('hmdb:name'),
'formula': txt('hmdb:chemical_formula'),
'avg_mw': txt('hmdb:average_molecular_weight'),
'smiles': txt('hmdb:smiles')})
pd.DataFrame(records).to_csv('hmdb_properties.csv', index=False)
print(f"Exported {len(records)} metabolites")When to use: Get all metabolites detected in a specific biofluid (e.g., urine for clinical screening).
biofluid = 'Urine'
urine_mets = []
for met in root.findall('hmdb:metabolite', NS):
for bf in met.findall(
'hmdb:biological_properties/hmdb:biospecimen_locations/hmdb:biospecimen', NS):
if bf.text and bf.text == biofluid:
urine_mets.append({
'hmdb_id': met.find('hmdb:accession', NS).text,
'name': met.find('hmdb:name', NS).text,
})
break
print(f"Metabolites in {biofluid}: {len(urine_mets)}")When to use: Summarize metabolite chemical classes for a hit list or the full database.
from collections import Counter
class_counts = Counter()
for met in root.findall('hmdb:metabolite', NS):
tax = met.find('hmdb:taxonomy', NS)
if tax is not None:
sc = tax.find('hmdb:super_class', NS)
if sc is not None and sc.text:
class_counts[sc.text] += 1
print(dict(class_counts.most_common(10)))| Problem | Cause | Solution |
|---|---|---|
find() returns None for known elements | Missing XML namespace | Always pass NS = {'hmdb': 'http://www.hmdb.ca'} |
MemoryError parsing full XML | ~8-12 GB needed in memory | Use ET.iterparse() with elem.clear() |
| Slow startup (>120s) | Parsing 6 GB XML | Parse once, build index dict; avoid re-parsing |
| Metabolite not found by name | Case sensitivity or synonym | Normalize to lowercase; try HMDB ID directly |
| Empty spectra for a metabolite | Not all entries have spectra (~10-15% coverage) | Check coverage table; use METLIN or MassBank for more spectra |
| Missing concentration data | Limited to well-studied clinical metabolites (~20%) | Cross-reference with MetaboAnalyst or literature |
| Duplicate entries for same compound | Secondary accessions (HMDB00XXXXX vs HMDB0000XXXX) | Use accession (primary), not secondary_accessions |
ET.iterparse missing data | Premature elem.clear() | Only clear after extracting all needed fields from the element |
Self-contained entry. The original reference file hmdb_data_fields.md (268 lines, field catalog with XML element names and descriptions) is consolidated into the Key Concepts "HMDB XML Entry Structure" table and the "Data Field Completeness" table above. The field catalog's per-element XML tag names are demonstrated in Core API code blocks. Omitted from original: web interface descriptions (not programmatic).
© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/proteomics-protein-engineering/hmdb-database of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Hmdb Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Hmdb Database this skilljaechang-hits/SciAgent-Skills | 371 | 1 repos | ~6.1k | Automated safety check: Pass | CC-BY-4.0 | |
| Bio Ensembl RESTGPTomics/bioSkills | 1.2k | 2 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Pride FetchClawBio/ClawBio | 1.2k | — | ~4.2k | Automated safety check: Pass | MIT | |
| Ensembl Databaseaipoch/medical-research-skills | 2k | — | ~1.5k | Automated safety check: Pass | MIT | |
| UniProt Database Accessdavila7/claude-code-templates | 32k | 14 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Singlecell Portalaipoch/medical-research-skills | 2k | — | ~1.2k | Automated safety check: Pass | MIT |
GPTomics/bioSkills
Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…
ClawBio/ClawBio
Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.
aipoch/medical-research-skills
Access Ensembl REST API for vertebrate genomic data; use when you need gene/ID lookups, sequence retrieval, variant effect prediction (VEP), or homology/assembly coordinate mapping.
davila7/claude-code-templates
Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.
aipoch/medical-research-skills
Programmatically query public single-cell study metadata from the Broad Institute Single Cell Portal REST API when you need to search and filter datasets by organism, tissue, disease, or cell type…
aipoch/medical-research-skills
Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need…
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping. Hmdb Database is an agent skill from jaechang-hits/SciAgent-Skills. Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping.
Hmdb Database fits situations like: tasks that involve Bioinformatics; tasks that involve REST APIs.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/hmdb-database in jaechang-hits/SciAgent-Skills) into .claude/skills/hmdb-database in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/hmdb-database in jaechang-hits/SciAgent-Skills) into .agents/skills/hmdb-database in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hmdb-database, .gemini/skills/hmdb-database, .github/skills/hmdb-database and .opencode/skills/hmdb-database in your project.
Going by SKILL.md and its folder, Hmdb Database needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 3 domains. In commands or code: hmdb.ca; the agent is likely to contact it when it follows the instructions. As links in the text: doi.org and metaboanalyst.ca. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Hmdb Database is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.1k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Hmdb Database: Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars), Pride Fetch (ClawBio/ClawBio, 1.2k stars), Ensembl Database (aipoch/medical-research-skills, 2k stars) and UniProt Database Access (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 371 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.