Agent skill

Hmdb Database

by jaechang-hits in jaechang-hits/SciAgent-Skills

Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping.

CC-BY-4.0Auto-check passedResearch & Science

Install Hmdb Database

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills hmdb-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/hmdb-database .claude/skills/hmdb-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hmdb-database
GitHub stars
371
Used in
1 other repo
Token cost
~6.1k tokens
SKILL.md length
1,078 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping.

  • Works in 6 steps: XML Setup and Metabolite Lookup → Chemical Properties → Biological Context → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 10 more sections
  • Calls pip; reaches hmdb.ca

What it does

Hmdb Database is an agent skill from jaechang-hits/SciAgent-Skills. Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping. No REST API — uses ~6 GB XML download. Use drugbank-database-access for drugs; pubchem-compound-search for live lookups.

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics and REST APIs. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve REST APIs

Example prompts

  • “/hmdb-database”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. XML Setup and Metabolite Lookup
  2. Chemical Properties
  3. Biological Context
  4. Disease and Biomarker Queries
  5. Spectral Data
  6. Cross-Database Mapping

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • hmdb.ca

    Also links to:

    • doi.org
    • metaboanalyst.ca

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hmdb Database loads about 6.1k tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 1,078 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 1,078 words, ~6,126 tokens.

Download SKILL.mdSave it as .claude/skills/hmdb-database/SKILL.md (or your agent's skills folder).
name
hmdb-database
description
Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping. No REST API — uses ~6 GB XML download. Use drugbank-database-access for drugs; pubchem-compound-search for live lookups.
license
CC-BY-4.0

HMDB Database — Local XML Access

Overview

Query the Human Metabolome Database (HMDB, 220,000+ metabolite entries) by parsing locally downloaded XML with Python's ElementTree. Covers metabolite lookup, chemical properties, biological context (pathways, enzymes, biofluids), disease/biomarker associations, spectral data for metabolite identification, and cross-database ID mapping to KEGG, PubChem, ChEBI, and DrugBank.

When to Use

  • Looking up metabolite information (description, chemical class, cellular location) by HMDB ID or name
  • Retrieving chemical properties (molecular weight, formula, SMILES, InChI, logP, PSA) for metabolomics analysis
  • Finding pathway associations and enzyme links for a set of metabolites
  • Identifying biofluid/tissue locations of metabolites (blood, urine, CSF, saliva)
  • Querying disease associations and normal/abnormal concentration ranges for biomarker discovery
  • Extracting NMR or MS spectral peak lists for metabolite identification
  • Mapping HMDB IDs to KEGG, PubChem, ChEBI, DrugBank, or other databases
  • For drug-specific data (interactions, targets, pharmacology) use drugbank-database-access instead
  • For live compound property queries without downloading use pubchem-compound-search instead

Prerequisites

  • HMDB XML download: Register at https://hmdb.ca/downloads — download hmdb_metabolites.xml.zip (~6 GB uncompressed)
  • Python packages: lxml (faster XPath) or standard xml.etree.ElementTree, pandas
  • No public REST API: HMDB has no programmatic REST API. All access is via local XML parsing or the web interface
  • R package (optional): hmdbQuery on CRAN provides some query wrappers but is limited and outdated
  • Rate limits: N/A for local XML parsing. The web interface has no documented rate limits but is not intended for scraping
bash
pip install lxml pandas

Quick Start

python
import xml.etree.ElementTree as ET

NS = {'hmdb': 'http://www.hmdb.ca'}

tree = ET.parse('hmdb_metabolites.xml')  # 60-120s for full XML
root = tree.getroot()

# Build lookup index (HMDB ID + lowercase name -> element)
metabolite_index = {}
for met in root.findall('hmdb:metabolite', NS):
    accession = met.find('hmdb:accession', NS)
    name = met.find('hmdb:name', NS)
    if accession is not None and accession.text:
        metabolite_index[accession.text] = met
    if name is not None and name.text:
        metabolite_index[name.text.lower()] = met

def find_metabolite(query):
    """Find metabolite by HMDB ID or name (case-insensitive)."""
    return metabolite_index.get(query) or metabolite_index.get(query.lower())

met = find_metabolite('HMDB0000122')  # Glucose
name = met.find('hmdb:name', NS).text
formula = met.find('hmdb:chemical_formula', NS).text
print(f"{name}: {formula}")
# Glucose: C6H12O6

Core API

1. XML Setup and Metabolite Lookup
python
import xml.etree.ElementTree as ET

NS = {'hmdb': 'http://www.hmdb.ca'}
tree = ET.parse('hmdb_metabolites.xml')
root = tree.getroot()
print(f"Total metabolite entries: {len(root.findall('hmdb:metabolite', NS))}")
# Total metabolite entries: ~220000+

For memory-constrained environments, use iterparse:

python
met_names = {}
for event, elem in ET.iterparse('hmdb_metabolites.xml', events=('end',)):
    if elem.tag == '{http://www.hmdb.ca}metabolite':
        acc = elem.find('{http://www.hmdb.ca}accession')
        name = elem.find('{http://www.hmdb.ca}name')
        if acc is not None and name is not None and acc.text and name.text:
            met_names[acc.text] = name.text
        elem.clear()
print(f"Parsed {len(met_names)} metabolites via iterparse")
2. Chemical Properties
python
def get_chemical_properties(met_element):
    """Extract chemical properties from a metabolite entry."""
    def txt(path):
        el = met_element.find(path, NS)
        return el.text if el is not None and el.text else None
    # Fields: accession, name, chemical_formula, average_molecular_weight,
    # monisotopic_molecular_weight, smiles, inchi, inchikey, state, iupac_name
    return {tag: txt(f'hmdb:{tag}') for tag in [
        'accession', 'name', 'chemical_formula', 'average_molecular_weight',
        'monisotopic_molecular_weight', 'smiles', 'inchi', 'inchikey', 'state']}

props = get_chemical_properties(find_metabolite('HMDB0000158'))  # L-Tyrosine
print(f"{props['name']}: MW={props['average_molecular_weight']}, SMILES={props['smiles']}")
python
# Extract taxonomy / chemical classification (ClassyFire ontology)
def get_classification(met_element):
    tax = met_element.find('hmdb:taxonomy', NS)
    if tax is None:
        return {}
    def txt(tag):
        el = tax.find(f'hmdb:{tag}', NS)
        return el.text if el is not None and el.text else None
    return {'kingdom': txt('kingdom'), 'super_class': txt('super_class'),
            'class': txt('class'), 'sub_class': txt('sub_class'),
            'direct_parent': txt('direct_parent')}

print(get_classification(find_metabolite('HMDB0000122')))
# {'kingdom': 'Organic compounds', 'super_class': 'Organooxygen compounds', ...}
3. Biological Context
python
def get_pathways(met_element):
    """Extract metabolic pathway associations."""
    pathways = []
    for pw in met_element.findall('hmdb:pathways/hmdb:pathway', NS):
        name = pw.find('hmdb:name', NS)
        smpdb_id = pw.find('hmdb:smpdb_id', NS)
        kegg_id = pw.find('hmdb:kegg_map_id', NS)
        pathways.append({
            'name': name.text if name is not None else None,
            'smpdb_id': smpdb_id.text if smpdb_id is not None else None,
            'kegg_map_id': kegg_id.text if kegg_id is not None else None,
        })
    return pathways

for pw in get_pathways(find_metabolite('HMDB0000122'))[:5]:
    print(f"  {pw['smpdb_id']}: {pw['name']}")
python
def get_biolocations(met_element):
    """Extract biofluid, tissue, and cellular locations."""
    bp = 'hmdb:biological_properties/hmdb:'
    def texts(path):
        return [el.text for el in met_element.findall(path, NS) if el.text]
    return {'biofluids': texts(f'{bp}biospecimen_locations/hmdb:biospecimen'),
            'tissues': texts(f'{bp}tissue_locations/hmdb:tissue'),
            'cellular': texts(f'{bp}cellular_locations/hmdb:cellular')}

locs = get_biolocations(find_metabolite('HMDB0000122'))
print(f"Glucose biofluids: {locs['biofluids']}")
python
def get_enzymes(met_element):
    """Extract associated enzymes/proteins with UniProt IDs."""
    return [{'name': (p.find('hmdb:name', NS).text
                      if p.find('hmdb:name', NS) is not None else None),
             'uniprot_id': (p.find('hmdb:uniprot_id', NS).text
                            if p.find('hmdb:uniprot_id', NS) is not None else None),
             'gene_name': (p.find('hmdb:gene_name', NS).text
                           if p.find('hmdb:gene_name', NS) is not None else None)}
            for p in met_element.findall('hmdb:protein_associations/hmdb:protein', NS)]

enz = get_enzymes(find_metabolite('HMDB0000122'))
print(f"Glucose-associated proteins: {len(enz)}")
for e in enz[:3]:
    print(f"  {e['gene_name']} ({e['uniprot_id']}): {e['name']}")
4. Disease and Biomarker Queries
python
def get_diseases(met_element):
    """Extract disease associations with OMIM IDs and PubMed references."""
    diseases = []
    for d in met_element.findall('hmdb:diseases/hmdb:disease', NS):
        name = d.find('hmdb:name', NS)
        omim = d.find('hmdb:omim_id', NS)
        pmids = [r.find('hmdb:pubmed_id', NS).text
                 for r in d.findall('hmdb:references/hmdb:reference', NS)
                 if r.find('hmdb:pubmed_id', NS) is not None and r.find('hmdb:pubmed_id', NS).text]
        diseases.append({'name': name.text if name is not None else None,
                         'omim_id': omim.text if omim is not None else None,
                         'pubmed_ids': pmids})
    return diseases

diseases = get_diseases(find_metabolite('HMDB0000122'))
print(f"Glucose disease associations: {len(diseases)}")
for d in diseases[:3]:
    print(f"  {d['name']} (OMIM: {d['omim_id']})")
python
def get_concentrations(met_element, biospecimen='Blood'):
    """Extract normal/abnormal concentration data filtered by biospecimen."""
    result = {'normal': [], 'abnormal': []}
    for ctype, key in [('normal_concentrations', 'normal'),
                       ('abnormal_concentrations', 'abnormal')]:
        for c in met_element.findall(f'hmdb:{ctype}/hmdb:concentration', NS):
            bio = c.find('hmdb:biospecimen', NS)
            if bio is None or bio.text != biospecimen:
                continue
            def txt(tag):
                el = c.find(f'hmdb:{tag}', NS)
                return el.text if el is not None else None
            result[key].append({'value': txt('concentration_value'),
                                'units': txt('concentration_units'),
                                'condition': txt('subject_condition')})
    return result

conc = get_concentrations(find_metabolite('HMDB0000122'), 'Blood')
print(f"Glucose blood: {len(conc['normal'])} normal, {len(conc['abnormal'])} abnormal")
5. Spectral Data
python
def get_ms_spectra(met_element):
    """Extract MS/MS spectral peak lists (m/z + intensity)."""
    spectra = []
    for spec in met_element.findall('hmdb:spectra/hmdb:spectrum', NS):
        stype = spec.find('hmdb:type', NS)
        if stype is None or 'MS' not in (stype.text or ''):
            continue
        peaks = [{'mz': float(p.find('hmdb:mass_charge', NS).text),
                  'intensity': float(p.find('hmdb:intensity', NS).text or 0)}
                 for p in spec.findall('hmdb:ms_ms_peaks/hmdb:ms_ms_peak', NS)
                 if p.find('hmdb:mass_charge', NS) is not None
                 and p.find('hmdb:mass_charge', NS).text]
        spectra.append({'type': stype.text, 'num_peaks': len(peaks), 'peaks': peaks})
    return spectra

spectra = get_ms_spectra(find_metabolite('HMDB0000122'))
print(f"Glucose MS spectra: {len(spectra)}")
python
def get_nmr_spectra(met_element):
    """Extract NMR spectral peak lists (1H, 13C). Same pattern as MS above."""
    spectra = []
    for spec in met_element.findall('hmdb:spectra/hmdb:spectrum', NS):
        stype = spec.find('hmdb:type', NS)
        if stype is None or 'NMR' not in (stype.text or ''):
            continue
        nucleus = spec.find('hmdb:nucleus', NS)
        shifts = [float(p.find('hmdb:chemical_shift', NS).text)
                  for p in spec.findall('hmdb:nmr_one_d_peaks/hmdb:nmr_one_d_peak', NS)
                  if p.find('hmdb:chemical_shift', NS) is not None
                  and p.find('hmdb:chemical_shift', NS).text]
        spectra.append({'type': stype.text,
                        'nucleus': nucleus.text if nucleus is not None else None,
                        'num_peaks': len(shifts), 'chemical_shifts': shifts})
    return spectra

nmr = get_nmr_spectra(find_metabolite('HMDB0000122'))
print(f"Glucose NMR spectra: {len(nmr)}")
6. Cross-Database Mapping
python
def get_external_ids(met_element):
    """Extract cross-database identifiers (KEGG, PubChem, ChEBI, DrugBank, CAS, etc.)."""
    fields = {'kegg_id': 'KEGG', 'pubchem_compound_id': 'PubChem', 'chebi_id': 'ChEBI',
              'drugbank_id': 'DrugBank', 'chemspider_id': 'ChemSpider',
              'cas_registry_number': 'CAS', 'biocyc_id': 'BioCyc', 'pdb_id': 'PDB',
              'foodb_id': 'FooDB', 'metlin_id': 'METLIN'}
    ids = {}
    for tag, label in fields.items():
        el = met_element.find(f'hmdb:{tag}', NS)
        if el is not None and el.text:
            ids[label] = el.text
    return ids

ids = get_external_ids(find_metabolite('HMDB0000122'))
print(f"Glucose cross-refs: {ids}")
# {'KEGG': 'C00031', 'PubChem': '5793', 'ChEBI': '17234', ...}
python
# Build cross-reference table for a metabolite list
import pandas as pd

queries = ['HMDB0000122', 'HMDB0000158', 'HMDB0000167', 'HMDB0000148']
rows = []
for q in queries:
    m = find_metabolite(q)
    if m is None: continue
    ids = get_external_ids(m)
    rows.append({'name': m.find('hmdb:name', NS).text, 'hmdb_id': q,
                 'kegg': ids.get('KEGG'), 'pubchem': ids.get('PubChem'),
                 'chebi': ids.get('ChEBI')})
print(pd.DataFrame(rows).to_string(index=False))

Key Concepts

HMDB XML Entry Structure
SectionXPathContent
Identityhmdb:accession, hmdb:name, hmdb:iupac_namePrimary identifiers
Chemicalhmdb:chemical_formula, hmdb:smiles, hmdb:inchiStructure descriptors
Propertieshmdb:average_molecular_weight, hmdb:statePhysical properties
Taxonomyhmdb:taxonomy/hmdb:kingdom, class, etc.Chemical classification
Biologicalhmdb:biological_propertiesBiofluids, tissues, cellular locations
Pathwayshmdb:pathways/hmdb:pathwaySMPDB + KEGG pathway links
Enzymeshmdb:protein_associations/hmdb:proteinAssociated proteins/enzymes
Diseaseshmdb:diseases/hmdb:diseaseDisease associations + OMIM IDs
Concentrationshmdb:normal_concentrations, hmdb:abnormal_concentrationsBiomarker reference ranges
Spectrahmdb:spectra/hmdb:spectrumNMR, MS/MS peak lists
External IDshmdb:kegg_id, hmdb:pubchem_compound_id, etc.Cross-database identifiers
Ontologyhmdb:ontologyPhysiological/disposition/process roles
Data Field Completeness

Not all entries have all fields populated. Coverage varies by metabolite class:

Field CategoryApproximate CoverageNotes
Accession, name, formula~100%Always present
SMILES, InChI, MW~90%Missing for some lipids and complex metabolites
Taxonomy/classification~85%Chemical ontology from ClassyFire
Biofluid locations~60%Best for common human metabolites
Pathways~40%Curated SMPDB + KEGG links
Protein associations~35%Enzyme-metabolite relationships
Disease associations~25%Primarily for biomarker metabolites
Normal concentrations~20%Reference ranges for clinical metabolites
MS/MS spectra~15%Experimental spectral libraries
NMR spectra~10%1H and 13C chemical shift data
File Format Comparison
FormatFileSizeBest For
Full XMLhmdb_metabolites.xml~6 GBComplete data access (all fields)
SDFstructures.sdf~200 MBChemical structures + basic properties
CSVVarious exports~50-500 MBTabular data (properties, concentrations)
FASTAhmdb_proteins.fasta~50 MBProtein sequence lookups

Common Workflows

Workflow 1: Metabolite Identification from Mass Spec

Goal: Match an observed m/z value to candidate metabolites using molecular weight. Assumes root, NS from Quick Start.

python
import pandas as pd

observed_mz = 180.063  # [M+H]+ for glucose
adduct_mass = 1.00728  # H+ adduct
target_mw = observed_mz - adduct_mass
tolerance_da = 0.01

candidates = []
for met in root.findall('hmdb:metabolite', NS):
    mw_el = met.find('hmdb:monisotopic_molecular_weight', NS)
    if mw_el is None or not mw_el.text:
        continue
    mw = float(mw_el.text)
    if abs(mw - target_mw) <= tolerance_da:
        candidates.append({
            'hmdb_id': met.find('hmdb:accession', NS).text,
            'name': met.find('hmdb:name', NS).text,
            'mw': mw, 'delta_da': abs(mw - target_mw),
            'formula': (met.find('hmdb:chemical_formula', NS).text
                        if met.find('hmdb:chemical_formula', NS) is not None else None),
        })

df = pd.DataFrame(candidates).sort_values('delta_da')
print(f"Candidates within {tolerance_da} Da of {target_mw:.3f}: {len(df)}")
print(df.head(10).to_string(index=False))
Workflow 2: Biomarker Discovery for a Disease

Goal: Find all metabolites associated with a disease and their concentration changes. Assumes root, NS from Quick Start.

python
import pandas as pd

disease_query = 'diabetes'
biomarkers = []
for met in root.findall('hmdb:metabolite', NS):
    for d in met.findall('hmdb:diseases/hmdb:disease', NS):
        dname = d.find('hmdb:name', NS)
        if dname is None or not dname.text:
            continue
        if disease_query.lower() not in dname.text.lower():
            continue
        abnormal = [c for c in met.findall(
            'hmdb:abnormal_concentrations/hmdb:concentration', NS)
            if c.find('hmdb:subject_condition', NS) is not None
            and disease_query.lower() in
            (c.find('hmdb:subject_condition', NS).text or '').lower()]
        biomarkers.append({
            'hmdb_id': met.find('hmdb:accession', NS).text,
            'metabolite': met.find('hmdb:name', NS).text,
            'disease': dname.text,
            'abnormal_measurements': len(abnormal),
        })

df = pd.DataFrame(biomarkers).drop_duplicates(['hmdb_id', 'disease'])
print(f"Metabolites linked to '{disease_query}': {len(df)}")
print(df.sort_values('abnormal_measurements', ascending=False).head(15).to_string(index=False))
Workflow 3: Pathway Enrichment from Metabolite List

Goal: Given a list of identified metabolites, find over-represented pathways. Assumes metabolite_index from Quick Start.

python
from collections import Counter
import pandas as pd

hit_ids = ['HMDB0000122', 'HMDB0000158', 'HMDB0000167',
           'HMDB0000148', 'HMDB0000064', 'HMDB0000161']

hit_pathways = Counter()
for hid in hit_ids:
    met = metabolite_index.get(hid)
    if met is None:
        continue
    for pw in met.findall('hmdb:pathways/hmdb:pathway', NS):
        pw_name = pw.find('hmdb:name', NS)
        if pw_name is not None and pw_name.text:
            hit_pathways[pw_name.text] += 1

enriched = [(pw, count) for pw, count in hit_pathways.most_common() if count >= 2]
df = pd.DataFrame(enriched, columns=['Pathway', 'Hit_Count'])
print(f"Pathways with 2+ hits from {len(hit_ids)} metabolites:")
print(df.to_string(index=False))

Key Parameters

ParameterFunction/EndpointDefaultDescription
NS (namespace dict)All XPath queries{'hmdb': 'http://www.hmdb.ca'}Required for all find/findall calls
tolerance_daMW matching0.01Mass tolerance in Daltons for metabolite ID
biospecimenget_concentrations()'Blood'Filter: Blood, Urine, Cerebrospinal Fluid (CSF), Saliva, etc.
eventsET.iterparse()('end',)Parse events; use ('end',) to fire on closing tags
adduct massMS identificationvaries1.00728 [M+H]+, 22.9892 [M+Na]+, -1.00728 [M-H]-
target_typeSpectral queries'MS' or 'NMR'Filter spectra by type string
Show full SKILL.md (458 more words)Show less

Best Practices

  1. Build an in-memory index on startup: Parse once (60-120s), build dict by HMDB ID + lowercase name. Never re-parse inside a loop
  2. Always pass the namespace dict: Every find()/findall() needs NS = {'hmdb': 'http://www.hmdb.ca'}. Omitting it returns empty results
  3. Use iterparse for memory constraints: Full XML requires ~8-12 GB RAM. Use elem.clear() in iterparse to process incrementally
  4. Guard against None and empty text: Many optional fields are absent. Always check el is not None and el.text before accessing
  5. Use monoisotopic weight for MS matching: monisotopic_molecular_weight is correct for mass spec; average_molecular_weight for other calculations
  6. Pre-filter by chemical class for large searches: Use taxonomy/classification to narrow searches before iterating all 220K+ entries

Common Recipes

Recipe: Export All Metabolite Properties to CSV

When to use: Create a flat table of all metabolites with key properties.

python
import pandas as pd
records = []
for met in root.findall('hmdb:metabolite', NS):
    def txt(p):
        el = met.find(p, NS)
        return el.text if el is not None and el.text else None
    records.append({'hmdb_id': txt('hmdb:accession'), 'name': txt('hmdb:name'),
                    'formula': txt('hmdb:chemical_formula'),
                    'avg_mw': txt('hmdb:average_molecular_weight'),
                    'smiles': txt('hmdb:smiles')})
pd.DataFrame(records).to_csv('hmdb_properties.csv', index=False)
print(f"Exported {len(records)} metabolites")
Recipe: Find Metabolites by Biofluid

When to use: Get all metabolites detected in a specific biofluid (e.g., urine for clinical screening).

python
biofluid = 'Urine'
urine_mets = []
for met in root.findall('hmdb:metabolite', NS):
    for bf in met.findall(
            'hmdb:biological_properties/hmdb:biospecimen_locations/hmdb:biospecimen', NS):
        if bf.text and bf.text == biofluid:
            urine_mets.append({
                'hmdb_id': met.find('hmdb:accession', NS).text,
                'name': met.find('hmdb:name', NS).text,
            })
            break
print(f"Metabolites in {biofluid}: {len(urine_mets)}")
Recipe: Chemical Class Distribution

When to use: Summarize metabolite chemical classes for a hit list or the full database.

python
from collections import Counter
class_counts = Counter()
for met in root.findall('hmdb:metabolite', NS):
    tax = met.find('hmdb:taxonomy', NS)
    if tax is not None:
        sc = tax.find('hmdb:super_class', NS)
        if sc is not None and sc.text:
            class_counts[sc.text] += 1
print(dict(class_counts.most_common(10)))

Troubleshooting

ProblemCauseSolution
find() returns None for known elementsMissing XML namespaceAlways pass NS = {'hmdb': 'http://www.hmdb.ca'}
MemoryError parsing full XML~8-12 GB needed in memoryUse ET.iterparse() with elem.clear()
Slow startup (>120s)Parsing 6 GB XMLParse once, build index dict; avoid re-parsing
Metabolite not found by nameCase sensitivity or synonymNormalize to lowercase; try HMDB ID directly
Empty spectra for a metaboliteNot all entries have spectra (~10-15% coverage)Check coverage table; use METLIN or MassBank for more spectra
Missing concentration dataLimited to well-studied clinical metabolites (~20%)Cross-reference with MetaboAnalyst or literature
Duplicate entries for same compoundSecondary accessions (HMDB00XXXXX vs HMDB0000XXXX)Use accession (primary), not secondary_accessions
ET.iterparse missing dataPremature elem.clear()Only clear after extracting all needed fields from the element

Bundled Resources

Self-contained entry. The original reference file hmdb_data_fields.md (268 lines, field catalog with XML element names and descriptions) is consolidated into the Key Concepts "HMDB XML Entry Structure" table and the "Data Field Completeness" table above. The field catalog's per-element XML tag names are demonstrated in Core API code blocks. Omitted from original: web interface descriptions (not programmatic).

  • drugbank-database-access -- Drug-specific data (interactions, targets, pharmacology) with similar local XML parsing pattern
  • pubchem-compound-search -- Live compound property lookups without downloading; PubChemPy REST API
  • kegg-database -- Pathway and compound database with REST API for cross-referencing HMDB pathway hits
  • matchms-spectral-matching -- Spectral similarity matching against reference libraries; complementary to HMDB spectral data
  • pyopenms-mass-spectrometry -- Full LC-MS/MS processing pipeline; use HMDB for metabolite identification step

References

© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/proteomics-protein-engineering/hmdb-database of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Hmdb Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hmdb Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hmdb Database this skilljaechang-hits/SciAgent-Skills3711 repos~6.1kAutomated safety check: PassCC-BY-4.0
Bio Ensembl RESTGPTomics/bioSkills1.2k2 repos~3.6kAutomated safety check: PassMIT
Pride FetchClawBio/ClawBio1.2k—~4.2kAutomated safety check: PassMIT
Ensembl Databaseaipoch/medical-research-skills2k—~1.5kAutomated safety check: PassMIT
UniProt Database Accessdavila7/claude-code-templates32k14 repos~1.7kAutomated safety check: PassMIT
Singlecell Portalaipoch/medical-research-skills2k—~1.2kAutomated safety check: PassMIT

Similar skills

  • Bio Ensembl REST

    GPTomics/bioSkills

    Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Pride Fetch

    ClawBio/ClawBio

    Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.

    1.2k GitHub stars~4.2k tokensUpdated today
    Research & ScienceAuto-check passed
  • Ensembl Database

    aipoch/medical-research-skills

    Access Ensembl REST API for vertebrate genomic data; use when you need gene/ID lookups, sequence retrieval, variant effect prediction (VEP), or homology/assembly coordinate mapping.

    2k GitHub stars~1.5k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    32k GitHub starsUsed in 14 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • Singlecell Portal

    aipoch/medical-research-skills

    Programmatically query public single-cell study metadata from the Broad Institute Single Cell Portal REST API when you need to search and filter datasets by organism, tissue, disease, or cell type…

    2k GitHub stars~1.2k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • Ena Database

    aipoch/medical-research-skills

    Access the European Nucleotide Archive (ENA) via REST APIs and FTP/Aspera to search and retrieve sequences, raw reads (FASTQ), assemblies, and metadata when you have accession IDs or need…

    2k GitHub stars~1.9k tokensUpdated 22 days ago
    Backend & APIsAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    371 GitHub stars~4k tokensUpdated 10 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    371 GitHub stars~3.2k tokensUpdated 10 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    371 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    371 GitHub stars~6.9k tokensUpdated 10 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub stars~2.3k tokensUpdated 10 days ago
    Auto-check passed

Questions about Hmdb Database

What does Hmdb Database do?

Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping. Hmdb Database is an agent skill from jaechang-hits/SciAgent-Skills. Parse HMDB (Human Metabolome Database) local XML for metabolite info, chemical properties, biological context, disease links, spectra, and cross-DB mapping.

When should I use Hmdb Database?

Hmdb Database fits situations like: tasks that involve Bioinformatics; tasks that involve REST APIs.

How do I install Hmdb Database in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/hmdb-database in jaechang-hits/SciAgent-Skills) into .claude/skills/hmdb-database in your project. Claude Code loads it when a task matches its description.

How do I install Hmdb Database in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/hmdb-database in jaechang-hits/SciAgent-Skills) into .agents/skills/hmdb-database in your project. Codex loads it when a task matches its description.

Can I use Hmdb Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill hmdb-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hmdb-database, .gemini/skills/hmdb-database, .github/skills/hmdb-database and .opencode/skills/hmdb-database in your project.

What does Hmdb Database need to run?

Going by SKILL.md and its folder, Hmdb Database needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Hmdb Database access the network?

SKILL.md names 3 domains. In commands or code: hmdb.ca; the agent is likely to contact it when it follows the instructions. As links in the text: doi.org and metaboanalyst.ca. This is read from the text; nothing was executed.

Is Hmdb Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hmdb Database use?

Hmdb Database is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hmdb Database use?

About 6.1k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hmdb Database?

Skills that share tags, products or a category with Hmdb Database: Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars), Pride Fetch (ClawBio/ClawBio, 1.2k stars), Ensembl Database (aipoch/medical-research-skills, 2k stars) and UniProt Database Access (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hmdb Database?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 371 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.