Agent skill

Metabolomics Workbench Database

by jaechang-hits in jaechang-hits/SciAgent-Skills

Query Metabolomics Workbench REST API (4,200+ NIH studies) for metabolite ID, study discovery, RefMet standardization, m/z precursor searches, and gene/protein annotations.

CC-BY-4.0Auto-check passedResearch & Science

Install Metabolomics Workbench Database

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill metabolomics-workbench-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills metabolomics-workbench-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/metabolomics-workbench-database .claude/skills/metabolomics-workbench-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
metabolomics-workbench-database
GitHub stars
374
Used in
1 other repo
Token cost
~5.3k tokens
SKILL.md length
1,021 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Query Metabolomics Workbench REST API (4,200+ NIH studies) for metabolite ID, study discovery, RefMet standardization, m/z precursor searches, and gene/protein annotations.

  • Works in 6 steps: Use refmet/match first for free-text… → moverz is TSV-only. Parse with… → Don't append /json to study/.../summary.… → …
  • Tasks that involve REST APIs
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 9 more sections
  • Calls pip; reaches metabolomicsworkbench.org

What it does

Metabolomics Workbench Database is an agent skill from jaechang-hits/SciAgent-Skills. Query Metabolomics Workbench REST API (4,200+ NIH studies) for metabolite ID, study discovery, RefMet standardization, m/z precursor searches, and gene/protein annotations. Quirks: compound inputitem rejects name (use pubchemcid/keggid/inchikey/etc.); free-text → compound is a two-step refmet/match→refmet/name flow; moverz endpoint returns TSV text, not JSON. Use hmdb-database for local XML; pubchem-compound-search for general compound lookup.

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering REST APIs, Bioinformatics and CSV and tabular files. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.

When your agent uses it

  • Tasks that involve REST APIs
  • Tasks that involve Bioinformatics
  • Tasks that involve CSV and tabular files

Example prompts

  • “/metabolomics-workbench-database”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Use refmet/match first for free-text input. compound/name/... is rejected by the server (name is not an allowed input_item).
  2. moverz is TSV-only. Parse with pd.read_csv(io.StringIO(r.text), sep="\t") — never call .json() on the response.
  3. Don't append /json to study/.../summary. The default is JSON; the suffix flips the response to TSV.
  4. refmet/name/{x}/all needs the canonical RefMet name. If you have user text, run it through refmet/match first.
  5. Compound results paged by formula come keyed '1','2',... — iterate dict.values() or pass to pd.DataFrame.from_dict(orient="index").
  6. Always time.sleep(0.3) in batch loops — no rate limit is published but the server is shared.

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • metabolomicsworkbench.org

    Also links to:

    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Metabolomics Workbench Database loads about 5.3k tokens when it runs. Until then it costs about 121 tokens; SKILL.md has 1,021 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~121
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 1,021 words, ~5,312 tokens.

Download SKILL.mdSave it as .claude/skills/metabolomics-workbench-database/SKILL.md (or your agent's skills folder).
name
metabolomics-workbench-database
description
Query Metabolomics Workbench REST API (4,200+ NIH studies) for metabolite ID, study discovery, RefMet standardization, m/z precursor searches, and gene/protein annotations. Quirks: compound input_item rejects `name` (use pubchem_cid/kegg_id/inchi_key/etc.); free-text → compound is a two-step refmet/match→refmet/name flow; moverz endpoint returns TSV text, not JSON. Use hmdb-database for local XML; pubchem-compound-search for general compound lookup.
license
CC-BY-4.0

Metabolomics Workbench Database — REST API Access

Overview

The Metabolomics Workbench (MW) REST API at https://www.metabolomicsworkbench.org/rest/ exposes 4,200+ metabolomics studies hosted at UCSD under NIH Common Fund sponsorship. URL pattern is /{context}/{input_item}/{input_value}/{output_item}/{format}. Contexts include compound, refmet, moverz, study, analysis, metabolite, gene, protein. Notable quirks discovered live:

  • compound/name/{x} is rejected — name is not an allowed input_item. Use pubchem_cid, kegg_id, inchi_key, hmdb_id, regno, lm_id, formula, smiles, or abbrev. For free-text input, go through refmet/match/{x} first.
  • refmet/name/{x}/all requires the exact RefMet name (e.g. Glucose, not D-glucose); use refmet/match/{x} for fuzzy normalisation first.
  • moverz/{REFMET|LIPIDS|MB}/{mz}/{ion}/{tol}/txt returns TSV text (no JSON variant).
  • The metstat/filter/... endpoint shown in older examples returns [] — replace with study/{context}/{value}/summary (or /metabolites) + client-side filtering.

No authentication required.

When to Use

  • Searching metabolite records by PubChem CID, KEGG ID, InChIKey, HMDB ID, formula, or SMILES
  • Discovering studies by species, disease, last_name, institute, analysis_type, or polarity
  • Standardising metabolite names to RefMet nomenclature for cross-study integration
  • Identifying unknown compounds from MS m/z values with adduct-aware matching (moverz)
  • Retrieving experimental metabolite tables (analyses, abundances) from published studies
  • Querying gene/protein annotations linked to metabolomics pathways
  • Downloading raw mwTab files for local analysis
  • For local 220K-metabolite XML parsing with NMR/MS spectra use hmdb-database instead
  • For live 110M-compound property lookups use pubchem-compound-search instead

Prerequisites

  • Python packages: requests, pandas
  • No API key required: publicly accessible
  • Rate limits: MW does not enforce strict limits; add time.sleep(0.3) between bulk requests
  • Base URL: https://www.metabolomicsworkbench.org/rest
bash
pip install requests pandas

Quick Start

python
import requests

BASE = "https://www.metabolomicsworkbench.org/rest"

# Two-step free-text → compound (the API rejects compound/name/...)
def lookup_by_name(name):
    # 1) Normalise to RefMet name
    r = requests.get(f"{BASE}/refmet/match/{name}", timeout=30)
    r.raise_for_status()
    refmet = r.json()
    if not refmet.get("refmet_name"):
        return None
    # 2) Pull full compound record by RefMet name (or by pubchem_cid)
    r2 = requests.get(f"{BASE}/refmet/name/{refmet['refmet_name']}/all", timeout=30)
    rec = r2.json() if r2.json() else {}
    return rec if isinstance(rec, dict) else None

c = lookup_by_name("glucose")
print(f"{c['name']}: formula={c['formula']}, PubChem CID={c['pubchem_cid']}, "
      f"InChIKey={c['inchi_key']}")
# Glucose: formula=C6H12O6, PubChem CID=5793, InChIKey=WQZGKKKJIJFFOK-GASJEMHNSA-N

Core API

Module 1: Compound Queries

compound/{input_item}/{input_value}/all/json — input_item must be one of regno, formula, inchi_key, lm_id, pubchem_cid, hmdb_id, kegg_id, smiles, abbrev. The legacy name input is rejected by the server.

python
import requests

BASE = "https://www.metabolomicsworkbench.org/rest"

# By PubChem CID
r = requests.get(f"{BASE}/compound/pubchem_cid/5793/all/json", timeout=30)
glucose = r.json()
print(f"PubChem 5793 -> {glucose['name']}, formula={glucose['formula']}, "
      f"HMDB={glucose.get('hmdb_id')}, KEGG={glucose.get('kegg_id')}")

# By KEGG ID
r = requests.get(f"{BASE}/compound/kegg_id/C00031/all/json", timeout=30)
print("KEGG C00031 ->", r.json()["name"])

# By InChIKey
r = requests.get(f"{BASE}/compound/inchi_key/WQZGKKKJIJFFOK-GASJEMHNSA-N/all/json", timeout=30)
print("InChIKey -> regno:", r.json()["regno"])
python
# Compound by formula returns a paged dict (multiple matches)
import requests
BASE = "https://www.metabolomicsworkbench.org/rest"
r = requests.get(f"{BASE}/compound/formula/C6H12O6/all/json", timeout=30)
matches = r.json()
print(f"Compounds with formula C6H12O6: {len(matches)}")
for k in list(matches)[:3]:
    print(f"  regno={matches[k]['regno']}  name={matches[k]['name']}")
Module 2: Study Discovery

study/{input_item}/{input_value}/{output} — input_item includes study_id, study_title, last_name, institute, analysis_id, metabolite_id, kegg_id, refmet_name. output includes summary, metabolites, factors, data, available_studies, species, disease. summary for study_id returns a dict (keyed by accession when multiple); for last_name/institute it returns a list.

python
import requests, pandas as pd

BASE = "https://www.metabolomicsworkbench.org/rest"

# Single-study summary — `study/study_id/{id}/summary` returns a flat dict
# (keys: study_id, study_title, species, institute, analysis_type, ...)
r = requests.get(f"{BASE}/study/study_id/ST000001/summary", timeout=30)
s = r.json()
print(f"{s['study_id']}: {s['study_title'][:60]}")
print(f"  Species : {s.get('species')}  Institute: {s.get('institute')}")
print(f"  Submit  : {s.get('submission_date')}")
python
# Studies that detected a metabolite — `study/refmet_name/{x}/summary` returns
# a thin index of (refmet_name, kegg_id, study_id) rows. Chain study_id → full summary
# to get title and species.
import requests, pandas as pd
BASE = "https://www.metabolomicsworkbench.org/rest"
r = requests.get(f"{BASE}/study/refmet_name/Glucose/summary", timeout=60)
d = r.json()
rows = list(d.values()) if isinstance(d, dict) else d
print(f"Studies referencing 'Glucose': {len(rows)}")
print(pd.DataFrame(rows).head(5).to_string(index=False))
# refmet_name kegg_id  study_id
#     Glucose  C00031  ST000001
#     Glucose  C00031  ST000002
#     ...
Module 3: RefMet Standardisation

refmet/match/{user_text} is a fuzzy normaliser — returns the standard RefMet record (no pubchem_cid/kegg_id though). refmet/name/{exact_refmet_name}/all returns the full record including IDs. Use them as a two-step pipeline.

python
import requests

BASE = "https://www.metabolomicsworkbench.org/rest"

def normalise_to_refmet(user_text):
    r = requests.get(f"{BASE}/refmet/match/{user_text}", timeout=30)
    r.raise_for_status()
    m = r.json()
    if not m or not m.get("refmet_name"):
        return None
    return m["refmet_name"]

def refmet_full(refmet_name):
    r = requests.get(f"{BASE}/refmet/name/{refmet_name}/all", timeout=30)
    r.raise_for_status()
    rec = r.json()
    return rec if isinstance(rec, dict) and rec else None

name = normalise_to_refmet("alpha-D-glucose")  # -> 'Glucose'
print(f"Normalised: {name}")
rec = refmet_full(name)
print(f"  PubChem CID : {rec['pubchem_cid']}")
print(f"  InChIKey    : {rec['inchi_key']}")
print(f"  Super class : {rec['super_class']} / {rec['main_class']} / {rec['sub_class']}")
Module 4: Study Filtering (replaces broken metstat)

The older metstat/filter/... endpoint returns []. Use the study context endpoints with client-side filtering instead.

python
import requests, pandas as pd

BASE = "https://www.metabolomicsworkbench.org/rest"

def study_ids_for_metabolite(refmet_name):
    """Return the study_id list that report a given RefMet name."""
    r = requests.get(f"{BASE}/study/refmet_name/{refmet_name}/summary", timeout=60)
    r.raise_for_status()
    d = r.json()
    rows = list(d.values()) if isinstance(d, dict) else d
    return sorted({row["study_id"] for row in rows if row.get("study_id")})

def study_summary(study_id):
    """Pull full summary (title, species, institute, dates) for one study_id.
    Response is a flat dict with keys study_id/study_title/species/institute/..."""
    return requests.get(f"{BASE}/study/study_id/{study_id}/summary", timeout=30).json()

# Find glucose studies, then enrich the first few
ids = study_ids_for_metabolite("Glucose")
print(f"Studies referencing 'Glucose': {len(ids)}")
rows = [study_summary(sid) for sid in ids[:5]]
df = pd.DataFrame(rows)
print(df[["study_id", "study_title", "species"]].head().to_string(index=False))
Module 5: m/z Precursor Search (moverz)

moverz/{REFMET|LIPIDS|MB}/{mz}/{ion}/{tolerance}/txt returns tab-separated text (not JSON). The first DB selector (REFMET, LIPIDS, or MB) is required — mz as the first segment is rejected.

python
import requests, io, pandas as pd

BASE = "https://www.metabolomicsworkbench.org/rest"

def moverz_search(db, mz, ion, tolerance=0.005):
    """Search precursor m/z in REFMET / LIPIDS / MB and return a DataFrame.
    Response is TSV text — no JSON variant."""
    assert db in {"REFMET", "LIPIDS", "MB"}
    r = requests.get(f"{BASE}/moverz/{db}/{mz}/{ion}/{tolerance}/txt", timeout=30)
    r.raise_for_status()
    return pd.read_csv(io.StringIO(r.text), sep="\t")

df = moverz_search("REFMET", 180.063, "M+H", 0.005)
print(f"Candidates at m/z 180.063 [M+H]+ (REFMET): {len(df)}")
print(df.head(5).to_string(index=False))
python
# Same query against the LIPIDS database
import requests, io, pandas as pd
BASE = "https://www.metabolomicsworkbench.org/rest"
r = requests.get(f"{BASE}/moverz/LIPIDS/760.585/M+H/0.01/txt", timeout=30)
df_lipids = pd.read_csv(io.StringIO(r.text), sep="\t")
print(f"Lipid candidates at m/z 760.585: {len(df_lipids)}")
print(df_lipids.head(3).to_string(index=False))
Module 6: Genes and Proteins
python
import requests

BASE = "https://www.metabolomicsworkbench.org/rest"

r = requests.get(f"{BASE}/gene/gene_symbol/HMGCR/all", timeout=30)
g = r.json()
print(f"{g['gene_symbol']} -> MGP: {g.get('mgp_id')}  ({g.get('gene_name', '')[:60]})")

# Protein by UniProt accession
r2 = requests.get(f"{BASE}/protein/uniprot_id/P04035/all", timeout=30)
p = r2.json()
print(f"UniProt P04035: {p.get('protein_name', '')[:60]}  organism={p.get('organism')}")

Key Concepts

Allowed input_item Values per Context
ContextValid input_itemNotes
compoundregno, formula, inchi_key, lm_id, pubchem_cid, hmdb_id, kegg_id, smiles, abbrevname is rejected — go via refmet/match first
refmetmatch, name, formula, exactmass, inchi_key, pubchem_cid, regnomatch is fuzzy; name requires the canonical RefMet name
studystudy_id, study_title, last_name, institute, analysis_id, metabolite_id, kegg_id, refmet_namesummary for study_id is a dict keyed by accession; for last_name/institute it's a list
moverz(path) REFMET / LIPIDS / MBFirst segment is the DB, not mz
genegene_id, gene_symbol, gene_name, mgp_idReturns a dict
proteinmgp_id, gene_id, uniprot_id, gene_symbolReturns a dict
Output Type Conventions
  • output=summary returns a dict when the input identifier is unique (e.g. study_id), a list when it isn't (e.g. last_name).
  • Appending /json to study/.../summary flips the response to TSV. Omit the format suffix — JSON is the default.
  • moverz only emits /txt (TSV); there is no JSON variant.
refmet/match vs refmet/name
  • refmet/match/{user_text} — fuzzy. Always returns a single dict with refmet_name, formula, exactmass, classification. Does not include pubchem_cid/inchi_key.
  • refmet/name/{exact_refmet_name}/all — requires the canonical RefMet name. Returns the full record including IDs. Returns an empty list if the name isn't canonical.

Common Workflows

Workflow 1: Free-Text → Full Compound Record
python
import requests, pandas as pd

BASE = "https://www.metabolomicsworkbench.org/rest"

def resolve_to_compound(user_text):
    # 1) Normalise via refmet/match
    rm = requests.get(f"{BASE}/refmet/match/{user_text}", timeout=30).json()
    if not rm.get("refmet_name"):
        return None
    name = rm["refmet_name"]
    # 2) Fetch full RefMet record (includes pubchem_cid / inchi_key)
    full = requests.get(f"{BASE}/refmet/name/{name}/all", timeout=30).json()
    if not isinstance(full, dict) or not full:
        return None
    # 3) Optionally pull the matching compound record via PubChem CID
    cid = full.get("pubchem_cid")
    compound = None
    if cid:
        compound = requests.get(f"{BASE}/compound/pubchem_cid/{cid}/all/json",
                                timeout=30).json()
    return {
        "input": user_text,
        "refmet_name": name,
        "formula": full.get("formula"),
        "pubchem_cid": cid,
        "hmdb_id": compound.get("hmdb_id") if compound else None,
        "kegg_id": compound.get("kegg_id") if compound else None,
        "inchi_key": full.get("inchi_key"),
    }

queries = ["glucose", "alpha-D-glucose", "L-tyrosine", "cholesterol"]
df = pd.DataFrame([resolve_to_compound(q) for q in queries])
print(df.to_string(index=False))
df.to_csv("name_resolution.csv", index=False)
Workflow 2: Annotate MS Hit List

Goal: Take a list of measured m/z values and assign RefMet candidate compounds.

python
import requests, io, pandas as pd, time

BASE = "https://www.metabolomicsworkbench.org/rest"

def annotate_peaks(mz_values, ion="M+H", tolerance=0.005):
    out = []
    for mz in mz_values:
        r = requests.get(f"{BASE}/moverz/REFMET/{mz}/{ion}/{tolerance}/txt", timeout=30)
        if r.status_code != 200 or not r.text.strip():
            time.sleep(0.3); continue
        df = pd.read_csv(io.StringIO(r.text), sep="\t")
        for _, row in df.iterrows():
            out.append({
                "query_mz": mz,
                "matched_mz": row["Matched m/z"],
                "delta": row["Delta"],
                "name": row["Name"],
                "formula": row["Formula"],
                "ion": row["Ion"],
                "main_class": row.get("Main class"),
            })
        time.sleep(0.3)
    return pd.DataFrame(out)

peaks = [180.063, 166.086, 90.055]   # glucose, phenylalanine, alanine
df_ann = annotate_peaks(peaks)
print(df_ann.head(10).to_string(index=False))
df_ann.to_csv("ms_annotations.csv", index=False)
Workflow 3: Find Studies Detecting a Metabolite
python
import requests, pandas as pd

BASE = "https://www.metabolomicsworkbench.org/rest"

def studies_with(refmet_name, enrich_n=20):
    """Return a DataFrame: study_id rows that report the metabolite, enriched with
    title + species for the first `enrich_n` IDs (via study/study_id/.../summary)."""
    r = requests.get(f"{BASE}/study/refmet_name/{refmet_name}/summary", timeout=60)
    r.raise_for_status()
    d = r.json()
    rows = list(d.values()) if isinstance(d, dict) else d
    ids = sorted({row["study_id"] for row in rows if row.get("study_id")})
    enriched = []
    for sid in ids[:enrich_n]:
        enriched.append(requests.get(
            f"{BASE}/study/study_id/{sid}/summary", timeout=30).json())
    return pd.DataFrame(enriched)

df = studies_with("Glucose", enrich_n=20)
print(f"Glucose-detecting studies (first 20 enriched): {len(df)}")
print(df.groupby("species").size().sort_values(ascending=False).head(8))
Show full SKILL.md (415 more words)Show less

Key Parameters

ParameterEndpointDefaultRange / OptionsEffect
contextpathrequiredcompound, refmet, moverz, study, analysis, metabolite, gene, proteinAPI context selector
input_itempathrequireddepends on context (see "Allowed input_item Values per Context")Identifier type
input_valuepathrequiredstringThe actual identifier or value
output_itempathallall, summary, metabolites, factors, data, etc.What aspect to return
formatpath(varies)json, txtmoverz only emits txt; do NOT append /json to study/.../summary
mz / ion / tolerancemoverz pathrequiredfloat / M+H, M-H, M+Na, M+K, etc. / floatMass tolerance in Da

Best Practices

  1. Use refmet/match first for free-text input. compound/name/... is rejected by the server (name is not an allowed input_item).
  2. moverz is TSV-only. Parse with pd.read_csv(io.StringIO(r.text), sep="\t") — never call .json() on the response.
  3. Don't append /json to study/.../summary. The default is JSON; the suffix flips the response to TSV.
  4. refmet/name/{x}/all needs the canonical RefMet name. If you have user text, run it through refmet/match first.
  5. Compound results paged by formula come keyed '1','2',... — iterate dict.values() or pass to pd.DataFrame.from_dict(orient="index").
  6. Always time.sleep(0.3) in batch loops — no rate limit is published but the server is shared.

Common Recipes

Recipe: Cross-Database ID Mapping (PubChem ↔ KEGG ↔ HMDB)
python
import requests, pandas as pd

BASE = "https://www.metabolomicsworkbench.org/rest"

def cross_refs(refmet_name):
    rm = requests.get(f"{BASE}/refmet/name/{refmet_name}/all", timeout=30).json()
    if not isinstance(rm, dict) or not rm:
        return None
    cid = rm.get("pubchem_cid")
    if not cid:
        return {"refmet": refmet_name, "pubchem_cid": None}
    c = requests.get(f"{BASE}/compound/pubchem_cid/{cid}/all/json", timeout=30).json()
    return {"refmet": refmet_name,
            "pubchem_cid": cid,
            "kegg_id": c.get("kegg_id"),
            "hmdb_id": c.get("hmdb_id"),
            "inchi_key": rm.get("inchi_key")}

df = pd.DataFrame([cross_refs(n) for n in ["Glucose", "L-Tyrosine", "Cholesterol"]])
print(df.to_string(index=False))
Recipe: Pull a Study's Metabolite Table
python
import requests, pandas as pd

BASE = "https://www.metabolomicsworkbench.org/rest"
r = requests.get(f"{BASE}/study/study_id/ST000001/metabolites", timeout=60)
r.raise_for_status()
rows = list(r.json().values())
df = pd.DataFrame(rows)
print(f"ST000001 metabolites: {len(df)}")
print(df[["analysis_id", "analysis_summary", "metabolite_name", "refmet_name"]].head(5).to_string(index=False))
Recipe: Gene → Metabolomics Pathways
python
import requests

BASE = "https://www.metabolomicsworkbench.org/rest"
g = requests.get(f"{BASE}/gene/gene_symbol/HMGCR/all", timeout=30).json()
print({k: g.get(k) for k in ["gene_symbol", "gene_id", "mgp_id",
                              "gene_name", "gene_synonyms"]})

Troubleshooting

ProblemCauseSolution
"This input item (name) is not allowed..."compound/name/... is rejectedUse one of the allowed input_item values (pubchem_cid, kegg_id, inchi_key, hmdb_id, etc.); for free text, go through refmet/match/{x} first
JSONDecodeError on moverz responsemoverz returns TSV text, not JSONParse with pd.read_csv(io.StringIO(r.text), sep="\t")
Empty list from refmet/name/{x}/allNeed the canonical RefMet name (case-sensitive)Normalise via refmet/match/{user_text} first, then plug refmet_name into refmet/name/.../all
metstat/filter/... returns []Endpoint syntax is non-functionalUse `study/{refmet_name
study/.../summary returns TSV instead of JSONThe /json suffix flips to TSVDrop the trailing /json — JSON is the default
Compound query by formula returns a dict, not a listServer pages multiple matches as {'1': {...}, '2': {...}}Iterate dict.values() (or pd.DataFrame.from_dict(d, orient='index'))
refmet/match response lacks pubchem_cidmatch returns the lightweight recordUse refmet/name/{refmet_name}/all for the full record
  • hmdb-database — Local HMDB XML (220K metabolites, NMR/MS spectra, disease links) for offline queries
  • pubchem-compound-search — General compound property lookups (110M+ compounds) via PubChemPy
  • kegg-database — Pathway and orthology data complementary to MW's study/metabolite hits
  • chembl-database-bioactivity — Bioactivity data for the same compounds

References

© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/proteomics-protein-engineering/metabolomics-workbench-database of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Metabolomics Workbench Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Metabolomics Workbench Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Metabolomics Workbench Database this skilljaechang-hits/SciAgent-Skills3741 repos~5.3kAutomated safety check: PassCC-BY-4.0
Pride FetchClawBio/ClawBio1.2k—~4.2kAutomated safety check: PassMIT
Bio Ensembl RESTGPTomics/bioSkills1.2k2 repos~3.6kAutomated safety check: PassMIT
Ensembl Databaseaipoch/medical-research-skills1.9k—~1.5kAutomated safety check: PassMIT
UniProt Database Accessdavila7/claude-code-templates33k14 repos~1.7kAutomated safety check: PassMIT
Bulkrna Cosinor RhythmTianGzlab/OmicsClaw161—~840Automated safety check: PassApache-2.0

Similar skills

  • Pride Fetch

    ClawBio/ClawBio

    Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.

    1.2k GitHub stars~4.2k tokensUpdated today
    Research & ScienceAuto-check passed
  • Bio Ensembl REST

    GPTomics/bioSkills

    Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Research & ScienceAuto-check passed
  • Ensembl Database

    aipoch/medical-research-skills

    Access Ensembl REST API for vertebrate genomic data; use when you need gene/ID lookups, sequence retrieval, variant effect prediction (VEP), or homology/assembly coordinate mapping.

    1.9k GitHub stars~1.5k tokensUpdated 23 days ago
    Research & ScienceAuto-check passed
  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    33k GitHub starsUsed in 14 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Ukb Ppp Region Fetch

    ClawBio/ClawBio

    Fetch a regional slice of plasma pQTL summary statistics from the UK Biobank Pharma Proteomics Project (UKB-PPP; Sun 2023 Nature) for a specific (protein, ancestry) measurement.

    1.2k GitHub stars~4.6k tokensUpdated today
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 11 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 11 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 11 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 11 days ago
    Auto-check passed

Questions about Metabolomics Workbench Database

What does Metabolomics Workbench Database do?

Query Metabolomics Workbench REST API (4,200+ NIH studies) for metabolite ID, study discovery, RefMet standardization, m/z precursor searches, and gene/protein annotations. Metabolomics Workbench Database is an agent skill from jaechang-hits/SciAgent-Skills. Query Metabolomics Workbench REST API (4,200+ NIH studies) for metabolite ID, study discovery, RefMet standardization, m/z precursor searches, and gene/protein annotations.

When should I use Metabolomics Workbench Database?

Metabolomics Workbench Database fits situations like: tasks that involve REST APIs; tasks that involve Bioinformatics; tasks that involve CSV and tabular files.

How do I install Metabolomics Workbench Database in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill metabolomics-workbench-database -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/metabolomics-workbench-database in jaechang-hits/SciAgent-Skills) into .claude/skills/metabolomics-workbench-database in your project. Claude Code loads it when a task matches its description.

How do I install Metabolomics Workbench Database in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill metabolomics-workbench-database -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/metabolomics-workbench-database in jaechang-hits/SciAgent-Skills) into .agents/skills/metabolomics-workbench-database in your project. Codex loads it when a task matches its description.

Can I use Metabolomics Workbench Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill metabolomics-workbench-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/metabolomics-workbench-database, .gemini/skills/metabolomics-workbench-database, .github/skills/metabolomics-workbench-database and .opencode/skills/metabolomics-workbench-database in your project.

What does Metabolomics Workbench Database need to run?

Going by SKILL.md and its folder, Metabolomics Workbench Database needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Metabolomics Workbench Database access the network?

SKILL.md names 2 domains. In commands or code: metabolomicsworkbench.org; the agent is likely to contact it when it follows the instructions. As links in the text: doi.org. This is read from the text; nothing was executed.

Is Metabolomics Workbench Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Metabolomics Workbench Database use?

Metabolomics Workbench Database is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Metabolomics Workbench Database use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Metabolomics Workbench Database?

Skills that share tags, products or a category with Metabolomics Workbench Database: Pride Fetch (ClawBio/ClawBio, 1.2k stars), Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars), Ensembl Database (aipoch/medical-research-skills, 1.9k stars) and UniProt Database Access (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Metabolomics Workbench Database?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.