Agent skill

Pubchem Compound Search

by jaechang-hits in jaechang-hits/SciAgent-Skills

Query PubChem (110M+ compounds) directly via the PUG-REST/JSON API with plain requests — no SDK install required.

CC-BY-4.0Auto-check passedResearch & Science

Install Pubchem Compound Search

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill pubchem-compound-search -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills pubchem-compound-search --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pubchem-compound-search .claude/skills/pubchem-compound-search && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pubchem-compound-search
GitHub stars
374
Used in
1 other repo
Token cost
~7k tokens
SKILL.md length
1,552 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Query PubChem (110M+ compounds) directly via the PUG-REST/JSON API with plain requests — no SDK install required.

  • Works in 8 steps: Always go through cids/JSON first when… → Batch properties, never loop them.… → URL-encode every SMILES. Use… → …
  • Tasks that involve Drug discovery and cheminformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 9 more sections
  • Calls pip; reaches pubchem.ncbi.nlm.nih.gov

What it does

Pubchem Compound Search is an agent skill from jaechang-hits/SciAgent-Skills. Query PubChem (110M+ compounds) directly via the PUG-REST/JSON API with plain requests — no SDK install required. Search by name/CID/SMILES/InChIKey/formula, retrieve properties (MW, XLogP, TPSA, H-bond counts), do similarity/substructure searches with async ListKey polling, fetch synonyms, descriptions, assay summaries, and download SDF/PNG. For local cheminformatics use rdkit; for bioactivity-centric workflows use chembl-database-bioactivity.

Its SKILL.md is about 7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Drug discovery and cheminformatics. It works with RDKit. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.

When your agent uses it

  • Tasks that involve Drug discovery and cheminformatics

Example prompts

  • “/pubchem-compound-search”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Always go through cids/JSON first when starting from a name or external identifier. The name→CID resolution and the CID→property lookup…
  2. Batch properties, never loop them. compound/cid/2244,3672,2157,.../property/MolecularWeight,XLogP,.../JSON accepts up to ~200 CIDs and any…
  3. URL-encode every SMILES. Use urllib.parse.quote(smiles, safe=""). Bare SMILES with =, #, (, ), [, ] will sometimes work but breaks…
  4. Treat similarity/substructure/formula as async. Branch on "Waiting" in response and poll listkey rather than re-issuing the search…
  5. Throttle ≤ 5 req/sec, ≤ 400/min. Insert time.sleep(0.25) in any tight loop. HTTP 503 means you tripped the limit — wait 10s and reduce…
  6. Cast MolecularWeight to float. It's returned as a string ("180.16") for full decimal fidelity. Comparing strings against numeric…
  7. For 100+ CIDs use POST. GET URLs over ~2000 chars get truncated by some HTTP proxies. PubChem also accepts POST with cid in the form body…
  8. 2025 SMILES property rename. PubChem renamed two SMILES properties in the 2025 PUG-REST schema: old IsomericSMILES (with stereo) → SMILES…

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pubchem.ncbi.nlm.nih.gov

    Also links to:

    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pubchem Compound Search loads about 7k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 1,552 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 1,552 words, ~7,042 tokens.

Download SKILL.mdSave it as .claude/skills/pubchem-compound-search/SKILL.md (or your agent's skills folder).
name
pubchem-compound-search
description
Query PubChem (110M+ compounds) directly via the PUG-REST/JSON API with plain `requests` — no SDK install required. Search by name/CID/SMILES/InChIKey/formula, retrieve properties (MW, XLogP, TPSA, H-bond counts), do similarity/substructure searches with async ListKey polling, fetch synonyms, descriptions, assay summaries, and download SDF/PNG. For local cheminformatics use rdkit; for bioactivity-centric workflows use chembl-database-bioactivity.
license
CC-BY-4.0

Overview

PubChem (NCBI) is the largest freely available chemical database — 110M+ compounds, 280M+ substances, and millions of bioassay records. Its PUG-REST JSON API is the canonical programmatic surface, and every example here uses it directly via plain requests. The Python pubchempy wrapper is not required; the PUG-REST URL grammar is small enough that direct calls are more transparent, easier to retry/cache, and avoid sandbox dependency issues (the library is not in TOOL_STATUS.md).

The URL pattern is fixed and predictable:

https://pubchem.ncbi.nlm.nih.gov/rest/pug/<input>/<operation>/<output>
  • <input> = compound/{name,cid,smiles,inchikey,formula}/<value>
  • <operation> = cids, property/<list>, synonyms, description, assaysummary, JSON (full record), SDF, PNG
  • <output> = JSON, CSV, TXT, SDF, PNG

For long-running operations (similarity, substructure, formula) the API returns HTTP 202 + {"Waiting": {"ListKey": "..."}}; poll compound/listkey/{key}/cids/JSON until it returns IdentifierList. The skill handles this pattern in Module 4.

When to Use

  • Looking up a compound by name, SMILES, InChIKey, or formula to get its PubChem CID
  • Retrieving molecular properties (molecular weight, XLogP, TPSA, H-bond donor/acceptor counts, rotatable bonds, formula, IUPAC name) for one or many CIDs in a single request
  • Finding structurally similar compounds via Tanimoto similarity (async ListKey poll)
  • Searching for compounds containing a substructure / pharmacophore motif
  • Fetching every synonym / trade name / CAS number for a CID
  • Pulling assay summary tables (active/inactive screening results) for a compound
  • Converting between identifier formats (name ↔ CID ↔ SMILES ↔ InChI ↔ InChIKey) in one API call
  • Downloading 2D SDF or PNG structures for figures or downstream RDKit work
  • For local cheminformatics (fingerprints, descriptors, 3D conformers, scaffold extraction) use rdkit
  • For deeper bioactivity / target-binding data (IC50, Ki against specific targets) use chembl-database-bioactivity

Prerequisites

  • Python packages: requests, pandas — both already in standard environments
  • No API key required — PubChem is fully public
  • Rate limits: max 5 requests/second and 400 requests/minute per IP. Throttle with time.sleep(0.25) in loops; return code 503 means you tripped the limit.

If you are inside a pixi/conda environment that already provides requests and pandas, skip the install and invoke scripts with pixi run python ....

bash
pip install requests pandas

Quick Start

python
import requests

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

# name → CID
cid = requests.get(f"{BASE}/compound/name/aspirin/cids/JSON").json()["IdentifierList"]["CID"][0]

# CID → properties (single call, many fields)
r = requests.get(
    f"{BASE}/compound/cid/{cid}/property/"
    "MolecularWeight,XLogP,TPSA,HBondDonorCount,HBondAcceptorCount,SMILES,IUPACName/JSON")
p = r.json()["PropertyTable"]["Properties"][0]
print(f"CID {cid} — {p['IUPACName']}")
print(f"  MW={p['MolecularWeight']}  XLogP={p['XLogP']}  TPSA={p['TPSA']}")
print(f"  HBD={p['HBondDonorCount']}  HBA={p['HBondAcceptorCount']}")
print(f"  SMILES={p['SMILES']}")

Core API

Module 1: Identifier lookup

Resolve any external identifier to a PubChem CID via /compound/{namespace}/{value}/cids/JSON. Namespaces: name, cid, smiles, inchikey, inchi, formula.

python
import requests
from urllib.parse import quote

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

# By name (returns all matching CIDs as a list)
cids = requests.get(f"{BASE}/compound/name/caffeine/cids/JSON").json()["IdentifierList"]["CID"]
print(f"caffeine CIDs: {cids}")

# By canonical SMILES (URL-encode!)
smi = quote("CC(=O)OC1=CC=CC=C1C(=O)O", safe="")
cid = requests.get(f"{BASE}/compound/smiles/{smi}/cids/JSON").json()["IdentifierList"]["CID"][0]
print(f"aspirin SMILES → CID {cid}")

# By InChIKey (exact match, fastest if you already have one)
ikey = "BSYNRYMUTXBXSQ-UHFFFAOYSA-N"
cid = requests.get(f"{BASE}/compound/inchikey/{ikey}/cids/JSON").json()["IdentifierList"]["CID"][0]
print(f"InChIKey → CID {cid}")
Module 2: Property retrieval

/compound/cid/{cid_or_csv}/property/<csv-list>/JSON returns all requested properties in one round trip. CIDs and property names are both CSV-joinable — batch up to ~200 CIDs and many properties at once.

python
import requests

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

# Full property set for a single compound (ibuprofen CID 3672)
url = (f"{BASE}/compound/cid/3672/property/"
       "MolecularWeight,XLogP,TPSA,HBondDonorCount,HBondAcceptorCount,"
       "RotatableBondCount,SMILES,InChIKey,IUPACName,MolecularFormula/JSON")
p = requests.get(url).json()["PropertyTable"]["Properties"][0]
print(f"{p['IUPACName']}  formula={p['MolecularFormula']}")
print(f"  MW={p['MolecularWeight']} XLogP={p['XLogP']} TPSA={p['TPSA']}")
print(f"  HBD={p['HBondDonorCount']} HBA={p['HBondAcceptorCount']} RotB={p['RotatableBondCount']}")
python
import requests, pandas as pd

# Batch: 4 CIDs, 3 properties — one request, one round trip
cids = "2244,3672,2157,2662"   # aspirin, ibuprofen, naproxen, celecoxib
r = requests.get(
    f"https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/cid/{cids}/property/"
    "MolecularWeight,XLogP,TPSA/JSON")
df = pd.DataFrame(r.json()["PropertyTable"]["Properties"])
print(df.to_string(index=False))
Module 3: Synonyms and description

Synonyms (trade names, CAS numbers, alternative spellings) and curated descriptions live at /compound/{ns}/{value}/{synonyms|description}/JSON.

python
import requests

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

# Full synonym list (aspirin has ~700)
info = requests.get(f"{BASE}/compound/cid/2244/synonyms/JSON").json()["InformationList"]["Information"][0]
print(f"aspirin synonyms: {len(info['Synonym'])}")
for s in info["Synonym"][:8]:
    print(f"  {s}")
python
import requests

# Curated descriptions (NCBI MeSH, CAMEO, etc.)
r = requests.get("https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/name/aspirin/description/JSON")
for item in r.json()["InformationList"]["Information"]:
    if "Description" in item:
        print(f"[{item.get('DescriptionSourceName','?')}]")
        print(f"  {item['Description'][:200]}…")
        print()
Module 4: Similarity & substructure search (async ListKey pattern)

Structure searches return HTTP 202 + {"Waiting": {"ListKey": "..."}}. Poll /compound/listkey/{key}/cids/JSON every ~2s until it returns IdentifierList. Wrap this in a helper since it's used everywhere.

python
import requests, time
from urllib.parse import quote

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

def poll_listkey(listkey, max_polls=10, interval=2.0):
    """Block until PubChem finishes async search; return CID list."""
    for _ in range(max_polls):
        time.sleep(interval)
        j = requests.get(f"{BASE}/compound/listkey/{listkey}/cids/JSON", timeout=20).json()
        if "IdentifierList" in j:
            return j["IdentifierList"]["CID"]
    raise TimeoutError(f"ListKey {listkey} did not complete")

# Tanimoto similarity (90% threshold, max 20 hits) — starting from aspirin SMILES
smi = quote("CC(=O)OC1=CC=CC=C1C(=O)O", safe="")
init = requests.get(
    f"{BASE}/compound/similarity/smiles/{smi}/JSON?Threshold=90&MaxRecords=20").json()
cids = poll_listkey(init["Waiting"]["ListKey"]) if "Waiting" in init \
       else init["IdentifierList"]["CID"]
print(f"aspirin @90% similarity: {len(cids)} hits, sample={cids[:5]}")
python
import requests
from urllib.parse import quote

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

# Substructure search — all compounds containing a sulfonamide group
smi = quote("S(=O)(=O)N", safe="")
init = requests.get(
    f"{BASE}/compound/substructure/smiles/{smi}/JSON?MaxRecords=20").json()
cids = poll_listkey(init["Waiting"]["ListKey"]) if "Waiting" in init \
       else init["IdentifierList"]["CID"]
print(f"sulfonamide-containing CIDs: {len(cids)}, sample={cids[:5]}")
Module 5: Assay summary (bioactivity data)

/compound/cid/{cid}/assaysummary/JSON returns a Table of every PubChem BioAssay the compound appears in (assay AID, target, outcome, micromolar activity if available).

python
import requests, pandas as pd

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

r = requests.get(f"{BASE}/compound/cid/2244/assaysummary/JSON", timeout=30)
rows = r.json().get("Table", {}).get("Row", [])
cols = r.json().get("Table", {}).get("Columns", {}).get("Column", [])
print(f"aspirin appears in {len(rows)} bioassays")

# First few columns + rows as a DataFrame
df = pd.DataFrame([row["Cell"] for row in rows[:5]], columns=cols)
print(df.iloc[:, :6].to_string(index=False))
Module 6: Structure file download (SDF / PNG)

/compound/cid/{cid}/SDF returns 2D MOL/SDF; /compound/cid/{cid}/PNG returns a structure image (use ?image_size=large for higher resolution).

python
import requests

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"
cid = 2519   # caffeine

# 2D SDF for downstream RDKit / OpenBabel
sdf = requests.get(f"{BASE}/compound/cid/{cid}/SDF", timeout=15).text
with open("caffeine.sdf", "w") as f:
    f.write(sdf)
print(f"caffeine.sdf: {len(sdf)} chars, ends with M  END={'M  END' in sdf}")

# PNG structure image
png = requests.get(f"{BASE}/compound/cid/{cid}/PNG?image_size=large", timeout=15).content
with open("caffeine.png", "wb") as f:
    f.write(png)
print(f"caffeine.png: {len(png)} bytes")

Key Concepts

Async ListKey pattern

Similarity, substructure, and formula searches are asynchronous — the API kicks off a background job and returns HTTP 202 with {"Waiting": {"ListKey": "<id>"}}. Poll /compound/listkey/{id}/cids/JSON every ~2 seconds until the response contains IdentifierList. Most searches finish in 5–15s; tighten polling for tiny searches, loosen for very large ones. Use the poll_listkey helper from Module 4 everywhere.

A small fraction of fast searches return IdentifierList directly on the first call (no Waiting field); check for both possibilities.

Property name reference
API nameMeaning
MolecularWeightMolecular weight (g/mol, string)
MolecularFormulaHill-system formula
SMILESIsomeric SMILES, with stereochemistry (2025+ name; was IsomericSMILES)
ConnectivitySMILESConnectivity-only SMILES, no stereo (2025+ name; was CanonicalSMILES)
IUPACNameCurated IUPAC name
InChI / InChIKeyIUPAC InChI / InChIKey
XLogPComputed logP (octanol/water)
TPSATopological polar surface area (Ų)
HBondDonorCountNumber of H-bond donors
HBondAcceptorCountNumber of H-bond acceptors
RotatableBondCountNumber of rotatable bonds
HeavyAtomCountNon-hydrogen atom count
ChargeFormal charge

CSV-join any subset in a single /property/<csv>/JSON URL. Note that MolecularWeight returns as a string; cast to float before arithmetic.

Response envelope shapes
  • cids/JSON → {"IdentifierList": {"CID": [int, ...]}}
  • property/.../JSON → {"PropertyTable": {"Properties": [{...}, ...]}} (one dict per CID, in input order)
  • synonyms/JSON → {"InformationList": {"Information": [{"CID": int, "Synonym": [str, ...]}]}}
  • description/JSON → {"InformationList": {"Information": [{"CID": int, "Description": str, "DescriptionSourceName": str, ...}, ...]}}
  • assaysummary/JSON → {"Table": {"Columns": {"Column": [...]}, "Row": [{"Cell": [...]}, ...]}}
  • async-init (similarity, substructure, formula) → 202 with {"Waiting": {"ListKey": "..."}}
  • listkey poll → {"IdentifierList": {"CID": [...]}} when ready

Common Workflows

Workflow 1: Property comparison across a drug panel

Goal: side-by-side physicochemical comparison of a small molecule set.

python
import requests, pandas as pd, time
from urllib.parse import quote

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

drugs = ["aspirin", "ibuprofen", "naproxen", "celecoxib"]

# Resolve names → CIDs in a loop (rate-limited)
cids = []
for d in drugs:
    cid = requests.get(f"{BASE}/compound/name/{quote(d)}/cids/JSON",
                       timeout=15).json()["IdentifierList"]["CID"][0]
    cids.append(cid)
    time.sleep(0.25)

# Single batched property pull
cid_csv = ",".join(str(c) for c in cids)
r = requests.get(
    f"{BASE}/compound/cid/{cid_csv}/property/"
    "MolecularWeight,XLogP,TPSA,HBondDonorCount,HBondAcceptorCount/JSON")
df = pd.DataFrame(r.json()["PropertyTable"]["Properties"])
df["Name"] = drugs
df = df[["Name", "CID", "MolecularWeight", "XLogP", "TPSA",
         "HBondDonorCount", "HBondAcceptorCount"]]
print(df.to_string(index=False))
Workflow 2: Lead compound → similar analogs → properties

Goal: starting from a kinase inhibitor (gefitinib), find 85%-similar analogs and pull their properties.

python
import requests, time
from urllib.parse import quote

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

def poll_listkey(listkey, max_polls=10, interval=2.0):
    for _ in range(max_polls):
        time.sleep(interval)
        j = requests.get(f"{BASE}/compound/listkey/{listkey}/cids/JSON",
                         timeout=20).json()
        if "IdentifierList" in j:
            return j["IdentifierList"]["CID"]
    raise TimeoutError("listkey timeout")

# 1. lead → CID → canonical SMILES
ref_cid = requests.get(
    f"{BASE}/compound/name/gefitinib/cids/JSON").json()["IdentifierList"]["CID"][0]
ref_smi = requests.get(
    f"{BASE}/compound/cid/{ref_cid}/property/SMILES/JSON"
    ).json()["PropertyTable"]["Properties"][0]["SMILES"]
print(f"gefitinib CID={ref_cid}  SMILES={ref_smi}")

# 2. similarity search
smi_q = quote(ref_smi, safe="")
init = requests.get(
    f"{BASE}/compound/similarity/smiles/{smi_q}/JSON?Threshold=85&MaxRecords=15"
    ).json()
sim_cids = poll_listkey(init["Waiting"]["ListKey"]) if "Waiting" in init \
           else init["IdentifierList"]["CID"]
print(f"  {len(sim_cids)} analogs @85% Tanimoto")

# 3. batch-pull properties for top 5 analogs
cid_csv = ",".join(str(c) for c in sim_cids[:5])
r = requests.get(
    f"{BASE}/compound/cid/{cid_csv}/property/"
    "MolecularWeight,XLogP,TPSA,RotatableBondCount/JSON")
for row in r.json()["PropertyTable"]["Properties"]:
    print(f"  CID {row['CID']}: MW={row['MolecularWeight']} XLogP={row['XLogP']}")
Workflow 3: Pharmacophore screen via substructure → bioactivity check

Goal: find compounds with a sulfonamide motif and check which have bioactivity records.

python
import requests, time
from urllib.parse import quote

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

def poll_listkey(listkey, max_polls=10, interval=2.0):
    for _ in range(max_polls):
        time.sleep(interval)
        j = requests.get(f"{BASE}/compound/listkey/{listkey}/cids/JSON",
                         timeout=20).json()
        if "IdentifierList" in j:
            return j["IdentifierList"]["CID"]
    raise TimeoutError("listkey timeout")

smi = quote("S(=O)(=O)N", safe="")
init = requests.get(
    f"{BASE}/compound/substructure/smiles/{smi}/JSON?MaxRecords=10").json()
cids = poll_listkey(init["Waiting"]["ListKey"]) if "Waiting" in init \
       else init["IdentifierList"]["CID"]
print(f"sulfonamide CIDs: {cids}")

# Bioactivity row counts for the first few hits
for cid in cids[:3]:
    rows = requests.get(f"{BASE}/compound/cid/{cid}/assaysummary/JSON",
                        timeout=30).json().get("Table", {}).get("Row", [])
    print(f"  CID {cid}: {len(rows)} assay rows")
    time.sleep(0.3)

Key Parameters

ParameterEndpoint / ModuleDefaultRange / OptionsEffect
<namespace>/compound/{ns}/<value>/...—name, cid, smiles, inchikey, inchi, formulaInput identifier type
<property csv>/.../property/<csv>/JSON—any subset of property names (see table)Which properties to return (one DB call per request)
Thresholdsimilarity (M4)900–100Tanimoto cutoff (percent)
MaxRecordssimilarity / substructure / formula(server-side default)1–10000Cap on async result list
image_size/compound/cid/{cid}/PNGmediumsmall, large, WxH (e.g. 500x500)PNG output resolution
record_type/compound/cid/{cid}/SDF2d2d, 3dSDF dimensionality (?record_type=3d)
MaxAssayResults/compound/cid/{cid}/assaysummary/JSON—intLimit assay rows when compound has thousands of records
Show full SKILL.md (703 more words)Show less

Best Practices

  1. Always go through cids/JSON first when starting from a name or external identifier. The name→CID resolution and the CID→property lookup are separate calls; doing both at once via name → property works but throws away the canonical CID list that downstream queries need.

  2. Batch properties, never loop them. compound/cid/2244,3672,2157,.../property/MolecularWeight,XLogP,.../JSON accepts up to ~200 CIDs and any number of properties — one round trip instead of N. Looping get_compounds per name is the most common rate-limit trap.

  3. URL-encode every SMILES. Use urllib.parse.quote(smiles, safe=""). Bare SMILES with =, #, (, ), [, ] will sometimes work but breaks unpredictably on +, /, \, or query-string-looking substrings.

  4. Treat similarity/substructure/formula as async. Branch on "Waiting" in response and poll listkey rather than re-issuing the search. Re-issuing creates a new ListKey and wastes the server's job slot.

  5. Throttle ≤ 5 req/sec, ≤ 400/min. Insert time.sleep(0.25) in any tight loop. HTTP 503 means you tripped the limit — wait 10s and reduce concurrency.

  6. Cast MolecularWeight to float. It's returned as a string ("180.16") for full decimal fidelity. Comparing strings against numeric thresholds is a silent bug.

  7. For 100+ CIDs use POST. GET URLs over ~2000 chars get truncated by some HTTP proxies. PubChem also accepts POST with cid in the form body: requests.post(f"{BASE}/compound/cid/property/MolecularWeight/JSON", data={"cid": cid_csv}).

  8. 2025 SMILES property rename. PubChem renamed two SMILES properties in the 2025 PUG-REST schema: old IsomericSMILES (with stereo) → SMILES, and old CanonicalSMILES (connectivity only, no stereo) → ConnectivitySMILES. The URL path still accepts the legacy names as input (e.g. /property/CanonicalSMILES/JSON returns 200), but the response JSON is keyed with the new names. So a request succeeds and only the parse step breaks with KeyError. Use SMILES / ConnectivitySMILES in new code and read those keys.

Common Recipes

Recipe 1 — Lipinski's Rule of Five check
python
import requests
from urllib.parse import quote

BASE = "https://pubchem.ncbi.nlm.nih.gov/rest/pug"

def check_lipinski(name):
    cid = requests.get(f"{BASE}/compound/name/{quote(name)}/cids/JSON"
                       ).json()["IdentifierList"]["CID"][0]
    p = requests.get(
        f"{BASE}/compound/cid/{cid}/property/"
        "MolecularWeight,XLogP,HBondDonorCount,HBondAcceptorCount/JSON"
        ).json()["PropertyTable"]["Properties"][0]
    mw, xlogp = float(p["MolecularWeight"]), p.get("XLogP", 0) or 0
    hbd, hba  = p["HBondDonorCount"], p["HBondAcceptorCount"]
    rules = {"MW ≤ 500": mw <= 500, "XLogP ≤ 5": xlogp <= 5,
             "HBD ≤ 5": hbd <= 5,   "HBA ≤ 10": hba <= 10}
    v = sum(1 for ok in rules.values() if not ok)
    return rules, v

rules, v = check_lipinski("metformin")
print(f"violations: {v}/4 ({'PASS' if v <= 1 else 'FAIL'})")
for r, ok in rules.items(): print(f"  {'✓' if ok else '✗'} {r}")
Recipe 2 — Pull every synonym / trade name
python
import requests

r = requests.get("https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/name/aspirin/synonyms/JSON")
syns = r.json()["InformationList"]["Information"][0]["Synonym"]
print(f"{len(syns)} synonyms")
for s in syns[:10]:
    print(f"  {s}")
Recipe 3 — Download a structure PNG
python
import requests

cid = 2519   # caffeine
png = requests.get(
    f"https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/cid/{cid}/PNG?image_size=large"
    ).content
with open("caffeine.png", "wb") as f:
    f.write(png)
print(f"wrote caffeine.png ({len(png)} bytes)")
Recipe 4 — Robust session with retry / 429 handling
python
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

s = requests.Session()
s.headers.update({"Accept": "application/json"})
s.mount("https://", HTTPAdapter(max_retries=Retry(
    total=4, backoff_factor=1.0,
    status_forcelist=[429, 500, 502, 503, 504],
    allowed_methods=["GET", "POST"])))

r = s.get(
    "https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/cid/2244/property/MolecularWeight/JSON",
    timeout=15)
r.raise_for_status()
print(r.json()["PropertyTable"]["Properties"][0]["MolecularWeight"])

Expected Outputs

  • CID lookup (/.../cids/JSON): {"IdentifierList": {"CID": [...]}} — list of integer CIDs, ordered by relevance for name searches.
  • Properties (/.../property/.../JSON): {"PropertyTable": {"Properties": [{"CID": ..., "MolecularWeight": "180.16", ...}, ...]}}. One row per CID in input order; MolecularWeight is a string.
  • Synonyms: {"InformationList": {"Information": [{"CID": 2244, "Synonym": ["aspirin", "ACETYLSALICYLIC ACID", "50-78-2", ...]}]}}.
  • Async search init: HTTP 202 with {"Waiting": {"ListKey": "12345..."}}. Poll /compound/listkey/{key}/cids/JSON until it returns IdentifierList.
  • Assay summary: {"Table": {"Columns": {"Column": [...col names...]}, "Row": [{"Cell": [...]}, ...]}}.
  • SDF: plain text starting with the CID header line; ends with M END followed by SDF property blocks.
  • PNG: binary image/png, ~2–4 KB at default size, ~10–20 KB at image_size=large.

Troubleshooting

ProblemCauseSolution
HTTP 404 PUGREST.NotFoundName / SMILES / formula matched no recordTry a CAS number or InChIKey; check spelling in the PubChem web UI; canonical SMILES from RDKit often resolves where input SMILES doesn't
HTTP 202 stuck in {"Waiting":...} for a similarity/substructure callAsync job still runningPoll /compound/listkey/{key}/cids/JSON every 2s up to ~30s; reduce MaxRecords if it never completes
HTTP 503 PUGREST.ServerBusyTripped the 5-req/s or 400-req/min rate limitInsert time.sleep(0.25) in loops; use the Retry session in Recipe 4; reduce concurrency
HTTP 400 on a SMILES URLSMILES wasn't URL-encodedWrap in urllib.parse.quote(smi, safe="") — #, +, / and \ all break path parsing
KeyError: 'CanonicalSMILES' / KeyError: 'IsomericSMILES'Requested old name; URL returns 200 but 2025 JSON is keyed ConnectivitySMILES / SMILESRead p["ConnectivitySMILES"] (connectivity, no stereo) or p["SMILES"] (with stereo); update the property CSV to the new names
TypeError: '>' not supported between instances of 'str' and 'int'MolecularWeight is a stringfloat(p["MolecularWeight"]) before any arithmetic comparison
Batch cid/2244,3672,... returns only some rowsURL exceeded server limitSwitch to requests.post(url, data={"cid": "2244,3672,..."}); same URL minus the value, body carries the CSV
Empty assaysummary TableCID has no bioassay recordsNot all compounds are assayed; verify on the PubChem web page
XLogP is None for a valid CIDProperty not computed for that compoundGuard with p.get("XLogP", 0) or 0 before arithmetic
  • chembl-database-bioactivity — IC50 / Ki / Kd target-binding data, deeper than PubChem's assay summaries
  • rdkit-cheminformatics — local SMILES/MOL manipulation, fingerprints, descriptors, scaffold extraction
  • pdb-database — protein structures co-crystallized with the small molecules found via PubChem CIDs

References

© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/structural-biology-drug-discovery/pubchem-compound-search of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pubchem Compound Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pubchem Compound Search compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pubchem Compound Search this skilljaechang-hits/SciAgent-Skills3741 repos~7kAutomated safety check: PassCC-BY-4.0
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT
Edu Chem Reactionwy51ai/edulab1.4k—~1.2kAutomated safety check: PassApache-2.0
RDKit Cheminformatics Practicesaiming-lab/AutoResearchClaw15k—~708Automated safety check: PassMIT
Rowanlamm-mit/scienceclaw2464 repos~3.1kAutomated safety check: WarnProprietary

Similar skills

  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 10 days ago
    Research & ScienceAuto-check passed
  • Edu Chem Reaction

    wy51ai/edulab

    把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • RDKit Cheminformatics Practices

    aiming-lab/AutoResearchClaw

    Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.

    15k GitHub stars~708 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Rowan

    lamm-mit/scienceclaw

    Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.

    246 GitHub starsUsed in 4 repos~3.1k tokens
    Research & ScienceAuto-check: warnings
  • Coot Rdkit

    pemsley/coot

    RDKit molecular manipulation and visualization within Coot's Python environment.

    168 GitHub stars~981 tokensUpdated 3 days ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Works with

Questions about Pubchem Compound Search

What does Pubchem Compound Search do?

Query PubChem (110M+ compounds) directly via the PUG-REST/JSON API with plain requests — no SDK install required. Pubchem Compound Search is an agent skill from jaechang-hits/SciAgent-Skills. Query PubChem (110M+ compounds) directly via the PUG-REST/JSON API with plain requests — no SDK install required.

When should I use Pubchem Compound Search?

Pubchem Compound Search fits situations like: tasks that involve Drug discovery and cheminformatics.

How do I install Pubchem Compound Search in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill pubchem-compound-search -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/pubchem-compound-search in jaechang-hits/SciAgent-Skills) into .claude/skills/pubchem-compound-search in your project. Claude Code loads it when a task matches its description.

How do I install Pubchem Compound Search in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill pubchem-compound-search -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/pubchem-compound-search in jaechang-hits/SciAgent-Skills) into .agents/skills/pubchem-compound-search in your project. Codex loads it when a task matches its description.

Can I use Pubchem Compound Search in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill pubchem-compound-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pubchem-compound-search, .gemini/skills/pubchem-compound-search, .github/skills/pubchem-compound-search and .opencode/skills/pubchem-compound-search in your project.

What does Pubchem Compound Search need to run?

Going by SKILL.md and its folder, Pubchem Compound Search needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Pubchem Compound Search access the network?

SKILL.md names 2 domains. In commands or code: pubchem.ncbi.nlm.nih.gov; the agent is likely to contact it when it follows the instructions. As links in the text: doi.org. This is read from the text; nothing was executed.

Is Pubchem Compound Search safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pubchem Compound Search use?

Pubchem Compound Search is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pubchem Compound Search use?

About 7k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pubchem Compound Search?

Skills that share tags, products or a category with Pubchem Compound Search: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars), Edu Chem Reaction (wy51ai/edulab, 1.4k stars) and RDKit Cheminformatics Practices (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pubchem Compound Search?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.