DiffDock Molecular Docking
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required.
$ npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills chembl-database-bioactivity --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/chembl-database-bioactivity .claude/skills/chembl-database-bioactivity && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "chembl-database-bioactivity" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/chembl-database-bioactivity into .claude/skills/chembl-database-bioactivity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chembl-database-bioactivity", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/chembl-database-bioactivityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills chembl-database-bioactivity --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/structural-biology-drug-discovery/chembl-database-bioactivity .agents/skills/chembl-database-bioactivity && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "chembl-database-bioactivity" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/chembl-database-bioactivity into .agents/skills/chembl-database-bioactivity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chembl-database-bioactivity", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills chembl-database-bioactivity --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/structural-biology-drug-discovery/chembl-database-bioactivity .cursor/skills/chembl-database-bioactivity && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "chembl-database-bioactivity" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/chembl-database-bioactivity into .cursor/skills/chembl-database-bioactivity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chembl-database-bioactivity", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/structural-biology-drug-discovery/chembl-database-bioactivity--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills chembl-database-bioactivity --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/structural-biology-drug-discovery/chembl-database-bioactivity .gemini/skills/chembl-database-bioactivity && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "chembl-database-bioactivity" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/chembl-database-bioactivity into .gemini/skills/chembl-database-bioactivity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chembl-database-bioactivity", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills chembl-database-bioactivityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/structural-biology-drug-discovery/chembl-database-bioactivity .github/skills/chembl-database-bioactivity && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "chembl-database-bioactivity" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/chembl-database-bioactivity into .github/skills/chembl-database-bioactivity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chembl-database-bioactivity", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills chembl-database-bioactivity --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/structural-biology-drug-discovery/chembl-database-bioactivity .opencode/skills/chembl-database-bioactivity && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "chembl-database-bioactivity" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/chembl-database-bioactivity into .opencode/skills/chembl-database-bioactivity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chembl-database-bioactivity", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
chembl-database-bioactivityQuery ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required.
Chembl Database Bioactivity is an agent skill from jaechang-hits/SciAgent-Skills. Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required. Search compounds, retrieve IC50/Ki/EC50 bioactivities, find target inhibitors, run SAR, access drug mechanism/indication data.
Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Drug discovery and cheminformatics and Protein structure and design. It works with Django. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-SA-4.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
ebi.ac.ukAlso links to:
chembl.gitbook.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Chembl Database Bioactivity loads about 6.4k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 1,177 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-SA-4.0 licence (© jaechang-hits). 1,177 words, ~6,426 tokens.
.claude/skills/chembl-database-bioactivity/SKILL.md (or your agent's skills folder).Why no SDK? The
chembl_webresource_clientpackage is convenient sugar over a public, no-auth REST/JSON API athttps://www.ebi.ac.uk/chembl/api/data/. When the SDK is unavailable, every operation can be reproduced with plainrequestsand URL parameters. This SKILL.md uses the REST path throughout so the code runs in any environment withrequestsinstalled. Django-style filter syntax (field__icontains=…,field__lte=…,field__range=a,b) works as URL query parameters.
ChEMBL is EMBL-EBI's bioactive molecule database: 2M+ compounds, 19M+ bioactivity measurements (IC50, Ki, EC50, Kd, …), 13K+ targets. The REST API at https://www.ebi.ac.uk/chembl/api/data/ returns JSON (append .json) or XML/YAML, requires no authentication, and supports Django-style query filters via URL parameters plus cursor-style pagination via page_meta.next.
rdkit-cheminformatics insteadpubchem-compound-searchrequests (only requirement). Optional: pandas for tabular analysis.time.sleep(0.2-0.5) between requests in batch loops; back off on HTTP 429.pip install requests
# Optional, for DataFrame work:
pip install pandasimport requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# Retrieve a molecule by ChEMBL ID
r = requests.get(f"{BASE}/molecule/CHEMBL25.json", timeout=15)
r.raise_for_status()
aspirin = r.json()
print(f"{aspirin['pref_name']}: MW={aspirin['molecule_properties']['mw_freebase']}")
# ASPIRIN: MW=180.16
# Search targets by full name (acronyms like 'EGFR' don't match pref_name — use full term)
r = requests.get(
f"{BASE}/target.json",
params={"pref_name__icontains": "epidermal growth factor receptor",
"target_type": "SINGLE PROTEIN", "limit": 5},
timeout=15,
)
targets = r.json()["targets"]
print(f"EGFR-like targets: {len(targets)}, first={targets[0]['target_chembl_id']}")
# Potent bioactivities: EGFR (CHEMBL203) IC50 <= 100 nM
r = requests.get(
f"{BASE}/activity.json",
params={"target_chembl_id": "CHEMBL203",
"standard_type": "IC50",
"standard_value__lte": 100,
"standard_units": "nM",
"limit": 5},
timeout=30,
)
data = r.json()
print(f"EGFR IC50 ≤ 100 nM records: {data['page_meta']['total_count']}")The SDK's field__operator=value syntax maps 1:1 to URL query parameters. Use & to combine filters.
| Operator | URL pattern | Example URL fragment |
|---|---|---|
__exact | field=value | target_type=SINGLE+PROTEIN |
__iexact | field__iexact=value | pref_name__iexact=aspirin |
__contains / __icontains | field__icontains=value | pref_name__icontains=kinase |
__startswith / __endswith | field__startswith=Epi | pref_name__endswith=nib |
__gt / __gte / __lt / __lte | field__lte=100 | standard_value__lte=100 |
__range | field__range=lo,hi | molecule_properties__mw_freebase__range=300,500 |
__in | field__in=a,b,c | standard_type__in=IC50,Ki,Kd |
__isnull | field__isnull=False (Python False/True strings) | pchembl_value__isnull=False |
__regex | field__regex=… | pref_name__regex=^EGF.*kinase$ |
__search | field__search=… | description__search=apoptosis |
When passed via requests.get(..., params={...}), the library handles URL encoding automatically (including the commas in __range and __in).
All endpoints accept .json, .xml, or .yaml suffix. JSON is the default below.
| Endpoint URL | Returns | Key fields |
|---|---|---|
/molecule/{chembl_id}.json | Compound by ID | pref_name, molecule_chembl_id, molecule_properties, molecule_structures |
/molecule.json?<filters> | Compound search | paginated molecules[] |
/target/{chembl_id}.json | Target by ID | pref_name, target_type, organism, target_components |
/target.json?<filters> | Target search | paginated targets[] |
/activity.json?<filters> | Bioactivity records | paginated activities[] |
/assay.json?<filters> | Assay details | paginated assays[] |
/drug.json?<filters> | Approved drug info | paginated drugs[]; supports /drug/{chembl_id}.json |
/mechanism.json?<filters> | Mechanism of action | paginated mechanisms[] |
/drug_indication.json?<filters> | Therapeutic indications | paginated drug_indications[] |
/similarity/{smiles}/{tanimoto}.json | Tanimoto similarity (0–100) | paginated molecules[] with similarity field |
/substructure/{smiles}.json | Substructure search | paginated molecules[] |
/image/{chembl_id}.svg | SVG structure image | binary SVG (NOT JSON) |
/molecule_form/{chembl_id}.json | Parent/salt forms | molecule_forms[] |
/protein_class.json | Protein classification hierarchy | hierarchical browse |
/document.json?<filters> | Literature source records | paginated documents[] |
{
"page_meta": {
"limit": 20,
"offset": 0,
"total_count": 12145,
"next": "/chembl/api/data/activity.json?...&offset=20",
"previous": null
},
"activities": [ /* or molecules[], targets[], etc. */ ]
}Walk page_meta.next (a relative URL — prefix with https://www.ebi.ac.uk) until it becomes null.
Properties accessible via molecule_properties on each record:
| Field | Description |
|---|---|
mw_freebase | Molecular weight (free base) |
full_mwt | Full molecular weight (including salts) |
alogp | Calculated LogP |
hba | Hydrogen bond acceptors |
hbd | Hydrogen bond donors |
psa | Polar surface area |
rtb | Rotatable bonds |
num_ro5_violations | Lipinski rule-of-5 violations |
ro3_pass | Rule of 3 compliance |
cx_most_apka / cx_most_bpka | Most acidic / basic pKa |
| Field | Description |
|---|---|
target_chembl_id | ChEMBL target identifier |
pref_name | Preferred (full) target name — acronyms like "EGFR" do NOT match; use the spelled-out term |
target_type | SINGLE PROTEIN, PROTEIN COMPLEX, ORGANISM, … |
organism | Target organism (e.g., Homo sapiens) |
tax_id | NCBI taxonomy ID |
target_components[] | Components (UniProt accession, sequence, …) |
| Field | Description |
|---|---|
standard_type | Activity type: IC50, Ki, Kd, EC50, … |
standard_value | Numerical activity value |
standard_units | Units: nM, uM, … |
pchembl_value | Normalized -log10 activity (>6 = potent) |
activity_comment | Activity annotations |
data_validity_comment | Data quality flags (check before analysis) |
potential_duplicate | Duplicate flag |
import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# By ChEMBL ID
r = requests.get(f"{BASE}/molecule/CHEMBL25.json", timeout=15)
aspirin = r.json()
print(f"{aspirin['pref_name']}: MW={aspirin['molecule_properties']['mw_freebase']}")
# By name (case-insensitive substring)
r = requests.get(f"{BASE}/molecule.json",
params={"pref_name__icontains": "imatinib", "limit": 5},
timeout=15)
for mol in r.json()["molecules"]:
print(f" {mol['molecule_chembl_id']} {mol.get('pref_name')!r}")
# By Lipinski-compliant property ranges
r = requests.get(f"{BASE}/molecule.json",
params={"molecule_properties__mw_freebase__range": "300,500",
"molecule_properties__alogp__lte": 5,
"molecule_properties__hba__lte": 10,
"molecule_properties__hbd__lte": 5,
"limit": 3},
timeout=15)
print(f"Lipinski-compliant total: {r.json()['page_meta']['total_count']}")import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# By ChEMBL ID
r = requests.get(f"{BASE}/target/CHEMBL203.json", timeout=15)
egfr = r.json()
print(f"{egfr['pref_name']} ({egfr['organism']}) — type={egfr['target_type']}")
# Search by full name (NOT acronym) + type
r = requests.get(f"{BASE}/target.json",
params={"pref_name__icontains": "kinase",
"target_type": "SINGLE PROTEIN", "limit": 5},
timeout=15)
d = r.json()
print(f"Kinase SINGLE_PROTEIN targets: total={d['page_meta']['total_count']}")
for t in d["targets"][:5]:
print(f" {t['target_chembl_id']:12s} {t.get('pref_name')!r} ({t['organism']})")
# By organism
r = requests.get(f"{BASE}/target.json",
params={"organism": "Homo sapiens", "limit": 3},
timeout=15)
print(f"Human targets: total={r.json()['page_meta']['total_count']}")import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# Potent inhibitors for a target (EGFR = CHEMBL203)
r = requests.get(f"{BASE}/activity.json",
params={"target_chembl_id": "CHEMBL203",
"standard_type": "IC50",
"standard_value__lte": 100,
"standard_units": "nM",
"limit": 5},
timeout=30)
data = r.json()
print(f"EGFR IC50≤100nM: total={data['page_meta']['total_count']}")
for act in data["activities"][:5]:
print(f" {act['molecule_chembl_id']:14s} IC50={act['standard_value']} nM "
f"pChEMBL={act.get('pchembl_value')}")
# All pChEMBL-tagged activities for a compound
r = requests.get(f"{BASE}/activity.json",
params={"molecule_chembl_id": "CHEMBL25",
"pchembl_value__isnull": "False",
"limit": 5},
timeout=30)
print(f"Aspirin pChEMBL activities: total={r.json()['page_meta']['total_count']}")
# Multiple activity types (CHEMBL240 = D2 dopamine receptor)
r = requests.get(f"{BASE}/activity.json",
params={"target_chembl_id": "CHEMBL240",
"standard_type__in": "IC50,Ki,Kd",
"limit": 5},
timeout=30)
print(f"D2 receptor IC50/Ki/Kd: total={r.json()['page_meta']['total_count']}")import requests
from urllib.parse import quote
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# Similarity search (Tanimoto ≥ 85%)
# Path-style endpoint: /similarity/{smiles}/{threshold}
# The SMILES MUST be URL-encoded (it contains '/', '(', ')' etc.)
aspirin_smiles = quote("CC(=O)Oc1ccccc1C(=O)O", safe="")
r = requests.get(f"{BASE}/similarity/{aspirin_smiles}/85.json",
params={"limit": 5}, timeout=30)
data = r.json()
print(f"Similar to aspirin (≥85% Tanimoto): total={data['page_meta']['total_count']}")
for m in data["molecules"][:5]:
print(f" {m['molecule_chembl_id']} similarity={m.get('similarity')}")# Substructure search
benzimidazole = quote("c1ccc2[nH]cnc2c1", safe="")
r = requests.get(f"{BASE}/substructure/{benzimidazole}.json",
params={"limit": 3}, timeout=30)
print(f"Benzimidazole substructure total: {r.json()['page_meta']['total_count']}")import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# Drug record (max clinical phase, ATC class, etc.)
r = requests.get(f"{BASE}/drug/CHEMBL941.json", timeout=15) # imatinib
drug = r.json()
print(f"Imatinib max_phase={drug.get('max_phase')}")
# Mechanisms of action — note: not every drug has mechanism records.
# Imatinib (CHEMBL941) returns 0 mechanism rows; sunitinib (CHEMBL535) has many.
r = requests.get(f"{BASE}/mechanism.json",
params={"molecule_chembl_id": "CHEMBL535"}, timeout=15)
for m in r.json()["mechanisms"]:
print(f" {m['mechanism_of_action']} → target {m.get('target_chembl_id')}")
# Therapeutic indications
r = requests.get(f"{BASE}/drug_indication.json",
params={"molecule_chembl_id": "CHEMBL941", "limit": 5},
timeout=15)
for ind in r.json()["drug_indications"]:
print(f" {ind.get('mesh_heading')!r} max_phase_for_ind={ind.get('max_phase_for_ind')}")# SVG molecular structure image — direct binary response, NOT JSON
# Do NOT call /image/{cid}.json — that endpoint raises JSONDecodeError.
r = requests.get(f"{BASE}/image/CHEMBL25.svg", timeout=15)
r.raise_for_status()
with open("aspirin.svg", "w") as f:
f.write(r.text)
print(f"Saved aspirin.svg ({len(r.text)} bytes, looks_svg={'<svg' in r.text})")Note: pref_name__icontains matches the spelled-out name. Acronyms like 'EGFR' or 'BRAF' return 0 results — use 'epidermal growth factor receptor' or 'B-raf' (with the hyphen).
import requests, pandas as pd, time
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# Step 1: Resolve the target by full name
r = requests.get(f"{BASE}/target.json",
params={"pref_name__icontains": "B-raf",
"target_type": "SINGLE PROTEIN", "limit": 5},
timeout=15)
targets = r.json()["targets"]
human_braf = next(t for t in targets if t["organism"] == "Homo sapiens")
target_id = human_braf["target_chembl_id"]
print(f"Using {target_id} — {human_braf['pref_name']}")
# Step 2: Paginate all potent IC50 activities (cap at 500 for demo)
url = (f"{BASE}/activity.json"
f"?target_chembl_id={target_id}"
f"&standard_type=IC50"
f"&standard_value__lte=100"
f"&standard_units=nM"
f"&pchembl_value__isnull=False"
f"&limit=200")
records = []
while url and len(records) < 500:
r = requests.get(url, timeout=30)
r.raise_for_status()
data = r.json()
records.extend(data["activities"])
nxt = data["page_meta"].get("next")
url = f"https://www.ebi.ac.uk{nxt}" if nxt else None
time.sleep(0.2)
df = pd.DataFrame(records)
df["standard_value"] = pd.to_numeric(df["standard_value"])
print(f"Retrieved {len(df)} potent {target_id} compounds")
print(df[["molecule_chembl_id", "standard_value", "pchembl_value"]].head(10))import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# Sunitinib (CHEMBL535) — has documented mechanisms + indications.
# Imatinib (CHEMBL941) sometimes returns 0 mechanism rows depending on ChEMBL release.
chembl_id = "CHEMBL535"
# Molecule record
m = requests.get(f"{BASE}/molecule/{chembl_id}.json", timeout=15).json()
print(f"Name: {m['pref_name']}")
print(f"MW : {m['molecule_properties']['mw_freebase']}")
# Mechanisms
mechs = requests.get(f"{BASE}/mechanism.json",
params={"molecule_chembl_id": chembl_id},
timeout=15).json()["mechanisms"]
for mc in mechs:
print(f" Mechanism: {mc['mechanism_of_action']}")
# Indications
inds = requests.get(f"{BASE}/drug_indication.json",
params={"molecule_chembl_id": chembl_id, "limit": 5},
timeout=15).json()["drug_indications"]
for ind in inds:
print(f" Indication: {ind.get('mesh_heading')} "
f"(Phase {ind.get('max_phase_for_ind')})")
# Bioactivity record count
total = requests.get(f"{BASE}/activity.json",
params={"molecule_chembl_id": chembl_id,
"pchembl_value__isnull": "False",
"limit": 1},
timeout=30).json()["page_meta"]["total_count"]
print(f"Total bioactivity records (pChEMBL-tagged): {total}")import requests, pandas as pd, time
from urllib.parse import quote
BASE = "https://www.ebi.ac.uk/chembl/api/data"
# Step 1: Similar compounds to a lead (e.g., quinoline scaffold)
lead_smiles = "c1ccc2c(c1)cc(nc2N)c3ccc(cc3)NC(=O)c4ccccc4"
r = requests.get(f"{BASE}/similarity/{quote(lead_smiles, safe='')}/80.json",
params={"limit": 20}, timeout=30)
analogs = r.json()["molecules"]
print(f"Analogs found: {len(analogs)}")
# Step 2: Collect bioactivities for each analog
records = []
for compound in analogs[:20]:
cid = compound["molecule_chembl_id"]
acts = requests.get(f"{BASE}/activity.json",
params={"molecule_chembl_id": cid,
"standard_type": "IC50",
"pchembl_value__isnull": "False",
"limit": 20},
timeout=30).json()["activities"]
for act in acts:
records.append({
"chembl_id": cid,
"target": act.get("target_pref_name"),
"IC50_nM": act.get("standard_value"),
"pchembl": act.get("pchembl_value"),
"mw": (compound.get("molecule_properties") or {}).get("mw_freebase"),
"alogp": (compound.get("molecule_properties") or {}).get("alogp"),
})
time.sleep(0.2)
df = pd.DataFrame(records)
if not df.empty:
df["IC50_nM"] = pd.to_numeric(df["IC50_nM"])
print(df.groupby("target")["IC50_nM"].describe())import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"
r = requests.get(f"{BASE}/molecule.json",
params={"molecule_properties__mw_freebase__range": "300,500",
"molecule_properties__alogp__lte": 5,
"molecule_properties__hba__lte": 10,
"molecule_properties__hbd__lte": 5,
"molecule_properties__num_ro5_violations": 0,
"limit": 1},
timeout=15)
print(f"Drug-like candidates: {r.json()['page_meta']['total_count']}")import requests, pandas as pd, time
BASE = "https://www.ebi.ac.uk/chembl/api/data"
url = (f"{BASE}/activity.json"
f"?target_chembl_id=CHEMBL203"
f"&standard_type=IC50"
f"&pchembl_value__isnull=False"
f"&limit=500")
all_acts = []
while url:
r = requests.get(url, timeout=60)
r.raise_for_status()
data = r.json()
all_acts.extend(data["activities"])
nxt = data["page_meta"].get("next")
url = f"https://www.ebi.ac.uk{nxt}" if nxt else None
time.sleep(0.3)
df = pd.DataFrame(all_acts)
df.to_csv("egfr_activities.csv", index=False)
print(f"Exported {len(df)} records → egfr_activities.csv")import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
def chembl_session(retries=3, backoff=1.0):
s = requests.Session()
s.headers.update({"Accept": "application/json"})
s.mount("https://", HTTPAdapter(max_retries=Retry(
total=retries, backoff_factor=backoff,
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET"])))
return s
session = chembl_session()
r = session.get("https://www.ebi.ac.uk/chembl/api/data/molecule/CHEMBL25.json", timeout=15)
print(r.json()["pref_name"])import requests
r = requests.get("https://www.ebi.ac.uk/chembl/api/data/image/CHEMBL25.svg", timeout=15)
r.raise_for_status()
with open("aspirin.svg", "w") as f:
f.write(r.text)| Parameter | Endpoint | Default | Description |
|---|---|---|---|
limit | all list endpoints | 20 | Page size; max 1000 |
offset | all list endpoints | 0 | Pagination offset (or follow page_meta.next) |
format | all endpoints | json (via .json suffix) | Also .xml, .yaml |
pref_name__icontains | /target, /molecule | — | Substring on full name; acronyms don't match, use full term |
target_chembl_id | /activity | — | E.g., CHEMBL203 (EGFR), CHEMBL240 (D2 receptor) |
molecule_chembl_id | /activity, /mechanism, /drug_indication | — | E.g., CHEMBL25 (aspirin) |
standard_type | /activity | — | IC50, Ki, Kd, EC50 |
standard_value__lte | /activity | — | Max activity value (paired with standard_units) |
pchembl_value__isnull | /activity | — | "False" to require pChEMBL-tagged data |
target_type | /target | — | SINGLE PROTEIN, PROTEIN COMPLEX, ORGANISM, … |
{tanimoto} (path) | /similarity/{smiles}/{tanimoto} | — | 0–100 Tanimoto threshold |
{smiles} (path) | /similarity, /substructure | — | URL-encoded SMILES (urllib.parse.quote(s, safe="")) |
| Problem | Cause | Solution |
|---|---|---|
pref_name__icontains=EGFR (or BRAF) returns 0 | ChEMBL stores spelled-out names; acronyms don't match | Use "epidermal growth factor receptor"; for BRAF use "B-raf" with the hyphen |
mechanism.json?molecule_chembl_id=CHEMBL941 returns empty | Not every drug has mechanism rows in every release (e.g., imatinib has 0 in current data) | Use CHEMBL535 (sunitinib) or CHEMBL192 (sildenafil) as known-populated examples |
JSONDecodeError on /image/{cid}.json | The image endpoint is binary, not JSON | Always use .svg or .png suffix: /image/{cid}.svg |
404 on /molecule/{id} | Invalid ChEMBL ID format | IDs must include the prefix: CHEMBL25, not 25 |
| 400 on similarity search | Unencoded SMILES (/ collides with URL path) | URL-encode: urllib.parse.quote(smiles, safe="") |
Empty next page but total_count higher | Reached internal limit (typically 10000 with offset pagination) | Narrow filters (date range, target class) and re-paginate; or use the ChEMBL FTP downloads for >100K records |
HTTP 429 Too Many Requests | Burst pace | Add time.sleep(0.3); mount a Retry adapter (see Recipe) |
Mixed units in activity records | Different assays report in nM / µM / % inhibition | Filter standard_units="nM" and prefer pchembl_value for cross-assay comparison |
data_validity_comment is non-empty | Curation flag (e.g., "Potential transcription error", "Outside typical range") | Drop these rows before SAR/regression analysis |
| Duplicate activity records | Same measurement reported in multiple sources | Check potential_duplicate=True and dedupe |
pchembl_value for cross-study comparisons — it normalizes IC50/Ki/EC50 to a comparable -log10 scale.data_validity_comment before computing aggregates — flagged rows can skew distributions.standard_units="nM" in activity queries to avoid mixing nM with µM.page_meta.next for pagination instead of incrementing offset manually — the URL already carries the right cursor./similarity/{smiles}/..., /substructure/{smiles}) with urllib.parse.quote(smi, safe="").Session with retry adapter for batch work (see Recipe) — ChEMBL handles a fair amount of traffic and occasionally returns 502/503.pref_name__icontains — EGFR, BRAF, HER2 all return 0 hits. Use the spelled-out term or filter via target_components__accession=<UniProt> instead.rdkit-cheminformatics — SMILES manipulation, fingerprints, descriptorsdatamol-cheminformatics — molecular preprocessing & featurizationpubchem-compound-search — alternative compound database (NIH; broader coverage but less bioactivity depth)pdb-database — 3D structures of ChEMBL targets via RCSB PDB REST APIopentargets-database — links ChEMBL drug-target evidence to disease associationschembl_webresource_client PyPI package; this SKILL.md uses the underlying REST API directly so no SDK install is needed.© jaechang-hits, CC-BY-SA-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/structural-biology-drug-discovery/chembl-database-bioactivity of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Chembl Database Bioactivity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Chembl Database Bioactivity this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~6.4k | Automated safety check: Pass | CC-BY-SA-4.0 | |
| DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3k | Automated safety check: Notes | MIT | |
| Biopipelineslocbp-uzh/biopipelines | 109 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Pdb Databasedavila7/claude-code-templates | 33k | 9 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Tooluniverseynulihao/AgentSkillOS | 618 | 2 repos | ~2.5k | Automated safety check: Pass | None | |
| Chai1JimLiu/science-skills | 228 | 4 repos | ~1.2k | Automated safety check: Pass | Apache-2.0 |
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
locbp-uzh/biopipelines
Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…
davila7/claude-code-templates
Access RCSB PDB for 3D protein/nucleic acid structures. An agent skill from davila7/claude-code-templates.
ynulihao/AgentSkillOS
A skill your agent uses when working with scientific research tools and workflows across bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery.
JimLiu/science-skills
Structure prediction for protein, nucleic-acid, and small-molecule complexes with the Chai-1 foundation model (Chai Discovery 2024, github.com/chaidiscovery/chai-lab).
aipoch/medical-research-skills
Access the RCSB Protein Data Bank (PDB) to search, download, and programmatically retrieve 3D macromolecular structures and metadata; use when you need structure discovery (text/sequence/3D…
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required. Chembl Database Bioactivity is an agent skill from jaechang-hits/SciAgent-Skills. Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required.
Chembl Database Bioactivity fits situations like: tasks that involve Drug discovery and cheminformatics; tasks that involve Protein structure and design.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/chembl-database-bioactivity in jaechang-hits/SciAgent-Skills) into .claude/skills/chembl-database-bioactivity in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/chembl-database-bioactivity in jaechang-hits/SciAgent-Skills) into .agents/skills/chembl-database-bioactivity in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chembl-database-bioactivity, .gemini/skills/chembl-database-bioactivity, .github/skills/chembl-database-bioactivity and .opencode/skills/chembl-database-bioactivity in your project.
Going by SKILL.md and its folder, Chembl Database Bioactivity needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 2 domains. In commands or code: ebi.ac.uk; the agent is likely to contact it when it follows the instructions. As links in the text: chembl.gitbook.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Chembl Database Bioactivity is published under the CC-BY-SA-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Chembl Database Bioactivity: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars), Pdb Database (davila7/claude-code-templates, 33k stars) and Tooluniverse (ynulihao/AgentSkillOS, 618 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.