Agent skill

Chembl Database Bioactivity

by jaechang-hits in jaechang-hits/SciAgent-Skills

Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required.

CC-BY-SA-4.0Auto-check passedResearch & Science

Install Chembl Database Bioactivity

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills chembl-database-bioactivity --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/chembl-database-bioactivity .claude/skills/chembl-database-bioactivity && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chembl-database-bioactivity
GitHub stars
374
Used in
1 other repo
Token cost
~6.4k tokens
SKILL.md length
1,177 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
CC-BY-SA-4.0

At a glance

Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required.

  • Works in 5 steps: Molecule Queries → Target Queries → Bioactivity Data → …
  • Tasks that involve Drug discovery and cheminformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 7 more sections
  • Calls pip; reaches ebi.ac.uk

What it does

Chembl Database Bioactivity is an agent skill from jaechang-hits/SciAgent-Skills. Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required. Search compounds, retrieve IC50/Ki/EC50 bioactivities, find target inhibitors, run SAR, access drug mechanism/indication data.

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Drug discovery and cheminformatics and Protein structure and design. It works with Django. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-SA-4.0.

When your agent uses it

  • Tasks that involve Drug discovery and cheminformatics
  • Tasks that involve Protein structure and design

Example prompts

  • “/chembl-database-bioactivity”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Molecule Queries
  2. Target Queries
  3. Bioactivity Data
  4. Structure-Based Search
  5. Drug and Mechanism Data

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ebi.ac.uk

    Also links to:

    • chembl.gitbook.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chembl Database Bioactivity loads about 6.4k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 1,177 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-SA-4.0 licence (© jaechang-hits). 1,177 words, ~6,426 tokens.

Download SKILL.mdSave it as .claude/skills/chembl-database-bioactivity/SKILL.md (or your agent's skills folder).
name
chembl-database-bioactivity
description
Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain `requests` — no SDK install required. Search compounds, retrieve IC50/Ki/EC50 bioactivities, find target inhibitors, run SAR, access drug mechanism/indication data.
license
CC-BY-SA-3.0

ChEMBL Database — Bioactivity Queries

Why no SDK? The chembl_webresource_client package is convenient sugar over a public, no-auth REST/JSON API at https://www.ebi.ac.uk/chembl/api/data/. When the SDK is unavailable, every operation can be reproduced with plain requests and URL parameters. This SKILL.md uses the REST path throughout so the code runs in any environment with requests installed. Django-style filter syntax (field__icontains=…, field__lte=…, field__range=a,b) works as URL query parameters.

Overview

ChEMBL is EMBL-EBI's bioactive molecule database: 2M+ compounds, 19M+ bioactivity measurements (IC50, Ki, EC50, Kd, …), 13K+ targets. The REST API at https://www.ebi.ac.uk/chembl/api/data/ returns JSON (append .json) or XML/YAML, requires no authentication, and supports Django-style query filters via URL parameters plus cursor-style pagination via page_meta.next.

When to Use

  • Finding compounds by name, ChEMBL ID, or physicochemical properties
  • Querying bioactivity data (IC50, Ki, EC50) for specific targets
  • Performing similarity or substructure searches using SMILES
  • Retrieving drug mechanisms of action and clinical indications
  • Identifying inhibitors, agonists, or bioactive molecules for a target
  • Analyzing structure-activity relationships (SAR) across compound series
  • Filtering molecules by Lipinski rule-of-5 or other drug-likeness criteria
  • For general cheminformatics (SMILES manipulation, fingerprints, descriptors) use rdkit-cheminformatics instead
  • For an alternative compound database (NIH, broader coverage) use pubchem-compound-search

Prerequisites

  • Python packages: requests (only requirement). Optional: pandas for tabular analysis.
  • No API key required: ChEMBL is freely accessible.
  • Rate limits: No published hard limit. The infrastructure is shared — add time.sleep(0.2-0.5) between requests in batch loops; back off on HTTP 429.
bash
pip install requests
# Optional, for DataFrame work:
pip install pandas

Quick Start

python
import requests

BASE = "https://www.ebi.ac.uk/chembl/api/data"

# Retrieve a molecule by ChEMBL ID
r = requests.get(f"{BASE}/molecule/CHEMBL25.json", timeout=15)
r.raise_for_status()
aspirin = r.json()
print(f"{aspirin['pref_name']}: MW={aspirin['molecule_properties']['mw_freebase']}")
# ASPIRIN: MW=180.16

# Search targets by full name (acronyms like 'EGFR' don't match pref_name — use full term)
r = requests.get(
    f"{BASE}/target.json",
    params={"pref_name__icontains": "epidermal growth factor receptor",
            "target_type": "SINGLE PROTEIN", "limit": 5},
    timeout=15,
)
targets = r.json()["targets"]
print(f"EGFR-like targets: {len(targets)}, first={targets[0]['target_chembl_id']}")

# Potent bioactivities: EGFR (CHEMBL203) IC50 <= 100 nM
r = requests.get(
    f"{BASE}/activity.json",
    params={"target_chembl_id": "CHEMBL203",
            "standard_type": "IC50",
            "standard_value__lte": 100,
            "standard_units": "nM",
            "limit": 5},
    timeout=30,
)
data = r.json()
print(f"EGFR IC50 ≤ 100 nM records: {data['page_meta']['total_count']}")

Key Concepts

Filter Operators (Django-style, as URL parameters)

The SDK's field__operator=value syntax maps 1:1 to URL query parameters. Use & to combine filters.

OperatorURL patternExample URL fragment
__exactfield=valuetarget_type=SINGLE+PROTEIN
__iexactfield__iexact=valuepref_name__iexact=aspirin
__contains / __icontainsfield__icontains=valuepref_name__icontains=kinase
__startswith / __endswithfield__startswith=Epipref_name__endswith=nib
__gt / __gte / __lt / __ltefield__lte=100standard_value__lte=100
__rangefield__range=lo,himolecule_properties__mw_freebase__range=300,500
__infield__in=a,b,cstandard_type__in=IC50,Ki,Kd
__isnullfield__isnull=False (Python False/True strings)pchembl_value__isnull=False
__regexfield__regex=…pref_name__regex=^EGF.*kinase$
__searchfield__search=…description__search=apoptosis

When passed via requests.get(..., params={...}), the library handles URL encoding automatically (including the commas in __range and __in).

Core Endpoints

All endpoints accept .json, .xml, or .yaml suffix. JSON is the default below.

Endpoint URLReturnsKey fields
/molecule/{chembl_id}.jsonCompound by IDpref_name, molecule_chembl_id, molecule_properties, molecule_structures
/molecule.json?<filters>Compound searchpaginated molecules[]
/target/{chembl_id}.jsonTarget by IDpref_name, target_type, organism, target_components
/target.json?<filters>Target searchpaginated targets[]
/activity.json?<filters>Bioactivity recordspaginated activities[]
/assay.json?<filters>Assay detailspaginated assays[]
/drug.json?<filters>Approved drug infopaginated drugs[]; supports /drug/{chembl_id}.json
/mechanism.json?<filters>Mechanism of actionpaginated mechanisms[]
/drug_indication.json?<filters>Therapeutic indicationspaginated drug_indications[]
/similarity/{smiles}/{tanimoto}.jsonTanimoto similarity (0–100)paginated molecules[] with similarity field
/substructure/{smiles}.jsonSubstructure searchpaginated molecules[]
/image/{chembl_id}.svgSVG structure imagebinary SVG (NOT JSON)
/molecule_form/{chembl_id}.jsonParent/salt formsmolecule_forms[]
/protein_class.jsonProtein classification hierarchyhierarchical browse
/document.json?<filters>Literature source recordspaginated documents[]
Response Shape
jsonc
{
  "page_meta": {
    "limit": 20,
    "offset": 0,
    "total_count": 12145,
    "next": "/chembl/api/data/activity.json?...&offset=20",
    "previous": null
  },
  "activities": [ /* or molecules[], targets[], etc. */ ]
}

Walk page_meta.next (a relative URL — prefix with https://www.ebi.ac.uk) until it becomes null.

Molecular Properties

Properties accessible via molecule_properties on each record:

FieldDescription
mw_freebaseMolecular weight (free base)
full_mwtFull molecular weight (including salts)
alogpCalculated LogP
hbaHydrogen bond acceptors
hbdHydrogen bond donors
psaPolar surface area
rtbRotatable bonds
num_ro5_violationsLipinski rule-of-5 violations
ro3_passRule of 3 compliance
cx_most_apka / cx_most_bpkaMost acidic / basic pKa
Target Information Fields
FieldDescription
target_chembl_idChEMBL target identifier
pref_namePreferred (full) target name — acronyms like "EGFR" do NOT match; use the spelled-out term
target_typeSINGLE PROTEIN, PROTEIN COMPLEX, ORGANISM, …
organismTarget organism (e.g., Homo sapiens)
tax_idNCBI taxonomy ID
target_components[]Components (UniProt accession, sequence, …)
Bioactivity Data Fields
FieldDescription
standard_typeActivity type: IC50, Ki, Kd, EC50, …
standard_valueNumerical activity value
standard_unitsUnits: nM, uM, …
pchembl_valueNormalized -log10 activity (>6 = potent)
activity_commentActivity annotations
data_validity_commentData quality flags (check before analysis)
potential_duplicateDuplicate flag

Core API

1. Molecule Queries
python
import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"

# By ChEMBL ID
r = requests.get(f"{BASE}/molecule/CHEMBL25.json", timeout=15)
aspirin = r.json()
print(f"{aspirin['pref_name']}: MW={aspirin['molecule_properties']['mw_freebase']}")

# By name (case-insensitive substring)
r = requests.get(f"{BASE}/molecule.json",
                 params={"pref_name__icontains": "imatinib", "limit": 5},
                 timeout=15)
for mol in r.json()["molecules"]:
    print(f"  {mol['molecule_chembl_id']}  {mol.get('pref_name')!r}")

# By Lipinski-compliant property ranges
r = requests.get(f"{BASE}/molecule.json",
                 params={"molecule_properties__mw_freebase__range": "300,500",
                         "molecule_properties__alogp__lte": 5,
                         "molecule_properties__hba__lte": 10,
                         "molecule_properties__hbd__lte": 5,
                         "limit": 3},
                 timeout=15)
print(f"Lipinski-compliant total: {r.json()['page_meta']['total_count']}")
2. Target Queries
python
import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"

# By ChEMBL ID
r = requests.get(f"{BASE}/target/CHEMBL203.json", timeout=15)
egfr = r.json()
print(f"{egfr['pref_name']} ({egfr['organism']}) — type={egfr['target_type']}")

# Search by full name (NOT acronym) + type
r = requests.get(f"{BASE}/target.json",
                 params={"pref_name__icontains": "kinase",
                         "target_type": "SINGLE PROTEIN", "limit": 5},
                 timeout=15)
d = r.json()
print(f"Kinase SINGLE_PROTEIN targets: total={d['page_meta']['total_count']}")
for t in d["targets"][:5]:
    print(f"  {t['target_chembl_id']:12s} {t.get('pref_name')!r}  ({t['organism']})")

# By organism
r = requests.get(f"{BASE}/target.json",
                 params={"organism": "Homo sapiens", "limit": 3},
                 timeout=15)
print(f"Human targets: total={r.json()['page_meta']['total_count']}")
3. Bioactivity Data
python
import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"

# Potent inhibitors for a target (EGFR = CHEMBL203)
r = requests.get(f"{BASE}/activity.json",
                 params={"target_chembl_id": "CHEMBL203",
                         "standard_type": "IC50",
                         "standard_value__lte": 100,
                         "standard_units": "nM",
                         "limit": 5},
                 timeout=30)
data = r.json()
print(f"EGFR IC50≤100nM: total={data['page_meta']['total_count']}")
for act in data["activities"][:5]:
    print(f"  {act['molecule_chembl_id']:14s} IC50={act['standard_value']} nM "
          f"pChEMBL={act.get('pchembl_value')}")

# All pChEMBL-tagged activities for a compound
r = requests.get(f"{BASE}/activity.json",
                 params={"molecule_chembl_id": "CHEMBL25",
                         "pchembl_value__isnull": "False",
                         "limit": 5},
                 timeout=30)
print(f"Aspirin pChEMBL activities: total={r.json()['page_meta']['total_count']}")

# Multiple activity types (CHEMBL240 = D2 dopamine receptor)
r = requests.get(f"{BASE}/activity.json",
                 params={"target_chembl_id": "CHEMBL240",
                         "standard_type__in": "IC50,Ki,Kd",
                         "limit": 5},
                 timeout=30)
print(f"D2 receptor IC50/Ki/Kd: total={r.json()['page_meta']['total_count']}")
python
import requests
from urllib.parse import quote
BASE = "https://www.ebi.ac.uk/chembl/api/data"

# Similarity search (Tanimoto ≥ 85%)
# Path-style endpoint: /similarity/{smiles}/{threshold}
# The SMILES MUST be URL-encoded (it contains '/', '(', ')' etc.)
aspirin_smiles = quote("CC(=O)Oc1ccccc1C(=O)O", safe="")
r = requests.get(f"{BASE}/similarity/{aspirin_smiles}/85.json",
                 params={"limit": 5}, timeout=30)
data = r.json()
print(f"Similar to aspirin (≥85% Tanimoto): total={data['page_meta']['total_count']}")
for m in data["molecules"][:5]:
    print(f"  {m['molecule_chembl_id']}  similarity={m.get('similarity')}")
python
# Substructure search
benzimidazole = quote("c1ccc2[nH]cnc2c1", safe="")
r = requests.get(f"{BASE}/substructure/{benzimidazole}.json",
                 params={"limit": 3}, timeout=30)
print(f"Benzimidazole substructure total: {r.json()['page_meta']['total_count']}")
5. Drug and Mechanism Data
python
import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"

# Drug record (max clinical phase, ATC class, etc.)
r = requests.get(f"{BASE}/drug/CHEMBL941.json", timeout=15)   # imatinib
drug = r.json()
print(f"Imatinib max_phase={drug.get('max_phase')}")

# Mechanisms of action — note: not every drug has mechanism records.
# Imatinib (CHEMBL941) returns 0 mechanism rows; sunitinib (CHEMBL535) has many.
r = requests.get(f"{BASE}/mechanism.json",
                 params={"molecule_chembl_id": "CHEMBL535"}, timeout=15)
for m in r.json()["mechanisms"]:
    print(f"  {m['mechanism_of_action']} → target {m.get('target_chembl_id')}")

# Therapeutic indications
r = requests.get(f"{BASE}/drug_indication.json",
                 params={"molecule_chembl_id": "CHEMBL941", "limit": 5},
                 timeout=15)
for ind in r.json()["drug_indications"]:
    print(f"  {ind.get('mesh_heading')!r}  max_phase_for_ind={ind.get('max_phase_for_ind')}")
python
# SVG molecular structure image — direct binary response, NOT JSON
# Do NOT call /image/{cid}.json — that endpoint raises JSONDecodeError.
r = requests.get(f"{BASE}/image/CHEMBL25.svg", timeout=15)
r.raise_for_status()
with open("aspirin.svg", "w") as f:
    f.write(r.text)
print(f"Saved aspirin.svg ({len(r.text)} bytes, looks_svg={'<svg' in r.text})")

Common Workflows

Workflow 1: Find Inhibitors for a Target

Note: pref_name__icontains matches the spelled-out name. Acronyms like 'EGFR' or 'BRAF' return 0 results — use 'epidermal growth factor receptor' or 'B-raf' (with the hyphen).

python
import requests, pandas as pd, time
BASE = "https://www.ebi.ac.uk/chembl/api/data"

# Step 1: Resolve the target by full name
r = requests.get(f"{BASE}/target.json",
                 params={"pref_name__icontains": "B-raf",
                         "target_type": "SINGLE PROTEIN", "limit": 5},
                 timeout=15)
targets = r.json()["targets"]
human_braf = next(t for t in targets if t["organism"] == "Homo sapiens")
target_id = human_braf["target_chembl_id"]
print(f"Using {target_id} — {human_braf['pref_name']}")

# Step 2: Paginate all potent IC50 activities (cap at 500 for demo)
url = (f"{BASE}/activity.json"
       f"?target_chembl_id={target_id}"
       f"&standard_type=IC50"
       f"&standard_value__lte=100"
       f"&standard_units=nM"
       f"&pchembl_value__isnull=False"
       f"&limit=200")
records = []
while url and len(records) < 500:
    r = requests.get(url, timeout=30)
    r.raise_for_status()
    data = r.json()
    records.extend(data["activities"])
    nxt = data["page_meta"].get("next")
    url = f"https://www.ebi.ac.uk{nxt}" if nxt else None
    time.sleep(0.2)

df = pd.DataFrame(records)
df["standard_value"] = pd.to_numeric(df["standard_value"])
print(f"Retrieved {len(df)} potent {target_id} compounds")
print(df[["molecule_chembl_id", "standard_value", "pchembl_value"]].head(10))
Workflow 2: Analyze a Known Drug
python
import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"

# Sunitinib (CHEMBL535) — has documented mechanisms + indications.
# Imatinib (CHEMBL941) sometimes returns 0 mechanism rows depending on ChEMBL release.
chembl_id = "CHEMBL535"

# Molecule record
m = requests.get(f"{BASE}/molecule/{chembl_id}.json", timeout=15).json()
print(f"Name: {m['pref_name']}")
print(f"MW  : {m['molecule_properties']['mw_freebase']}")

# Mechanisms
mechs = requests.get(f"{BASE}/mechanism.json",
                     params={"molecule_chembl_id": chembl_id},
                     timeout=15).json()["mechanisms"]
for mc in mechs:
    print(f"  Mechanism: {mc['mechanism_of_action']}")

# Indications
inds = requests.get(f"{BASE}/drug_indication.json",
                    params={"molecule_chembl_id": chembl_id, "limit": 5},
                    timeout=15).json()["drug_indications"]
for ind in inds:
    print(f"  Indication: {ind.get('mesh_heading')}  "
          f"(Phase {ind.get('max_phase_for_ind')})")

# Bioactivity record count
total = requests.get(f"{BASE}/activity.json",
                     params={"molecule_chembl_id": chembl_id,
                             "pchembl_value__isnull": "False",
                             "limit": 1},
                     timeout=30).json()["page_meta"]["total_count"]
print(f"Total bioactivity records (pChEMBL-tagged): {total}")
Workflow 3: SAR Study
python
import requests, pandas as pd, time
from urllib.parse import quote
BASE = "https://www.ebi.ac.uk/chembl/api/data"

# Step 1: Similar compounds to a lead (e.g., quinoline scaffold)
lead_smiles = "c1ccc2c(c1)cc(nc2N)c3ccc(cc3)NC(=O)c4ccccc4"
r = requests.get(f"{BASE}/similarity/{quote(lead_smiles, safe='')}/80.json",
                 params={"limit": 20}, timeout=30)
analogs = r.json()["molecules"]
print(f"Analogs found: {len(analogs)}")

# Step 2: Collect bioactivities for each analog
records = []
for compound in analogs[:20]:
    cid = compound["molecule_chembl_id"]
    acts = requests.get(f"{BASE}/activity.json",
                        params={"molecule_chembl_id": cid,
                                "standard_type": "IC50",
                                "pchembl_value__isnull": "False",
                                "limit": 20},
                        timeout=30).json()["activities"]
    for act in acts:
        records.append({
            "chembl_id": cid,
            "target": act.get("target_pref_name"),
            "IC50_nM": act.get("standard_value"),
            "pchembl": act.get("pchembl_value"),
            "mw":    (compound.get("molecule_properties") or {}).get("mw_freebase"),
            "alogp": (compound.get("molecule_properties") or {}).get("alogp"),
        })
    time.sleep(0.2)

df = pd.DataFrame(records)
if not df.empty:
    df["IC50_nM"] = pd.to_numeric(df["IC50_nM"])
    print(df.groupby("target")["IC50_nM"].describe())

Common Recipes

Recipe: Virtual Screening Filter (Lipinski rule-of-5)
python
import requests
BASE = "https://www.ebi.ac.uk/chembl/api/data"
r = requests.get(f"{BASE}/molecule.json",
                 params={"molecule_properties__mw_freebase__range": "300,500",
                         "molecule_properties__alogp__lte": 5,
                         "molecule_properties__hba__lte": 10,
                         "molecule_properties__hbd__lte": 5,
                         "molecule_properties__num_ro5_violations": 0,
                         "limit": 1},
                 timeout=15)
print(f"Drug-like candidates: {r.json()['page_meta']['total_count']}")
Recipe: Paginate Activities to CSV
python
import requests, pandas as pd, time
BASE = "https://www.ebi.ac.uk/chembl/api/data"

url = (f"{BASE}/activity.json"
       f"?target_chembl_id=CHEMBL203"
       f"&standard_type=IC50"
       f"&pchembl_value__isnull=False"
       f"&limit=500")
all_acts = []
while url:
    r = requests.get(url, timeout=60)
    r.raise_for_status()
    data = r.json()
    all_acts.extend(data["activities"])
    nxt = data["page_meta"].get("next")
    url = f"https://www.ebi.ac.uk{nxt}" if nxt else None
    time.sleep(0.3)

df = pd.DataFrame(all_acts)
df.to_csv("egfr_activities.csv", index=False)
print(f"Exported {len(df)} records → egfr_activities.csv")
Recipe: Robust Session with Retries
python
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

def chembl_session(retries=3, backoff=1.0):
    s = requests.Session()
    s.headers.update({"Accept": "application/json"})
    s.mount("https://", HTTPAdapter(max_retries=Retry(
        total=retries, backoff_factor=backoff,
        status_forcelist=[429, 500, 502, 503, 504],
        allowed_methods=["GET"])))
    return s

session = chembl_session()
r = session.get("https://www.ebi.ac.uk/chembl/api/data/molecule/CHEMBL25.json", timeout=15)
print(r.json()["pref_name"])
Recipe: Download SVG Structure Image
python
import requests
r = requests.get("https://www.ebi.ac.uk/chembl/api/data/image/CHEMBL25.svg", timeout=15)
r.raise_for_status()
with open("aspirin.svg", "w") as f:
    f.write(r.text)
Show full SKILL.md (508 more words)Show less

Key Parameters

ParameterEndpointDefaultDescription
limitall list endpoints20Page size; max 1000
offsetall list endpoints0Pagination offset (or follow page_meta.next)
formatall endpointsjson (via .json suffix)Also .xml, .yaml
pref_name__icontains/target, /molecule—Substring on full name; acronyms don't match, use full term
target_chembl_id/activity—E.g., CHEMBL203 (EGFR), CHEMBL240 (D2 receptor)
molecule_chembl_id/activity, /mechanism, /drug_indication—E.g., CHEMBL25 (aspirin)
standard_type/activity—IC50, Ki, Kd, EC50
standard_value__lte/activity—Max activity value (paired with standard_units)
pchembl_value__isnull/activity—"False" to require pChEMBL-tagged data
target_type/target—SINGLE PROTEIN, PROTEIN COMPLEX, ORGANISM, …
{tanimoto} (path)/similarity/{smiles}/{tanimoto}—0–100 Tanimoto threshold
{smiles} (path)/similarity, /substructure—URL-encoded SMILES (urllib.parse.quote(s, safe=""))

Troubleshooting

ProblemCauseSolution
pref_name__icontains=EGFR (or BRAF) returns 0ChEMBL stores spelled-out names; acronyms don't matchUse "epidermal growth factor receptor"; for BRAF use "B-raf" with the hyphen
mechanism.json?molecule_chembl_id=CHEMBL941 returns emptyNot every drug has mechanism rows in every release (e.g., imatinib has 0 in current data)Use CHEMBL535 (sunitinib) or CHEMBL192 (sildenafil) as known-populated examples
JSONDecodeError on /image/{cid}.jsonThe image endpoint is binary, not JSONAlways use .svg or .png suffix: /image/{cid}.svg
404 on /molecule/{id}Invalid ChEMBL ID formatIDs must include the prefix: CHEMBL25, not 25
400 on similarity searchUnencoded SMILES (/ collides with URL path)URL-encode: urllib.parse.quote(smiles, safe="")
Empty next page but total_count higherReached internal limit (typically 10000 with offset pagination)Narrow filters (date range, target class) and re-paginate; or use the ChEMBL FTP downloads for >100K records
HTTP 429 Too Many RequestsBurst paceAdd time.sleep(0.3); mount a Retry adapter (see Recipe)
Mixed units in activity recordsDifferent assays report in nM / µM / % inhibitionFilter standard_units="nM" and prefer pchembl_value for cross-assay comparison
data_validity_comment is non-emptyCuration flag (e.g., "Potential transcription error", "Outside typical range")Drop these rows before SAR/regression analysis
Duplicate activity recordsSame measurement reported in multiple sourcesCheck potential_duplicate=True and dedupe

Best Practices

  • Use pchembl_value for cross-study comparisons — it normalizes IC50/Ki/EC50 to a comparable -log10 scale.
  • Always check data_validity_comment before computing aggregates — flagged rows can skew distributions.
  • Pin standard_units="nM" in activity queries to avoid mixing nM with µM.
  • Follow page_meta.next for pagination instead of incrementing offset manually — the URL already carries the right cursor.
  • URL-encode SMILES in path-style endpoints (/similarity/{smiles}/..., /substructure/{smiles}) with urllib.parse.quote(smi, safe="").
  • Use a Session with retry adapter for batch work (see Recipe) — ChEMBL handles a fair amount of traffic and occasionally returns 502/503.
  • For >100K records prefer the ChEMBL FTP downloads over paginated API calls.
  • Be deliberate about acronyms in pref_name__icontains — EGFR, BRAF, HER2 all return 0 hits. Use the spelled-out term or filter via target_components__accession=<UniProt> instead.
  • rdkit-cheminformatics — SMILES manipulation, fingerprints, descriptors
  • datamol-cheminformatics — molecular preprocessing & featurization
  • pubchem-compound-search — alternative compound database (NIH; broader coverage but less bioactivity depth)
  • pdb-database — 3D structures of ChEMBL targets via RCSB PDB REST API
  • opentargets-database — links ChEMBL drug-target evidence to disease associations

References

© jaechang-hits, CC-BY-SA-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/structural-biology-drug-discovery/chembl-database-bioactivity of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Chembl Database Bioactivity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chembl Database Bioactivity compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chembl Database Bioactivity this skilljaechang-hits/SciAgent-Skills3741 repos~6.4kAutomated safety check: PassCC-BY-SA-4.0
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT
Pdb Databasedavila7/claude-code-templates33k9 repos~2.3kAutomated safety check: PassMIT
Tooluniverseynulihao/AgentSkillOS6182 repos~2.5kAutomated safety check: PassNone
Chai1JimLiu/science-skills2284 repos~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Pdb Database

    davila7/claude-code-templates

    Access RCSB PDB for 3D protein/nucleic acid structures. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 9 repos~2.3k tokens
    Research & ScienceAuto-check passed
  • Tooluniverse

    ynulihao/AgentSkillOS

    A skill your agent uses when working with scientific research tools and workflows across bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery.

    618 GitHub starsUsed in 2 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Chai1

    JimLiu/science-skills

    Structure prediction for protein, nucleic-acid, and small-molecule complexes with the Chai-1 foundation model (Chai Discovery 2024, github.com/chaidiscovery/chai-lab).

    228 GitHub starsUsed in 4 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Pdb Database

    aipoch/medical-research-skills

    Access the RCSB Protein Data Bank (PDB) to search, download, and programmatically retrieve 3D macromolecular structures and metadata; use when you need structure discovery (text/sequence/3D…

    1.9k GitHub stars~1.6k tokensUpdated 23 days ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 11 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 11 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 11 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 11 days ago
    Auto-check passed

Works with

Questions about Chembl Database Bioactivity

What does Chembl Database Bioactivity do?

Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required. Chembl Database Bioactivity is an agent skill from jaechang-hits/SciAgent-Skills. Query ChEMBL (2M+ compounds, 19M+ bioactivity measurements, 13K+ targets) via the public REST/JSON API with plain requests — no SDK install required.

When should I use Chembl Database Bioactivity?

Chembl Database Bioactivity fits situations like: tasks that involve Drug discovery and cheminformatics; tasks that involve Protein structure and design.

How do I install Chembl Database Bioactivity in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/chembl-database-bioactivity in jaechang-hits/SciAgent-Skills) into .claude/skills/chembl-database-bioactivity in your project. Claude Code loads it when a task matches its description.

How do I install Chembl Database Bioactivity in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/chembl-database-bioactivity in jaechang-hits/SciAgent-Skills) into .agents/skills/chembl-database-bioactivity in your project. Codex loads it when a task matches its description.

Can I use Chembl Database Bioactivity in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill chembl-database-bioactivity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chembl-database-bioactivity, .gemini/skills/chembl-database-bioactivity, .github/skills/chembl-database-bioactivity and .opencode/skills/chembl-database-bioactivity in your project.

What does Chembl Database Bioactivity need to run?

Going by SKILL.md and its folder, Chembl Database Bioactivity needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Chembl Database Bioactivity access the network?

SKILL.md names 2 domains. In commands or code: ebi.ac.uk; the agent is likely to contact it when it follows the instructions. As links in the text: chembl.gitbook.io. This is read from the text; nothing was executed.

Is Chembl Database Bioactivity safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Chembl Database Bioactivity use?

Chembl Database Bioactivity is published under the CC-BY-SA-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chembl Database Bioactivity use?

About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Chembl Database Bioactivity?

Skills that share tags, products or a category with Chembl Database Bioactivity: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars), Pdb Database (davila7/claude-code-templates, 33k stars) and Tooluniverse (ynulihao/AgentSkillOS, 618 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chembl Database Bioactivity?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.