Agent skill

Chembl Database

by google-deepmind in google-deepmind/science-skills

Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.

Apache-2.0Auto-check passedResearch & Science

Install Chembl Database

skills CLI
$ npx skills add google-deepmind/science-skills --skill chembl-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-deepmind/science-skills chembl-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/chembl_database .claude/skills/chembl-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chembl-database
GitHub stars
3.2k
Used in
2 other repos
Token cost
~2.9k tokens
SKILL.md length
1,087 words
Files
4 (incl. scripts, references)
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.

  • Works in 10 steps: Check API Status → Molecule Queries → Target Queries → …
  • The user asks about compounds
  • SKILL.md covers Prerequisites, Core Rules, Utility Script and Common Options, plus 2 more sections
  • Runs Python scripts from its folder; calls uv

What it does

Chembl Database is an agent skill from google-deepmind/science-skills. Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures. Use when the user asks about compounds, targets, IC50/Ki values, drug mechanisms, or structure searches.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/api_endpoints.md` and `scripts/chembl_api.py`).

It sits in Research & Science. It works with Bash. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.

When your agent uses it

  • The user asks about compounds
  • Drug mechanisms
  • Structure searches

Example prompts

  • “/chembl-database”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Check API Status
  2. Molecule Queries
  3. Target Queries
  4. Bioactivity Data
  5. Drug Information
  6. Structure-Based Searches
  7. Compound Image
  8. Cross-Referencing with Other Databases
  9. Pagination
  10. Other Endpoints

What it can do on your machine

Read from SKILL.md and the folder at commit 6883275. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • chembl.gitbook.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chembl Database loads about 2.9k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 1,087 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google-deepmind/science-skills at commit 6883275, republished under its Apache-2.0 licence (© google-deepmind). 1,087 words, ~2,875 tokens.

Download SKILL.mdSave it as .claude/skills/chembl-database/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
chembl-database
description
Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures. Use when the user asks about compounds, targets, IC50/Ki values, drug mechanisms, or structure searches.

ChEMBL Database Query

Prerequisites

  1. uv: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH.
  2. User Notification: If .licenses/chembl_database_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://chembl.gitbook.io/chembl-interface-documentation/about, then (2) create the file recording the notification text and timestamp.

Core Rules

  • [!IMPORTANT] Use the Utility Scripts: You MUST ALWAYS use the provided utility script scripts/chembl_api.py for all ChEMBL API interactions, including checking status. NEVER use curl or custom Python requests to query the ChEMBL API directly. This ensures rate limit is enfoced and also retries on network errors.

  • Output to File (Required): The --output flag is required for every subcommand. All JSON results are written to the specified file. After running the command, read the output file with jq or your own code to extract the data. List results are typically wrapped in a JSON array keyed by the endpoint name (e.g., molecules, activities).

  • Notification: If this skill is used, ensure this is mentioned in the output.

Utility Script

All ChEMBL API queries use one script with subcommands:

bash
uv run scripts/chembl_api.py <subcommand> --output <file> [options]

1. Check API Status
bash
uv run scripts/chembl_api.py status --output /tmp/status.json

2. Molecule Queries

Fetch by ChEMBL ID: bash uv run scripts/chembl_api.py molecule --id CHEMBL25 --output /tmp/mol.json

Search by name: bash uv run scripts/chembl_api.py molecule --search "aspirin" --limit 3 --output /tmp/mol_search.json

Batch fetch: bash uv run scripts/chembl_api.py molecule --ids "CHEMBL25;CHEMBL1642" --limit 10 --output /tmp/mol_batch.json

Filter by properties: bash uv run scripts/chembl_api.py molecule --filter molecule_properties__mw_freebase__lte=500 --limit 5 --output /tmp/mol_filter.json

Filter by range: bash uv run scripts/chembl_api.py molecule --filter molecule_properties__mw_freebase__range=150,200 --limit 5 --output /tmp/mol_range.json

Download SDF structure file: bash uv run scripts/chembl_api.py molecule --id CHEMBL25 --dl_format sdf --output /tmp/aspirin.sdf

Tip: SDF/MOL files can be passed directly to tools like PyMOL or RDKit for 3D visualization and analysis.


3. Target Queries

Search for targets: bash uv run scripts/chembl_api.py target --search "EGFR" --limit 5 --output /tmp/targets.json

Fetch by ID: bash uv run scripts/chembl_api.py target --id CHEMBL203 --output /tmp/egfr.json


4. Bioactivity Data

Fetch activity by ID: bash uv run scripts/chembl_api.py activity --id 31863 --output /tmp/act.json

Search activities: bash uv run scripts/chembl_api.py activity --search "EGFR" --limit 5 --output /tmp/act_search.json

Filter activities for a target: bash uv run scripts/chembl_api.py activity --filter target_chembl_id=CHEMBL203 standard_type=IC50 --limit 10 --output /tmp/egfr_ic50.json

Normalize bioactivity units to nM: bash uv run scripts/chembl_api.py activity --filter target_chembl_id=CHEMBL203 standard_type=IC50 --limit 5 --normalize --output /tmp/egfr_normalized.json

Important: Bioactivity values come in various units (nM, µM, pM). Use --normalize to convert all values to nM for consistent comparison. Each record will include normalized_value_nM and normalization_note.


5. Drug Information

Fetch drug details: bash uv run scripts/chembl_api.py drug --id CHEMBL25 --output /tmp/drug.json

Drug indications: bash uv run scripts/chembl_api.py drug_indication --filter molecule_chembl_id=CHEMBL25 --limit 10 --output /tmp/indications.json

Filter indications by phase: bash uv run scripts/chembl_api.py drug_indication --filter molecule_chembl_id=CHEMBL25 max_phase_for_ind=4.0 --limit 10 --output /tmp/approved_indications.json

Drug warnings: bash uv run scripts/chembl_api.py drug_warning --limit 5 --output /tmp/warnings.json

Mechanisms of action: bash uv run scripts/chembl_api.py mechanism --filter molecule_chembl_id=CHEMBL25 --limit 5 --output /tmp/mech.json


6. Structure-Based Searches

Note: Both similarity and substructure searches are performed server-side on ChEMBL's pre-indexed database. They do not require a local RDKit installation.

Similarity search (SMILES + threshold): bash uv run scripts/chembl_api.py similarity --smiles "CC(=O)Oc1ccccc1C(=O)O" --similarity 85 --limit 5 --output /tmp/similar.json

Substructure search (SMILES): bash uv run scripts/chembl_api.py substructure --smiles "c1ccccc1" --limit 5 --output /tmp/substruct.json


7. Compound Image

Download a 2D structure image (SVG by default, scalable for publication):

bash
uv run scripts/chembl_api.py image --id CHEMBL25 --output /tmp/chembl25.svg

Options:

  • --dimensions: Image size in pixels (max 500, default 500).
  • --engine: Rendering engine (default: rdkit).
  • --img_format: Output format — svg (default, vector) or png (raster).

8. Cross-Referencing with Other Databases

ChEMBL integrates with UniProt, Ensembl, PubChem, and other databases. Common cross-referencing patterns:

Find a ChEMBL target from a UniProt accession: bash uv run scripts/chembl_api.py target --filter target_components__accession=P00533 --limit 5 --output /tmp/uniprot_target.json

Resolve any ChEMBL ID to its entity type: bash uv run scripts/chembl_api.py chembl_id_lookup --id CHEMBL203 --output /tmp/lookup.json

Look up cross-reference sources: bash uv run scripts/chembl_api.py xref_source --limit 10 --output /tmp/xrefs.json

Tip: Use the target_component endpoint to find UniProt accessions, gene names, and protein sequences for any ChEMBL target.


Show full SKILL.md (423 more words)Show less
9. Pagination

All list endpoints support --limit and --offset for pagination:

bash
# First page: 2 results starting at offset 0
uv run scripts/chembl_api.py molecule --limit 2 --offset 0 --output /tmp/page1.json

# Second page: next 2 results starting at offset 2
uv run scripts/chembl_api.py molecule --limit 2 --offset 2 --output /tmp/page2.json

The response includes page_meta with total_count, limit, offset, next, and previous links. Use successive --offset values to page through large result sets.


10. Other Endpoints

All remaining endpoints follow the same pattern:

bash
uv run scripts/chembl_api.py <subcommand> --output <file> [--id ID | --ids ID1;ID2 | --search QUERY] [--limit N] [--offset N] [--filter KEY=VAL ...]

Key subcommands at a glance:

  • molecule (searchable: true): Molecules/compounds — the primary entry point
  • target (searchable: true): Drug targets (proteins, organisms, etc.)
  • activity (searchable: true): Bioactivity data (IC50, Ki, EC50, etc.)
  • drug (searchable: false): Approved drugs
  • mechanism (searchable: false): Mechanisms of action
  • assay (searchable: true): Assay descriptions
  • similarity (searchable: false): Similarity search (special)
  • substructure (searchable: false): Substructure search (special)
  • image (searchable: false): Compound image download (special)

Full subcommand list:

  • activity_supp (searchable: false): Supplementary activity data
  • assay_class (searchable: false): Assay classifications
  • atc_class (searchable: false): ATC drug classifications
  • binding_site (searchable: false): Binding site information
  • biotherapeutic (searchable: false): Biotherapeutic molecules
  • cell_line (searchable: false): Cell line details
  • chembl_id_lookup (searchable: true): ChEMBL ID resolution
  • chembl_release (searchable: false): Database release info
  • compound_record (searchable: false): Compound records
  • compound_structural_alert (searchable: false): Structural alerts
  • document (searchable: true): Literature documents
  • document_similarity (searchable: false): Document similarity
  • drug_indication (searchable: false): Drug indications
  • drug_warning (searchable: false): Drug safety warnings
  • go_slim (searchable: false): GO slim terms
  • metabolism (searchable: false): Metabolism data
  • molecule_form (searchable: false): Molecule forms (salts/parents)
  • organism (searchable: false): Organisms
  • protein_classification (searchable: true): Protein classifications
  • source (searchable: false): Data sources
  • target_component (searchable: false): Target protein components
  • target_relation (searchable: false): Target relationships
  • tissue (searchable: false): Tissue types
  • xref_source (searchable: false): Cross-reference sources
  • status (searchable: false): API status check (special)

Common Options

  • --output FILE: Required. Output file path for JSON results.
  • --id ID: Fetch a single record by ID.
  • --ids ID1;ID2;...: Batch fetch multiple records.
  • --search QUERY: Free-text search (only for searchable endpoints, marked ✓).
  • --limit N: Max results to return (default: 5).
  • --offset N: Pagination offset.
  • --filter KEY=VAL: Filter parameters (can specify multiple).
  • --normalize: (activity only) Normalize values to nM.
  • --dl_format sdf|mol: (molecule only) Download structure file.

Reference

Workflow

  1. Use status --output /tmp/status.json to verify the API is available.
  2. Search for targets, molecules, or drugs using the relevant subcommand.
  3. Read the output JSON file to extract IDs and data.
  4. Use IDs from search results to fetch detailed records.
  5. Query activity with filters to get bioactivity data for targets/molecules. Use --normalize when comparing values across studies.
  6. Use similarity or substructure for server-side structure-based queries.
  7. Download compound images with image or structure files with molecule --dl_format sdf.
  8. Use target --filter target_components__accession=<UniProt> to cross- reference with UniProt.

© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/chembl_database of google-deepmind/science-skills.

  • SKILL.md
  • references/api_endpoints.md
  • references/citation.bib
  • scripts/chembl_api.py

Open the folder on GitHubat commit 6883275

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in google-deepmind/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Chembl Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chembl Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chembl Database this skillgoogle-deepmind/science-skills3.2k2 repos~2.9kAutomated safety check: PassApache-2.0
Tbtoolsxuzhougeng/wispterm440—~2.3kAutomated safety check: PassMIT
bioSkills InstallerGPTomics/bioSkills1.2k1 repos~789Automated safety check: PassMIT
Academic Research Mappertinyfish-io/tinyfish-cookbook2.2k—~3.5kAutomated safety check: PassMIT
Deep Researchteam-attention/hoyeon173—~6kAutomated safety check: PassMIT
Ds BaselineOpenLAIR/dr-claw1.2k—~6.9kAutomated safety check: PassMIT

Similar skills

  • Tbtools

    xuzhougeng/wispterm

    A skill your agent uses when the user asks about TBtools, TBtools-II, TBtools RPC API, TBtools CLI, or bioinformatics operations available through TBtools such as sequence manipulation, BLAST…

    440 GitHub stars~2.3k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Research & ScienceAuto-check passed
  • Academic Research Mapper

    tinyfish-io/tinyfish-cookbook

    Map the research landscape for any technical or academic topic by searching arXiv, Semantic Scholar, and Google Scholar in parallel.

    2.2k GitHub stars~3.5k tokensUpdated 6 days ago
    Research & ScienceAuto-check passed
  • Deep Research

    team-attention/hoyeon

    Deep web research skill using parallel subagents + chromux browser-explorer + Gemini.

    173 GitHub stars~6k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Ds Baseline

    OpenLAIR/dr-claw

    A skill your agent uses when a quest needs to attach, import, reproduce, repair, verify, compare, or publish a baseline and its metrics.

    1.2k GitHub stars~6.9k tokensUpdated 20 days ago
    Research & ScienceAuto-check passed
  • Launch the interactive web dashboard to visualize a codebase's knowledge graph

    136 GitHub starsUsed in 2 repos~1.9k tokens
    Knowledge ManagementAuto-check passed

More from google-deepmind/science-skills

All 40 skills in this repo
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Auto-check: notes
  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Auto-check: notes
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Auto-check: notes
  • Gnomad Database

    google-deepmind/science-skills

    Query the Genome Aggregation Database (gnomAD). An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~775 tokens
    Auto-check passed

Works with

Questions about Chembl Database

What does Chembl Database do?

Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures. Chembl Database is an agent skill from google-deepmind/science-skills. Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.

When should I use Chembl Database?

Chembl Database fits situations like: the user asks about compounds; drug mechanisms; structure searches.

How do I install Chembl Database in Claude Code?

Run `npx skills add google-deepmind/science-skills --skill chembl-database -a claude-code`. Or copy the skill folder (skills/chembl_database in google-deepmind/science-skills) into .claude/skills/chembl-database in your project. Claude Code loads it when a task matches its description.

How do I install Chembl Database in Codex?

Run `npx skills add google-deepmind/science-skills --skill chembl-database -a codex`. Or copy the skill folder (skills/chembl_database in google-deepmind/science-skills) into .agents/skills/chembl-database in your project. Codex loads it when a task matches its description.

Can I use Chembl Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill chembl-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chembl-database, .gemini/skills/chembl-database, .github/skills/chembl-database and .opencode/skills/chembl-database in your project.

What does Chembl Database need to run?

Going by SKILL.md and its folder, Chembl Database needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Chembl Database access the network?

SKILL.md names 1 domain. As links in the text: chembl.gitbook.io. This is read from the text; nothing was executed.

Is Chembl Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Chembl Database use?

Chembl Database is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chembl Database use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Chembl Database?

Skills that share tags, products or a category with Chembl Database: Tbtools (xuzhougeng/wispterm, 440 stars), bioSkills Installer (GPTomics/bioSkills, 1.2k stars), Academic Research Mapper (tinyfish-io/tinyfish-cookbook, 2.2k stars) and Deep Research (team-attention/hoyeon, 173 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chembl Database?

google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,216 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on September 15, 2026.

Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.