Agent skill

Zinc Database

by aipoch in aipoch/medical-research-skills

Access the ZINC (230M+ purchasable compounds) database when you need to look up compounds by ZINC ID/SMILES, run similarity/analog searches, or download 3D ready-to-dock structures for virtual…

MITAuto-check passedResearch & Science

Install Zinc Database

skills CLI
$ npx skills add aipoch/medical-research-skills --skill zinc-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills zinc-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Evidence Insight/zinc-database' .claude/skills/zinc-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
zinc-database
GitHub stars
2k
Token cost
~1.9k tokens
SKILL.md length
431 words
Files
3 (incl. references)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

Access the ZINC (230M+ purchasable compounds) database when you need to look up compounds by ZINC ID/SMILES, run similarity/analog searches, or download 3D ready-to-dock structures for virtual…

  • Works in 5 steps: Build a virtual screening library by… → Retrieve compounds by identifier (ZINC… → Search by structure (SMILES) to find… → …
  • Tasks that involve Drug discovery and cheminformatics
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 1 more section
  • Calls curl; reaches cartblanche22.docking.org

What it does

Zinc Database is an agent skill from aipoch/medical-research-skills. Access the ZINC (230M+ purchasable compounds) database when you need to look up compounds by ZINC ID/SMILES, run similarity/analog searches, or download 3D ready-to-dock structures for virtual screening and drug discovery.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/api_reference.md` and `zinc-database_audit_result_v1.json`).

It sits in Research & Science, covering Drug discovery and cheminformatics. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Tasks that involve Drug discovery and cheminformatics

Example prompts

  • “/zinc-database”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Build a virtual screening library by sampling purchasable compounds (e.g., fragment/lead-like/drug-like subsets).
  2. Retrieve compounds by identifier (ZINC ID) for follow-up analysis, procurement, or reporting.
  3. Search by structure (SMILES) to find exact matches or analogs via similarity thresholds.
  4. Validate supplier availability by querying supplier/catalog identifiers and mapping them to ZINC entries.
  5. Download docking-ready 3D structures (e.g., MOL2/SDF/DB2) organized by ZINC tranches for docking pipelines.

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • cartblanche22.docking.org

    Also links to:

    • files.docking.org
    • zinc.docking.org
    • wiki.docking.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Zinc Database loads about 1.9k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 431 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 431 words, ~1,916 tokens.

Download SKILL.mdSave it as .claude/skills/zinc-database/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
zinc-database
description
Access the ZINC (230M+ purchasable compounds) database when you need to look up compounds by ZINC ID/SMILES, run similarity/analog searches, or download 3D ready-to-dock structures for virtual screening and drug discovery.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

Use this skill when you need to:

  1. Build a virtual screening library by sampling purchasable compounds (e.g., fragment/lead-like/drug-like subsets).
  2. Retrieve compounds by identifier (ZINC ID) for follow-up analysis, procurement, or reporting.
  3. Search by structure (SMILES) to find exact matches or analogs via similarity thresholds.
  4. Validate supplier availability by querying supplier/catalog identifiers and mapping them to ZINC entries.
  5. Download docking-ready 3D structures (e.g., MOL2/SDF/DB2) organized by ZINC tranches for docking pipelines.

Key Features

  • ZINC22 access (CartBlanche22 web + API) for large-scale purchasable chemical space.
  • Lookup by ZINC ID (single or batch).
  • SMILES search with optional similarity/analog expansion via distance parameters.
  • Supplier/catalog queries to cross-reference vendor codes and catalogs.
  • Random sampling for benchmarking, diversity sampling, and screening set generation.
  • Property-aware filtering using tranche codes (H-bond donors, LogP, MW, reactivity phase).
  • 3D structure downloads from the ZINC22 files library (tranche-organized).

Dependencies

  • curl (tested with 7.70+)
  • Python >=3.9
  • pandas>=2.0.0 (parsing tabular API output)
  • (optional) requests>=2.31.0 (if replacing curl with native HTTP)
  • (optional) rdkit>=2023.09.1 (structure validation, fingerprints, downstream cheminformatics)

Example Usage

The following example is a complete runnable script that:

  1. queries by ZINC ID, 2) runs a SMILES similarity search, 3) samples random compounds, and 4) parses tranche properties.
python
#!/usr/bin/env python3
import subprocess
from io import StringIO
import re
import pandas as pd

BASE = "https://cartblanche22.docking.org"

def curl_get(url: str) -> str:
    r = subprocess.run(["curl", "-sS", url], capture_output=True, text=True)
    r.check_returncode()
    return r.stdout

def query_by_zinc_id(zinc_id: str, output_fields="zinc_id,smiles,catalogs,tranche") -> pd.DataFrame:
    # Common pattern used by CartBlanche22: <endpoint>.txt:<field>=<value>&output_fields=...
    url = f"{BASE}/substances.txt:zinc_id={zinc_id}&output_fields={output_fields}"
    txt = curl_get(url)
    return pd.read_csv(StringIO(txt), sep="\t")

def search_by_smiles(smiles: str, dist: int = 0, adist: int = 0,
                     output_fields="zinc_id,smiles,tranche") -> pd.DataFrame:
    url = (
        f"{BASE}/smiles.txt:smiles={smiles}"
        f"&dist={dist}&adist={adist}&output_fields={output_fields}"
    )
    txt = curl_get(url)
    return pd.read_csv(StringIO(txt), sep="\t")

def random_compounds(count: int = 100, subset: str | None = None,
                     output_fields="zinc_id,smiles,tranche") -> pd.DataFrame:
    url = f"{BASE}/substance/random.txt:count={count}&output_fields={output_fields}"
    if subset:
        url += f"&subset={subset}"
    txt = curl_get(url)
    return pd.read_csv(StringIO(txt), sep="\t")

def parse_tranche(tranche: str):
    """
    Tranche format: H##P###M###-phase
      H##   = H-bond donors
      P###  = LogP * 10
      M###  = molecular weight (Da)
      phase = reactivity classification
    Example: H05P035M400-0
    """
    m = re.match(r"H(\d+)P(\d+)M(\d+)-(\d+)", str(tranche))
    if not m:
        return None
    return {
        "h_donors": int(m.group(1)),
        "logP": int(m.group(2)) / 10.0,
        "mw": int(m.group(3)),
        "phase": int(m.group(4)),
    }

def main():
    # 1) Lookup by ZINC ID
    df_id = query_by_zinc_id("ZINC000000000001")
    print("By ZINC ID:")
    print(df_id.head(), "\n")

    # 2) SMILES exact / similarity search (example: benzene)
    df_smiles = search_by_smiles("c1ccccc1", dist=3, output_fields="zinc_id,smiles,tranche")
    print("SMILES similarity search (dist=3):")
    print(df_smiles.head(), "\n")

    # 3) Random sampling (lead-like)
    df_rand = random_compounds(count=50, subset="lead-like", output_fields="zinc_id,smiles,tranche")
    df_rand["tranche_props"] = df_rand["tranche"].apply(parse_tranche)
    print("Random lead-like sample with parsed tranche:")
    print(df_rand.head(), "\n")

    # 4) Simple tranche-based filtering example
    # Keep compounds with MW <= 350 and logP <= 3.5 when tranche parsing is available
    props = df_rand["tranche_props"].dropna().apply(pd.Series)
    filtered = df_rand.loc[props.index].copy()
    filtered = filtered.join(props)
    filtered = filtered[(filtered["mw"] <= 350) & (filtered["logP"] <= 3.5)]
    print(f"Filtered (mw<=350, logP<=3.5): {len(filtered)} rows")
    print(filtered[["zinc_id", "smiles", "tranche", "mw", "logP"]].head())

if __name__ == "__main__":
    main()

Implementation Details

Data Sources and Access Points
Core Query Patterns

CartBlanche22 commonly exposes endpoints in the form:

  • .../substances.txt:zinc_id=<ID1,ID2,...>&output_fields=...
  • .../smiles.txt:smiles=<SMILES>&dist=<n>&adist=<n>&output_fields=...
  • .../catitems.txt:catitem_id=<SUPPLIER_CODE>
  • .../substance/random.txt:count=<N>&subset=<subset>&output_fields=...

Returned data is typically tab-separated text; request only needed columns via output_fields to reduce payload.

Show full SKILL.md (171 more words)Show less
Similarity Parameters (dist, adist)
  • dist: similarity/analog expansion control (often used as a threshold-like knob; smaller values yield closer analogs).
  • adist: alternative distance parameter for broader expansion.
  • Practical guidance:
    • Start with exact match (dist=0, adist=0).
    • Expand gradually (e.g., dist=1..3 for close analogs; higher values for broader exploration).
Output Fields

Commonly useful fields (availability depends on endpoint/data):

  • zinc_id: ZINC identifier
  • smiles: SMILES representation
  • sub_id: internal substance identifier
  • supplier_code: vendor catalog number
  • catalogs: supplier/catalog list
  • tranche: encoded property bin (H donors, LogP, MW, phase)

Example:

bash
curl "https://cartblanche22.docking.org/substances.txt:zinc_id=ZINC000000000001&output_fields=zinc_id,smiles,catalogs,tranche"
Tranche Encoding (Property Binning)

ZINC tranches encode coarse physicochemical properties:

  • Format: H##P###M###-phase
    • H##: H-bond donors
    • P###: LogP × 10
    • M###: molecular weight (Da)
    • phase: reactivity classification

Use tranche parsing to implement fast, server-side-friendly filtering workflows (e.g., lead-like/drug-like constraints) before downloading 3D structures.

3D Structure Downloads (Docking-Ready)

For docking workflows, use the ZINC22 files library:

Files are organized by tranche and provided in formats such as MOL2, SDF, and DB2.GZ (for DOCK). For large batch downloads, prefer tranche-based retrieval and parallel download tools (e.g., wget, aria2c) while respecting server load.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in scientific-skills/Evidence Insight/zinc-database of aipoch/medical-research-skills.

  • SKILL.md
  • references/api_reference.md
  • zinc-database_audit_result_v1.json

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Zinc Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Zinc Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Zinc Database this skillaipoch/medical-research-skills2k—~1.9kAutomated safety check: PassMIT
MolecodeAtomFlow-AI/MoleCode305—~1.9kAutomated safety check: PassMIT
Drug DiscoveryTommy-yw/RunbookHermes5461 repos~2.3kAutomated safety check: PassMIT
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone
Edu Chem Reactionwy51ai/edulab1.4k—~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • Molecode

    AtomFlow-AI/MoleCode

    A skill your agent uses for deterministic molecule understanding, graph-level editing, generation, and validation with MoleCode — an explicit Mermaid graph in which every atom and bond is a typed…

    305 GitHub stars~1.9k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Drug Discovery

    Tommy-yw/RunbookHermes

    Pharmaceutical research assistant for drug discovery workflows.

    546 GitHub starsUsed in 1 repo~2.3k tokens
    Research & ScienceAuto-check passed
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Edu Chem Reaction

    wy51ai/edulab

    把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。

    1.4k GitHub stars~1.2k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 20 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 20 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 20 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 20 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 20 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 20 days ago
    Auto-check passed

Questions about Zinc Database

What does Zinc Database do?

Access the ZINC (230M+ purchasable compounds) database when you need to look up compounds by ZINC ID/SMILES, run similarity/analog searches, or download 3D ready-to-dock structures for virtual…. Zinc Database is an agent skill from aipoch/medical-research-skills. Access the ZINC (230M+ purchasable compounds) database when you need to look up compounds by ZINC ID/SMILES, run similarity/analog searches, or download 3D ready-to-dock structures for virtual screening and drug discovery.

When should I use Zinc Database?

Zinc Database fits situations like: tasks that involve Drug discovery and cheminformatics.

How do I install Zinc Database in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill zinc-database -a claude-code`. Or copy the skill folder (scientific-skills/Evidence Insight/zinc-database in aipoch/medical-research-skills) into .claude/skills/zinc-database in your project. Claude Code loads it when a task matches its description.

How do I install Zinc Database in Codex?

Run `npx skills add aipoch/medical-research-skills --skill zinc-database -a codex`. Or copy the skill folder (scientific-skills/Evidence Insight/zinc-database in aipoch/medical-research-skills) into .agents/skills/zinc-database in your project. Codex loads it when a task matches its description.

Can I use Zinc Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill zinc-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/zinc-database, .gemini/skills/zinc-database, .github/skills/zinc-database and .opencode/skills/zinc-database in your project.

What does Zinc Database need to run?

Going by SKILL.md and its folder, Zinc Database needs the command-line tools its instructions call (curl). Our summary lists: Python 3.

Does Zinc Database access the network?

SKILL.md names 4 domains. In commands or code: cartblanche22.docking.org; the agent is likely to contact it when it follows the instructions. As links in the text: files.docking.org, zinc.docking.org and wiki.docking.org. This is read from the text; nothing was executed.

Is Zinc Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Zinc Database use?

Zinc Database is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Zinc Database use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.1k tokens, read only when the agent opens those files.

What are the alternatives to Zinc Database?

Skills that share tags, products or a category with Zinc Database: Molecode (AtomFlow-AI/MoleCode, 305 stars), Drug Discovery (Tommy-yw/RunbookHermes, 546 stars), DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars) and Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Zinc Database?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,973 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.