Agent skill

Pharma ML Tools

by DrugClaw in DrugClaw/DrugClaw

Pharmaceutical machine-learning workflow guide for library profiling, molecular featurization, benchmark dataset fetch, medicinal-chemistry filtering, and optional pose-generation handoff.

Apache-2.0Auto-check passedResearch & Science

Install Pharma ML Tools

skills CLI
$ npx skills add DrugClaw/DrugClaw --skill pharma-ml-tools -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DrugClaw/DrugClaw pharma-ml-tools --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DrugClaw/DrugClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pharma/pharma-ml-tools .claude/skills/pharma-ml-tools && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pharma-ml-tools
GitHub stars
125
Token cost
~1.3k tokens
SKILL.md length
406 words
Files
5
Skills in repo
25
Repo updated
First seen
Licence
Apache-2.0

At a glance

Pharmaceutical machine-learning workflow guide for library profiling, molecular featurization, benchmark dataset fetch, medicinal-chemistry filtering, and optional pose-generation handoff.

  • Works in 6 steps: Normalize the input table first and… → Run datamol_library_profile.py before… → Use medchem_screen.py before large… → …
  • The user asks for datamol
  • SKILL.md covers Environment Check, Bundled Assets, Preferred Workflow and Library Profiling With Datamol, plus 6 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Pharma ML Tools is an agent skill from DrugClaw/DrugClaw. Pharmaceutical machine-learning workflow guide for library profiling, molecular featurization, benchmark dataset fetch, medicinal-chemistry filtering, and optional pose-generation handoff. Use when the user asks for datamol, molfeat, PyTDC, medchem, compound-library triage, dataset preparation, or chemistry-ML baselines beyond simple descriptor calculation.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `templates/datamol_library_profile.py`, `templates/medchem_screen.py` and `templates/molfeat_featurize.py`).

It sits in Research & Science, covering Drug discovery and cheminformatics. The repository describes itself as: 💊 AI Research Assistant for Accelerated Drug Discovery. 🦞. The licence is Apache-2.0.

When your agent uses it

  • The user asks for datamol
  • Compound-library triage
  • Dataset preparation
  • Chemistry-ML baselines beyond simple descriptor calculation

Example prompts

  • “/pharma-ml-tools”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Normalize the input table first and identify the exact SMILES column.
  2. Run datamol_library_profile.py before building models so duplicates, invalid structures, and scaffold concentration are visible.
  3. Use medchem_screen.py before large docking or QSAR jobs to flag problematic chemotypes.
  4. Use molfeat_featurize.py when the user needs model-ready features rather than only descriptor summaries.
  5. Use pytdc_dataset_fetch.py when the user needs reproducible public benchmark datasets rather than ad hoc CSV collection.
  6. Keep outputs under a dedicated directory such as ./pharma_ml/.

What it can do on your machine

Read from SKILL.md and the folder at commit 960a6e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pharma ML Tools loads about 1.3k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 406 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from DrugClaw/DrugClaw at commit 960a6e0, republished under its Apache-2.0 licence (© DrugClaw). 406 words, ~1,260 tokens.

Download SKILL.mdSave it as .claude/skills/pharma-ml-tools/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
pharma-ml-tools
description
Pharmaceutical machine-learning workflow guide for library profiling, molecular featurization, benchmark dataset fetch, medicinal-chemistry filtering, and optional pose-generation handoff. Use when the user asks for datamol, molfeat, PyTDC, medchem, compound-library triage, dataset preparation, or chemistry-ML baselines beyond simple descriptor calculation.
source
drugclaw
updated_at
2026-03-11

Pharma ML Tools

Use this skill when the user asks for compound-library profiling, chemistry ML feature generation, medicinal-chemistry screening, or benchmark dataset preparation.

Typical triggers:

  • standardize and profile a compound library before QSAR or screening
  • featurize molecules with molfeat for downstream ML
  • pull benchmark-ready ADME, toxicity, DTI, or DDI datasets with PyTDC
  • apply medicinal-chemistry rules or alert filters before prioritization
  • compare scaffolds, duplicates, or diversity in a virtual-screening library
  • prepare a docking or QSAR campaign with better compound hygiene

Environment Check

bash
which python3 || true
python3 - <<'PY'
mods = ["pandas", "numpy", "datamol", "molfeat", "medchem"]
for name in mods:
    try:
        __import__(name)
        print(f"{name}: ok")
    except Exception as exc:
        print(f"{name}: missing ({exc})")
try:
    import tdc
    print("PyTDC: ok")
except Exception as exc:
    print(f"PyTDC: missing ({exc})")
PY

If a requested module is missing, say so explicitly. Do not claim the screen, featurization, or dataset pull completed.

Bundled Assets

  • templates/datamol_library_profile.py
  • templates/molfeat_featurize.py
  • templates/pytdc_dataset_fetch.py
  • templates/medchem_screen.py

Preferred Workflow

  1. Normalize the input table first and identify the exact SMILES column.
  2. Run datamol_library_profile.py before building models so duplicates, invalid structures, and scaffold concentration are visible.
  3. Use medchem_screen.py before large docking or QSAR jobs to flag problematic chemotypes.
  4. Use molfeat_featurize.py when the user needs model-ready features rather than only descriptor summaries.
  5. Use pytdc_dataset_fetch.py when the user needs reproducible public benchmark datasets rather than ad hoc CSV collection.
  6. Keep outputs under a dedicated directory such as ./pharma_ml/.

Library Profiling With Datamol

bash
python3 templates/datamol_library_profile.py \
  --input libraries/kinase_hits.csv \
  --smiles-column smiles \
  --id-column compound_id \
  --output pharma_ml/kinase_hits_profile.csv \
  --summary pharma_ml/kinase_hits_profile.json

Use this first for:

  • canonical SMILES and InChIKey generation
  • invalid structure detection
  • scaffold counts
  • molecular-property summaries before modeling

Molfeat Featurization

bash
python3 templates/molfeat_featurize.py \
  --input libraries/kinase_hits.csv \
  --smiles-column smiles \
  --id-column compound_id \
  --featurizer ecfp \
  --output pharma_ml/kinase_hits_ecfp.csv \
  --summary pharma_ml/kinase_hits_ecfp.json

Supported baseline featurizers in the bundled template:

  • ecfp
  • maccs
  • rdkit2d

Use this for local QSAR, ranking, clustering, or embedding handoff.

PyTDC Benchmark Datasets

bash
python3 templates/pytdc_dataset_fetch.py \
  --task adme \
  --dataset Caco2_Wang \
  --split-method scaffold \
  --out-dir pharma_ml/caco2_wang

Good use cases:

  • ADME or toxicity baselines
  • DTI or DDI dataset retrieval
  • reproducible train/valid/test splits for benchmarking
Show full SKILL.md (153 more words)Show less

Medicinal-Chemistry Screening

bash
python3 templates/medchem_screen.py \
  --input libraries/kinase_hits.csv \
  --smiles-column smiles \
  --id-column compound_id \
  --output pharma_ml/kinase_hits_medchem.csv \
  --summary pharma_ml/kinase_hits_medchem.json

Use this for:

  • Rule-of-Five and lead-like checks
  • alert-oriented library triage
  • quick pass/fail summaries before wet-lab nomination

Treat these filters as prioritization heuristics, not hard truth.

DiffDock Boundary

If the user asks for diffusion docking or deep pose generation, acknowledge that this runtime already includes docking-tools for Vina-style workflows, but DiffDock-class workflows require a heavier environment with PyTorch Geometric, model weights, and usually GPU acceleration. Do not pretend that support is bundled unless the environment is confirmed.

Output Expectations

Good answers should mention:

  • the exact input file and SMILES column
  • which template ran
  • valid versus invalid molecule counts
  • whether outputs are profiling, features, dataset splits, or medchem filters
  • what files were written
  • any module, network, or dataset-license caveats

For public APIs such as PubChem, ChEMBL, openFDA, ClinicalTrials.gov, or OpenAlex, activate pharma-db-tools. For RDKit descriptors, ADMET heuristics, DrugBank, QSAR, or structure-aware affinity, activate chem-tools. For docking and pose-level workflows, activate docking-tools.

© DrugClaw, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/pharma/pharma-ml-tools of DrugClaw/DrugClaw.

  • SKILL.md
  • templates/datamol_library_profile.py
  • templates/medchem_screen.py
  • templates/molfeat_featurize.py
  • templates/pytdc_dataset_fetch.py

Open the folder on GitHubat commit 960a6e0

Compare with similar skills

Pharma ML Tools next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pharma ML Tools compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pharma ML Tools this skillDrugClaw/DrugClaw125—~1.3kAutomated safety check: PassApache-2.0
MolecodeAtomFlow-AI/MoleCode306—~1.9kAutomated safety check: PassMIT
Drug DiscoveryTommy-yw/RunbookHermes5461 repos~2.3kAutomated safety check: PassMIT
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone
Edu Chem Reactionwy51ai/edulab1.4k—~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • Molecode

    AtomFlow-AI/MoleCode

    A skill your agent uses for deterministic molecule understanding, graph-level editing, generation, and validation with MoleCode — an explicit Mermaid graph in which every atom and bond is a typed…

    306 GitHub stars~1.9k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Drug Discovery

    Tommy-yw/RunbookHermes

    Pharmaceutical research assistant for drug discovery workflows.

    546 GitHub starsUsed in 1 repo~2.3k tokens
    Research & ScienceAuto-check passed
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Edu Chem Reaction

    wy51ai/edulab

    把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。

    1.4k GitHub stars~1.2k tokensUpdated 11 days ago
    Research & ScienceAuto-check passed
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed

More from DrugClaw/DrugClaw

All 25 skills in this repo
  • Bio DB Tools

    DrugClaw/DrugClaw

    Query public biology databases and APIs including UniProt, RCSB PDB, AlphaFold DB, ClinVar, dbSNP, gnomAD, Ensembl, GEO, InterPro, KEGG, OpenTargets, Reactome, and STRING.

    125 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Grn Tools

    DrugClaw/DrugClaw

    Gene regulatory network workflow guide for transcriptomics and single-cell expression matrices using Arboreto, GRNBoost2, and GENIE3.

    125 GitHub stars~692 tokensUpdated 6 mo ago
    Auto-check passed
  • Knowledge Graph Tools

    DrugClaw/DrugClaw

    Drug-discovery knowledge-graph workflow guide for assembling drug-target-disease-pathway relationship graphs from OpenTargets GraphQL, ChEMBL REST, STRING PPI, and Reactome pathway APIs, then…

    125 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Literature Review Tools

    DrugClaw/DrugClaw

    Research-literature workflow guide for evidence-matrix assembly, citation-table normalization, structured review synthesis, and research-gap mapping.

    125 GitHub stars~923 tokensUpdated 6 mo ago
    Auto-check passed
  • Medical Data Tools

    DrugClaw/DrugClaw

    Medical data workflow guide for DICOM metadata inspection and basic de-identification, physiological signal analysis with NeuroKit2, and cohort-table profiling for clinical research datasets.

    125 GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • Omics Tools

    DrugClaw/DrugClaw

    Omics and single-cell workflow guide for AnnData, Scanpy-style dataset profiling, PyDESeq2-oriented count checks, pysam alignment inspection, and pyOpenMS mass-spectrometry summaries.

    125 GitHub stars~1.1k tokensUpdated 6 mo ago
    Auto-check passed

Questions about Pharma ML Tools

What does Pharma ML Tools do?

Pharmaceutical machine-learning workflow guide for library profiling, molecular featurization, benchmark dataset fetch, medicinal-chemistry filtering, and optional pose-generation handoff. Pharma ML Tools is an agent skill from DrugClaw/DrugClaw. Pharmaceutical machine-learning workflow guide for library profiling, molecular featurization, benchmark dataset fetch, medicinal-chemistry filtering, and optional pose-generation handoff.

When should I use Pharma ML Tools?

Pharma ML Tools fits situations like: the user asks for datamol; compound-library triage; dataset preparation; chemistry-ML baselines beyond simple descriptor calculation.

How do I install Pharma ML Tools in Claude Code?

Run `npx skills add DrugClaw/DrugClaw --skill pharma-ml-tools -a claude-code`. Or copy the skill folder (skills/pharma/pharma-ml-tools in DrugClaw/DrugClaw) into .claude/skills/pharma-ml-tools in your project. Claude Code loads it when a task matches its description.

How do I install Pharma ML Tools in Codex?

Run `npx skills add DrugClaw/DrugClaw --skill pharma-ml-tools -a codex`. Or copy the skill folder (skills/pharma/pharma-ml-tools in DrugClaw/DrugClaw) into .agents/skills/pharma-ml-tools in your project. Codex loads it when a task matches its description.

Can I use Pharma ML Tools in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DrugClaw/DrugClaw --skill pharma-ml-tools -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pharma-ml-tools, .gemini/skills/pharma-ml-tools, .github/skills/pharma-ml-tools and .opencode/skills/pharma-ml-tools in your project.

What does Pharma ML Tools need to run?

Going by SKILL.md and its folder, Pharma ML Tools needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Pharma ML Tools access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Pharma ML Tools safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pharma ML Tools use?

Pharma ML Tools is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pharma ML Tools use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pharma ML Tools?

Skills that share tags, products or a category with Pharma ML Tools: Molecode (AtomFlow-AI/MoleCode, 306 stars), Drug Discovery (Tommy-yw/RunbookHermes, 546 stars), DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars) and Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pharma ML Tools?

DrugClaw (a GitHub organization) maintains it in DrugClaw/DrugClaw, which has 125 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on March 23, 2026.

Source: DrugClaw/DrugClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.