Official agent skill

Drug Discovery Pipeline

by NVIDIA in NVIDIA/skills

NOTE: molecule and target inputs and your NGCAPIKEY are transmitted to external NVIDIA-hosted API endpoints on every call.

OfficialApache-2.0Auto-check: notesResearch & Science

Install Drug Discovery Pipeline

skills CLI
$ npx skills add NVIDIA/skills --skill drug-discovery-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills drug-discovery-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bionemo-drug-discovery-pipeline .claude/skills/drug-discovery-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
drug-discovery-pipeline
GitHub stars
3.6k
Used in
1 other repo
Token cost
~1.8k tokens
SKILL.md length
246 words
Files
7
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

NOTE: molecule and target inputs and your NGCAPIKEY are transmitted to external NVIDIA-hosted API endpoints on every call.

  • Works in 3 steps: Generate molecules with GenMol → Dock molecules with DiffDock → Predict binding affinity with Boltz2
  • The user wants to generate and screen small molecule drug candidates
  • SKILL.md covers Overview, Before you start, Step 1: Generate molecules… and Step 2: Dock molecules with…, plus 3 more sections
  • Reaches health.api.nvidia.com; needs NGC_API_KEY

What it does

Drug Discovery Pipeline is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. NOTE: molecule and target inputs and your NGCAPIKEY are transmitted to external NVIDIA-hosted API endpoints on every call. Use local NIM containers for confidential or proprietary data. Run a complete computational drug discovery pipeline using NVIDIA BioNeMo NIMs: generate drug-like molecules with GenMol, dock them to a protein target with DiffDock, then predict binding affinity with Boltz2. Use this skill whenever the user wants to generate and screen small molecule drug candidates, perform hit discovery…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `BENCHMARK.md`, `evals/config.yml` and `evals/evals.json`).

It sits in Research & Science, covering Drug discovery and cheminformatics. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user wants to generate and screen small molecule drug candidates
  • Perform hit discovery
  • Optimize leads against a protein target
  • Do virtual screening combining molecule generation

Example prompts

  • “/drug-discovery-pipeline”

Requirements

  • Python 3
  • Docker
  • A credential in NGC_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, Write, AskUserQuestion

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Generate molecules with GenMol
  2. Dock molecules with DiffDock
  3. Predict binding affinity with Boltz2

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • health.api.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NGC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Drug Discovery Pipeline loads about 1.8k tokens when it runs. Until then it costs about 236 tokens; SKILL.md has 246 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~236
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 246 words, ~1,840 tokens.

Download SKILL.mdSave it as .claude/skills/drug-discovery-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
drug-discovery-pipeline
description
NOTE: molecule and target inputs and your NGC_API_KEY are transmitted to external NVIDIA-hosted API endpoints on every call. Use local NIM containers for confidential or proprietary data. Run a complete computational drug discovery pipeline using NVIDIA BioNeMo NIMs: generate drug-like molecules with GenMol, dock them to a protein target with DiffDock, then predict binding affinity with Boltz2. Use this skill whenever the user wants to generate and screen small molecule drug candidates, perform hit discovery, optimize leads against a protein target, or do virtual screening combining molecule generation, docking, and affinity prediction. Triggers on: drug discovery pipeline, hit discovery, lead optimization, virtual screening, molecule generation, molecular docking, binding affinity, GenMol, DiffDock, Boltz2, SMILES, SAFE notation, NIM microservice. This is a multi-step pipeline composing three BioNeMo NIMs.
allowed-tools
Bash, Read, Write, AskUserQuestion
license
Apache-2.0 AND CC-BY-4.0

Drug Discovery Pipeline

Screen drug candidates end-to-end using three BioNeMo NIMs in sequence:

Step 1: GenMol    →  Step 2: DiffDock  →  Step 3: Boltz2
(Generate mols)      (Dock to target)      (Predict affinity)

Overview

This pipeline is used for:

  • De novo hit discovery: generate drug-like molecules and screen them against a target
  • Lead optimization: start from a known scaffold and generate improved analogs, then dock and score
  • Virtual screening: dock a library of candidates and filter by docking confidence + affinity

Before you start

Confirm with the user:

  1. Target protein: PDB file or sequence of the binding target
  2. Starting point: de novo (no scaffold) or scaffold decoration (known core)?
  3. Scoring: drug-likeness (QED) or lipophilicity (LogP)?
  4. API mode: hosted or local Docker?

For local Docker, do not assume all NIMs are running on localhost:8000 at the same time. Either run one container at a time and hand files/results between steps, or start each NIM on a distinct host port and set the per-step URLs.


Step 1: Generate molecules with GenMol

GenMol requires SAFE notation input (not raw SMILES). Use the safe-mol package.

python
import requests, json, os
import safe as sf                          # pip install safe-mol
from pathlib import Path

NGC_API_KEY = os.getenv("NGC_API_KEY")
HOSTED = True

if HOSTED:
    genmol_url = "https://health.api.nvidia.com/v1/biology/nvidia/genmol/generate"
    headers = {"Content-Type": "application/json",
               "Authorization": f"Bearer {NGC_API_KEY}"}
else:
    genmol_url = "http://localhost:8000/generate"
    headers = {"Content-Type": "application/json"}

# De novo generation (no scaffold):
safe_input = "[*{20-30}]"

# Scaffold decoration (known core):
# scaffold_smiles = "c1ccccc1"
# safe_input = sf.encode(scaffold_smiles) + ".[*{5-10}]"

payload = {
    "smiles": safe_input,              # field is named 'smiles' but takes SAFE notation
    "num_molecules": 30,               # request more to compensate for post-generation filtering
    "scoring": "QED",                  # QED or LogP
    "unique": True,
    "temperature": "1.0",             # NOTE: must be string, not float
    "noise": "1.0",                   # NOTE: must be string, not float
}

r = requests.post(genmol_url, headers=headers, json=payload)
r.raise_for_status()
molecules = r.json()["molecules"]
molecules_sorted = sorted(molecules, key=lambda x: x["score"], reverse=True)
top_20 = molecules_sorted[:20]

print(f"Generated {len(molecules)} valid molecules (requested 30)")
print("Top 5 by QED score:")
for m in top_20[:5]:
    print(f"  {m['smiles'][:50]}  score={m['score']:.4f}")

Step 2: Dock molecules with DiffDock

Prepare the protein and dock each candidate:

python
# Load protein (ATOM records only)
receptor_pdb_raw = Path("target.pdb").read_text()
receptor_pdb = "\n".join(line for line in receptor_pdb_raw.splitlines()
                          if line.startswith("ATOM"))

if HOSTED:
    diffdock_url = "https://health.api.nvidia.com/v1/biology/mit/diffdock"
else:
    diffdock_url = "http://localhost:8000/molecular-docking/diffdock/generate"

docking_results = []

for i, mol in enumerate(top_20):
    payload = {
        "protein": receptor_pdb,
        "ligand": mol["smiles"],
        "ligand_file_type": "txt",     # "txt" for SMILES input
        "num_poses": 5,
        "time_divisions": 20,
        "steps": 18,
        "save_trajectory": False,
    }

    r = requests.post(diffdock_url, headers=headers, json=payload)
    r.raise_for_status()
    result = r.json()

    best_conf = result["position_confidence"][0]  # rank 1 pose
    best_pose = result["ligand_positions"][0]

    docking_results.append({
        "smiles": mol["smiles"],
        "qed_score": mol["score"],
        "docking_confidence": best_conf,
        "best_pose_sdf": best_pose,
    })
    print(f"  Mol {i+1:2d}: QED={mol['score']:.3f}  docking_conf={best_conf:.4f}")

# Rank by docking confidence
docking_results.sort(key=lambda x: x["docking_confidence"], reverse=True)
print(f"\nTop 3 by docking confidence:")
for d in docking_results[:3]:
    print(f"  {d['smiles'][:50]}  conf={d['docking_confidence']:.4f}")

Step 3: Predict binding affinity with Boltz2

For the top docking candidates, predict structure-based binding affinity:

python
if HOSTED:
    boltz_url = "https://health.api.nvidia.com/v1/biology/mit/boltz2/predict"
else:
    boltz_url = "http://localhost:8000/biology/mit/boltz2/predict"

# Use the target protein sequence (not PDB)
target_sequence = "<YOUR_TARGET_PROTEIN_SEQUENCE>"

affinity_results = []
for d in docking_results[:5]:  # score top 5 docking hits
    payload = {
        "polymers": [
            {"id": "A", "molecule_type": "protein", "sequence": target_sequence}
        ],
        "ligands": [
            {"id": "L1", "smiles": d["smiles"], "predict_affinity": True}
        ],
        "recycling_steps": 3,
        "sampling_steps": 50,
        "diffusion_samples": 1,
        "output_format": "mmcif",
    }

    r = requests.post(boltz_url, headers=headers, json=payload)
    r.raise_for_status()
    result = r.json()

    aff = result["affinities"]["L1"]
    pic50 = aff["affinity_pic50"][0]
    prob_binding = aff["affinity_probability_binary"][0]

    affinity_results.append({
        **d,
        "pic50": pic50,
        "probability_binding": prob_binding,
    })
    print(f"  {d['smiles'][:40]}  pIC50={pic50:.2f}  P(bind)={prob_binding:.3f}")

# Final ranking by pIC50
affinity_results.sort(key=lambda x: x["pic50"], reverse=True)

Interpreting results

  • GenMol QED score: 0–1; >0.5 is drug-like
  • DiffDock confidence: higher = more reliable binding pose prediction
  • Boltz2 pIC50: predicted -log10(IC50); >6 = sub-micromolar, >8 = very potent
  • P(bind): probability of binary binding; >0.7 = likely binder

Quick reference — skill dependencies

StepSkillKey endpoint
Molecule generationgenmol-nim/biology/nvidia/genmol/generate
Dockingdiffdock-nim/molecular-docking/diffdock/generate
Affinity predictionboltz2-nim/biology/mit/boltz2/predict

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files in skills/bionemo-drug-discovery-pipeline of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/config.yml
  • evals/evals.json
  • evals/files/egfr.pdb
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Drug Discovery Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Drug Discovery Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Drug Discovery Pipeline this skillNVIDIA/skills3.6k1 repos~1.8kAutomated safety check: NotesApache-2.0
MolecodeAtomFlow-AI/MoleCode306—~1.9kAutomated safety check: PassMIT
Drug DiscoveryTommy-yw/RunbookHermes5461 repos~2.3kAutomated safety check: PassMIT
Complexa Binder DesignNVIDIA-BioNeMo/bionemo-agent-toolkit479—~3.1kAutomated safety check: NotesApache-2.0
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone

Similar skills

  • Molecode

    AtomFlow-AI/MoleCode

    A skill your agent uses for deterministic molecule understanding, graph-level editing, generation, and validation with MoleCode — an explicit Mermaid graph in which every atom and bond is a typed…

    306 GitHub stars~1.9k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Drug Discovery

    Tommy-yw/RunbookHermes

    Pharmaceutical research assistant for drug discovery workflows.

    546 GitHub starsUsed in 1 repo~2.3k tokens
    Research & ScienceAuto-check passed
  • Complexa Binder Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided…

    479 GitHub stars~3.1k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Deep Researcher Research

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses when asked to run deep research or Deep Researcher Agent research through a reachable NVIDIA Deep Researcher Agent Blueprint backend.

    886 GitHub stars~4.4k tokensUpdated yesterday
    Research & ScienceAuto-check: notes

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Drug Discovery Pipeline

What does Drug Discovery Pipeline do?

NOTE: molecule and target inputs and your NGCAPIKEY are transmitted to external NVIDIA-hosted API endpoints on every call. Drug Discovery Pipeline is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. NOTE: molecule and target inputs and your NGCAPIKEY are transmitted to external NVIDIA-hosted API endpoints on every call.

When should I use Drug Discovery Pipeline?

Drug Discovery Pipeline fits situations like: the user wants to generate and screen small molecule drug candidates; perform hit discovery; optimize leads against a protein target; do virtual screening combining molecule generation.

How do I install Drug Discovery Pipeline in Claude Code?

Run `npx skills add NVIDIA/skills --skill drug-discovery-pipeline -a claude-code`. Or copy the skill folder (skills/bionemo-drug-discovery-pipeline in NVIDIA/skills) into .claude/skills/drug-discovery-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Drug Discovery Pipeline in Codex?

Run `npx skills add NVIDIA/skills --skill drug-discovery-pipeline -a codex`. Or copy the skill folder (skills/bionemo-drug-discovery-pipeline in NVIDIA/skills) into .agents/skills/drug-discovery-pipeline in your project. Codex loads it when a task matches its description.

Can I use Drug Discovery Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill drug-discovery-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/drug-discovery-pipeline, .gemini/skills/drug-discovery-pipeline, .github/skills/drug-discovery-pipeline and .opencode/skills/drug-discovery-pipeline in your project.

What does Drug Discovery Pipeline need to run?

Going by SKILL.md and its folder, Drug Discovery Pipeline needs credentials named NGC_API_KEY. Our summary lists: Python 3; Docker; A credential in NGC_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, AskUserQuestion.

Does Drug Discovery Pipeline access the network?

SKILL.md names 1 domain. In commands or code: health.api.nvidia.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Drug Discovery Pipeline safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Drug Discovery Pipeline use?

Drug Discovery Pipeline is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Drug Discovery Pipeline use?

About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Drug Discovery Pipeline?

Skills that share tags, products or a category with Drug Discovery Pipeline: Molecode (AtomFlow-AI/MoleCode, 306 stars), Drug Discovery (Tommy-yw/RunbookHermes, 546 stars), Complexa Binder Design (NVIDIA-BioNeMo/bionemo-agent-toolkit, 479 stars) and DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Drug Discovery Pipeline?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.