Agent skill

Drug Docking Analysis

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Post-docking analysis of virtual screening results including score distributions, enrichment metrics (ROC AUC, enrichment factors), and ligand efficiency calculations.

MITAuto-check passedResearch & Science

Install Drug Docking Analysis

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill drug-docking-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills drug-docking-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/drug-docking-analysis .claude/skills/drug-docking-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
drug-docking-analysis
GitHub stars
175
Token cost
~2.2k tokens
SKILL.md length
935 words
Files
54 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Post-docking analysis of virtual screening results including score distributions, enrichment metrics (ROC AUC, enrichment factors), and ligand efficiency calculations.

  • Works in 5 steps: Prerequisites: produce a… → Basic analysis (no labels) → With enrichment analysis (labeled library) → …
  • Tasks that involve Drug discovery and cheminformatics
  • SKILL.md covers Goal, Instructions, Outputs and Constraints

What it does

Drug Docking Analysis is an agent skill from learningmatter-mit/AtomisticSkills. Post-docking analysis of virtual screening results including score distributions, enrichment metrics (ROC AUC, enrichment factors), and ligand efficiency calculations.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 58 other files, including scripts (for example `examples/README.md` and `examples/cdk2-htvs/output/analysis_summary.json`).

It sits in Research & Science, covering Drug discovery and cheminformatics. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

When your agent uses it

  • Tasks that involve Drug discovery and cheminformatics

Example prompts

  • “/drug-docking-analysis”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Prerequisites: produce a docking_ranked.csv
  2. Basic analysis (no labels)
  3. With enrichment analysis (labeled library)
  4. Aggregate microstates before ranking
  5. Interpret results

What it can do on your machine

Read from SKILL.md and the folder at commit 7f2d86d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Drug Docking Analysis loads about 2.2k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 935 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 7f2d86d, republished under its MIT licence (© learningmatter-mit). 935 words, ~2,160 tokens.

Download SKILL.mdSave it as .claude/skills/drug-docking-analysis/SKILL.md (or your agent's skills folder). This skill also uses 53 other files; get the full folder from GitHub.
name
drug-docking-analysis
description
Post-docking analysis of virtual screening results including score distributions, enrichment metrics (ROC AUC, enrichment factors), and ligand efficiency calculations.
metadata.category
drug-discovery
metadata.venv
cpu

drug-docking-analysis

Goal

To analyze virtual screening docking results by computing score distributions, ligand efficiency metrics, and (when labeled actives/inactives are available) enrichment statistics. This skill sits between drug-docking-vina and downstream refinement stages, providing quantitative assessment of docking campaign quality.

Important: this skill does not assess pose quality. A compound can receive an excellent Vina score with a physically implausible pose (internal clashes, strained torsions, mis-assigned bond orders). Always pair this analysis with drug-pose-validation (PoseBusters) before acting on the top-ranked compounds.

Outputs:

  • Score distribution KDE plot
  • Score vs. molecular weight scatter (visualizes Vina's size/lipophilicity bias)
  • Ligand efficiency (LE, BEI, SEI) distributions
  • ROC curve with AUC (when labels available)
  • Enrichment factor bar chart at 1%, 2%, 5%, 10%, 20% (when labels available)
  • Enriched results CSV with per-compound efficiency metrics

Instructions

0. Prerequisites: produce a docking_ranked.csv

This skill expects a ranked CSV produced by drug-docking-vina's collect_results.py. If you are starting from raw drug-docking-vina JSON output, run the collect step first:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/../drug-docking-vina/scripts/collect_results.py \
  --results docking/results/docking_results.json \
  --library_csv library/library_master.csv \
  --output_dir docking/analysis/

The collect step joins the docking scores with the library CSV to pull SMILES, labels, and (when present) parent_compound_id / microstate_id columns. See drug-docking-vina SKILL.md step 5 for details.

1. Basic analysis (no labels)

When you have docking results but no active/inactive labels:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/analyze_docking.py \
  --docking_csv docking/docking_ranked.csv \
  --output_dir docking/analysis/

This produces score KDE, score vs. MW, and ligand efficiency plots.

2. With enrichment analysis (labeled library)

When your library has known actives and inactives:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/analyze_docking.py \
  --docking_csv docking/docking_ranked.csv \
  --library_csv library/library_master.csv \
  --active_label active \
  --inactive_label inactive \
  --output_dir docking/analysis/

The --library_csv must have compound_id and label columns. If labels are already in the docking CSV, the library CSV is not needed.

3. Aggregate microstates before ranking

If the library contained enumerated protomers/tautomers (each parent compound appearing multiple times under different compound_id values), you must collapse microstates back to best-score-per-parent before computing any ranking metric. Without aggregation, compounds with more enumerated forms get extra chances to rank high and inflate the apparent library size, biasing enrichment.

Pass --parent_id_col parent_compound_id to opt in:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/analyze_docking.py \
  --docking_csv docking/docking_ranked.csv \
  --parent_id_col parent_compound_id \
  --output_dir docking/analysis/

collect_results.py propagates parent_compound_id and microstate_id from the library CSV automatically when those columns exist, so if you prepared ligands with drug-ligand-prep in its enumeration mode you should use this flag.

See examples/README.md for a side-by-side comparison of aggregated vs. unaggregated analysis on a small synthetic library.

4. Interpret results

Score distribution: A healthy screen typically shows a unimodal distribution with the bulk of scores in the -6 to -8 kcal/mol range and a tail extending toward -9 or beyond; the tail is where you look for hits. A bimodal distribution often means the library contains structurally distinct subsets (e.g. actives + fillers with very different MW profiles) and should be investigated before ranking.

Score vs. MW: Vina scores are biased toward larger, more lipophilic molecules. This plot reveals the bias. If actives cluster in a different MW range than inactives, the enrichment may be driven by size rather than binding complementarity, and you should re-rank by a size-normalized metric (LE or BEI) or apply an MW cutoff.

Ligand efficiency (LE): LE = -score / heavy_atom_count. Normalizes for molecular size. LE > 0.3 kcal/mol/HA is a common rule-of-thumb threshold for drug-like efficiency, but note that this threshold was derived from calibrated experimental binding free energies (Hopkins et al., 2004), not from Vina scores. Vina scores correlate with binding affinity but are not on the same kcal/mol scale as true free energies, so applying the 0.3 threshold directly to Vina-derived LE is common practice but not rigorously justified. Treat the dashed line on the plot as a visual reference, not a hard cutoff. BEI (-score*1000/MW) and SEI (-score*1000/TPSA) provide alternative normalizations; TPSA here is RDKit's Ertl 2D method via Descriptors.TPSA.

Enrichment: ROC AUC > 0.7 indicates decent discrimination. EF1% > 5x indicates meaningful early enrichment. Interpret in context: AUC and EF depend heavily on the decoy set composition, so report the decoy source alongside the metrics and prefer relative comparisons between protocol variants over absolute enrichment claims.

Pose quality: This skill reports score-based metrics only. None of these plots tell you whether the docked poses are physically reasonable. Always run drug-pose-validation (PoseBusters) on the top-ranked compounds before committing them to MD or experimental validation; a great score with a clashing or strained pose is a false positive waiting to happen.

Show full SKILL.md (257 more words)Show less

Outputs

  • docking_analysis.csv: per-compound enriched CSV with columns compound_id, label, docking_score, mw, clogp, tpsa, heavy_atoms, le, bei, sei, smiles. Rows come from the input docking CSV (after microstate aggregation if --parent_id_col is set) and are ordered in the same order as the input.
  • analysis_summary.json: summary metrics (n_compounds, microstate_aggregation, score_*, le_*, and, when labels are available, an enrichment sub-dict with AUC and EF at 1/2/5/10/20%).
  • plots/score_kde.*, plots/score_hist_kde.*: score distribution KDEs (PNG and PDF).
  • plots/score_vs_mw.*: score vs. molecular weight scatter with active/inactive coloring when labels are available.
  • plots/le_distribution.*: ligand efficiency distribution with the 0.3 kcal/mol/HA reference line.
  • plots/roc_curve.* and plots/enrichment_factors.*: only written when labels are available.

Constraints

  • Environment: Requires cpu.
  • Dependencies: rdkit, numpy, scipy, matplotlib.
  • Input format: CSV with at minimum compound_id, best_affinity, and smiles columns. Optional passthrough columns: label, parent_compound_id, microstate_id. Use collect_results.py in drug-docking-vina to produce a CSV in the expected shape.
  • Join key: join between the docking CSV and the optional --library_csv is on compound_id, not on SMILES. The SMILES string in a docked row may reflect a specific protonation or tautomer microstate and will not necessarily match the canonical SMILES in the parent library. Always use compound_id as the key.
  • Score convention: Assumes Vina-style scores where more negative = better binding.
  • Enrichment caveats: ROC AUC and EF depend on the decoy set. ChEMBL-confirmed inactives (structurally similar to actives) give lower AUC than property-matched decoys (DUD-E). Always report the decoy source.
  • No pose-quality signal: score-based metrics only. Run drug-pose-validation alongside this skill to catch physically implausible poses that happen to score well.

Author: Matthew Cox Contact: GitHub @mcox3406

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 53 other files (scripts) in skills/drug-docking-analysis of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/README.md
  • examples/cdk2-htvs/inputs/docking_ranked.csv
  • examples/cdk2-htvs/output/analysis_summary.json
  • examples/cdk2-htvs/output/docking_analysis.csv
  • examples/cdk2-htvs/output/plots/enrichment_factors.pdf
  • examples/cdk2-htvs/output/plots/enrichment_factors.png
  • examples/cdk2-htvs/output/plots/le_distribution.pdf
  • examples/cdk2-htvs/output/plots/le_distribution.png
  • examples/cdk2-htvs/output/plots/roc_curve.pdf
  • examples/cdk2-htvs/output/plots/roc_curve.png
  • examples/cdk2-htvs/output/plots/score_hist_kde.pdf
  • examples/cdk2-htvs/output/plots/score_hist_kde.png
  • examples/cdk2-htvs/output/plots/score_kde.pdf
  • examples/cdk2-htvs/output/plots/score_kde.png
  • examples/cdk2-htvs/output/plots/score_vs_mw.pdf
  • … and 38 more

Open the folder on GitHubat commit 7f2d86d

Compare with similar skills

Drug Docking Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Drug Docking Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Drug Docking Analysis this skilllearningmatter-mit/AtomisticSkills175—~2.2kAutomated safety check: PassMIT
MolecodeAtomFlow-AI/MoleCode305—~1.9kAutomated safety check: PassMIT
Drug DiscoveryTommy-yw/RunbookHermes5461 repos~2.3kAutomated safety check: PassMIT
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone
Edu Chem Reactionwy51ai/edulab1.4k—~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • Molecode

    AtomFlow-AI/MoleCode

    A skill your agent uses for deterministic molecule understanding, graph-level editing, generation, and validation with MoleCode — an explicit Mermaid graph in which every atom and bond is a typed…

    305 GitHub stars~1.9k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Drug Discovery

    Tommy-yw/RunbookHermes

    Pharmaceutical research assistant for drug discovery workflows.

    546 GitHub starsUsed in 1 repo~2.3k tokens
    Research & ScienceAuto-check passed
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Edu Chem Reaction

    wy51ai/edulab

    把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。

    1.4k GitHub stars~1.2k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    175 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    175 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    175 GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    175 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    175 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    175 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Questions about Drug Docking Analysis

What does Drug Docking Analysis do?

Post-docking analysis of virtual screening results including score distributions, enrichment metrics (ROC AUC, enrichment factors), and ligand efficiency calculations. Drug Docking Analysis is an agent skill from learningmatter-mit/AtomisticSkills. Post-docking analysis of virtual screening results including score distributions, enrichment metrics (ROC AUC, enrichment factors), and ligand efficiency calculations.

When should I use Drug Docking Analysis?

Drug Docking Analysis fits situations like: tasks that involve Drug discovery and cheminformatics.

How do I install Drug Docking Analysis in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill drug-docking-analysis -a claude-code`. Or copy the skill folder (skills/drug-docking-analysis in learningmatter-mit/AtomisticSkills) into .claude/skills/drug-docking-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Drug Docking Analysis in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill drug-docking-analysis -a codex`. Or copy the skill folder (skills/drug-docking-analysis in learningmatter-mit/AtomisticSkills) into .agents/skills/drug-docking-analysis in your project. Codex loads it when a task matches its description.

Can I use Drug Docking Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill drug-docking-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/drug-docking-analysis, .gemini/skills/drug-docking-analysis, .github/skills/drug-docking-analysis and .opencode/skills/drug-docking-analysis in your project.

What does Drug Docking Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Drug Docking Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Drug Docking Analysis access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Drug Docking Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Drug Docking Analysis use?

Drug Docking Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Drug Docking Analysis use?

About 2.2k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Drug Docking Analysis?

Skills that share tags, products or a category with Drug Docking Analysis: Molecode (AtomFlow-AI/MoleCode, 305 stars), Drug Discovery (Tommy-yw/RunbookHermes, 546 stars), DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars) and Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Drug Docking Analysis?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 175 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 6, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.