DiffDock Molecular Docking
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
Pythonic RDKit wrapper with sensible defaults for drug discovery.
$ npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills datamol-cheminformatics --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/datamol-cheminformatics .claude/skills/datamol-cheminformatics && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "datamol-cheminformatics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/datamol-cheminformatics into .claude/skills/datamol-cheminformatics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datamol-cheminformatics", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/datamol-cheminformaticsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills datamol-cheminformatics --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/structural-biology-drug-discovery/datamol-cheminformatics .agents/skills/datamol-cheminformatics && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "datamol-cheminformatics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/datamol-cheminformatics into .agents/skills/datamol-cheminformatics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datamol-cheminformatics", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills datamol-cheminformatics --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/structural-biology-drug-discovery/datamol-cheminformatics .cursor/skills/datamol-cheminformatics && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "datamol-cheminformatics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/datamol-cheminformatics into .cursor/skills/datamol-cheminformatics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datamol-cheminformatics", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/structural-biology-drug-discovery/datamol-cheminformatics--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills datamol-cheminformatics --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/structural-biology-drug-discovery/datamol-cheminformatics .gemini/skills/datamol-cheminformatics && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "datamol-cheminformatics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/datamol-cheminformatics into .gemini/skills/datamol-cheminformatics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datamol-cheminformatics", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills datamol-cheminformaticsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/structural-biology-drug-discovery/datamol-cheminformatics .github/skills/datamol-cheminformatics && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "datamol-cheminformatics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/datamol-cheminformatics into .github/skills/datamol-cheminformatics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datamol-cheminformatics", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills datamol-cheminformatics --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/structural-biology-drug-discovery/datamol-cheminformatics .opencode/skills/datamol-cheminformatics && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "datamol-cheminformatics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/datamol-cheminformatics into .opencode/skills/datamol-cheminformatics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datamol-cheminformatics", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
datamol-cheminformaticsPythonic RDKit wrapper with sensible defaults for drug discovery.
Datamol Cheminformatics is an agent skill from jaechang-hits/SciAgent-Skills. Pythonic RDKit wrapper with sensible defaults for drug discovery. SMILES parsing, standardization, descriptors, fingerprints, similarity, clustering, diversity selection, scaffold analysis, BRICS/RECAP fragmentation, 3D conformers, and visualization. Returns native rdkit.Chem.Mol. Prefer datamol for standard workflows; use RDKit directly for advanced control.
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Drug discovery and cheminformatics. It works with RDKit. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.datamol.iordkit.orggithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Datamol Cheminformatics loads about 4.4k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 857 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 857 words, ~4,422 tokens.
.claude/skills/datamol-cheminformatics/SKILL.md (or your agent's skills folder).Datamol provides a lightweight, Pythonic abstraction layer over RDKit for molecular cheminformatics. It simplifies common drug discovery operations — SMILES parsing, standardization, descriptors, fingerprints, clustering, scaffolds, conformers, and visualization — with sensible defaults, built-in parallelization, and cloud storage support via fsspec. All molecular objects are native rdkit.Chem.Mol instances, ensuring full RDKit compatibility.
uv pip install datamolimport datamol as dm
import numpy as np
import pandas as pdimport datamol as dm
# Parse and standardize
mol = dm.to_mol("CC(=O)Oc1ccccc1C(=O)O") # Aspirin
mol = dm.standardize_mol(mol)
print(dm.to_smiles(mol)) # Canonical SMILES
# Compute descriptors
desc = dm.descriptors.compute_many_descriptors(mol)
print(f"MW: {desc['mw']:.1f}, LogP: {desc['logp']:.2f}, TPSA: {desc['tpsa']:.1f}")
# Generate fingerprint
fp = dm.to_fp(mol, fp_type='ecfp', radius=2, n_bits=2048)
print(f"Fingerprint shape: {fp.shape}") # (2048,)Parsing molecules:
import datamol as dm
# From SMILES (returns None on failure)
mol = dm.to_mol("CCO")
if mol is None:
print("Invalid SMILES")
# Format conversions
smiles = dm.to_smiles(mol, isomeric=True) # Canonical SMILES
inchi = dm.to_inchi(mol)
inchikey = dm.to_inchikey(mol)
selfies = dm.to_selfies(mol)Standardization (always recommended for external data):
mol = dm.standardize_mol(
mol,
disconnect_metals=True,
normalize=True,
reionize=True
)
clean_smiles = dm.standardize_smiles("C(C)O") # From SMILES directlyFile I/O:
# Reading (supports local, S3, GCS, HTTP via fsspec)
df = dm.read_sdf("compounds.sdf", mol_column='mol')
df = dm.read_csv("data.csv", smiles_column="SMILES", mol_column="mol")
df = dm.read_excel("compounds.xlsx", sheet_name=0, mol_column="mol")
df = dm.open_df("file.sdf") # Auto-detect format
# Writing
dm.to_sdf(df, "output.sdf", mol_column="mol")
dm.to_smi(mols, "output.smi")
dm.to_xlsx(df, "output.xlsx", mol_columns=["mol"]) # Renders molecule images
# Remote files
df = dm.read_sdf("s3://bucket/compounds.sdf")
dm.to_sdf(mols, "s3://bucket/output.sdf")import datamol as dm
mol = dm.to_mol("c1ccc(cc1)CCN")
# Standard descriptor set (single molecule)
desc = dm.descriptors.compute_many_descriptors(mol)
# Returns dict: {'mw': 121.18, 'logp': 1.41, 'hbd': 1, 'hba': 1,
# 'tpsa': 26.02, 'n_aromatic_atoms': 6, ...}
# Batch computation (parallel)
mols = [dm.to_mol(s) for s in ["CCO", "c1ccccc1", "CC(=O)O"]]
desc_df = dm.descriptors.batch_compute_many_descriptors(
mols, n_jobs=-1, progress=True
)
print(desc_df.head())
# Specific descriptors
n_stereo = dm.descriptors.n_stereo_centers(mol)
n_aromatic = dm.descriptors.n_aromatic_atoms(mol)
aromatic_ratio = dm.descriptors.n_aromatic_atoms_proportion(mol)
n_rigid = dm.descriptors.n_rigid_bonds(mol)Drug-likeness filtering (Lipinski Rule of Five):
def is_druglike(mol):
desc = dm.descriptors.compute_many_descriptors(mol)
return (desc['mw'] <= 500 and desc['logp'] <= 5
and desc['hbd'] <= 5 and desc['hba'] <= 10)
druglike = [m for m in mols if is_druglike(m)]
print(f"Drug-like: {len(druglike)}/{len(mols)}")import datamol as dm
mol = dm.to_mol("c1ccc(cc1)CCN")
# Fingerprint types
fp_ecfp = dm.to_fp(mol, fp_type='ecfp', radius=2, n_bits=2048) # Morgan/ECFP
fp_maccs = dm.to_fp(mol, fp_type='maccs') # MACCS keys (167 bits)
fp_topo = dm.to_fp(mol, fp_type='topological') # Topological
fp_ap = dm.to_fp(mol, fp_type='atompair') # Atom pairs
# Pairwise distances (Tanimoto distance = 1 - similarity)
mols = [dm.to_mol(s) for s in ["CCO", "CCCO", "c1ccccc1"]]
dist_matrix = dm.pdist(mols, n_jobs=-1)
print(f"Distance vector shape: {dist_matrix.shape}")
# Distances between two sets
query = [dm.to_mol("CCO")]
library = [dm.to_mol(s) for s in ["CCCO", "c1ccccc1", "CC(=O)O"]]
distances = dm.cdist(query, library, n_jobs=-1)
print(f"Query-library distances: {distances.shape}")import datamol as dm
mols = [dm.to_mol(s) for s in smiles_list] # Assume smiles_list defined
# Butina clustering (suitable for ~1000 molecules, builds full distance matrix)
clusters = dm.cluster_mols(mols, cutoff=0.2, n_jobs=-1)
for i, cluster in enumerate(clusters[:5]):
print(f"Cluster {i}: {len(cluster)} molecules")
# Diversity selection (works for larger libraries)
diverse_mols = dm.pick_diverse(mols, npick=100)
print(f"Selected {len(diverse_mols)} diverse molecules")
# Cluster centroids
centroids = dm.pick_centroids(mols, npick=50)
print(f"Selected {len(centroids)} centroids")Murcko scaffold extraction:
import datamol as dm
from collections import Counter
mol = dm.to_mol("c1ccc(cc1)CCN")
scaffold = dm.to_scaffold_murcko(mol)
print(f"Scaffold: {dm.to_smiles(scaffold)}")
# Scaffold frequency analysis
scaffolds = [dm.to_scaffold_murcko(m) for m in mols]
scaffold_smiles = [dm.to_smiles(s) for s in scaffolds]
counts = Counter(scaffold_smiles)
print(f"Top scaffolds: {counts.most_common(5)}")
# Scaffold-based train/test split (for ML)
scaffold_to_mols = {}
for mol, scaf in zip(mols, scaffold_smiles):
scaffold_to_mols.setdefault(scaf, []).append(mol)
scaffolds_list = list(scaffold_to_mols.keys())
split_idx = int(0.8 * len(scaffolds_list))
train_mols = [m for s in scaffolds_list[:split_idx] for m in scaffold_to_mols[s]]
test_mols = [m for s in scaffolds_list[split_idx:] for m in scaffold_to_mols[s]]Fragmentation:
mol = dm.to_mol("CC(=O)Oc1ccccc1C(=O)O") # Aspirin
# BRICS (16 bond types, retrosynthetic)
brics_frags = dm.fragment.brics(mol)
print(f"BRICS fragments: {brics_frags}") # Set of fragment SMILES with [1*] attachment points
# RECAP (11 bond types, combinatorial)
recap_frags = dm.fragment.recap(mol)
# MMPA (matched molecular pair analysis)
mmpa_frags = dm.fragment.mmpa_frag(mol)import datamol as dm
mol = dm.to_mol("c1ccc(cc1)CCN")
# Generate conformers
mol_3d = dm.conformers.generate(
mol,
n_confs=50, # Number to generate
rms_cutoff=0.5, # Filter similar (Angstroms)
minimize_energy=True, # UFF minimization
method='ETKDGv3' # Embedding method
)
print(f"Generated {mol_3d.GetNumConformers()} conformers")
# Access coordinates
conf = mol_3d.GetConformer(0)
positions = conf.GetPositions() # Nx3 array
print(f"Atom positions shape: {positions.shape}")
# Cluster conformers by RMSD
clusters = dm.conformers.cluster(mol_3d, rms_cutoff=1.0)
centroids = dm.conformers.return_centroids(mol_3d, clusters)
# Solvent accessible surface area
sasa = dm.conformers.sasa(mol_3d, n_jobs=-1)
print(f"SASA values: {sasa[:3]}")| Use Datamol when... | Use RDKit directly when... |
|---|---|
| Standard SMILES ↔ Mol conversions | Custom fingerprint definitions |
| Batch processing with parallelization | Low-level atom/bond manipulation |
| Quick descriptor computation | Substructure query optimization |
| File I/O (SDF, CSV, Excel, cloud) | Reaction enumeration (large-scale) |
| Clustering & diversity selection | Custom force field parameters |
| Scaffold analysis | Advanced stereochemistry handling |
rdkit.Chem.Mol objects — fully compatible with RDKit functionsmol column containing Mol objectsFunctions supporting n_jobs parameter: dm.read_sdf, dm.descriptors.batch_compute_many_descriptors, dm.cluster_mols, dm.pdist, dm.cdist, dm.conformers.sasa. Use n_jobs=-1 for all cores, progress=True for progress bars.
import datamol as dm
# 1. Load and standardize
df = dm.read_sdf("compounds.sdf")
df['mol'] = df['mol'].apply(lambda m: dm.standardize_mol(m) if m else None)
df = df[df['mol'].notna()]
print(f"Loaded {len(df)} valid molecules")
# 2. Compute descriptors and filter by drug-likeness
desc_df = dm.descriptors.batch_compute_many_descriptors(
df['mol'].tolist(), n_jobs=-1, progress=True
)
druglike = (desc_df['mw'] <= 500) & (desc_df['logp'] <= 5) & (desc_df['hbd'] <= 5) & (desc_df['hba'] <= 10)
filtered_df = df[druglike.values].reset_index(drop=True)
print(f"Drug-like compounds: {len(filtered_df)}")
# 3. Select diverse subset
diverse = dm.pick_diverse(filtered_df['mol'].tolist(), npick=100)
# 4. Visualize
dm.viz.to_image(diverse[:20], legends=[dm.to_smiles(m) for m in diverse[:20]],
n_cols=5, mol_size=(300, 300), outfile="diverse_hits.png")import datamol as dm
import numpy as np
# Query actives and screening library
actives = [dm.to_mol(s) for s in active_smiles] # Known actives
library = [dm.to_mol(s) for s in library_smiles] # Screening library
# Calculate distances (Tanimoto)
distances = dm.cdist(actives, library, n_jobs=-1)
min_distances = distances.min(axis=0) # Best match to any active
similarities = 1 - min_distances
# Rank and select top hits
top_idx = np.argsort(similarities)[::-1][:100]
top_hits = [library[i] for i in top_idx]
top_scores = [similarities[i] for i in top_idx]
print(f"Top hit similarity: {top_scores[0]:.3f}")
# Visualize top hits
dm.viz.to_image(top_hits[:20],
legends=[f"Sim: {s:.3f}" for s in top_scores[:20]],
outfile="screening_hits.png")import datamol as dm
# Group compounds by scaffold
scaffolds = [dm.to_scaffold_murcko(m) for m in mols]
scaffold_smiles = [dm.to_smiles(s) for s in scaffolds]
sar_df = pd.DataFrame({
'mol': mols, 'scaffold': scaffold_smiles, 'activity': activities
})
# Analyze each scaffold series
for scaffold, group in sar_df.groupby('scaffold'):
if len(group) >= 3:
print(f"Scaffold: {scaffold} | N={len(group)} | "
f"Activity: {group['activity'].min():.2f}–{group['activity'].max():.2f}")
dm.viz.to_image(group['mol'].tolist(), align=True,
legends=[f"Act: {a:.2f}" for a in group['activity']])| Function | Parameter | Default | Description |
|---|---|---|---|
dm.to_fp | fp_type | 'ecfp' | Fingerprint type: ecfp, maccs, topological, atompair |
dm.to_fp | radius | 2 | Morgan radius (ecfp only); radius=2 ≈ ECFP4 |
dm.to_fp | n_bits | 2048 | Fingerprint length (ecfp, topological) |
dm.cluster_mols | cutoff | 0.2 | Tanimoto distance threshold (0=identical, 1=different) |
dm.pick_diverse | npick | required | Number of diverse molecules to select |
dm.conformers.generate | n_confs | None | Number of conformers (None = auto) |
dm.conformers.generate | rms_cutoff | None | RMSD filter threshold (Angstroms) |
dm.conformers.generate | method | 'ETKDGv3' | Embedding: ETKDGv3, ETKDGv2, ETKDG |
dm.standardize_mol | disconnect_metals | False | Remove metal-ligand bonds |
dm.read_sdf | sanitize | True | Apply molecule sanitization |
dm.read_sdf | remove_hs | True | Remove explicit hydrogens |
dm.viz.to_image | align | False | Align molecules by MCS |
dm.viz.to_image | use_svg | False | Output SVG (True) or PNG (False) |
dm.standardize_mol() with disconnect_metals=True, normalize=True, reionize=True before any analysis. Different SMILES representations of the same molecule will produce different fingerprintsdm.to_mol() returns None for invalid SMILES. Filter these before batch operations to avoid crashesn_jobs=-1, progress=True to batch operations. Sequential processing of 10,000+ molecules is unnecessarily slowdm.cluster_mols) builds a full distance matrix. Use for ≤~1,000 molecules. For larger sets, use dm.pick_diverse() or hierarchical methodss3fs or gcsfs for cloud supportWhen to use: Clean a list of SMILES strings before any downstream analysis.
import datamol as dm
smiles_list = ["CC(=O)Oc1ccccc1C(=O)O", "c1ccccc1", "invalid_smiles", "CC(N)C(=O)O"]
mols = [dm.to_mol(s) for s in smiles_list]
valid = [(s, m) for s, m in zip(smiles_list, mols) if m is not None]
standardized = [(s, dm.standardize_mol(m)) for s, m in valid]
print(f"Valid: {len(valid)}/{len(smiles_list)}")
for orig, mol in standardized:
print(f" {orig} → {dm.to_smiles(mol)}")When to use: Compare a small compound set against each other or a reference library.
import datamol as dm
import numpy as np
smiles = ["CC(=O)Oc1ccccc1C(=O)O", "c1ccc(cc1)C(=O)O", "CC(N)C(=O)O", "c1ccccc1"]
mols = [dm.to_mol(s) for s in smiles]
fps = [dm.to_fp(m) for m in mols]
# Pairwise Tanimoto similarity
n = len(fps)
sim_matrix = np.zeros((n, n))
for i in range(n):
for j in range(n):
sim_matrix[i, j] = dm.similarity.tanimoto(fps[i], fps[j])
print(f"Similarity matrix shape: {sim_matrix.shape}")
print(f"Most similar pair: {np.unravel_index(np.argsort(sim_matrix.ravel())[-3], (n, n))}")| Problem | Cause | Solution |
|---|---|---|
dm.to_mol() returns None | Invalid or non-canonical SMILES | Try dm.standardize_smiles() first; check for kekulization issues |
| MemoryError during clustering | Full distance matrix for large set | Use dm.pick_diverse() instead of dm.cluster_mols for >1000 molecules |
| Slow conformer generation | Too many conformers or large molecule | Reduce n_confs, increase rms_cutoff, or limit molecule size |
| Remote file access fails | Missing fsspec backend | Install s3fs (AWS), gcsfs (GCP), or adlfs (Azure) |
| Descriptor computation fails | Molecule has no conformer | Standardize first; some 3D descriptors need dm.conformers.generate() |
dm.to_xlsx missing images | openpyxl not installed | uv pip install openpyxl |
| Inconsistent fingerprints | Different SMILES for same molecule | Standardize all molecules before fingerprint computation |
| Scaffold extraction returns full molecule | No ring system in molecule | Murcko scaffolds require at least one ring; acyclic molecules return themselves |
| Reaction product is None | Reactant doesn't match SMARTS pattern | Verify reactant matches reaction template; check atom mapping |
Import error for dm.viz | Missing visualization dependencies | uv pip install Pillow cairosvg |
© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/structural-biology-drug-discovery/datamol-cheminformatics of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Datamol Cheminformatics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Datamol Cheminformatics this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3k | Automated safety check: Notes | MIT | |
| Biopipelineslocbp-uzh/biopipelines | 109 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Edu Chem Reactionwy51ai/edulab | 1.4k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| RDKit Cheminformatics Practicesaiming-lab/AutoResearchClaw | 15k | — | ~708 | Automated safety check: Pass | MIT | |
| Rowanlamm-mit/scienceclaw | 246 | 4 repos | ~3.1k | Automated safety check: Warn | Proprietary |
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
locbp-uzh/biopipelines
Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…
wy51ai/edulab
把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。
aiming-lab/AutoResearchClaw
Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.
lamm-mit/scienceclaw
Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.
pemsley/coot
RDKit molecular manipulation and visualization within Coot's Python environment.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Pythonic RDKit wrapper with sensible defaults for drug discovery. Datamol Cheminformatics is an agent skill from jaechang-hits/SciAgent-Skills. Pythonic RDKit wrapper with sensible defaults for drug discovery.
Datamol Cheminformatics fits situations like: tasks that involve Drug discovery and cheminformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/datamol-cheminformatics in jaechang-hits/SciAgent-Skills) into .claude/skills/datamol-cheminformatics in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/datamol-cheminformatics in jaechang-hits/SciAgent-Skills) into .agents/skills/datamol-cheminformatics in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datamol-cheminformatics, .gemini/skills/datamol-cheminformatics, .github/skills/datamol-cheminformatics and .opencode/skills/datamol-cheminformatics in your project.
Going by SKILL.md and its folder, Datamol Cheminformatics needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md names 3 domains. As links in the text: docs.datamol.io, rdkit.org and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Datamol Cheminformatics is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Datamol Cheminformatics: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars), Edu Chem Reaction (wy51ai/edulab, 1.4k stars) and RDKit Cheminformatics Practices (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.