Agent skill

Datamol Cheminformatics

by jaechang-hits in jaechang-hits/SciAgent-Skills

Pythonic RDKit wrapper with sensible defaults for drug discovery.

Apache-2.0Auto-check passedResearch & Science

Install Datamol Cheminformatics

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills datamol-cheminformatics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/datamol-cheminformatics .claude/skills/datamol-cheminformatics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
datamol-cheminformatics
GitHub stars
374
Used in
1 other repo
Token cost
~4.4k tokens
SKILL.md length
857 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
Apache-2.0

At a glance

Pythonic RDKit wrapper with sensible defaults for drug discovery.

  • Works in 9 steps: Molecular I/O & Standardization → Descriptors & Properties → Fingerprints & Similarity → …
  • Tasks that involve Drug discovery and cheminformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 9 more sections
  • Calls uv

What it does

Datamol Cheminformatics is an agent skill from jaechang-hits/SciAgent-Skills. Pythonic RDKit wrapper with sensible defaults for drug discovery. SMILES parsing, standardization, descriptors, fingerprints, similarity, clustering, diversity selection, scaffold analysis, BRICS/RECAP fragmentation, 3D conformers, and visualization. Returns native rdkit.Chem.Mol. Prefer datamol for standard workflows; use RDKit directly for advanced control.

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Drug discovery and cheminformatics. It works with RDKit. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Drug discovery and cheminformatics

Example prompts

  • “/datamol-cheminformatics”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Molecular I/O & Standardization
  2. Descriptors & Properties
  3. Fingerprints & Similarity
  4. Clustering & Diversity Selection
  5. Scaffolds & Fragments
  6. 3D Conformers
  7. Drug Discovery Pipeline: Load → Filter → Cluster → Visualize
  8. Virtual Screening: Query → Similarity → Rank
  9. SAR Analysis: Group by Scaffold → Compare Activities

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.datamol.io
    • rdkit.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Datamol Cheminformatics loads about 4.4k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 857 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 857 words, ~4,422 tokens.

Download SKILL.mdSave it as .claude/skills/datamol-cheminformatics/SKILL.md (or your agent's skills folder).
name
datamol-cheminformatics
description
Pythonic RDKit wrapper with sensible defaults for drug discovery. SMILES parsing, standardization, descriptors, fingerprints, similarity, clustering, diversity selection, scaffold analysis, BRICS/RECAP fragmentation, 3D conformers, and visualization. Returns native rdkit.Chem.Mol. Prefer datamol for standard workflows; use RDKit directly for advanced control.
license
Apache-2.0

Datamol Cheminformatics Toolkit

Overview

Datamol provides a lightweight, Pythonic abstraction layer over RDKit for molecular cheminformatics. It simplifies common drug discovery operations — SMILES parsing, standardization, descriptors, fingerprints, clustering, scaffolds, conformers, and visualization — with sensible defaults, built-in parallelization, and cloud storage support via fsspec. All molecular objects are native rdkit.Chem.Mol instances, ensuring full RDKit compatibility.

When to Use

  • Parsing, validating, and standardizing molecular structures from SMILES, SDF, or other formats
  • Computing molecular descriptors and fingerprints for ML featurization
  • Similarity searching and diversity selection from compound libraries
  • Clustering compounds by structural similarity (Butina clustering)
  • Scaffold analysis and scaffold-based train/test splitting for ML
  • BRICS/RECAP molecular fragmentation for fragment-based design
  • 3D conformer generation and analysis
  • Visualizing molecules as grids with alignment and highlighting
  • Batch processing molecular datasets with parallelization
  • For quick gene lookups use gget instead; for advanced substructure queries or custom fingerprints, use RDKit directly

Prerequisites

bash
uv pip install datamol
python
import datamol as dm
import numpy as np
import pandas as pd

Quick Start

python
import datamol as dm

# Parse and standardize
mol = dm.to_mol("CC(=O)Oc1ccccc1C(=O)O")  # Aspirin
mol = dm.standardize_mol(mol)
print(dm.to_smiles(mol))  # Canonical SMILES

# Compute descriptors
desc = dm.descriptors.compute_many_descriptors(mol)
print(f"MW: {desc['mw']:.1f}, LogP: {desc['logp']:.2f}, TPSA: {desc['tpsa']:.1f}")

# Generate fingerprint
fp = dm.to_fp(mol, fp_type='ecfp', radius=2, n_bits=2048)
print(f"Fingerprint shape: {fp.shape}")  # (2048,)

Core API

1. Molecular I/O & Standardization

Parsing molecules:

python
import datamol as dm

# From SMILES (returns None on failure)
mol = dm.to_mol("CCO")
if mol is None:
    print("Invalid SMILES")

# Format conversions
smiles = dm.to_smiles(mol, isomeric=True)  # Canonical SMILES
inchi = dm.to_inchi(mol)
inchikey = dm.to_inchikey(mol)
selfies = dm.to_selfies(mol)

Standardization (always recommended for external data):

python
mol = dm.standardize_mol(
    mol,
    disconnect_metals=True,
    normalize=True,
    reionize=True
)
clean_smiles = dm.standardize_smiles("C(C)O")  # From SMILES directly

File I/O:

python
# Reading (supports local, S3, GCS, HTTP via fsspec)
df = dm.read_sdf("compounds.sdf", mol_column='mol')
df = dm.read_csv("data.csv", smiles_column="SMILES", mol_column="mol")
df = dm.read_excel("compounds.xlsx", sheet_name=0, mol_column="mol")
df = dm.open_df("file.sdf")  # Auto-detect format

# Writing
dm.to_sdf(df, "output.sdf", mol_column="mol")
dm.to_smi(mols, "output.smi")
dm.to_xlsx(df, "output.xlsx", mol_columns=["mol"])  # Renders molecule images

# Remote files
df = dm.read_sdf("s3://bucket/compounds.sdf")
dm.to_sdf(mols, "s3://bucket/output.sdf")
2. Descriptors & Properties
python
import datamol as dm

mol = dm.to_mol("c1ccc(cc1)CCN")

# Standard descriptor set (single molecule)
desc = dm.descriptors.compute_many_descriptors(mol)
# Returns dict: {'mw': 121.18, 'logp': 1.41, 'hbd': 1, 'hba': 1,
#                'tpsa': 26.02, 'n_aromatic_atoms': 6, ...}

# Batch computation (parallel)
mols = [dm.to_mol(s) for s in ["CCO", "c1ccccc1", "CC(=O)O"]]
desc_df = dm.descriptors.batch_compute_many_descriptors(
    mols, n_jobs=-1, progress=True
)
print(desc_df.head())

# Specific descriptors
n_stereo = dm.descriptors.n_stereo_centers(mol)
n_aromatic = dm.descriptors.n_aromatic_atoms(mol)
aromatic_ratio = dm.descriptors.n_aromatic_atoms_proportion(mol)
n_rigid = dm.descriptors.n_rigid_bonds(mol)

Drug-likeness filtering (Lipinski Rule of Five):

python
def is_druglike(mol):
    desc = dm.descriptors.compute_many_descriptors(mol)
    return (desc['mw'] <= 500 and desc['logp'] <= 5
            and desc['hbd'] <= 5 and desc['hba'] <= 10)

druglike = [m for m in mols if is_druglike(m)]
print(f"Drug-like: {len(druglike)}/{len(mols)}")
3. Fingerprints & Similarity
python
import datamol as dm

mol = dm.to_mol("c1ccc(cc1)CCN")

# Fingerprint types
fp_ecfp = dm.to_fp(mol, fp_type='ecfp', radius=2, n_bits=2048)  # Morgan/ECFP
fp_maccs = dm.to_fp(mol, fp_type='maccs')    # MACCS keys (167 bits)
fp_topo = dm.to_fp(mol, fp_type='topological')  # Topological
fp_ap = dm.to_fp(mol, fp_type='atompair')    # Atom pairs

# Pairwise distances (Tanimoto distance = 1 - similarity)
mols = [dm.to_mol(s) for s in ["CCO", "CCCO", "c1ccccc1"]]
dist_matrix = dm.pdist(mols, n_jobs=-1)
print(f"Distance vector shape: {dist_matrix.shape}")

# Distances between two sets
query = [dm.to_mol("CCO")]
library = [dm.to_mol(s) for s in ["CCCO", "c1ccccc1", "CC(=O)O"]]
distances = dm.cdist(query, library, n_jobs=-1)
print(f"Query-library distances: {distances.shape}")
4. Clustering & Diversity Selection
python
import datamol as dm

mols = [dm.to_mol(s) for s in smiles_list]  # Assume smiles_list defined

# Butina clustering (suitable for ~1000 molecules, builds full distance matrix)
clusters = dm.cluster_mols(mols, cutoff=0.2, n_jobs=-1)
for i, cluster in enumerate(clusters[:5]):
    print(f"Cluster {i}: {len(cluster)} molecules")

# Diversity selection (works for larger libraries)
diverse_mols = dm.pick_diverse(mols, npick=100)
print(f"Selected {len(diverse_mols)} diverse molecules")

# Cluster centroids
centroids = dm.pick_centroids(mols, npick=50)
print(f"Selected {len(centroids)} centroids")
5. Scaffolds & Fragments

Murcko scaffold extraction:

python
import datamol as dm
from collections import Counter

mol = dm.to_mol("c1ccc(cc1)CCN")
scaffold = dm.to_scaffold_murcko(mol)
print(f"Scaffold: {dm.to_smiles(scaffold)}")

# Scaffold frequency analysis
scaffolds = [dm.to_scaffold_murcko(m) for m in mols]
scaffold_smiles = [dm.to_smiles(s) for s in scaffolds]
counts = Counter(scaffold_smiles)
print(f"Top scaffolds: {counts.most_common(5)}")

# Scaffold-based train/test split (for ML)
scaffold_to_mols = {}
for mol, scaf in zip(mols, scaffold_smiles):
    scaffold_to_mols.setdefault(scaf, []).append(mol)
scaffolds_list = list(scaffold_to_mols.keys())
split_idx = int(0.8 * len(scaffolds_list))
train_mols = [m for s in scaffolds_list[:split_idx] for m in scaffold_to_mols[s]]
test_mols = [m for s in scaffolds_list[split_idx:] for m in scaffold_to_mols[s]]

Fragmentation:

python
mol = dm.to_mol("CC(=O)Oc1ccccc1C(=O)O")  # Aspirin

# BRICS (16 bond types, retrosynthetic)
brics_frags = dm.fragment.brics(mol)
print(f"BRICS fragments: {brics_frags}")  # Set of fragment SMILES with [1*] attachment points

# RECAP (11 bond types, combinatorial)
recap_frags = dm.fragment.recap(mol)

# MMPA (matched molecular pair analysis)
mmpa_frags = dm.fragment.mmpa_frag(mol)
6. 3D Conformers
python
import datamol as dm

mol = dm.to_mol("c1ccc(cc1)CCN")

# Generate conformers
mol_3d = dm.conformers.generate(
    mol,
    n_confs=50,           # Number to generate
    rms_cutoff=0.5,       # Filter similar (Angstroms)
    minimize_energy=True,  # UFF minimization
    method='ETKDGv3'      # Embedding method
)
print(f"Generated {mol_3d.GetNumConformers()} conformers")

# Access coordinates
conf = mol_3d.GetConformer(0)
positions = conf.GetPositions()  # Nx3 array
print(f"Atom positions shape: {positions.shape}")

# Cluster conformers by RMSD
clusters = dm.conformers.cluster(mol_3d, rms_cutoff=1.0)
centroids = dm.conformers.return_centroids(mol_3d, clusters)

# Solvent accessible surface area
sasa = dm.conformers.sasa(mol_3d, n_jobs=-1)
print(f"SASA values: {sasa[:3]}")

Key Concepts

Datamol vs RDKit Decision Guide
Use Datamol when...Use RDKit directly when...
Standard SMILES ↔ Mol conversionsCustom fingerprint definitions
Batch processing with parallelizationLow-level atom/bond manipulation
Quick descriptor computationSubstructure query optimization
File I/O (SDF, CSV, Excel, cloud)Reaction enumeration (large-scale)
Clustering & diversity selectionCustom force field parameters
Scaffold analysisAdvanced stereochemistry handling
Key Data Types
  • All molecules are native rdkit.Chem.Mol objects — fully compatible with RDKit functions
  • Fingerprints are numpy arrays (dense bit vectors)
  • DataFrames use pandas with a mol column containing Mol objects
  • Distance matrices use Tanimoto distance (0 = identical, 1 = completely different)
Parallelization

Functions supporting n_jobs parameter: dm.read_sdf, dm.descriptors.batch_compute_many_descriptors, dm.cluster_mols, dm.pdist, dm.cdist, dm.conformers.sasa. Use n_jobs=-1 for all cores, progress=True for progress bars.

Common Workflows

1. Drug Discovery Pipeline: Load → Filter → Cluster → Visualize
python
import datamol as dm

# 1. Load and standardize
df = dm.read_sdf("compounds.sdf")
df['mol'] = df['mol'].apply(lambda m: dm.standardize_mol(m) if m else None)
df = df[df['mol'].notna()]
print(f"Loaded {len(df)} valid molecules")

# 2. Compute descriptors and filter by drug-likeness
desc_df = dm.descriptors.batch_compute_many_descriptors(
    df['mol'].tolist(), n_jobs=-1, progress=True
)
druglike = (desc_df['mw'] <= 500) & (desc_df['logp'] <= 5) & (desc_df['hbd'] <= 5) & (desc_df['hba'] <= 10)
filtered_df = df[druglike.values].reset_index(drop=True)
print(f"Drug-like compounds: {len(filtered_df)}")

# 3. Select diverse subset
diverse = dm.pick_diverse(filtered_df['mol'].tolist(), npick=100)

# 4. Visualize
dm.viz.to_image(diverse[:20], legends=[dm.to_smiles(m) for m in diverse[:20]],
                n_cols=5, mol_size=(300, 300), outfile="diverse_hits.png")
2. Virtual Screening: Query → Similarity → Rank
python
import datamol as dm
import numpy as np

# Query actives and screening library
actives = [dm.to_mol(s) for s in active_smiles]  # Known actives
library = [dm.to_mol(s) for s in library_smiles]  # Screening library

# Calculate distances (Tanimoto)
distances = dm.cdist(actives, library, n_jobs=-1)
min_distances = distances.min(axis=0)  # Best match to any active
similarities = 1 - min_distances

# Rank and select top hits
top_idx = np.argsort(similarities)[::-1][:100]
top_hits = [library[i] for i in top_idx]
top_scores = [similarities[i] for i in top_idx]
print(f"Top hit similarity: {top_scores[0]:.3f}")

# Visualize top hits
dm.viz.to_image(top_hits[:20],
    legends=[f"Sim: {s:.3f}" for s in top_scores[:20]],
    outfile="screening_hits.png")
3. SAR Analysis: Group by Scaffold → Compare Activities
python
import datamol as dm

# Group compounds by scaffold
scaffolds = [dm.to_scaffold_murcko(m) for m in mols]
scaffold_smiles = [dm.to_smiles(s) for s in scaffolds]

sar_df = pd.DataFrame({
    'mol': mols, 'scaffold': scaffold_smiles, 'activity': activities
})

# Analyze each scaffold series
for scaffold, group in sar_df.groupby('scaffold'):
    if len(group) >= 3:
        print(f"Scaffold: {scaffold} | N={len(group)} | "
              f"Activity: {group['activity'].min():.2f}–{group['activity'].max():.2f}")
        dm.viz.to_image(group['mol'].tolist(), align=True,
            legends=[f"Act: {a:.2f}" for a in group['activity']])

Key Parameters

FunctionParameterDefaultDescription
dm.to_fpfp_type'ecfp'Fingerprint type: ecfp, maccs, topological, atompair
dm.to_fpradius2Morgan radius (ecfp only); radius=2 ≈ ECFP4
dm.to_fpn_bits2048Fingerprint length (ecfp, topological)
dm.cluster_molscutoff0.2Tanimoto distance threshold (0=identical, 1=different)
dm.pick_diversenpickrequiredNumber of diverse molecules to select
dm.conformers.generaten_confsNoneNumber of conformers (None = auto)
dm.conformers.generaterms_cutoffNoneRMSD filter threshold (Angstroms)
dm.conformers.generatemethod'ETKDGv3'Embedding: ETKDGv3, ETKDGv2, ETKDG
dm.standardize_moldisconnect_metalsFalseRemove metal-ligand bonds
dm.read_sdfsanitizeTrueApply molecule sanitization
dm.read_sdfremove_hsTrueRemove explicit hydrogens
dm.viz.to_imagealignFalseAlign molecules by MCS
dm.viz.to_imageuse_svgFalseOutput SVG (True) or PNG (False)

Best Practices

  1. Always standardize molecules from external sources — call dm.standardize_mol() with disconnect_metals=True, normalize=True, reionize=True before any analysis. Different SMILES representations of the same molecule will produce different fingerprints
  2. Check for None after parsing — dm.to_mol() returns None for invalid SMILES. Filter these before batch operations to avoid crashes
  3. Use parallel processing for datasets — pass n_jobs=-1, progress=True to batch operations. Sequential processing of 10,000+ molecules is unnecessarily slow
  4. Choose fingerprints by use case — ECFP (Morgan): general structural similarity; MACCS: fast, smaller space; Atom pairs: distance-sensitive. ECFP with radius=2, n_bits=2048 is the most common default
  5. Mind clustering scale limits — Butina clustering (dm.cluster_mols) builds a full distance matrix. Use for ≤~1,000 molecules. For larger sets, use dm.pick_diverse() or hierarchical methods
  6. Use scaffold splitting for ML — random splits leak similar structures into train/test. Always use scaffold-based splitting for molecular property prediction models
  7. Leverage fsspec for cloud data — all I/O functions accept S3, GCS, and HTTP paths directly. Install s3fs or gcsfs for cloud support
Show full SKILL.md (260 more words)Show less

Common Recipes

Recipe: Batch SMILES Validation and Standardization

When to use: Clean a list of SMILES strings before any downstream analysis.

python
import datamol as dm

smiles_list = ["CC(=O)Oc1ccccc1C(=O)O", "c1ccccc1", "invalid_smiles", "CC(N)C(=O)O"]
mols = [dm.to_mol(s) for s in smiles_list]
valid = [(s, m) for s, m in zip(smiles_list, mols) if m is not None]
standardized = [(s, dm.standardize_mol(m)) for s, m in valid]
print(f"Valid: {len(valid)}/{len(smiles_list)}")
for orig, mol in standardized:
    print(f"  {orig} → {dm.to_smiles(mol)}")
Recipe: Pairwise Similarity Matrix

When to use: Compare a small compound set against each other or a reference library.

python
import datamol as dm
import numpy as np

smiles = ["CC(=O)Oc1ccccc1C(=O)O", "c1ccc(cc1)C(=O)O", "CC(N)C(=O)O", "c1ccccc1"]
mols = [dm.to_mol(s) for s in smiles]
fps = [dm.to_fp(m) for m in mols]

# Pairwise Tanimoto similarity
n = len(fps)
sim_matrix = np.zeros((n, n))
for i in range(n):
    for j in range(n):
        sim_matrix[i, j] = dm.similarity.tanimoto(fps[i], fps[j])
print(f"Similarity matrix shape: {sim_matrix.shape}")
print(f"Most similar pair: {np.unravel_index(np.argsort(sim_matrix.ravel())[-3], (n, n))}")

Troubleshooting

ProblemCauseSolution
dm.to_mol() returns NoneInvalid or non-canonical SMILESTry dm.standardize_smiles() first; check for kekulization issues
MemoryError during clusteringFull distance matrix for large setUse dm.pick_diverse() instead of dm.cluster_mols for >1000 molecules
Slow conformer generationToo many conformers or large moleculeReduce n_confs, increase rms_cutoff, or limit molecule size
Remote file access failsMissing fsspec backendInstall s3fs (AWS), gcsfs (GCP), or adlfs (Azure)
Descriptor computation failsMolecule has no conformerStandardize first; some 3D descriptors need dm.conformers.generate()
dm.to_xlsx missing imagesopenpyxl not installeduv pip install openpyxl
Inconsistent fingerprintsDifferent SMILES for same moleculeStandardize all molecules before fingerprint computation
Scaffold extraction returns full moleculeNo ring system in moleculeMurcko scaffolds require at least one ring; acyclic molecules return themselves
Reaction product is NoneReactant doesn't match SMARTS patternVerify reactant matches reaction template; check atom mapping
Import error for dm.vizMissing visualization dependenciesuv pip install Pillow cairosvg
  • rdkit-cheminformatics — full RDKit API for advanced operations not covered by datamol's simplified interface
  • pubchem-compound-search — retrieve compound data by name, CID, or structure from PubChem
  • scikit-learn-machine-learning — ML model training using datamol-generated features
  • matplotlib-scientific-plotting — custom publication-quality molecular property plots

References

© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/structural-biology-drug-discovery/datamol-cheminformatics of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Datamol Cheminformatics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Datamol Cheminformatics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Datamol Cheminformatics this skilljaechang-hits/SciAgent-Skills3741 repos~4.4kAutomated safety check: PassApache-2.0
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT
Edu Chem Reactionwy51ai/edulab1.4k—~1.2kAutomated safety check: PassApache-2.0
RDKit Cheminformatics Practicesaiming-lab/AutoResearchClaw15k—~708Automated safety check: PassMIT
Rowanlamm-mit/scienceclaw2464 repos~3.1kAutomated safety check: WarnProprietary

Similar skills

  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Edu Chem Reaction

    wy51ai/edulab

    把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。

    1.4k GitHub stars~1.2k tokensUpdated today
    Research & ScienceAuto-check passed
  • RDKit Cheminformatics Practices

    aiming-lab/AutoResearchClaw

    Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.

    15k GitHub stars~708 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Rowan

    lamm-mit/scienceclaw

    Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.

    246 GitHub starsUsed in 4 repos~3.1k tokens
    Research & ScienceAuto-check: warnings
  • Coot Rdkit

    pemsley/coot

    RDKit molecular manipulation and visualization within Coot's Python environment.

    168 GitHub stars~981 tokensUpdated 2 days ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 11 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 11 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 11 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 11 days ago
    Auto-check passed

Works with

Questions about Datamol Cheminformatics

What does Datamol Cheminformatics do?

Pythonic RDKit wrapper with sensible defaults for drug discovery. Datamol Cheminformatics is an agent skill from jaechang-hits/SciAgent-Skills. Pythonic RDKit wrapper with sensible defaults for drug discovery.

When should I use Datamol Cheminformatics?

Datamol Cheminformatics fits situations like: tasks that involve Drug discovery and cheminformatics.

How do I install Datamol Cheminformatics in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/datamol-cheminformatics in jaechang-hits/SciAgent-Skills) into .claude/skills/datamol-cheminformatics in your project. Claude Code loads it when a task matches its description.

How do I install Datamol Cheminformatics in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/datamol-cheminformatics in jaechang-hits/SciAgent-Skills) into .agents/skills/datamol-cheminformatics in your project. Codex loads it when a task matches its description.

Can I use Datamol Cheminformatics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill datamol-cheminformatics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datamol-cheminformatics, .gemini/skills/datamol-cheminformatics, .github/skills/datamol-cheminformatics and .opencode/skills/datamol-cheminformatics in your project.

What does Datamol Cheminformatics need to run?

Going by SKILL.md and its folder, Datamol Cheminformatics needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Datamol Cheminformatics access the network?

SKILL.md names 3 domains. As links in the text: docs.datamol.io, rdkit.org and github.com. This is read from the text; nothing was executed.

Is Datamol Cheminformatics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Datamol Cheminformatics use?

Datamol Cheminformatics is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Datamol Cheminformatics use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Datamol Cheminformatics?

Skills that share tags, products or a category with Datamol Cheminformatics: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars), Edu Chem Reaction (wy51ai/edulab, 1.4k stars) and RDKit Cheminformatics Practices (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Datamol Cheminformatics?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.