Build molecules with RDKit — from SMILES or a scaffold, analogues and series (each with its parent), properties (MW, cLogP, TPSA, Lipinski, Veber, QED, alerts), similarity and substructure search…

MITAuto-check passedResearch & Science

Install Rdkit

skills CLI
$ npx skills add autonomous-ai/openharness --skill rdkit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/openharness rdkit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/rdkit/skills/rdkit .claude/skills/rdkit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rdkit
GitHub stars
1.1k
Token cost
~3k tokens
SKILL.md length
1,240 words
Files
1
Skills in repo
99
Repo updated
First seen
Licence
MIT

At a glance

Build molecules with RDKit — from SMILES or a scaffold, analogues and series (each with its parent), properties (MW, cLogP, TPSA, Lipinski, Veber, QED, alerts), similarity and substructure search…

  • Any request that ends in a molecule
  • SKILL.md covers Build, write, verdict, A shape the user kept, SMILES, the parts that matter and SMARTS, for finding things, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • A chemical series

What it does

Rdkit is an agent skill from autonomous-ai/openharness. Build molecules with RDKit — from SMILES or a scaffold, analogues and series (each with its parent), properties (MW, cLogP, TPSA, Lipinski, Veber, QED, alerts), similarity and substructure search, conformer ensembles written as SDF for the pane. Use for any request that ends in a molecule, a property or a chemical series.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Drug discovery and cheminformatics. It works with RDKit. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.

When your agent uses it

  • Any request that ends in a molecule
  • A chemical series

Example prompts

  • “/rdkit”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 54a1f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rdkit loads about 3k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 1,240 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/openharness at commit 54a1f1b, republished under its MIT licence (© autonomous-ai). 1,240 words, ~3,009 tokens.

Download SKILL.mdSave it as .claude/skills/rdkit/SKILL.md (or your agent's skills folder).
name
rdkit
description
Build molecules with RDKit — from SMILES or a scaffold, analogues and series (each with its parent), properties (MW, cLogP, TPSA, Lipinski, Veber, QED, alerts), similarity and substructure search, conformer ensembles written as SDF for the pane. Use for any request that ends in a molecule, a property or a chemical series.

rdkit

RDKit is the cheminformatics toolkit: a molecule is a graph (Chem.Mol), written as SMILES, queried with SMARTS, given coordinates by distance geometry. Tools: $RDKIT_PYTHON (the pinned venv, with numpy and pandas), harness_rdkit on PYTHONPATH (build, embed, write, compare), $RDKIT_TOOLCHAIN/verdict.py (the pane header). Never install another RDKit.

Build, write, verdict

bash
"$RDKIT_PYTHON" molecules/hello.py                        # runs the script → out/<name>.sdf and the rest
"$RDKIT_PYTHON" "$RDKIT_TOOLCHAIN/verdict.py"             # judges the newest molecule → pane header
"$RDKIT_PYTHON" "$RDKIT_TOOLCHAIN/harness_rdkit.py" design "CCO" ethanol   # the same, without a script
python
from harness_rdkit import design, mol_from_smiles, embed_conformers, embed_3d, write_outputs, properties, similarity, substructure

design("CC(C)Cc1ccc(cc1)C(C)C(=O)O", "ibuprofen")         # parse → 10 conformers → minimise → describe → write
design("CC(C)(O)Cc1ccc(C(C)C(=O)O)cc1", "ibuprofen_oh", parent="ibuprofen")   # an analogue: name its parent

mol = mol_from_smiles("CN1C=NC2=C1C(=O)N(C)C(=O)N2C", "caffeine")   # the steps, when work happens between
confs = embed_conformers(mol, n=10, seed=7)               # ETKDGv3 + MMFF94, deduplicated, lowest first, aligned
write_outputs(confs, "caffeine", parent=False)            # False: not an analogue of anything here

What design/write_outputs write to out/, and what each is for:

filewhat it iswho reads it
<name>.sdfthe lowest-energy conformerthe pane (the artifact), the verdict, PyMOL/ChimeraX
<name>.conformers.sdfevery kept conformer, lowest first, with energy/delta_energythe pane's conformer player
<name>.svg, <name>.pngthe 2D depiction (the SVG drawn on the parent's core)the pane, reports
<name>.molecule.jsonidentifiers, properties, Lipinski/Veber, plain-language flags, per-atom Gasteiger charge, Crippen logP and hybridisation, functional groups, PAINS/Brenk alerts, conformer energies, the parent and the atoms that changedthe pane
series.jsonevery molecule designed here, in order, with its parent and key propertiesthe pane's series strip and table
properties.json, report.jsonthe newest molecule's properties; the verdict's inputthe verdict

The pane shows the molecule while it is being made (out/.progress.json: embedding, minimising, describing); never write these files by hand. properties(mol) returns the properties without writing; depict(mol, path) the PNG alone; describe(mol, name) the record and SVG without writing. read_smi("molecules/series.smi") reads a SMILES name list, table(mols) turns molecules into rows for pandas.DataFrame.

Parents. An analogue is compared with its parent: the maximum common substructure, the atoms that changed (+O, −C +O), the property deltas, a 2D depiction drawn on the parent's core and a 3D pose superposed on it. Pass parent="<name already designed>" (or a SMILES) whenever you make an analogue; left out, the parent is inferred as the earlier molecule this one is the smallest edit of, and parent=False says there is none. Design the lead first, then its analogues, so the series reads in order.

A shape the user kept

The pane can scan a non-ring single bond with native MMFF94, with the other internal coordinates fixed. The user chooses four connected atoms, moves along the energy curve, and keeps a pose and note in out/torsions/<id>/. This is a rigid, vacuum, single-bond experiment, not a geometry optimization or a free-energy calculation. Do not infer populations, kinetics, activity or binding. The input must have explicit hydrogen atoms, 3D coordinates, 4–200 atoms, one connected component and MMFF94 parameters. It uses 10°, 15° or 30° spacing and never silently falls back to UFF.

When continuing the user's experiment, read study.json for the method, selected angle, atom indices, full-precision coordinates and note. Load selected.sdf with Chem.SDMolSupplier(path, removeHs=False)[0], then author the requested follow-up under a new name. Keep the original study unchanged. source.mol is the exact input conformer; scan.sdf contains every sampled angle; energies.csv gives absolute and within-scan relative energies in kcal/mol. study.zip includes these files and standalone reproduce.py/harness_torsion.py; the recorded RDKit version is required to verify the calculation. JSON coordinates preserve full precision; MOL/SDF exchange files round them to four decimals. Unsaved pane scans are not files: ask the user to Keep study only when their chosen pose or note is needed for the next step.

SMILES, the parts that matter

  • Atoms are bare symbols; lowercase is aromatic (c1ccccc1 benzene). Bonds: - single (implicit), = double, # triple, / \ around a double bond for E/Z.
  • Branches in parentheses: CC(C)C isobutane. Rings close on matching digits: C1CCCCC1 cyclohexane, c1ccc2ccccc2c1 naphthalene. Reuse a digit once it is closed.
  • Brackets for anything not a default: charge [NH4+], [O-]; isotope [13C]; explicit H [nH] — pyrrole is c1cc[nH]c1, and forgetting the H is the commonest SMILES error there is.
  • Stereo: [C@H] / [C@@H] at a centre, F/C=C/F trans. Write it when it matters; RDKit will not guess, and an unspecified centre silently becomes a racemate.
  • Dot separates components: a salt is CC(=O)[O-].[Na+], and most calculations want the parent only.

Useful anchors: water O, ethanol CCO, benzene c1ccccc1, phenol Oc1ccccc1, aspirin CC(=O)Oc1ccccc1C(=O)O, paracetamol CC(=O)Nc1ccc(O)cc1, caffeine CN1C=NC2=C1C(=O)N(C)C(=O)N2C, ibuprofen CC(C)Cc1ccc(cc1)C(C)C(=O)O, naproxen COc1ccc2cc(ccc2c1)C(C)C(=O)O, glucose OC[C@H]1OC(O)[C@H](O)[C@@H](O)[C@@H]1O, penicillin G core CC1(C)S[C@@H]2[C@H](NC(=O)Cc3ccccc3)C(=O)N2[C@H]1C(=O)O.

SMARTS, for finding things

SMARTS is SMILES plus queries: [#6] any carbon, [C,N] either, [!c] not aromatic carbon, [R2] in two rings, [X3] three connections, [OX2H] a hydroxyl oxygen, * anything, ~ any bond.

python
substructure(mol, "[OX2H]")                  # hydroxyls → ((3,), (7,))
substructure(mol, "c1ccccc1")                # benzene rings
substructure(mol, "[CX3](=O)[OX2H1]")        # carboxylic acid

Groups worth keeping: carboxylic acid [CX3](=O)[OX2H1], amide [NX3][CX3](=[OX1]), primary amine [NX3;H2;!$(NC=O)], sulfonamide [SX4](=[OX1])(=[OX1])([NX3]), nitro [N+](=O)[O-], halogen [F,Cl,Br,I].

Show full SKILL.md (545 more words)Show less

Common tasks

Modify a scaffold — edit the SMILES where the substituent goes, or replace a group in place:

python
from rdkit import Chem
core = Chem.MolFromSmiles("CC(=O)Oc1ccccc1C(=O)O")
out  = Chem.ReplaceSubstructs(core, Chem.MolFromSmarts("[CX3](=O)[OX2H1]"),
                              Chem.MolFromSmiles("C(=O)NC"), replaceAll=True)[0]
Chem.SanitizeMol(out); print(Chem.MolToSmiles(out))

Enumerate analogues — one substituent list, one loop, a table, and the ones worth looking at written out, each with its parent:

python
import pandas as pd
from rdkit import Chem
from harness_rdkit import mol_from_smiles, properties, design
subs = {"H": "", "F": "F", "Cl": "Cl", "OMe": "OC", "CF3": "C(F)(F)F"}
mols = [mol_from_smiles(f"CC(C)Cc1ccc(C(C)C(=O)O)c({r})c1" if r else "CC(C)Cc1ccc(cc1)C(C)C(=O)O", name)
        for name, r in subs.items()]
print(pd.DataFrame([{"R": m.GetProp("_Name"), **properties(m)} for m in mols])[["R", "mw", "logp", "tpsa", "qed"]])
design("CC(C)Cc1ccc(cc1)C(C)C(=O)O", "ibuprofen")         # the lead first
design(Chem.MolToSmiles(mols[-1]), "analogue_cf3", parent="ibuprofen")   # then the analogue, into the series

Similarity search across a list — Morgan (ECFP4) Tanimoto; > 0.7 is a close analogue, < 0.3 unrelated:

python
from harness_rdkit import read_smi, similarity
query = "CC(=O)Oc1ccccc1C(=O)O"
hits = sorted(((similarity(query, m), m.GetProp("_Name")) for m in read_smi("molecules/library.smi")), reverse=True)
for score, name in hits[:10]: print(f"{score:.2f}  {name}")

3D and conformers — design already runs a small conformer search (conformers=10): ETKDGv3 embeddings, MMFF94 minimisation, duplicates dropped, the lowest within 10 kcal/mol kept, aligned, and written as the ensemble the pane plays. Raise it for a flexible molecule (conformers=30), lower it to 1 for a single seed-7 pose. embed_conformers(mol, n, seed) is the same search by hand; an unspecified stereocentre is pinned to one configuration across the ensemble and still reported as unspecified. Energies are vacuum force-field numbers — a guide to shape, not to populations.

Save for other tools: write_outputs writes the SDFs. Chem.MolToPDBFile(mol, "out/x.pdb"), Chem.MolToXYZFile(mol, "out/x.xyz"), Chem.MolToSmiles(mol) for the canonical string.

Pitfalls

  • Sanitization is not optional. Chem.MolFromSmiles returns None for anything that will not sanitize — a five-valent carbon, an unclosed ring, c1ccccc1 written with the wrong aromaticity. A None is a typo in the SMILES; fix the string, never sanitize=False your way past it.
  • Hydrogens. Properties are computed on the graph without them; 3D needs them. embed_3d adds them and gives you a new molecule — the original stays flat, which is what the depiction wants.
  • Stereo. An unspecified centre embeds as one arbitrary enantiomer and the SDF will look decided. Write the stereo into the SMILES, or say in your answer that the centre is unspecified.
  • Protonation. SMILES are written neutral by convention; a carboxylic acid is C(=O)O even though it is an anion at pH 7.4. cLogP and TPSA assume the neutral form — say so rather than "correcting" it.
  • Salts and mixtures. Strip them before computing: Chem.MolStandardize.rdMolStandardize.LargestFragmentChooser().
  • Macrocycles and cages sometimes fail to embed; embed_3d retries with random coordinates, and after that it is a different seed or useMacrocycleTorsions.
  • cLogP is Crippen's estimate, QED a 2012 desirability score, Lipinski a rule of thumb for oral absorption. They are guides, not measurements, and nothing here predicts activity, binding or safety.

Rules

  • molecules/ holds scripts and .smi lists, out/ holds everything produced. One design call per molecule, name matching the file you want in the pane.
  • Run the verdict after every script: "$RDKIT_PYTHON" "$RDKIT_TOOLCHAIN/verdict.py". It is ready when a script exists, the newest out/*.sdf parses with a real 3D conformer, and no error is open; Lipinski and Veber violations, PAINS/Brenk alerts and a strained energy are warnings, an unspecified stereocentre a note — none of them failures.
  • The pane opens the newest molecule and follows each new one: 3D (wire, stick, ball-and-stick, spacefill; hydrogens; colour by element, Gasteiger charge, Crippen lipophilicity or change from the parent; VDW/SAS/SES surfaces coloured by electrostatic or lipophilicity potential; hover for an atom's charge and hybridisation; click to measure distances, angles and dihedrals), the 2D depiction linked atom-for-atom with its functional groups, the properties with Lipinski/Veber badges and plain-language flags, the conformer player, the parent overlay, and the series as a strip and a comparison table. Writing the files through design/write_outputs is how you put something in it; point the user at what to look at ("colour by change", "the conformer tab") rather than describing the picture.

© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in store/agents/rdkit/skills/rdkit of autonomous-ai/openharness.

Open the folder on GitHubat commit 54a1f1b

Compare with similar skills

Rdkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rdkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rdkit this skillautonomous-ai/openharness1.1k—~3kAutomated safety check: PassMIT
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Edu Chem Reactionwy51ai/edulab1.4k—~1.2kAutomated safety check: PassApache-2.0
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT
RDKit Cheminformatics Practicesaiming-lab/AutoResearchClaw15k—~708Automated safety check: PassMIT
Rowanlamm-mit/scienceclaw2444 repos~3.1kAutomated safety check: WarnProprietary

Similar skills

  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Edu Chem Reaction

    wy51ai/edulab

    把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。

    1.4k GitHub stars~1.2k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed
  • RDKit Cheminformatics Practices

    aiming-lab/AutoResearchClaw

    Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.

    15k GitHub stars~708 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Rowan

    lamm-mit/scienceclaw

    Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.

    244 GitHub starsUsed in 4 repos~3.1k tokens
    Research & ScienceAuto-check: warnings
  • RDKit Cheminformatics

    davila7/claude-code-templates

    Guides molecular work with RDKit in Python: reading SMILES and SDF, sanitization, descriptors, fingerprints, substructure and similarity search, reactions and coordinates.

    32k GitHub starsUsed in 15 repos~5k tokens
    Research & ScienceAuto-check passed

More from autonomous-ai/openharness

All 99 skills in this repo
  • G-code Slicer Tool

    autonomous-ai/openharness

    Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Home Assistant Automation Builder

    autonomous-ai/openharness

    Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.

    1.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Score Music Composer

    autonomous-ai/openharness

    Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.

    1.1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • OrcaSlicer 3MF and G-code Workflow

    autonomous-ai/openharness

    Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sheets and Docs Report Builder

    autonomous-ai/openharness

    Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.

    1.1k GitHub stars~708 tokensUpdated today
    Auto-check passed
  • Bambu Labs

    autonomous-ai/openharness

    Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.

    1.1k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check: warnings

Works with

Questions about Rdkit

What does Rdkit do?

Build molecules with RDKit — from SMILES or a scaffold, analogues and series (each with its parent), properties (MW, cLogP, TPSA, Lipinski, Veber, QED, alerts), similarity and substructure search…. Rdkit is an agent skill from autonomous-ai/openharness. Build molecules with RDKit — from SMILES or a scaffold, analogues and series (each with its parent), properties (MW, cLogP, TPSA, Lipinski, Veber, QED, alerts), similarity and substructure search, conformer ensembles written as SDF for the pane.

When should I use Rdkit?

Rdkit fits situations like: any request that ends in a molecule; A chemical series.

How do I install Rdkit in Claude Code?

Run `npx skills add autonomous-ai/openharness --skill rdkit -a claude-code`. Or copy the skill folder (store/agents/rdkit/skills/rdkit in autonomous-ai/openharness) into .claude/skills/rdkit in your project. Claude Code loads it when a task matches its description.

How do I install Rdkit in Codex?

Run `npx skills add autonomous-ai/openharness --skill rdkit -a codex`. Or copy the skill folder (store/agents/rdkit/skills/rdkit in autonomous-ai/openharness) into .agents/skills/rdkit in your project. Codex loads it when a task matches its description.

Can I use Rdkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill rdkit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rdkit, .gemini/skills/rdkit, .github/skills/rdkit and .opencode/skills/rdkit in your project.

What does Rdkit need to run?

SKILL.md names no scripts, command-line tools or credentials: Rdkit is instructions for the agent only. Our summary lists: Python 3.

Does Rdkit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rdkit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Rdkit use?

Rdkit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rdkit use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Rdkit?

Skills that share tags, products or a category with Rdkit: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Edu Chem Reaction (wy51ai/edulab, 1.4k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars) and RDKit Cheminformatics Practices (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rdkit?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,137 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 7, 2026.

Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.