Agent skill

RDKit Descriptors and Fingerprints

by jinzhezenggroup in jinzhezenggroup/computational-chemistry-agent-skills

Computes RDKit physicochemical descriptors and molecular fingerprints from SMILES through a uv-run CLI script that skips and logs invalid molecules.

LGPL-3.0Auto-check passedResearch & Science

Install RDKit Descriptors and Fingerprints

skills CLI
$ npx skills add jinzhezenggroup/computational-chemistry-agent-skills --skill rdkit-repr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jinzhezenggroup/computational-chemistry-agent-skills rdkit-repr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jinzhezenggroup/computational-chemistry-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/molecular-representation/rdkit-repr .claude/skills/rdkit-repr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rdkit-repr
GitHub stars
148
Token cost
~2.3k tokens
SKILL.md length
513 words
Files
2 (incl. scripts)
Skills in repo
62
Repo updated
First seen
Licence
LGPL-3.0

At a glance

Computes RDKit physicochemical descriptors and molecular fingerprints from SMILES through a uv-run CLI script that skips and logs invalid molecules.

  • Works in 3 steps: Compute physicochemical descriptors → .csv → Compute molecular fingerprints → .npy or… → List available descriptors
  • Computing Lipinski or physicochemical descriptors for a SMILES dataset
  • SKILL.md covers Quick Start, Core Tasks, Descriptor Presets Reference and Output Format Notes, plus 2 more sections
  • Runs Python scripts from its folder; calls uv

What it does

The skill wraps `scripts/rdkit_helper.py` with three subcommands: `desc` writes descriptors to CSV, `fp` writes fingerprints to .npy or CSV, and `list-desc` shows the available descriptor names and presets. Input can be a single SMILES string, a CSV with a `smiles` column (configurable) or a .smi file. The default `physchem` preset has 25 descriptors, a Lipinski preset has 6, and you can instead name descriptors yourself; `--no-merge` keeps only the SMILES and the descriptors.

Several behaviors matter to the agent. The script prints an environment check for Python, RDKit, NumPy and pandas unless `--no-env` is given, writes invalid SMILES to a `*.skipped.csv` file instead of crashing, and ends each run by printing absolute output paths as RESULT lines. It has to be launched with `uv run` followed by the script path, because dependencies come from inline script metadata and `uv run python` bypasses that.

When your agent uses it

  • Computing Lipinski or physicochemical descriptors for a SMILES dataset
  • Generating fingerprints as a NumPy array for machine learning
  • Listing the RDKit descriptor names and presets that are available
  • Cleaning a molecule list where some SMILES strings are invalid

Example prompts

  • “Compute the physchem descriptors for molecules.csv and save them next to the input.”
  • “Generate fingerprints for data.smi as a .npy file.”
  • “Apply the Lipinski preset to my compounds file and tell me which rows were skipped.”

Requirements

  • `uv` installed
  • RDKit, pandas and NumPy, which `uv run` installs from inline metadata
  • Compatibility (from SKILL.md): Requires uv. Dependencies (rdkit, pandas, numpy) are declared as PEP 723 inline script metadata and are installed automatically when the script is invoked with `uv run <script_path>` (do NOT use `uv run python <script_path>` -- that bypasses the inline metadata and will not install dependencies automatically).

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Compute physicochemical descriptors → .csv
  2. Compute molecular fingerprints → .npy or .csv
  3. List available descriptors

What it can do on your machine

Read from SKILL.md and the folder at commit 5c19e75. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • rdkit.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires uv. Dependencies (rdkit, pandas, numpy) are declared as PEP 723 inline script metadata and are installed automatically when the script is invoked with `uv run <script_path>` (do NOT use `uv run python <script_path>` -- that bypasses the inline metadata and will not install dependencies automatically).

    From compatibility in the SKILL.md frontmatter.

Context cost

RDKit Descriptors and Fingerprints loads about 2.3k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 513 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jinzhezenggroup/computational-chemistry-agent-skills at commit 5c19e75, republished under its LGPL-3.0 licence (© jinzhezenggroup). 513 words, ~2,296 tokens.

Download SKILL.mdSave it as .claude/skills/rdkit-repr/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
rdkit-repr
description
A standardized CLI wrapper for RDKit molecular featurization workflows that handles physicochemical descriptor computation (outputs .csv) and molecular fingerprint extraction (outputs .npy or .csv), with built-in SMILES validation. USE WHEN you need to compute RDKit molecular descriptors or fingerprints from SMILES datasets (.csv/.smi), or when you want to list all available descriptor names and presets.
compatibility
Requires uv. Dependencies (rdkit, pandas, numpy) are declared as PEP 723 inline script metadata and are installed automatically when the script is invoked with `uv run <script_path>` (do NOT use `uv run python <script_path>` -- that bypasses the inline metadata and will not install dependencies automatically).
metadata.author
luzitian
metadata.version
1.0
metadata.repository
https://github.com/rdkit/rdkit

RDKit Molecular Featurization

This skill provides practical command patterns for RDKit descriptor and fingerprint extraction using the standardized CLI wrapper: <skill_path>/scripts/rdkit_helper.py.

Key behaviors (important for Agents):

  • The script prints environment detection (Python/RDKit/NumPy/Pandas) by default.
  • Bad/illegal SMILES are skipped and logged to *.skipped.csv (no crash).
  • Each run ends by printing absolute output paths like:
    • [RESULT] desc_csv=/abs/path.csv
    • [RESULT] fp_npy=/abs/path.npy
    • [RESULT] fp_csv=/abs/path.csv

Quick Start

Check CLI help:

bash
uv run <skill_path>/scripts/rdkit_helper.py --help

Check subcommand help:

bash
uv run <skill_path>/scripts/rdkit_helper.py desc --help
uv run <skill_path>/scripts/rdkit_helper.py fp --help
uv run <skill_path>/scripts/rdkit_helper.py list-desc --help

Disable environment printing (optional):

bash
uv run <skill_path>/scripts/rdkit_helper.py --no-env desc --smiles "CCO" --output out.csv

Core Tasks

1) Compute physicochemical descriptors → .csv

Single SMILES (default preset: physchem, 25 descriptors):

bash
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --smiles "CCO" \
    --output /tmp/CCO.desc.csv

From CSV (default SMILES column is smiles):

bash
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --file data.csv \
    --smiles-col smiles \
    --output data.desc.csv

From SMI:

bash
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --file molecules.smi \
    --output molecules.desc.csv

Choose a descriptor preset:

bash
# Lipinski drug-likeness (6 descriptors: MolWt, MolLogP, NumHDonors, ...)
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --file data.csv --preset lipinski --output data.lipinski.csv

# Extended physicochemical (25 descriptors, default)
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --file data.csv --preset physchem --output data.physchem.csv

# Topological / graph indices (56 descriptors: BalabanJ, BertzCT, Chi*, PEOE_VSA*, ...)
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --file data.csv --preset topological --output data.topo.csv

# All RDKit descriptors (~200 descriptors)
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --file data.csv --preset all --output data.all_desc.csv

Select specific descriptors (overrides --preset):

bash
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --file data.csv \
    --descriptors "MolWt,MolLogP,TPSA,NumHDonors,NumHAcceptors" \
    --output data.custom.csv

Suppress merging back original CSV columns (output only smiles + descriptors):

bash
uv run <skill_path>/scripts/rdkit_helper.py desc \
    --file data.csv --preset physchem --no-merge --output data.desc_only.csv

2) Compute molecular fingerprints → .npy or .csv

Available fingerprint types:

TypeDescriptionDefault bits
morgan2Morgan circular FP radius 2 (ECFP4-like), bit vector2048
morgan3Morgan circular FP radius 3 (ECFP6-like), bit vector2048
morgan2_countMorgan radius-2 count vector2048
rdkitRDKit path-based FP, bit vector2048
maccsMACCS 167 structural keys (bit vector, --nbits ignored)167
topologicalTopological torsion FP (count vector, hashed to --nbits)2048
atompairAtom-pair FP (count vector, hashed to --nbits)2048
layeredLayered substructure FP, bit vector2048
patternSMARTS pattern FP, bit vector2048

Single SMILES, output as NumPy array (.npy):

bash
uv run <skill_path>/scripts/rdkit_helper.py fp \
    --smiles "CCO" \
    --type morgan2 \
    --output /tmp/CCO.morgan2.npy

From CSV, Morgan ECFP4 (2048 bits):

bash
uv run <skill_path>/scripts/rdkit_helper.py fp \
    --file data.csv \
    --smiles-col smiles \
    --type morgan2 \
    --nbits 2048 \
    --output data.morgan2.npy

From SMI, MACCS keys (always 167 bits):

bash
uv run <skill_path>/scripts/rdkit_helper.py fp \
    --file molecules.smi \
    --type maccs \
    --output molecules.maccs.npy

Output as CSV (smiles + bit_0 … bit_N-1 columns):

bash
uv run <skill_path>/scripts/rdkit_helper.py fp \
    --file data.csv \
    --type rdkit \
    --nbits 1024 \
    --format csv \
    --output data.rdkfp.csv

Atom-pair fingerprint, 4096 bits:

bash
uv run <skill_path>/scripts/rdkit_helper.py fp \
    --file data.csv \
    --type atompair \
    --nbits 4096 \
    --output data.atompair.npy

3) List available descriptors

List all descriptors and built-in presets:

bash
uv run <skill_path>/scripts/rdkit_helper.py list-desc

List descriptors in a specific preset group:

bash
uv run <skill_path>/scripts/rdkit_helper.py list-desc --group lipinski
uv run <skill_path>/scripts/rdkit_helper.py list-desc --group physchem
uv run <skill_path>/scripts/rdkit_helper.py list-desc --group topological
uv run <skill_path>/scripts/rdkit_helper.py list-desc --group all

Descriptor Presets Reference

PresetCountTypical Use
lipinski6Quick drug-likeness screening (Ro5 filter)
physchem25General ML features: MW, logP, TPSA, ring counts, charge stats, …
topological56Graph/topology indices: Balaban J, Kappa, Chi, PEOE_VSA, EState_VSA, …
all~200Full RDKit descriptor set (includes fragment counts, MQN, etc.)

Show full SKILL.md (215 more words)Show less

Output Format Notes

desc output (CSV):

  • Columns: smiles, then one column per descriptor.
  • When --file is a .csv and --no-merge is not set, original CSV columns are appended.
  • Rows only contain valid SMILES (invalid ones are logged to *.skipped.csv).

fp output:

  • .npy (default): NumPy array of shape (N_valid, nbits), dtype uint8 (bit) or int32 (count).
  • .csv: smiles column followed by bit_0 … bit_{nbits-1} columns.
  • MACCS keys always produce 167 bits regardless of --nbits.

Agent Checklist

When using this skill for users:

  1. Confirm input format:
    • .csv requires a SMILES column (default smiles)
    • .smi uses the first token of each line as SMILES
  2. Quote SMILES containing special shell characters (brackets/parentheses):
    • Example: --smiles "[C@@H](O)(F)Cl"
  3. For CSV workflows, verify column names:
    • desc: --smiles-col
    • fp: --smiles-col
  4. Choose the right preset or fingerprint type for the downstream task:
    • Drug screening / Ro5: --preset lipinski
    • General ML featurization: --preset physchem or --type morgan2
    • Structural similarity search: --type morgan2 or --type rdkit
    • Substructure matching: --type maccs or --type pattern
  5. Watch for skipped SMILES:
    • Check *.skipped.csv and decide whether to fix or permanently drop them
  6. Always capture absolute output paths:
    • Look for [RESULT] ...=/abs/path in stdout
  7. If debugging is needed, enable full traceback:
    • RDKIT_HELPER_TRACE=1 uv run <skill_path>/scripts/rdkit_helper.py ...

References

© jinzhezenggroup, LGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in molecular-representation/rdkit-repr of jinzhezenggroup/computational-chemistry-agent-skills.

  • SKILL.md
  • scripts/rdkit_helper.py

Open the folder on GitHubat commit 5c19e75

Compare with similar skills

RDKit Descriptors and Fingerprints next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RDKit Descriptors and Fingerprints compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RDKit Descriptors and Fingerprints this skilljinzhezenggroup/computational-chemistry-agent-skills148—~2.3kAutomated safety check: PassLGPL-3.0
Nvmolkit UsageNVIDIA-BioNeMo/bionemo-agent-toolkit479—~4.4kAutomated safety check: PassApache-2.0
ADMET Prediction for Drug CandidatesGPTomics/bioSkills1.2k1 repos~5kAutomated safety check: PassMIT
Vaex Out-of-Core DataFramesdavila7/claude-code-templates33k12 repos~1.6kAutomated safety check: PassMIT
Statistical Data Analysislingzhi227/agent-research-skills390—~886Automated safety check: PassNone
Q-EDA Exploratory AnalysisTyrealQ/q-skills108—~1.1kAutomated safety check: PassMIT

Similar skills

  • Nvmolkit Usage

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF…

    479 GitHub stars~4.4k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Predicts absorption, distribution, metabolism, excretion and toxicity for drug candidates with ADMETlab 3.0, ADMET-AI, DeepChem and chemprop, plus druglikeness filters.

    1.2k GitHub starsUsed in 1 repo~5k tokens
    Research & ScienceAuto-check passed
  • Vaex Out-of-Core DataFrames

    davila7/claude-code-templates

    Processes tabular datasets too large for RAM with Vaex: lazy DataFrames, fast aggregations, big-data plots and ML pipelines over CSV, HDF5, Arrow and Parquet.

    33k GitHub starsUsed in 12 repos~1.6k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    390 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.

    108 GitHub stars~1.1k tokensUpdated 17 days ago
    Data & AnalyticsAuto-check passed
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed

More from jinzhezenggroup/computational-chemistry-agent-skills

All 62 skills in this repo
  • Quantum ESPRESSO DFT Task Builder

    jinzhezenggroup/computational-chemistry-agent-skills

    Turns a user-supplied atomic structure and DFT settings into a runnable Quantum ESPRESSO input file, stopping short of submitting the job.

    148 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • DP-GEN Simplify Workflow

    jinzhezenggroup/computational-chemistry-agent-skills

    Prepares, validates and runs DP-GEN simplify jobs that thin out repeated or redundant DeepMD datasets, generating param.json and machine.json for local or scheduler runs.

    148 GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed
  • LAMMPS with DeePMD-kit

    jinzhezenggroup/computational-chemistry-agent-skills

    Prepares and runs molecular dynamics simulations in LAMMPS with a DeePMD machine-learning potential, writing the input script and choosing NVE, NVT or NPT.

    148 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed
  • LAMMPS ReaxFF Setup

    jinzhezenggroup/computational-chemistry-agent-skills

    Prepares and explains LAMMPS input scripts for reactive molecular dynamics with the ReaxFF potential, including charge equilibration and ensemble choice.

    148 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • RDKit Conformer Generator

    jinzhezenggroup/computational-chemistry-agent-skills

    Generates 3D molecular conformers from SMILES strings or files with RDKit, keeps the lowest-energy one per molecule, and falls back to 2D coordinates when embedding fails.

    148 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Unimol

    jinzhezenggroup/computational-chemistry-agent-skills

    A standardized CLI wrapper for Uni-Mol molecular ML workflows that handles representation extraction (embeddings), model training (regression/classification), and property prediction with built-in…

    148 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed

Questions about RDKit Descriptors and Fingerprints

What does RDKit Descriptors and Fingerprints do?

Computes RDKit physicochemical descriptors and molecular fingerprints from SMILES through a uv-run CLI script that skips and logs invalid molecules. npy or CSV, and `list-desc` shows the available descriptor names and presets.smi file.

When should I use RDKit Descriptors and Fingerprints?

RDKit Descriptors and Fingerprints fits situations like: computing Lipinski or physicochemical descriptors for a SMILES dataset; generating fingerprints as a NumPy array for machine learning; listing the RDKit descriptor names and presets that are available; cleaning a molecule list where some SMILES strings are invalid.

How do I install RDKit Descriptors and Fingerprints in Claude Code?

Run `npx skills add jinzhezenggroup/computational-chemistry-agent-skills --skill rdkit-repr -a claude-code`. Or copy the skill folder (molecular-representation/rdkit-repr in jinzhezenggroup/computational-chemistry-agent-skills) into .claude/skills/rdkit-repr in your project. Claude Code loads it when a task matches its description.

How do I install RDKit Descriptors and Fingerprints in Codex?

Run `npx skills add jinzhezenggroup/computational-chemistry-agent-skills --skill rdkit-repr -a codex`. Or copy the skill folder (molecular-representation/rdkit-repr in jinzhezenggroup/computational-chemistry-agent-skills) into .agents/skills/rdkit-repr in your project. Codex loads it when a task matches its description.

Can I use RDKit Descriptors and Fingerprints in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jinzhezenggroup/computational-chemistry-agent-skills --skill rdkit-repr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rdkit-repr, .gemini/skills/rdkit-repr, .github/skills/rdkit-repr and .opencode/skills/rdkit-repr in your project.

What does RDKit Descriptors and Fingerprints need to run?

Going by SKILL.md and its folder, RDKit Descriptors and Fingerprints needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: `uv` installed; RDKit, pandas and NumPy, which `uv run` installs from inline metadata. Compatibility (from SKILL.md): Requires uv. Dependencies (rdkit, pandas, numpy) are declared as PEP 723 inline script metadata and are installed automatically when the script is invoked with `uv run <script_path>` (do NOT use `uv run python <script_path>` -- that bypasses the inline metadata and will not install dependencies automatically)..

Does RDKit Descriptors and Fingerprints access the network?

SKILL.md names 1 domain. As links in the text: rdkit.org. This is read from the text; nothing was executed.

Is RDKit Descriptors and Fingerprints safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does RDKit Descriptors and Fingerprints use?

RDKit Descriptors and Fingerprints is published under the LGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RDKit Descriptors and Fingerprints use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to RDKit Descriptors and Fingerprints?

Skills that share tags, products or a category with RDKit Descriptors and Fingerprints: Nvmolkit Usage (NVIDIA-BioNeMo/bionemo-agent-toolkit, 479 stars), ADMET Prediction for Drug Candidates (GPTomics/bioSkills, 1.2k stars), Vaex Out-of-Core DataFrames (davila7/claude-code-templates, 33k stars) and Statistical Data Analysis (lingzhi227/agent-research-skills, 390 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RDKit Descriptors and Fingerprints?

jinzhezenggroup (a GitHub organization) maintains it in jinzhezenggroup/computational-chemistry-agent-skills, which has 148 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on October 9, 2026.

Source: jinzhezenggroup/computational-chemistry-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.