DiffDock Molecular Docking
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
Calculates molecular fingerprints (ECFP/Morgan, FCFP, MACCS, RDKit, AtomPair, TopologicalTorsion, Avalon, MAP4, MHFP6) and physicochemical descriptors (Lipinski, QED, TPSA, Crippen LogP, 3D shape)…
$ npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-molecular-descriptors --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/chemoinformatics/molecular-descriptors .claude/skills/bio-molecular-descriptors && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-molecular-descriptors" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/molecular-descriptors into .claude/skills/bio-molecular-descriptors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-molecular-descriptors", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/molecular-descriptorsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-molecular-descriptors --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/chemoinformatics/molecular-descriptors .agents/skills/bio-molecular-descriptors && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-molecular-descriptors" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/molecular-descriptors into .agents/skills/bio-molecular-descriptors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-molecular-descriptors", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-molecular-descriptors --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/chemoinformatics/molecular-descriptors .cursor/skills/bio-molecular-descriptors && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-molecular-descriptors" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/molecular-descriptors into .cursor/skills/bio-molecular-descriptors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-molecular-descriptors", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path chemoinformatics/molecular-descriptors--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-molecular-descriptors --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/chemoinformatics/molecular-descriptors .gemini/skills/bio-molecular-descriptors && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-molecular-descriptors" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/molecular-descriptors into .gemini/skills/bio-molecular-descriptors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-molecular-descriptors", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-molecular-descriptorsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/chemoinformatics/molecular-descriptors .github/skills/bio-molecular-descriptors && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-molecular-descriptors" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/molecular-descriptors into .github/skills/bio-molecular-descriptors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-molecular-descriptors", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-molecular-descriptors --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/chemoinformatics/molecular-descriptors .opencode/skills/bio-molecular-descriptors && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-molecular-descriptors" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/molecular-descriptors into .opencode/skills/bio-molecular-descriptors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-molecular-descriptors", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-molecular-descriptorsCalculates molecular fingerprints (ECFP/Morgan, FCFP, MACCS, RDKit, AtomPair, TopologicalTorsion, Avalon, MAP4, MHFP6) and physicochemical descriptors (Lipinski, QED, TPSA, Crippen LogP, 3D shape)…
Bio Molecular Descriptors is an agent skill from GPTomics/bioSkills. Calculates molecular fingerprints (ECFP/Morgan, FCFP, MACCS, RDKit, AtomPair, TopologicalTorsion, Avalon, MAP4, MHFP6) and physicochemical descriptors (Lipinski, QED, TPSA, Crippen LogP, 3D shape) with explicit choice tables, bit vs count semantics, and partial-charge model selection. Use when featurizing molecules for similarity, QSAR, virtual screening, or ML, or selecting the correct fingerprint for a chemotype-aware task.
Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/calculate_descriptors.py` and `usage-guide.md`).
It sits in Research & Science, covering Drug discovery and cheminformatics. It works with RDKit. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
doi.orgpypi.orgdocs.openforcefield.orgautodock-vina.readthedocs.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Molecular Descriptors loads about 4.5k tokens when it runs. Until then it costs about 114 tokens; SKILL.md has 1,730 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,730 words, ~4,541 tokens.
.claude/skills/bio-molecular-descriptors/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: RDKit 2024.09+, numpy 1.26+, pandas 2.2+, map4 1.1+ (MAP4), mhfp 1.9+. Use mapchiral separately when the stereochemistry-aware MAP4C fingerprint is intended.
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturesIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Featurize molecules for similarity search, QSAR, virtual screening, or ML. Fingerprint performance is dataset- and objective-dependent: ECFP4 is a strong drug-like baseline, atom-pair and topological-torsion fingerprints expose longer-range topology, MAP4/MHFP6 target broader chemical-space searches, and 3D conformer-based descriptors are needed when shape and stereochemistry matter.
For canonicalization before featurization, see chemoinformatics/molecular-standardization. For 3D-only descriptors, see chemoinformatics/conformer-generation.
| Fingerprint | Type | Radius/Path | Bits | Use case | Fails when |
|---|---|---|---|---|---|
| Morgan (ECFP) | Circular | r=2 (ECFP4), r=3 (ECFP6) | 2048 typical | Drug-like similarity, ML default | Loses long-range topology; bit collisions at low nBits |
| FCFP | Functional Morgan | r=2 default | 2048 | Pharmacophore-aware similarity | Same caveats as ECFP; less specific |
| MACCS | Substructure key | 166 fixed bits | 167 | Quick fingerprint, drug-likeness | Too sparse for large diverse libraries |
| RDKit FP | Path/subgraph-based | paths and branched subgraphs up to 7 bonds by default | 2048 | RDKit-native ECFP alternative | Drug-like only; not optimal for scaffold hopping |
| AtomPair | Pair + topological distance | All atom pairs | 2048 | Long-range topological similarity | Slower than ECFP; harder to interpret |
| TopologicalTorsion | 4-atom torsion | All TT | 2048 | Path-pattern similarity | Like AP, slower than ECFP |
| Avalon | Substructure + atom pairs | Mixed | 512/1024 | Fast similarity | Less standard; older |
| MAP4 (MinHashed atom-pair) | MinHash atom-pair | r=1,2 | 1024/2048 | Biological + metabolite diversity | map4 library required; slower hash |
| MHFP6 (MinHash) | MinHash ECFP-like | r=3 (diam 6) | 2048 | Large-library nearest-neighbor with a compatible MinHash/LSH index | Different distance semantics from folded-bit Tanimoto |
| Pharm2D | 2D pharmacophore | feature pairs/triplets | sparse | Pharmacophore search | Sparse, slower |
Decision: For drug-like similarity ranking, start with ECFP4 2048 bit because it is fast and well characterized. MHFP6 outperformed ECFP4 for analog recovery in the benchmark reported by Probst and Reymond (2018), making it a candidate for large, diverse libraries. For scaffold hopping, benchmark ECFP4, AtomPair, TopologicalTorsion, and pharmacophore fingerprints on target-relevant actives and decoys; published comparisons do not support a universal AtomPair advantage (Gardiner et al. 2011; Riniker & Landrum 2013).
| Form | Use | Library impact |
|---|---|---|
| Bit (0/1) | Tanimoto similarity, BulkTanimotoSimilarity, RDKit fingerprint folding | Standard for similarity |
| Count (integer) | Some ML methods, RF on counts, neural fingerprints | Loses bit-level fast operations; richer signal |
| Sparse (dict) | Direct chemical interpretation (which fragments at which atoms) | Use for SHAP / atomic attribution |
from rdkit import Chem
from rdkit.Chem import rdFingerprintGenerator
mol = Chem.MolFromSmiles('CCO')
morgan = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
ecfp4_bit = morgan.GetFingerprint(mol)
ecfp4_count = morgan.GetCountFingerprint(mol)
ecfp4_sparse = morgan.GetSparseCountFingerprint(mol)ECFP-X notation: X is the diameter in bonds. RDKit's radius parameter is half of X.
| Notation | RDKit radius | Diameter | Captures |
|---|---|---|---|
| ECFP0 | 0 | 0 | Atom identity only |
| ECFP2 | 1 | 2 | Atom + immediate neighbors |
| ECFP4 | 2 | 4 | Atom + 2-bond environment |
| ECFP6 | 3 | 6 | Atom + 3-bond environment |
Trade-off: Larger radius captures more specific local environments but increases collisions at fixed nBits. ECFP4 2048 is a common baseline (Rogers & Hahn 2010; Wu et al. 2018). O'Boyle and Sayle (2016) showed that increasing folded-fingerprint length can improve virtual-screening performance, but they did not establish a universal 4096-bit setting or a 1-5% collision rate. Measure collision occupancy and model performance for the dataset; increase nBits or use an unhashed sparse representation when needed.
FCFP (Functional-Class) uses RDKit's Morgan feature invariants (donor, acceptor, aromatic, halogen, basic, and acidic) instead of atom identity. Hydrophobe is a family in BaseFeatures.fdef, but it is not one of the default Morgan feature-invariant classes. FCFP trades atom-specificity for functional-equivalence.
ecfp_generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
feature_invariants = rdFingerprintGenerator.GetMorganFeatureAtomInvGen()
fcfp_generator = rdFingerprintGenerator.GetMorganGenerator(
radius=2, fpSize=2048, atomInvariantsGenerator=feature_invariants)
ecfp4 = ecfp_generator.GetFingerprint(mol)
fcfp4 = fcfp_generator.GetFingerprint(mol)When to use FCFP4: Scaffold-hopping campaigns, pharmacophore-driven similarity, cross-target activity prediction.
When to use ECFP4: Within-series QSAR, lead optimization, when chemotype identity matters.
Conformer-dependent descriptors (asphericity, eccentricity, principal moments of inertia, RDF) require a generated 3D structure. A single conformer may be unrepresentative when the molecule is flexible; measure descriptor variation across a conformer ensemble when the downstream conclusion depends on 3D shape.
Goal: Compute 3D shape descriptors over a conformer ensemble rather than from a single (possibly unrepresentative) conformer.
Approach: Add explicit hydrogens, embed N conformers with ETKDGv3, MMFF-optimize them all, then evaluate the descriptor across each conformer for downstream averaging.
from rdkit.Chem import AllChem, Descriptors3D
mol = Chem.MolFromSmiles('CCCCO')
mol = Chem.AddHs(mol)
params = AllChem.ETKDGv3()
params.randomSeed = 42
conf_ids = AllChem.EmbedMultipleConfs(mol, numConfs=20, params=params)
if not conf_ids:
raise RuntimeError('ETKDGv3 failed to generate any conformers')
if not AllChem.MMFFHasAllMoleculeParams(mol):
raise ValueError('MMFF94 parameters are unavailable for this molecule')
optimization_results = AllChem.MMFFOptimizeMoleculeConfs(mol)
if any(status != 0 for status, _ in optimization_results):
raise RuntimeError('MMFF94 optimization did not converge for every conformer')
asphericities = [Descriptors3D.Asphericity(mol, confId=c) for c in conf_ids]Decision: For QSAR / ML, choose and document the conformer count using a convergence check on representative molecules. Report the aggregation rule, such as a simple mean or a Boltzmann-weighted average, and the energy model used for any weights.
| Method | Software | Cost | Accuracy | Use for |
|---|---|---|---|---|
| Gasteiger-Marsili | RDKit, Open Babel | Fast | Empirical, rough | Charge-aware preparation or models that explicitly require Gasteiger charges; Vina/Vinardo scoring itself does not require assigned atom charges |
| MMFF94 | RDKit | 0.1s/mol | Force-field consistent | MMFF energy, conformer ranking |
| AM1-BCC | antechamber (AmberTools) | ~10s/mol | Semi-empirical | MD setup, FEP, GAFF |
| RESP | psi4, Gaussian | minutes/mol | Restrained fit to a quantum-mechanical ESP; protocol-specific | Force-field workflows parameterized for that RESP protocol |
| OpenFF Recharge | openff-recharge | Workflow-dependent | Framework for generating/retrieving QC ESP data and fitting library charges, BCCs, RESP charges, or virtual sites | Developing or evaluating charge models; it is not one charge-assignment method |
from rdkit.Chem import AllChem
AllChem.ComputeGasteigerCharges(mol)
for atom in mol.GetAtoms():
print(atom.GetIdx(), atom.GetPropsAsDict().get('_GasteigerCharge', None))Critical: Charge method must match downstream. Gasteiger charges in an AMBER MD run violate the assumptions of the protein force field.
For libraries spanning drug-like molecules, natural products, peptides, and metabolites, compare ECFP4 with MAP4 or MHFP6 on task-relevant retrieval benchmarks. MAP4 and MHFP6 use MinHash with atom-pair or circular-substructure shingles, but no universal pairwise-similarity range establishes that ECFP4 is saturated for every mixed library.
from mhfp.encoder import MHFPEncoder
encoder = MHFPEncoder(2048)
mhfp6 = encoder.encode_mol(mol, radius=3)MHFP6 distance is Jaccard on MinHash, not standard Tanimoto. Use MHFPEncoder.distance(fp1, fp2).
| Descriptor | Source | Range | Drug-like cutoff |
|---|---|---|---|
| MolWt | RDKit Descriptors.MolWt | ~50-2000 Da | <=500 (Lipinski) |
| MolLogP (Crippen) | RDKit Descriptors.MolLogP | -5 to 8 | <=5 (Lipinski) |
| HBD | Lipinski.NumHDonors | 0-10 | <=5 (Lipinski) |
| HBA | Lipinski.NumHAcceptors | 0-15 | <=10 (Lipinski) |
| TPSA | Descriptors.TPSA (Ertl) | 0-200 A^2 | <=140 (Veber oral); <=90 (BBB+) |
| RotBonds | Lipinski.NumRotatableBonds | 0-15 | <=10 (Veber) |
| AromaticRings | Lipinski.NumAromaticRings | 0-6 | <=3-4 (Ritchie-Macdonald aromatic ring count) |
| HeavyAtoms | Descriptors.HeavyAtomCount | <=50 (lead-like) | |
| FractionCSP3 | Descriptors.FractionCSP3 | 0-1 | Descriptive; higher sp3 character was associated with clinical progression by Lovering et al. (2009), without a universal cutoff |
| QED | QED.qed | 0-1 | Higher is more similar to the reference property distributions; a project may use >=0.5 as a triage heuristic |
| SAscore | sascorer.calculateScore (external) | 1-10 | Lower is easier by the model; project cutoffs such as <=4 or >6 require dataset calibration |
Goal: Compute a standard physicochemical descriptor panel for drug-likeness filtering and QSAR features.
Approach: Combine RDKit Descriptors, Lipinski, and QED calls into a single dict so the caller gets MW, LogP, HBD/HBA, TPSA, rotatable bonds, aromatic rings, fraction sp3, and QED in one pass.
from rdkit.Chem import Descriptors, Lipinski, QED
def physchem(mol):
return {
'MolWt': Descriptors.MolWt(mol),
'MolLogP': Descriptors.MolLogP(mol),
'HBD': Lipinski.NumHDonors(mol),
'HBA': Lipinski.NumHAcceptors(mol),
'TPSA': Descriptors.TPSA(mol),
'RotBonds': Lipinski.NumRotatableBonds(mol),
'AromRings': Lipinski.NumAromaticRings(mol),
'FractionCSP3': Descriptors.FractionCSP3(mol),
'QED': QED.qed(mol),
}| Rule | Constraints | Source |
|---|---|---|
| Lipinski Ro5 | MW<=500, LogP<=5, HBD<=5, HBA<=10 | Lipinski 1997 |
| Veber | RotBonds<=10, TPSA<=140 | Veber 2002 (oral) |
| Ghose | 160<=MW<=480, -0.4<=LogP<=5.6, 40<=MR<=130, 20<=atoms<=70 | Ghose 1999 |
| Egan | LogP<=5.88, TPSA<=131.6 | Egan 2000 |
| Muegge | 200<=MW<=600, -2<=LogP<=5, TPSA<=150, rings<=7, C>4, heteroatoms>1, RotBonds<=15, HBD<=5, HBA<=10 | Muegge 2001 |
| Lead-like | MW<=350, LogP<=3 | Teague 1999 |
| Fragment Ro3 | MW<=300, LogP<=3, HBD<=3, HBA<=3, RotBonds<=3, TPSA<=60 A^2 | Congreve 2003 |
| Pfizer CNS MPO | Six desirability functions: ClogP, ClogD, MW, TPSA, HBD, and pKa | Wager 2010 |
Use case: Treat Ro5 and Veber criteria as risk indicators rather than universal hard cutoffs. Doak et al. (2014) analyze orally bioavailable drugs and candidates beyond the Rule of 5, but do not support the claim that approximately 30% of marketed oral drugs violate at least one rule. For CNS prioritization, implement the six-property Wager MPO desirability score rather than replacing it with three hard thresholds.
QED (Bickerton 2012) is a single-number drug-likeness measure (0-1) combining 8 properties (MW, LogP, HBD, HBA, PSA, RotBonds, AromaticRings, structural alerts) via desirability functions.
Caveat: QED summarizes desirability functions derived from property distributions of marketed oral drugs; it is not a supervised predictor trained specifically on FDA-approved drugs. It can under-rank fragment-like or natural-product-like molecules, so do not use it as the sole filter for those libraries.
| Symptom | Cause | Fix |
|---|---|---|
| Fingerprint changes between runs | Random seed not set for canonicalization | RDKit Morgan is deterministic; check if input differs (stereo, charges) |
| MACCS bit count != 166 | RDKit MACCS returns 167 bits (bit 0 unused) | Slice [1:] if comparing to literature 166-bit |
| Crippen LogP differs from XLogP | Different model | Use Descriptors.MolLogP for Crippen; XLogP3 requires external lib |
| 3D descriptor differs between calls | Different conformer | Set confId=0 explicitly; or average over ensemble |
| QED returns nan | Charged species or non-standard atom | Standardize (uncharge) before QED |
| Count-vector similarity differs from bit-vector similarity | Count multiplicities change the generalized Tanimoto calculation | RDKit supports Tanimoto on sparse count vectors; record the vector type and do not compare its threshold directly with a folded-bit threshold |
| MolWt off by ~1 from PubChem | Implicit H counted differently | Use Descriptors.ExactMolWt for monoisotopic; PubChem reports average |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in chemoinformatics/molecular-descriptors of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Molecular Descriptors next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Molecular Descriptors this skillGPTomics/bioSkills | 1.2k | 2 repos | ~4.5k | Automated safety check: Pass | MIT | |
| DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3k | Automated safety check: Notes | MIT | |
| Biopipelineslocbp-uzh/biopipelines | 109 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Edu Chem Reactionwy51ai/edulab | 1.4k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| RDKit Cheminformatics Practicesaiming-lab/AutoResearchClaw | 15k | — | ~708 | Automated safety check: Pass | MIT | |
| Rowanlamm-mit/scienceclaw | 246 | 4 repos | ~3.1k | Automated safety check: Warn | Proprietary |
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
locbp-uzh/biopipelines
Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…
wy51ai/edulab
把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。
aiming-lab/AutoResearchClaw
Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.
lamm-mit/scienceclaw
Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.
pemsley/coot
RDKit molecular manipulation and visualization within Coot's Python environment.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Works with
Categories
Calculates molecular fingerprints (ECFP/Morgan, FCFP, MACCS, RDKit, AtomPair, TopologicalTorsion, Avalon, MAP4, MHFP6) and physicochemical descriptors (Lipinski, QED, TPSA, Crippen LogP, 3D shape)…. Bio Molecular Descriptors is an agent skill from GPTomics/bioSkills. Calculates molecular fingerprints (ECFP/Morgan, FCFP, MACCS, RDKit, AtomPair, TopologicalTorsion, Avalon, MAP4, MHFP6) and physicochemical descriptors (Lipinski, QED, TPSA, Crippen LogP, 3D shape) with explicit choice tables, bit vs count semantics, and partial-charge model selection.
Bio Molecular Descriptors fits situations like: featurizing molecules for similarity; virtual screening; selecting the correct fingerprint for a chemotype-aware task.
Run `npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a claude-code`. Or copy the skill folder (chemoinformatics/molecular-descriptors in GPTomics/bioSkills) into .claude/skills/bio-molecular-descriptors in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a codex`. Or copy the skill folder (chemoinformatics/molecular-descriptors in GPTomics/bioSkills) into .agents/skills/bio-molecular-descriptors in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-molecular-descriptors -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-molecular-descriptors, .gemini/skills/bio-molecular-descriptors, .github/skills/bio-molecular-descriptors and .opencode/skills/bio-molecular-descriptors in your project.
Going by SKILL.md and its folder, Bio Molecular Descriptors needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 4 domains. As links in the text: doi.org, pypi.org, docs.openforcefield.org and autodock-vina.readthedocs.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Molecular Descriptors is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Molecular Descriptors: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars), Edu Chem Reaction (wy51ai/edulab, 1.4k stars) and RDKit Cheminformatics Practices (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.