DiffDock Molecular Docking
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and…
$ npx skills add GPTomics/bioSkills --skill bio-substructure-search -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-substructure-search --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/chemoinformatics/substructure-search .claude/skills/bio-substructure-search && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-substructure-search" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/substructure-search into .claude/skills/bio-substructure-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-substructure-search", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/substructure-searchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-substructure-search -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-substructure-search --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/chemoinformatics/substructure-search .agents/skills/bio-substructure-search && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-substructure-search" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/substructure-search into .agents/skills/bio-substructure-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-substructure-search", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-substructure-search -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-substructure-search --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/chemoinformatics/substructure-search .cursor/skills/bio-substructure-search && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-substructure-search" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/substructure-search into .cursor/skills/bio-substructure-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-substructure-search", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path chemoinformatics/substructure-search--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-substructure-search -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-substructure-search --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/chemoinformatics/substructure-search .gemini/skills/bio-substructure-search && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-substructure-search" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/substructure-search into .gemini/skills/bio-substructure-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-substructure-search", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-substructure-searchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-substructure-search -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/chemoinformatics/substructure-search .github/skills/bio-substructure-search && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-substructure-search" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/substructure-search into .github/skills/bio-substructure-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-substructure-search", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-substructure-search -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-substructure-search --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/chemoinformatics/substructure-search .opencode/skills/bio-substructure-search && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-substructure-search" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/substructure-search into .opencode/skills/bio-substructure-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-substructure-search", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-substructure-searchSearches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and…
Bio Substructure Search is an agent skill from GPTomics/bioSkills. Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and reactive/PAINS/REOS/Brenk filter catalogs. Use when filtering compounds by pharmacophore features, functional groups, scaffold matches, or screening for assay-interference / structural alerts.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/substructure_search.py` and `usage-guide.md`).
It sits in Research & Science. It works with RDKit. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
doi.orgdaylight.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Substructure Search loads about 4.3k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 1,575 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,575 words, ~4,324 tokens.
.claude/skills/bio-substructure-search/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: RDKit 2024.09+. SMARTS dialect follows Daylight specification with RDKit extensions.
Before using code patterns, verify installed versions match. If versions differ:
pip show rdkit then help(rdkit.Chem.MolFromSmarts) to check signaturesIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Search molecular collections for structural patterns using SMARTS. The choice of SMARTS dialect, atom/bond matching mode, and structural-alert catalog determines whether the search is correctly capturing the intended chemistry. PAINS (Baell & Holloway 2010) is the most-cited but most-misunderstood filter -- it identifies patterns of assay interference, not "bad molecules". Knowing when to apply each catalog and how to interpret hits is essential.
For SMARTS-based reactions (transforming matched substructures), see chemoinformatics/reaction-enumeration. For 3D pharmacophore matching, see chemoinformatics/pharmacophore-modeling.
| Token | Meaning | Example |
|---|---|---|
[#6] | Atom by atomic number | [#6] carbon (any hybridization) |
c | Lowercase = aromatic | c1ccccc1 benzene aromatic |
C | Uppercase = aliphatic only | C(=O)O carboxylic acid carbon |
[CX4] | Atom + connection count X | [CX4] sp3 carbon (4 connections) |
[CX3]=O | Carbonyl (CX3 = sp2 with 3 bonds) | matches ketone, aldehyde, ester C |
[#6;R] | Atom in ring | [#6;R] ring carbon |
[#6;!R] | Atom not in ring | [#6;!R] acyclic carbon |
[#6;r6] | Atom in 6-membered ring | [#6;r6] six-ring carbon |
[a] | Any aromatic atom | [a] |
[!#1] | Anything except H | [!#1] heavy atom |
[N;H2] | N with exactly 2 H; neighboring chemistry unconstrained | [NH2] also matches non-amine NH2 environments unless context is added |
[N+] | Positively charged N | [N+](=O)[O-] nitro |
[$(...)] | Recursive SMARTS | [$(c1ccccc1)] aromatic 6-ring atom |
[c]([F,Cl,Br,I]) | OR within brackets | aryl halide |
~ | Any bond type | c~c any aromatic-aromatic bond |
@ | Ring bond | c@c requires the matched bond to be in a ring |
- | Single bond explicit | C-C |
= | Double bond | C=O |
: | Aromatic bond explicit |
| Pattern | SMARTS | Notes |
|---|---|---|
| Hydroxyl (alcohol + phenol) | [OX2H] | OX2H avoids matching O- in OH- |
| Phenol only | [OX2H][c] | OH attached to aromatic carbon |
| Aliphatic OH only | [OX2H][CX4] | OH attached to sp3 C |
| Carboxylic acid | [CX3](=O)[OX2H1] | C(=O)OH |
| Carboxylate | [CX3](=O)[O-] | C(=O)O- (deprotonated) |
| Ester | [CX3](=O)[OX2][!H] | C(=O)O-R |
| Amide | [CX3](=[OX1])[NX3] | C(=O)N-R |
| Primary amine attached to carbon (excluding common amide-like N) | [NX3;H2;$(N-[#6]);!$(N-[C,S,P]=[O,S,N])] | Carbon-substituted -NH2; extend the exclusions for a project-specific amine definition |
| Secondary amine attached to two carbons | [NX3;H1;$(N(-[#6])-[#6]);!$(N-[C,S,P]=[O,S,N])] | Carbon-substituted -NH- excluding common amide-like N |
| Neutral tertiary amine attached to three carbons | [NX3;H0;+0;$(N(-[#6])(-[#6])-[#6]);!$(N-[C,S,P]=[O,S,N])] | Carbon-substituted -NR2 excluding common amide-like N |
| Quaternary amine | [NX4+] | -NR4+ |
| Nitro | [N+](=O)[O-] | -NO2 |
| Nitrile | [CX2]#[NX1] | -C#N |
| Sulfonamide | [SX4](=[OX1])(=[OX1])[NX3] | -S(=O)(=O)N |
| Aryl halide | [c][F,Cl,Br,I] | halogen on aromatic |
| Aliphatic halide | [CX4][F,Cl,Br,I] | halogen on sp3 C |
| Hydrogen bond donor | Use a named feature definition such as RDKit BaseFeatures.fdef or Lipinski.NumHDonors | [#7,#8;!H0] is only a simplified N/O-H query and is not a universal HBD model |
| Hydrogen bond acceptor | Use a named feature definition such as RDKit BaseFeatures.fdef or Lipinski.NumHAcceptors | No short universal SMARTS correctly captures every accepted HBA chemistry model |
| Michael acceptor | [CX3]=[CX3][CX3]=O | enone, acrylamide warhead |
| Aldehyde | [CX3H1](=O) | -CHO |
| Ketone | [CX3;H0](=[OX1])([#6])[#6] | Carbonyl carbon has two carbon substituents and no hydrogen |
Goal: Test whether a molecule contains a SMARTS pattern and enumerate the matching atom indices.
Approach: Parse the molecule with MolFromSmiles and the pattern with MolFromSmarts, gate with HasSubstructMatch, then call GetSubstructMatches and map each atom index back to the molecule for inspection.
from rdkit import Chem
mol = Chem.MolFromSmiles('c1ccc(O)cc1CCO')
pattern = Chem.MolFromSmarts('[OX2H]')
if mol.HasSubstructMatch(pattern):
matches = mol.GetSubstructMatches(pattern)
for match in matches:
atoms = [mol.GetAtomWithIdx(i).GetSymbol() for i in match]HasSubstructMatch returns bool, GetSubstructMatches returns tuple of tuples of atom indices.
[$(pattern)] matches an atom that also matches the entire pattern starting from itself. Critical for context-aware matching.
# Aromatic carbon attached to a carbonyl
pat = Chem.MolFromSmarts('[$(c[C](=O))]')
# Aniline-type N (aromatic carbon-N-H)
pat = Chem.MolFromSmarts('[$([NX3;H2][c])]')
# Neutral tertiary amine with three sp3-carbon neighbors
pat = Chem.MolFromSmarts('[$([NX3]([CX4])([CX4])[CX4])]')
# H-bond donor (per Lipinski, exclude quaternary)
hbd = Chem.MolFromSmarts('[#7,#8;!H0;!$([NX3+])]')
# For H-bond acceptors, use RDKit's maintained Lipinski/feature definitions
# instead of an ad hoc universal SMARTS.
from rdkit.Chem import Lipinski
n_acceptors = Lipinski.NumHAcceptors(mol)| Filter | Origin | Patterns | Use case | Failure mode |
|---|---|---|---|---|
| PAINS_A | Baell & Holloway 2010 | 16 | Most populated source-data patterns (>=150 analogues per pattern) | Many false positives in primary screens; legitimate medicines flagged |
| PAINS_B | Baell & Holloway 2010 | 55 | Intermediate source-data population (15-149 analogues per pattern) | Similar |
| PAINS_C | Baell & Holloway 2010 | 409 | Least populated source-data patterns (1-14 analogues per pattern) | Most permissive |
| BRENK | Brenk 2008 (DDS unsuitable) | 105 | Reactive / toxicity / undesirable | Useful for fragment / virtual library |
| NIH | NIH MLSMR | 180 in RDKit 2024.09 | Reactive groups, unstable | Legacy filter; verify count after toolkit upgrades |
| ZINC | ZINC clean-leads | 50 in RDKit 2024.09 | Drug-like cleanup | Verify definitions after toolkit upgrades |
| Glaxo / Eli Lilly | Vendor lists | varies | Internal "ugly" filters | Often unpublished |
| REOS | Walters & Murcko 2002 | property + structural | Drug-likeness combined filter | Hand-curated thresholds |
The PAINS A/B/C families encode pattern population in the original screening dataset, not increasing or decreasing external evidence strength.
| Scenario | Catalog | Reason |
|---|---|---|
| Hit validation from biochemical screen | PAINS_A | Identify assay-interference candidates |
| Library prep for HTS | PAINS_A + Brenk + ZINC | Remove clearly bad |
| Fragment library design | Brenk + ZINC | Remove reactive; PAINS less critical at fragments |
| Lead optimization | None mandatory | Filters can exclude valid leads |
| Natural product analog | None | Filters trained on synthetic chemistry |
| Covalent inhibitor design | Skip warhead filter | Warheads ARE the design |
Critical: Capuzzi et al. (2017) found PAINS alerts in 87 FDA-approved small-molecule drugs. PAINS is a flag for assay validation, not a killing filter.
Goal: Split a molecule list into PAINS-flagged and PAINS-clean sets using one or more PAINS catalog tiers.
Approach: Configure FilterCatalogParams with the requested catalog enums, build a FilterCatalog once, and for each molecule use GetFirstMatch to either bucket it as clean or record the matching pattern description.
from rdkit.Chem.FilterCatalog import FilterCatalog, FilterCatalogParams
def pains_filter(mols, catalogs=('PAINS_A',)):
params = FilterCatalogParams()
for cat in catalogs:
params.AddCatalog(getattr(FilterCatalogParams.FilterCatalogs, cat))
catalog = FilterCatalog(params)
flagged = []
clean = []
for mol in mols:
if mol is None:
continue
entry = catalog.GetFirstMatch(mol)
if entry is None:
clean.append(mol)
else:
flagged.append((mol, entry.GetDescription()))
return clean, flaggedAvailable catalog names: PAINS_A, PAINS_B, PAINS_C, PAINS (all), BRENK, NIH, ZINC, ALL.
For HTS triage, filter electrophilic warheads (acrylamide, chloroacetamide, etc.) unless designing covalent inhibitors.
Goal: Flag molecules containing electrophilic warheads or other reactive functional groups that would interfere with biochemical HTS.
Approach: Maintain a named SMARTS dictionary of reactive groups (acid halides, epoxides, Michael acceptors, etc.), then per molecule scan each pattern with HasSubstructMatch and return the first matching warhead name.
REACTIVE_SMARTS = {
'acid_anhydride': '[CX3](=O)O[CX3](=O)',
'acid_halide': '[CX3](=O)[F,Cl,Br,I]',
'alpha_halo_carbonyl': '[CX3](=O)C([F,Cl,Br,I])',
'aldehyde_reactive': '[CX3H1](=O)[#6;X4]', # aliphatic aldehydes
'epoxide': 'C1OC1',
'aziridine': 'C1NC1',
'isocyanate': '[NX2]=C=[OX1]',
'isothiocyanate': '[NX2]=C=[SX1]',
'beta_lactam': 'C1(=O)NCC1',
'sulfonyl_halide': '[SX4](=O)(=O)[F,Cl,Br,I]',
'Michael_acceptor': '[CX3]=[CX3][CX3]=O',
'vinyl_sulfone': '[SX4](=O)(=O)C=C',
}
def reactive_filter(mol, exclude_warheads=True):
if not exclude_warheads:
return False, None
for name, smarts in REACTIVE_SMARTS.items():
if mol.HasSubstructMatch(Chem.MolFromSmarts(smarts)):
return True, name
return False, NoneFor covalent-inhibitor design, see chemoinformatics/covalent-design; these warheads are the desired chemistry, not noise to filter.
Goal: Reduce a molecule library to those that match all required SMARTS patterns and none of the excluded ones.
Approach: Start from the full molecule list, iteratively intersect with each include SMARTS using HasSubstructMatch, then subtract any molecule matching an exclude SMARTS.
def filter_library(mols, include=None, exclude=None):
keep = list(mols)
if include:
for s in include:
p = Chem.MolFromSmarts(s)
keep = [m for m in keep if m and m.HasSubstructMatch(p)]
if exclude:
for s in exclude:
p = Chem.MolFromSmarts(s)
keep = [m for m in keep if m and not m.HasSubstructMatch(p)]
return keepAtom maps [C:1] track atoms through transformations. Used in reactions (reaction-enumeration skill) but also for substructure-based extraction:
# Find amide N with attached aryl
pat = Chem.MolFromSmarts('[CX3:1](=O)[NX3:2][c:3]')
match = mol.GetSubstructMatch(pat)
amide_C, amide_N, aryl_C = matchTrigger: Library contains natural products, polyphenols, flavonoids, quinones.
Mechanism: PAINS_A patterns target rhodanines, curcumins, polyhydroxylated polyphenols -- legitimate scaffolds in natural-product chemistry.
Symptom: Library hits flagged as PAINS but trace back to validated natural products with confirmed activity.
Fix: Use PAINS as a flag not a delete. Cross-check flagged compounds for orthogonal-assay confirmation (label-free e.g. SPR, ITC).
Trigger: SMARTS pattern with c (aromatic) for a heteroatom-rich ring; molecule parsed with different aromaticity model.
Mechanism: RDKit, OpenEye, ChemAxon differ on whether furan, thiazole, tropone, etc. are aromatic.
Symptom: Same pattern matches in one toolkit, not in another.
Fix: Re-canonicalize molecules within RDKit before applying SMARTS. Or use [#6]:[#6] instead of c:c (explicit element + bond type).
Trigger: SMARTS targets keto form C(=O) but molecule is enol C(O)=C.
Mechanism: Default canonical form differs by toolkit + standardization choice.
Symptom: Known matching molecule reports no match.
Fix: Use tautomer-aware match: enumerate tautomers and OR-match. Or canonicalize first via chemoinformatics/molecular-standardization. Or expand pattern with [$(C(=O)),$(C(O)=C)].
Trigger: SMARTS without /\@ stereo markers applied to mol with explicit stereo.
Mechanism: SMARTS matching is stereo-agnostic by default.
Symptom: Wrong stereoisomer is matched as well as right one.
Fix: mol.GetSubstructMatches(pattern, useChirality=True) to require chirality match.
Trigger: A query must distinguish an isolated benzene ring from a six-membered aromatic ring embedded in a fused system.
Mechanism: c1ccccc1 matches six-membered aromatic cycles and therefore does match benzene cycles within naphthalene. Extra ring-membership or fusion constraints are required to exclude fused systems.
Symptom: A nominal "benzene" query returns fused polyaromatics that the project intended to exclude.
Fix: Keep c1ccccc1 when any aromatic six-cycle is desired. When an isolated ring is required, add explicit ring-degree/fusion constraints and test the query against benzene, naphthalene, indole, and representative substituted controls.
Trigger: Deeply nested recursive SMARTS over a large library.
Mechanism: Each [$()] re-evaluates the inner pattern for every candidate atom.
Symptom: Search 10x-100x slower than expected.
Fix: Flatten recursion where possible; pre-filter with simpler pattern, then re-test with the recursive one.
| Symptom | Cause | Fix |
|---|---|---|
Chem.MolFromSmarts returns None | Invalid SMARTS grammar | Validate with Chem.MolFromSmarts(smi, mergeHs=False); check parens, brackets |
[OH] gives unexpected hydroxyl matches | Query does not state the intended valence/connectivity model | Use [OX2H] for neutral alcohol/phenol oxygen or a more specific context-aware pattern |
| Pattern matches but library is "empty" | Mol failed sanitize | Try Chem.SDMolSupplier(sanitize=False) then catch errors |
| Multiple matches per molecule | Single-match query expected | GetSubstructMatch returns first; GetSubstructMatches returns all |
| Match indices but no fragment | Match returns atom indices in pattern order | Map to original mol via mol.GetAtomWithIdx(i) |
| PAINS catalog initialization slow | Loading 1000+ patterns on every call | Build catalog once, reuse for batch |
| Stereo SMARTS not matching | useChirality=False (default) | mol.GetSubstructMatches(p, useChirality=True) |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in chemoinformatics/substructure-search of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Substructure Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Substructure Search this skillGPTomics/bioSkills | 1.2k | 2 repos | ~4.3k | Automated safety check: Pass | MIT | |
| DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3k | Automated safety check: Notes | MIT | |
| Edu Chem Reactionwy51ai/edulab | 1.4k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Biopipelineslocbp-uzh/biopipelines | 109 | — | ~2.4k | Automated safety check: Pass | MIT | |
| RDKit Cheminformatics Practicesaiming-lab/AutoResearchClaw | 15k | — | ~708 | Automated safety check: Pass | MIT | |
| Rowanlamm-mit/scienceclaw | 244 | 4 repos | ~3.1k | Automated safety check: Warn | Proprietary |
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
wy51ai/edulab
把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。
locbp-uzh/biopipelines
Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…
aiming-lab/AutoResearchClaw
Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.
lamm-mit/scienceclaw
Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.
pemsley/coot
RDKit molecular manipulation and visualization within Coot's Python environment.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Works with
Categories
Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and…. Bio Substructure Search is an agent skill from GPTomics/bioSkills. Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and reactive/PAINS/REOS/Brenk filter catalogs.
Bio Substructure Search fits situations like: filtering compounds by pharmacophore features; functional groups; scaffold matches; screening for assay-interference / structural alerts.
Run `npx skills add GPTomics/bioSkills --skill bio-substructure-search -a claude-code`. Or copy the skill folder (chemoinformatics/substructure-search in GPTomics/bioSkills) into .claude/skills/bio-substructure-search in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-substructure-search -a codex`. Or copy the skill folder (chemoinformatics/substructure-search in GPTomics/bioSkills) into .agents/skills/bio-substructure-search in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-substructure-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-substructure-search, .gemini/skills/bio-substructure-search, .github/skills/bio-substructure-search and .opencode/skills/bio-substructure-search in your project.
Going by SKILL.md and its folder, Bio Substructure Search needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: doi.org and daylight.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Substructure Search is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Substructure Search: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Edu Chem Reaction (wy51ai/edulab, 1.4k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars) and RDKit Cheminformatics Practices (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.