Agent skill

Bio Substructure Search

by GPTomics in GPTomics/bioSkills

Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and…

MITAuto-check passedResearch & Science

Install Bio Substructure Search

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-substructure-search -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-substructure-search --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/chemoinformatics/substructure-search .claude/skills/bio-substructure-search && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-substructure-search
GitHub stars
1.2k
Used in
2 other repos
Token cost
~4.3k tokens
SKILL.md length
1,575 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and…

  • Filtering compounds by pharmacophore features
  • SKILL.md covers Version Compatibility, SMARTS Grammar Essentials, Common SMARTS Patterns and Basic Substructure Match, plus 11 more sections
  • Runs Python scripts from its folder; calls pip
  • Functional groups

What it does

Bio Substructure Search is an agent skill from GPTomics/bioSkills. Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and reactive/PAINS/REOS/Brenk filter catalogs. Use when filtering compounds by pharmacophore features, functional groups, scaffold matches, or screening for assay-interference / structural alerts.

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/substructure_search.py` and `usage-guide.md`).

It sits in Research & Science. It works with RDKit. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Filtering compounds by pharmacophore features
  • Functional groups
  • Scaffold matches
  • Screening for assay-interference / structural alerts

Example prompts

  • “Use the bio-substructure-search skill to search molecular libraries for substructure matches using SMARTS patterns with explicit handling of…”
  • “/bio-substructure-search”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org
    • daylight.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Substructure Search loads about 4.3k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 1,575 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,575 words, ~4,324 tokens.

Download SKILL.mdSave it as .claude/skills/bio-substructure-search/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-substructure-search
description
Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and reactive/PAINS/REOS/Brenk filter catalogs. Use when filtering compounds by pharmacophore features, functional groups, scaffold matches, or screening for assay-interference / structural alerts.
tool_type
python
primary_tool
RDKit

Version Compatibility

Reference examples tested with: RDKit 2024.09+. SMARTS dialect follows Daylight specification with RDKit extensions.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show rdkit then help(rdkit.Chem.MolFromSmarts) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Search molecular collections for structural patterns using SMARTS. The choice of SMARTS dialect, atom/bond matching mode, and structural-alert catalog determines whether the search is correctly capturing the intended chemistry. PAINS (Baell & Holloway 2010) is the most-cited but most-misunderstood filter -- it identifies patterns of assay interference, not "bad molecules". Knowing when to apply each catalog and how to interpret hits is essential.

For SMARTS-based reactions (transforming matched substructures), see chemoinformatics/reaction-enumeration. For 3D pharmacophore matching, see chemoinformatics/pharmacophore-modeling.

SMARTS Grammar Essentials

TokenMeaningExample
[#6]Atom by atomic number[#6] carbon (any hybridization)
cLowercase = aromaticc1ccccc1 benzene aromatic
CUppercase = aliphatic onlyC(=O)O carboxylic acid carbon
[CX4]Atom + connection count X[CX4] sp3 carbon (4 connections)
[CX3]=OCarbonyl (CX3 = sp2 with 3 bonds)matches ketone, aldehyde, ester C
[#6;R]Atom in ring[#6;R] ring carbon
[#6;!R]Atom not in ring[#6;!R] acyclic carbon
[#6;r6]Atom in 6-membered ring[#6;r6] six-ring carbon
[a]Any aromatic atom[a]
[!#1]Anything except H[!#1] heavy atom
[N;H2]N with exactly 2 H; neighboring chemistry unconstrained[NH2] also matches non-amine NH2 environments unless context is added
[N+]Positively charged N[N+](=O)[O-] nitro
[$(...)]Recursive SMARTS[$(c1ccccc1)] aromatic 6-ring atom
[c]([F,Cl,Br,I])OR within bracketsaryl halide
~Any bond typec~c any aromatic-aromatic bond
@Ring bondc@c requires the matched bond to be in a ring
-Single bond explicitC-C
=Double bondC=O
:Aromatic bond explicit

Common SMARTS Patterns

PatternSMARTSNotes
Hydroxyl (alcohol + phenol)[OX2H]OX2H avoids matching O- in OH-
Phenol only[OX2H][c]OH attached to aromatic carbon
Aliphatic OH only[OX2H][CX4]OH attached to sp3 C
Carboxylic acid[CX3](=O)[OX2H1]C(=O)OH
Carboxylate[CX3](=O)[O-]C(=O)O- (deprotonated)
Ester[CX3](=O)[OX2][!H]C(=O)O-R
Amide[CX3](=[OX1])[NX3]C(=O)N-R
Primary amine attached to carbon (excluding common amide-like N)[NX3;H2;$(N-[#6]);!$(N-[C,S,P]=[O,S,N])]Carbon-substituted -NH2; extend the exclusions for a project-specific amine definition
Secondary amine attached to two carbons[NX3;H1;$(N(-[#6])-[#6]);!$(N-[C,S,P]=[O,S,N])]Carbon-substituted -NH- excluding common amide-like N
Neutral tertiary amine attached to three carbons[NX3;H0;+0;$(N(-[#6])(-[#6])-[#6]);!$(N-[C,S,P]=[O,S,N])]Carbon-substituted -NR2 excluding common amide-like N
Quaternary amine[NX4+]-NR4+
Nitro[N+](=O)[O-]-NO2
Nitrile[CX2]#[NX1]-C#N
Sulfonamide[SX4](=[OX1])(=[OX1])[NX3]-S(=O)(=O)N
Aryl halide[c][F,Cl,Br,I]halogen on aromatic
Aliphatic halide[CX4][F,Cl,Br,I]halogen on sp3 C
Hydrogen bond donorUse a named feature definition such as RDKit BaseFeatures.fdef or Lipinski.NumHDonors[#7,#8;!H0] is only a simplified N/O-H query and is not a universal HBD model
Hydrogen bond acceptorUse a named feature definition such as RDKit BaseFeatures.fdef or Lipinski.NumHAcceptorsNo short universal SMARTS correctly captures every accepted HBA chemistry model
Michael acceptor[CX3]=[CX3][CX3]=Oenone, acrylamide warhead
Aldehyde[CX3H1](=O)-CHO
Ketone[CX3;H0](=[OX1])([#6])[#6]Carbonyl carbon has two carbon substituents and no hydrogen

Basic Substructure Match

Goal: Test whether a molecule contains a SMARTS pattern and enumerate the matching atom indices.

Approach: Parse the molecule with MolFromSmiles and the pattern with MolFromSmarts, gate with HasSubstructMatch, then call GetSubstructMatches and map each atom index back to the molecule for inspection.

python
from rdkit import Chem

mol = Chem.MolFromSmiles('c1ccc(O)cc1CCO')
pattern = Chem.MolFromSmarts('[OX2H]')

if mol.HasSubstructMatch(pattern):
    matches = mol.GetSubstructMatches(pattern)
    for match in matches:
        atoms = [mol.GetAtomWithIdx(i).GetSymbol() for i in match]

HasSubstructMatch returns bool, GetSubstructMatches returns tuple of tuples of atom indices.

Recursive SMARTS for Context-Aware Patterns

[$(pattern)] matches an atom that also matches the entire pattern starting from itself. Critical for context-aware matching.

python
# Aromatic carbon attached to a carbonyl
pat = Chem.MolFromSmarts('[$(c[C](=O))]')

# Aniline-type N (aromatic carbon-N-H)
pat = Chem.MolFromSmarts('[$([NX3;H2][c])]')

# Neutral tertiary amine with three sp3-carbon neighbors
pat = Chem.MolFromSmarts('[$([NX3]([CX4])([CX4])[CX4])]')

# H-bond donor (per Lipinski, exclude quaternary)
hbd = Chem.MolFromSmarts('[#7,#8;!H0;!$([NX3+])]')

# For H-bond acceptors, use RDKit's maintained Lipinski/feature definitions
# instead of an ad hoc universal SMARTS.
from rdkit.Chem import Lipinski
n_acceptors = Lipinski.NumHAcceptors(mol)

Structural-Alert Filter Catalogs

FilterOriginPatternsUse caseFailure mode
PAINS_ABaell & Holloway 201016Most populated source-data patterns (>=150 analogues per pattern)Many false positives in primary screens; legitimate medicines flagged
PAINS_BBaell & Holloway 201055Intermediate source-data population (15-149 analogues per pattern)Similar
PAINS_CBaell & Holloway 2010409Least populated source-data patterns (1-14 analogues per pattern)Most permissive
BRENKBrenk 2008 (DDS unsuitable)105Reactive / toxicity / undesirableUseful for fragment / virtual library
NIHNIH MLSMR180 in RDKit 2024.09Reactive groups, unstableLegacy filter; verify count after toolkit upgrades
ZINCZINC clean-leads50 in RDKit 2024.09Drug-like cleanupVerify definitions after toolkit upgrades
Glaxo / Eli LillyVendor listsvariesInternal "ugly" filtersOften unpublished
REOSWalters & Murcko 2002property + structuralDrug-likeness combined filterHand-curated thresholds

The PAINS A/B/C families encode pattern population in the original screening dataset, not increasing or decreasing external evidence strength.

When to Apply Each Filter

ScenarioCatalogReason
Hit validation from biochemical screenPAINS_AIdentify assay-interference candidates
Library prep for HTSPAINS_A + Brenk + ZINCRemove clearly bad
Fragment library designBrenk + ZINCRemove reactive; PAINS less critical at fragments
Lead optimizationNone mandatoryFilters can exclude valid leads
Natural product analogNoneFilters trained on synthetic chemistry
Covalent inhibitor designSkip warhead filterWarheads ARE the design

Critical: Capuzzi et al. (2017) found PAINS alerts in 87 FDA-approved small-molecule drugs. PAINS is a flag for assay validation, not a killing filter.

PAINS Filter

Goal: Split a molecule list into PAINS-flagged and PAINS-clean sets using one or more PAINS catalog tiers.

Approach: Configure FilterCatalogParams with the requested catalog enums, build a FilterCatalog once, and for each molecule use GetFirstMatch to either bucket it as clean or record the matching pattern description.

python
from rdkit.Chem.FilterCatalog import FilterCatalog, FilterCatalogParams

def pains_filter(mols, catalogs=('PAINS_A',)):
    params = FilterCatalogParams()
    for cat in catalogs:
        params.AddCatalog(getattr(FilterCatalogParams.FilterCatalogs, cat))
    catalog = FilterCatalog(params)

    flagged = []
    clean = []
    for mol in mols:
        if mol is None:
            continue
        entry = catalog.GetFirstMatch(mol)
        if entry is None:
            clean.append(mol)
        else:
            flagged.append((mol, entry.GetDescription()))
    return clean, flagged

Available catalog names: PAINS_A, PAINS_B, PAINS_C, PAINS (all), BRENK, NIH, ZINC, ALL.

Reaction-Reactive Group Filter (custom)

For HTS triage, filter electrophilic warheads (acrylamide, chloroacetamide, etc.) unless designing covalent inhibitors.

Goal: Flag molecules containing electrophilic warheads or other reactive functional groups that would interfere with biochemical HTS.

Approach: Maintain a named SMARTS dictionary of reactive groups (acid halides, epoxides, Michael acceptors, etc.), then per molecule scan each pattern with HasSubstructMatch and return the first matching warhead name.

python
REACTIVE_SMARTS = {
    'acid_anhydride': '[CX3](=O)O[CX3](=O)',
    'acid_halide': '[CX3](=O)[F,Cl,Br,I]',
    'alpha_halo_carbonyl': '[CX3](=O)C([F,Cl,Br,I])',
    'aldehyde_reactive': '[CX3H1](=O)[#6;X4]',  # aliphatic aldehydes
    'epoxide': 'C1OC1',
    'aziridine': 'C1NC1',
    'isocyanate': '[NX2]=C=[OX1]',
    'isothiocyanate': '[NX2]=C=[SX1]',
    'beta_lactam': 'C1(=O)NCC1',
    'sulfonyl_halide': '[SX4](=O)(=O)[F,Cl,Br,I]',
    'Michael_acceptor': '[CX3]=[CX3][CX3]=O',
    'vinyl_sulfone': '[SX4](=O)(=O)C=C',
}

def reactive_filter(mol, exclude_warheads=True):
    if not exclude_warheads:
        return False, None
    for name, smarts in REACTIVE_SMARTS.items():
        if mol.HasSubstructMatch(Chem.MolFromSmarts(smarts)):
            return True, name
    return False, None

For covalent-inhibitor design, see chemoinformatics/covalent-design; these warheads are the desired chemistry, not noise to filter.

Show full SKILL.md (638 more words)Show less

Library Filtering with Multiple Patterns

Goal: Reduce a molecule library to those that match all required SMARTS patterns and none of the excluded ones.

Approach: Start from the full molecule list, iteratively intersect with each include SMARTS using HasSubstructMatch, then subtract any molecule matching an exclude SMARTS.

python
def filter_library(mols, include=None, exclude=None):
    keep = list(mols)
    if include:
        for s in include:
            p = Chem.MolFromSmarts(s)
            keep = [m for m in keep if m and m.HasSubstructMatch(p)]
    if exclude:
        for s in exclude:
            p = Chem.MolFromSmarts(s)
            keep = [m for m in keep if m and not m.HasSubstructMatch(p)]
    return keep

Atom Map Indices in SMARTS

Atom maps [C:1] track atoms through transformations. Used in reactions (reaction-enumeration skill) but also for substructure-based extraction:

python
# Find amide N with attached aryl
pat = Chem.MolFromSmarts('[CX3:1](=O)[NX3:2][c:3]')
match = mol.GetSubstructMatch(pat)

amide_C, amide_N, aryl_C = match

Per-Tool Failure Modes

PAINS -- false positive on natural product

Trigger: Library contains natural products, polyphenols, flavonoids, quinones.

Mechanism: PAINS_A patterns target rhodanines, curcumins, polyhydroxylated polyphenols -- legitimate scaffolds in natural-product chemistry.

Symptom: Library hits flagged as PAINS but trace back to validated natural products with confirmed activity.

Fix: Use PAINS as a flag not a delete. Cross-check flagged compounds for orthogonal-assay confirmation (label-free e.g. SPR, ITC).

Aromaticity dialect mismatch

Trigger: SMARTS pattern with c (aromatic) for a heteroatom-rich ring; molecule parsed with different aromaticity model.

Mechanism: RDKit, OpenEye, ChemAxon differ on whether furan, thiazole, tropone, etc. are aromatic.

Symptom: Same pattern matches in one toolkit, not in another.

Fix: Re-canonicalize molecules within RDKit before applying SMARTS. Or use [#6]:[#6] instead of c:c (explicit element + bond type).

Tautomer-sensitive pattern miss

Trigger: SMARTS targets keto form C(=O) but molecule is enol C(O)=C.

Mechanism: Default canonical form differs by toolkit + standardization choice.

Symptom: Known matching molecule reports no match.

Fix: Use tautomer-aware match: enumerate tautomers and OR-match. Or canonicalize first via chemoinformatics/molecular-standardization. Or expand pattern with [$(C(=O)),$(C(O)=C)].

Stereochemistry ignored

Trigger: SMARTS without /\@ stereo markers applied to mol with explicit stereo.

Mechanism: SMARTS matching is stereo-agnostic by default.

Symptom: Wrong stereoisomer is matched as well as right one.

Fix: mol.GetSubstructMatches(pattern, useChirality=True) to require chirality match.

Ring closure / fused-ring specificity

Trigger: A query must distinguish an isolated benzene ring from a six-membered aromatic ring embedded in a fused system.

Mechanism: c1ccccc1 matches six-membered aromatic cycles and therefore does match benzene cycles within naphthalene. Extra ring-membership or fusion constraints are required to exclude fused systems.

Symptom: A nominal "benzene" query returns fused polyaromatics that the project intended to exclude.

Fix: Keep c1ccccc1 when any aromatic six-cycle is desired. When an isolated ring is required, add explicit ring-degree/fusion constraints and test the query against benzene, naphthalene, indole, and representative substituted controls.

Recursive SMARTS performance

Trigger: Deeply nested recursive SMARTS over a large library.

Mechanism: Each [$()] re-evaluates the inner pattern for every candidate atom.

Symptom: Search 10x-100x slower than expected.

Fix: Flatten recursion where possible; pre-filter with simpler pattern, then re-test with the recursive one.

Common Errors

SymptomCauseFix
Chem.MolFromSmarts returns NoneInvalid SMARTS grammarValidate with Chem.MolFromSmarts(smi, mergeHs=False); check parens, brackets
[OH] gives unexpected hydroxyl matchesQuery does not state the intended valence/connectivity modelUse [OX2H] for neutral alcohol/phenol oxygen or a more specific context-aware pattern
Pattern matches but library is "empty"Mol failed sanitizeTry Chem.SDMolSupplier(sanitize=False) then catch errors
Multiple matches per moleculeSingle-match query expectedGetSubstructMatch returns first; GetSubstructMatches returns all
Match indices but no fragmentMatch returns atom indices in pattern orderMap to original mol via mol.GetAtomWithIdx(i)
PAINS catalog initialization slowLoading 1000+ patterns on every callBuild catalog once, reuse for batch
Stereo SMARTS not matchinguseChirality=False (default)mol.GetSubstructMatches(p, useChirality=True)

References

  • chemoinformatics/molecular-io - Parse molecules before searching
  • chemoinformatics/molecular-standardization - Canonicalize tautomers before SMARTS
  • chemoinformatics/similarity-searching - Fingerprint-based fuzzy matching
  • chemoinformatics/scaffold-analysis - Scaffold-based pattern derivation
  • chemoinformatics/reaction-enumeration - SMARTS for chemical transformations
  • chemoinformatics/admet-prediction - PAINS as ADMET filter
  • chemoinformatics/covalent-design - Warhead chemistry

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in chemoinformatics/substructure-search of GPTomics/bioSkills.

  • SKILL.md
  • examples/substructure_search.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Substructure Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Substructure Search compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Substructure Search this skillGPTomics/bioSkills1.2k2 repos~4.3kAutomated safety check: PassMIT
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Edu Chem Reactionwy51ai/edulab1.4k—~1.2kAutomated safety check: PassApache-2.0
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT
RDKit Cheminformatics Practicesaiming-lab/AutoResearchClaw15k—~708Automated safety check: PassMIT
Rowanlamm-mit/scienceclaw2444 repos~3.1kAutomated safety check: WarnProprietary

Similar skills

  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Edu Chem Reaction

    wy51ai/edulab

    把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。

    1.4k GitHub stars~1.2k tokensUpdated 11 days ago
    Research & ScienceAuto-check passed
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed
  • RDKit Cheminformatics Practices

    aiming-lab/AutoResearchClaw

    Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.

    15k GitHub stars~708 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Rowan

    lamm-mit/scienceclaw

    Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.

    244 GitHub starsUsed in 4 repos~3.1k tokens
    Research & ScienceAuto-check: warnings
  • Coot Rdkit

    pemsley/coot

    RDKit molecular manipulation and visualization within Coot's Python environment.

    168 GitHub stars~981 tokensUpdated yesterday
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Substructure Search

What does Bio Substructure Search do?

Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and…. Bio Substructure Search is an agent skill from GPTomics/bioSkills. Searches molecular libraries for substructure matches using SMARTS patterns with explicit handling of recursive SMARTS, ring membership, aromaticity dialect, vector binding, atom map indices, and reactive/PAINS/REOS/Brenk filter catalogs.

When should I use Bio Substructure Search?

Bio Substructure Search fits situations like: filtering compounds by pharmacophore features; functional groups; scaffold matches; screening for assay-interference / structural alerts.

How do I install Bio Substructure Search in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-substructure-search -a claude-code`. Or copy the skill folder (chemoinformatics/substructure-search in GPTomics/bioSkills) into .claude/skills/bio-substructure-search in your project. Claude Code loads it when a task matches its description.

How do I install Bio Substructure Search in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-substructure-search -a codex`. Or copy the skill folder (chemoinformatics/substructure-search in GPTomics/bioSkills) into .agents/skills/bio-substructure-search in your project. Codex loads it when a task matches its description.

Can I use Bio Substructure Search in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-substructure-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-substructure-search, .gemini/skills/bio-substructure-search, .github/skills/bio-substructure-search and .opencode/skills/bio-substructure-search in your project.

What does Bio Substructure Search need to run?

Going by SKILL.md and its folder, Bio Substructure Search needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Substructure Search access the network?

SKILL.md names 2 domains. As links in the text: doi.org and daylight.com. This is read from the text; nothing was executed.

Is Bio Substructure Search safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Substructure Search use?

Bio Substructure Search is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Substructure Search use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Substructure Search?

Skills that share tags, products or a category with Bio Substructure Search: DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Edu Chem Reaction (wy51ai/edulab, 1.4k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars) and RDKit Cheminformatics Practices (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Substructure Search?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.