Agent skill

Bio Molecular Standardization

by GPTomics in GPTomics/bioSkills

Standardizes molecular structures using the ChEMBL structure pipeline for normalization and parent selection plus RDKit rdMolStandardize for explicit custom steps such as tautomer canonicalization…

MITAuto-check passedResearch & Science

Install Bio Molecular Standardization

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-molecular-standardization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-molecular-standardization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/chemoinformatics/molecular-standardization .claude/skills/bio-molecular-standardization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-molecular-standardization
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.5k tokens
SKILL.md length
1,628 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Standardizes molecular structures using the ChEMBL structure pipeline for normalization and parent selection plus RDKit rdMolStandardize for explicit custom steps such as tautomer canonicalization…

  • Preparing libraries for QSAR training
  • SKILL.md covers Version Compatibility, Standardization Pipeline Stages, Pipeline Reconciliation and ChEMBL Structure Pipeline…, plus 9 more sections
  • Runs Python scripts from its folder; calls pip
  • Joining datasets across sources

What it does

Bio Molecular Standardization is an agent skill from GPTomics/bioSkills. Standardizes molecular structures using the ChEMBL structure pipeline for normalization and parent selection plus RDKit rdMolStandardize for explicit custom steps such as tautomer canonicalization, salt/solvent stripping, charge handling, stereochemistry handling, mixture selection, and isotope normalization. Explicitly compares ChEMBL, canSARchem, RDKit, and PubChem standardization choices. Use when preparing libraries for QSAR training, joining datasets across sources, deduplicating compound collections, or…

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/standardize_library.py` and `usage-guide.md`).

It sits in Research & Science, covering Drug discovery and cheminformatics and Database schema design. It works with RDKit. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Preparing libraries for QSAR training
  • Joining datasets across sources
  • Deduplicating compound collections
  • Building canonical compound registries

Example prompts

  • “Use the bio-molecular-standardization skill to standardiz molecular structures using the ChEMBL structure pipeline for normalization and parent…”
  • “/bio-molecular-standardization”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org
    • rdkit.org
    • openbabel.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Molecular Standardization loads about 4.5k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 1,628 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,628 words, ~4,537 tokens.

Download SKILL.mdSave it as .claude/skills/bio-molecular-standardization/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-molecular-standardization
description
Standardizes molecular structures using the ChEMBL structure pipeline for normalization and parent selection plus RDKit rdMolStandardize for explicit custom steps such as tautomer canonicalization, salt/solvent stripping, charge handling, stereochemistry handling, mixture selection, and isotope normalization. Explicitly compares ChEMBL, canSARchem, RDKit, and PubChem standardization choices. Use when preparing libraries for QSAR training, joining datasets across sources, deduplicating compound collections, or building canonical compound registries.
tool_type
python
primary_tool
RDKit

Version Compatibility

Reference examples tested with: RDKit 2024.09+ and chembl_structure_pipeline 1.2+. MolVS 0.1.1 is a legacy package; use RDKit's maintained rdMolStandardize module for custom pipelines.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Molecular Standardization

Convert raw molecular structures into a consistent form for ML training data, deduplication, registry, and cross-database joining. Skipping standardization can create data leakage when alternate representations of one compound enter different splits, distort QSAR inputs, and cause database join misses. The ChEMBL structure pipeline (Bento et al. 2020) is built on RDKit and applies ChEMBL-specific normalization and parent-selection rules. canSARchem (Dolciami et al. 2022) adds canonical-tautomer selection before parent extraction. RDKit's maintained rdMolStandardize module provides primitives for building an explicit custom pipeline.

For format-level I/O and aromaticity perception, see chemoinformatics/molecular-io. For descriptor calculation after standardization, see chemoinformatics/molecular-descriptors.

Standardization Pipeline Stages

StageRDKit ToolOperationCommon errors caught
1. SanitizationChem.SanitizeMolKekulize, assign aromaticity, fix valencesWrong valence on N/O
2. Salt strippingrdMolStandardize.FragmentRemover or LargestFragmentChooserRemove counterionsCl-, Na+, K+, OH-
3. Mixture choiceLargestFragmentChooserPick parent fragmentCo-crystals, hydrates
4. Charge neutralizationUnchargerNeutralize while preserving net chargePermanent charges preserved (quaternary N+)
5. Tautomer canonicalizationTautomerEnumerator.CanonicalizePick canonical tautomerKeto/enol; amide/imidate
6. Stereo standardizationChem.AssignStereochemistryConsistent stereo descriptorsLost wedges, ambiguous R/S
7. Isotope normalizationExplicitly set selected atom isotope labels to 0Remove 13C, 2H labelsTracer studies; preserve labels when scientifically meaningful
8. Output canonicalizationChem.MolToSmiles(canonical=True)Canonical SMILES + InChIKeyRound-trip stability

Pipeline Reconciliation

PipelineOriginTautomer canonicalizationSalt definitionUse case
ChEMBL pipelineEBI ChEMBLNot performed by standardize_mol or get_parent_molChEMBL salt list (extensive)ChEMBL-compatible registration
canSARchemICR Cancer Research UKCanonical tautomer BEFORE parent extractionExtended salt listCancer drug discovery
PubChem (OpenEye)NIH NCBIOpenEye QUACPAC tautomerPubChem salt listBioassay data, large-scale
RDKit rdMolStandardize defaultGreg LandrumRDKit TautomerEnumeratorRDKit defaultGeneral purpose, open source

Key difference (canSARchem vs ChEMBL):

  • ChEMBL standardizes the representation and extracts a parent, but does not canonicalize tautomers.
  • canSARchem canonicalizes the tautomer before parent extraction.

This difference matters when alternate tautomeric inputs must be registered as one parent. Do not describe ChEMBL output as tautomer-canonical unless an explicit tautomer step is added and documented.

ChEMBL Structure Pipeline (Reference Implementation)

ChEMBL's standardization is the most widely-used reference. The Python package chembl_structure_pipeline exposes the validated pipeline.

Goal: Apply the industry-reference ChEMBL standardization pipeline to a SMILES.

Approach: Parse SMILES with RDKit, run standardize_mol (sanitize, normalize, and standardize charges), then get_parent_mol (strip salts/counter-ions), and emit canonical SMILES. Add rdMolStandardize.TautomerEnumerator separately only when the project requires tautomer canonicalization.

python
from chembl_structure_pipeline import standardize_mol, get_parent_mol
from rdkit import Chem

def chembl_pipeline(smi):
    mol = Chem.MolFromSmiles(smi)
    if mol is None:
        return None, 'parse_failure'
    standardized = standardize_mol(mol)
    parent, exclude = get_parent_mol(standardized)
    if exclude:
        return None, 'excluded_by_chembl'
    return Chem.MolToSmiles(parent), 'ok'

standardize_mol: sanitize, normalize functional groups, and standardize charges; returns one RDKit molecule.

get_parent_mol: strip salts/counter-ions and choose the parent; returns (parent_mol, exclude_flag).

Output: canonical SMILES of the selected parent after the ChEMBL transformations, or an explicit excluded_by_chembl status when the parent carries ChEMBL's exclusion flag. Neutralizable acid/base sites may be normalized, but permanent or otherwise non-removable charges can remain; do not assume every emitted parent is neutral.

Full Standardization with rdMolStandardize

For more granular control or non-ChEMBL workflows.

Goal: Execute each standardization step explicitly to control salt stripping, charge handling, tautomer canonicalization, and isotope normalization.

Approach: Run the 8-stage pipeline (sanitize, largest fragment, normalize, uncharge, tautomer canonicalize, isotope strip, stereo standardize, canonical SMILES) sequentially with rdMolStandardize primitives.

python
from rdkit import Chem
from rdkit.Chem.MolStandardize import rdMolStandardize

def full_standardize(smi, keep_isotopes=False):
    mol = Chem.MolFromSmiles(smi)
    if mol is None:
        return None

    Chem.SanitizeMol(mol)

    largest = rdMolStandardize.LargestFragmentChooser(preferOrganic=True)
    mol = largest.choose(mol)

    normalizer = rdMolStandardize.Normalizer()
    mol = normalizer.normalize(mol)

    uncharger = rdMolStandardize.Uncharger(canonicalOrder=True)
    mol = uncharger.uncharge(mol)

    enumerator = rdMolStandardize.TautomerEnumerator()
    mol = enumerator.Canonicalize(mol)

    if not keep_isotopes:
        for atom in mol.GetAtoms():
            atom.SetIsotope(0)

    Chem.AssignStereochemistry(mol, cleanIt=True, force=True)
    return Chem.MolToSmiles(mol)

canonicalOrder=True makes the uncharger choose neutralization sites in canonical order when more than one equivalent site is available. It does not itself decide whether a permanent charge is retained; inspect charge-sensitive structures and keep force=False unless a documented policy requires otherwise.

Salt Stripping Edge Cases

Salt formActionExample
Mono-saltStrip counter-ion[Na+].CC(=O)[O-] -> CC(=O)O
Di-saltStrip both[Na+].[Na+].CC(=O)[O-].CC(=O)[O-] -> CC(=O)O
Mixed saltLargest organic fragmentCCO.CC(=O)O -> CCO (or CC(=O)O depending on rule)
Co-crystalHardest caseCC(=O)O.CCOC(C)=O -- both organic; default returns largest
HydrateStrip watersCC(=O)O.O -> CC(=O)O
SolvateStrip solventsCC(=O)O.CO -> CC(=O)O
Quaternary ammoniumPreserve charge[N+](C)(C)(C)C (permanent charge; do NOT neutralize)

LargestFragmentChooser(preferOrganic=True) prefers organic fragments over inorganic counter-ions even if smaller; for co-crystals, default rule picks largest organic fragment.

Tautomer Canonicalization (debated)

Tautomer canonicalization is the most controversial standardization step. There is no universally-correct canonical tautomer for many drug-like molecules.

Tautomer pairWhy the policy matters
Keto/enolCanonicalization can select a representation different from the experimentally relevant bound or solution form
Lactam/lactimHeterocycle scoring rules and toolkit versions may choose different representatives
Amidine/iminolProton placement changes donor/acceptor annotations and downstream matching
Phenol/keto (e.g., naphthol/naphthalenone)Aromaticity and functional-group perception can change with the selected representation
2H-pyrazole / 1H-pyrazoleNitrogen identity and donor/acceptor assignments depend on proton placement

Treat the enumerator's canonical result as a reproducible representation chosen by its configured scoring rules, not as a prediction of the dominant tautomer in vivo. Record the RDKit version and any custom transforms or scoring changes.

Practical rules:

  • Always apply consistent canonicalization across train + test for ML
  • For prospective prediction, predict for both tautomers if disagreement could matter
  • For library deduplication, canonical tautomer is the standard answer
  • For docking, use an ionization-aware preparation workflow. For Open Babel, the documented CLI is obabel input.sdf -O output.sdf -p 7.4; validate generated states because its rule-based protonation is not a substitute for project-specific pKa analysis.
python
from rdkit.Chem.MolStandardize import rdMolStandardize

def canonical_tautomer(smi):
    mol = Chem.MolFromSmiles(smi)
    enumerator = rdMolStandardize.TautomerEnumerator()
    canon = enumerator.Canonicalize(mol)
    return Chem.MolToSmiles(canon)

Stereochemistry Standardization

python
from rdkit import Chem

def standardize_stereo(mol, remove_undefined=False):
    Chem.AssignStereochemistry(mol, cleanIt=True, force=True)
    if remove_undefined:
        Chem.RemoveStereochemistry(mol)
    return mol

Cases:

  • Explicit stereo with @ / \ / / -> preserved
  • Wedge bonds in SDF -> re-perceived from 3D coords if present
  • Ambiguous stereo (no markers) -> left as-is, marked as undefined
  • Racemic (explicit "rac") -> keep as racemate

For ML, remove stereochemistry only when the endpoint, data curation, and model representation justify treating stereoisomers as equivalent; record that policy and test its effect. For docking and FEP, preserve the intended stereoisomer and reject unintended stereo changes.

Show full SKILL.md (678 more words)Show less

Standardization for ML Training (avoiding data leakage)

Goal: Build a standardized + deduplicated training set with replicate-averaged activity for QSAR or ADMET model training.

Approach: Standardize every SMILES through the ChEMBL pipeline, compute InChIKey as canonical identity, group by InChIKey, and mean-aggregate activities; report replicate count for confidence weighting.

python
import pandas as pd
from chembl_structure_pipeline import standardize_mol, get_parent_mol

def prepare_qsar_data(df, smiles_col='smiles', activity_col='pIC50'):
    standardized = []
    for i, row in df.iterrows():
        mol = Chem.MolFromSmiles(row[smiles_col])
        if mol is None:
            continue
        try:
            mol = standardize_mol(mol)
            mol, exclude = get_parent_mol(mol)
            if exclude:
                continue
            standardized.append({
                'smiles': Chem.MolToSmiles(mol),
                'inchikey': Chem.MolToInchiKey(mol),
                'activity': row[activity_col],
            })
        except Exception:
            continue

    df_std = pd.DataFrame(standardized)
    if df_std.empty:
        return pd.DataFrame(columns=['inchikey', 'smiles', 'activity', 'n_replicates'])
    df_std = df_std.groupby('inchikey').agg(
        smiles=('smiles', 'first'),
        activity=('activity', 'mean'),
        n_replicates=('activity', 'count'),
    ).reset_index()
    return df_std

Standard InChIKey may collapse some mobile-hydrogen tautomer representations, but this is not a substitute for an explicitly chosen tautomer policy. Replicate count signals measurement reliability.

Per-Tool Failure Modes

ChEMBL pipeline -- inorganic salt fails

Trigger: Molecule is genuinely an inorganic salt (e.g., NaCl, K2SO4).

Mechanism: get_parent_mol chooses largest organic; falls back to largest fragment for fully inorganic.

Symptom: Returns the salt itself (not a drug).

Fix: Pre-filter to compounds with ≥1 carbon atom.

Uncharger -- charge-state policy mismatch

Trigger: A molecule combines a non-removable charge, such as quaternary ammonium, with other neutralizable sites, or the desired physiological ionization state differs from a structure-normalization rule.

Mechanism: Uncharger adds or removes hydrogens from neutralizable acids and bases. It cannot remove a permanent charge that has no corresponding hydrogen edit; by default it may preserve an opposite neutralizable charge when a non-removable charge is present so that the total charge remains balanced. force=True instead neutralizes all sites that can be neutralized even if the remaining permanent charge leaves a nonzero total charge.

Symptom: The permanent charge remains, but other sites or the total charge differ from the protonation state intended for docking or modeling.

Fix: Choose force according to the documented total-charge policy, keep force=False when balanced countercharges should be preserved, and inspect/prepare physiological protonation states separately.

Tautomer enumerator -- combinatorial explosion

Trigger: Molecule with many tautomerizable groups (polyhydroxylated heterocycle).

Mechanism: TautomerEnumerator.Enumerate generates all possible tautomers; can produce thousands.

Symptom: OOM or hour-long compute on single molecule.

Fix: Use Canonicalize when only the configured canonical representation is needed. Before Enumerate, call enumerator.SetMaxTransforms(limit) (and, when appropriate, SetMaxTautomers(limit)) to cap the search.

Legacy MolVS -- import or compatibility failure

Trigger: Code still using legacy from molvs import Standardizer.

Mechanism: The standalone MolVS package is legacy and may not support current Python/RDKit versions. RDKit's maintained rdMolStandardize module remains available.

Symptom: ImportError or AttributeError on newer RDKit.

Fix: Migrate deliberately to from rdkit.Chem.MolStandardize import rdMolStandardize; compare outputs because RDKit functions are not drop-in aliases for every MolVS workflow.

Round-trip InChIKey mismatch

Trigger: Records were processed with different standardization settings or entered in different salt, charge, isotope, stereo, or tautomer forms.

Mechanism: The pipelines did not apply the same explicitly versioned transformations before identity generation.

Symptom: Apparently equivalent records produce different InChIKeys, or an expected database join fails.

Fix: Record and apply the same toolkit version, standardization stages, tautomer policy, and InChI options to both datasets; compare full standardized structures when results still differ.

Common Errors

SymptomCauseFix
ImportError from standalone molvsLegacy package incompatible with current environmentUse maintained rdkit.Chem.MolStandardize.rdMolStandardize APIs and validate output
standardize_mol raises or input parsing returns NoneInvalid or unsanitizable inputCapture the exception/input index and inspect sanitization deliberately; do not silently accept a partially sanitized structure
Stripped wrong fragmentLargestFragmentChooser ambiguityManually inspect; consider custom logic
Tautomer differs between datasetsDifferent tautomer rules or toolkit versionsPin and record the same TautomerEnumerator settings and version
Unexpected charge distribution with permanent ionsUncharger total-charge policy does not match the intended protonation workflowReview non-removable and neutralizable sites; choose force deliberately and prepare physiological states separately
Same InChIKey for apparently different recordsStandard-InChI normalization or a rare hash collisionCompare full InChI and standardized structures; InChIKey has no longer form
Pipeline slow on large libraryPer-molecule Python overheadProcess independent molecules in validated chunks or worker processes; chembl_structure_pipeline itself is a per-molecule API

References

  • chemoinformatics/molecular-io - Parse molecules before standardizing
  • chemoinformatics/molecular-descriptors - Apply descriptors to standardized molecules
  • chemoinformatics/similarity-searching - Standardize before comparing
  • chemoinformatics/substructure-search - Standardize before SMARTS matching
  • chemoinformatics/qsar-modeling - Mandatory upstream for QSAR

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in chemoinformatics/molecular-standardization of GPTomics/bioSkills.

  • SKILL.md
  • examples/standardize_library.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Molecular Standardization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Molecular Standardization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Molecular Standardization this skillGPTomics/bioSkills1.2k1 repos~4.5kAutomated safety check: PassMIT
Reactions StandardizationVectorSpaceLab/AREX-Skill331—~1.2kAutomated safety check: PassBSD-3-Clause
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT
Edu Chem Reactionwy51ai/edulab1.4k—~1.2kAutomated safety check: PassApache-2.0
RDKit Cheminformatics Practicesaiming-lab/AutoResearchClaw15k—~708Automated safety check: PassMIT

Similar skills

  • Reactions Standardization

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for RDKit reaction SMARTS/RXN workflows, product sanitization, MolStandardize cleanup/normalization/fragment/tautomer handling, R-group decomposition, stereochemistry/CIP…

    331 GitHub stars~1.2k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 10 days ago
    Research & ScienceAuto-check passed
  • Edu Chem Reaction

    wy51ai/edulab

    把一个化学反应做成自包含的微观 3D 交互演示网页:左/上为 Three.js 可交互分子动画 (拖滑块看断键·成键·原子重组,分步高亮),右为 KaTeX 反应方程 + 分步讲解 + 原子守恒计数 + 可选能量-反应进程曲线。支持三入口——给定文字反应/方程、随机出题、上传图片识别后演示。

    1.4k GitHub stars~1.2k tokensUpdated today
    Research & ScienceAuto-check passed
  • RDKit Cheminformatics Practices

    aiming-lab/AutoResearchClaw

    Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.

    15k GitHub stars~708 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Rowan

    lamm-mit/scienceclaw

    Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.

    246 GitHub starsUsed in 4 repos~3.1k tokens
    Research & ScienceAuto-check: warnings

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Molecular Standardization

What does Bio Molecular Standardization do?

Standardizes molecular structures using the ChEMBL structure pipeline for normalization and parent selection plus RDKit rdMolStandardize for explicit custom steps such as tautomer canonicalization…. Bio Molecular Standardization is an agent skill from GPTomics/bioSkills. Standardizes molecular structures using the ChEMBL structure pipeline for normalization and parent selection plus RDKit rdMolStandardize for explicit custom steps such as tautomer canonicalization, salt/solvent stripping, charge handling, stereochemistry handling, mixture selection, and isotope normalization.

When should I use Bio Molecular Standardization?

Bio Molecular Standardization fits situations like: preparing libraries for QSAR training; joining datasets across sources; deduplicating compound collections; building canonical compound registries.

How do I install Bio Molecular Standardization in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-molecular-standardization -a claude-code`. Or copy the skill folder (chemoinformatics/molecular-standardization in GPTomics/bioSkills) into .claude/skills/bio-molecular-standardization in your project. Claude Code loads it when a task matches its description.

How do I install Bio Molecular Standardization in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-molecular-standardization -a codex`. Or copy the skill folder (chemoinformatics/molecular-standardization in GPTomics/bioSkills) into .agents/skills/bio-molecular-standardization in your project. Codex loads it when a task matches its description.

Can I use Bio Molecular Standardization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-molecular-standardization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-molecular-standardization, .gemini/skills/bio-molecular-standardization, .github/skills/bio-molecular-standardization and .opencode/skills/bio-molecular-standardization in your project.

What does Bio Molecular Standardization need to run?

Going by SKILL.md and its folder, Bio Molecular Standardization needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Molecular Standardization access the network?

SKILL.md names 3 domains. As links in the text: doi.org, rdkit.org and openbabel.org. This is read from the text; nothing was executed.

Is Bio Molecular Standardization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Molecular Standardization use?

Bio Molecular Standardization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Molecular Standardization use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Molecular Standardization?

Skills that share tags, products or a category with Bio Molecular Standardization: Reactions Standardization (VectorSpaceLab/AREX-Skill, 331 stars), DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars), Biopipelines (locbp-uzh/biopipelines, 109 stars) and Edu Chem Reaction (wy51ai/edulab, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Molecular Standardization?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.