Hypothesis Generation
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures…
$ npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-similarity-searching --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/chemoinformatics/similarity-searching .claude/skills/bio-similarity-searching && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-similarity-searching" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/similarity-searching into .claude/skills/bio-similarity-searching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-similarity-searching", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/similarity-searchingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-similarity-searching --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/chemoinformatics/similarity-searching .agents/skills/bio-similarity-searching && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-similarity-searching" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/similarity-searching into .agents/skills/bio-similarity-searching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-similarity-searching", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-similarity-searching --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/chemoinformatics/similarity-searching .cursor/skills/bio-similarity-searching && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-similarity-searching" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/similarity-searching into .cursor/skills/bio-similarity-searching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-similarity-searching", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path chemoinformatics/similarity-searching--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-similarity-searching --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/chemoinformatics/similarity-searching .gemini/skills/bio-similarity-searching && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-similarity-searching" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/similarity-searching into .gemini/skills/bio-similarity-searching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-similarity-searching", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-similarity-searchingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/chemoinformatics/similarity-searching .github/skills/bio-similarity-searching && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-similarity-searching" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/similarity-searching into .github/skills/bio-similarity-searching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-similarity-searching", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-similarity-searching --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/chemoinformatics/similarity-searching .opencode/skills/bio-similarity-searching && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-similarity-searching" agent skill from https://github.com/GPTomics/bioSkills/tree/main/chemoinformatics/similarity-searching into .opencode/skills/bio-similarity-searching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-similarity-searching", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-similarity-searchingPerforms molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures…
Bio Similarity Searching is an agent skill from GPTomics/bioSkills. Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures, scaffold-hopping vs lead-optimization regimes, activity-cliff diagnosis, and large-library nearest-neighbor methods (BulkTanimoto, MHFP6 LSH forest, USRCAT). Use when ranking compounds by structural resemblance to a query, clustering libraries, finding analogs, or diagnosing activity cliffs.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/similarity_search.py` and `usage-guide.md`).
It sits in Research & Science. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
doi.orgrdkit.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Similarity Searching loads about 4.6k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,623 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,623 words, ~4,582 tokens.
.claude/skills/bio-similarity-searching/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: RDKit 2024.09+, scikit-learn 1.4+, mhfp 1.9+.
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturesIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Find structurally similar compounds and cluster libraries by similarity. The choice of similarity coefficient and fingerprint is task-aware: Tanimoto for symmetric similarity in lead optimization, Tversky for asymmetric "substructure-like" queries, Dice for higher sensitivity in low-similarity regimes, and MaxCommon Substructure (MCS) for scaffold-hopping. Tanimoto similarity above 0.7 is not a guarantee of activity preservation; activity cliffs (similar molecules with dissimilar activities) are common (Maggiora 2014).
For fingerprint choice, see chemoinformatics/molecular-descriptors. For 3D shape similarity, see chemoinformatics/shape-similarity.
| Coefficient | Formula | Range | Symmetric | Use case | Fails when |
|---|---|---|---|---|---|
| Tanimoto | c / (a + b - c) | 0-1 | Yes | Default for ECFP4 similarity, ranking analogs | Saturates at low similarity (drug vs natural product) |
| Dice | 2c / (a + b) | 0-1 | Yes | Bit or nonnegative sparse-count vectors when Dice semantics are intended | Thresholds depend on vector type; analog choice subjective |
| Cosine (Ochiai) | c / sqrt(a*b) | 0-1 | Yes | Count vectors, weighted similarity | Not standard for bit vectors |
| Tversky alpha,beta | c / (alpha*(a-c) + beta*(b-c) + c) | 0-1 | No when alpha != beta | Asymmetric "is A a substructure of B" queries | Parameter choice subjective; alpha=1,beta=0 = substructure-like |
| Hamming | (a + b - 2c) / nBits | 0-1 | Yes | Binary bit vectors when bit-wise disagreement matters | Does not preserve count magnitude |
| Russell-Rao | c / nBits | 0-1 | Yes | Sparse fingerprints | Biased by fingerprint density |
| Kulczynski | (c/a + c/b) / 2 | 0-1 | Yes | When fingerprints have very different bit-density | Less standard |
Where a = set bits in fp1, b = set bits in fp2, c = bits in common.
| Scenario | Coefficient | Why |
|---|---|---|
| Standard analog search (drug-like, ECFP4) | Tanimoto, start near 0.7 | Repository starting heuristic; calibrate against project actives and analog judgments |
| Sensitive search at lower similarity | Dice, threshold 0.45 | Dice is roughly 2*Tanimoto/(1+Tanimoto); more sensitive in middle range |
| Substructure-like ranking | Tversky alpha=1, beta=0 | Asymmetric: rewards compounds containing query features |
| Count fingerprints (neural, atom-environment) | Cosine | Bit-vector Tanimoto loses information |
| Activity-cliff diagnosis | Tanimoto + property difference | Detect ECFP4>=0.85 but |
| Cross-target / scaffold-hopping | FCFP4 Tanimoto OR AtomPair Tanimoto | Pharmacophore-equivalent matches different scaffolds |
| Metabolomics / natural products | MHFP6 Jaccard | ECFP4 saturates near 0.2 across diverse classes |
| 3D shape | Tanimoto on shape volume overlap | See shape-similarity skill |
| Threshold | Interpretation | Caveat |
|---|---|---|
| >=0.85 | Likely same scaffold + close analog | Activity cliffs still possible |
| 0.70-0.85 | Same series, R-group variation | Standard "similar" threshold |
| 0.55-0.70 | Related chemotype, different decoration | Useful for series expansion |
| 0.35-0.55 | Distant analog, possible scaffold hop | Many false positives |
| <0.35 | Mostly noise; use 3D shape or pharmacophore instead | ECFP4 not informative |
These bands are working defaults for ECFP4-like fingerprints, not transferable calibration. Inspect the target dataset's similarity distribution and known series before setting a cutoff. Maggiora's similarity principle states "similar molecules tend to have similar activity" -- but activity cliffs (Stumpfe & Bajorath 2012) violate this. Treat high ECFP4 similarity as a prioritization signal, not evidence that activity will be preserved.
| Goal | Workflow | Tools |
|---|---|---|
| Find analogs of a hit (lead opt) | ECFP4 Tanimoto >=0.7 search | RDKit BulkTanimotoSimilarity |
| Find scaffold hops | FCFP4 OR AtomPair Tanimoto >=0.5 + filter MCS | RDKit + rdFMCS |
| Cluster library by chemotype | Butina clustering at Tanimoto 0.6 cutoff | RDKit Butina.ClusterData |
| Diversity sampling | MaxMin selection on Tanimoto | RDKit rdSimDivPickers.MaxMinPicker |
| Nearest neighbors in >1M library | LSH forest with MHFP6 | mhfp.lsh_forest.LSHForestHelper |
| Activity cliff diagnosis | Tanimoto + pIC50 delta scatter | Custom analysis |
| 3D similarity (shape) | USRCAT / Open3DAlign / ROCS | shape-similarity skill |
Goal: Rank a library by ECFP4 Tanimoto similarity to a query molecule, returning hits above a threshold.
Approach: Generate ECFP4 fingerprints for all molecules once, then use BulkTanimotoSimilarity for O(N) lookup.
from rdkit import Chem, DataStructs
from rdkit.Chem import rdFingerprintGenerator
def precompute_fps(smiles_list, radius=2, nBits=2048):
generator = rdFingerprintGenerator.GetMorganGenerator(
radius=radius, fpSize=nBits)
fps = []
for smi in smiles_list:
mol = Chem.MolFromSmiles(smi)
if mol is None:
fps.append(None)
else:
fps.append(generator.GetFingerprint(mol))
return fps
def search(query_smi, library_fps, threshold=0.7):
qmol = Chem.MolFromSmiles(query_smi)
if qmol is None:
raise ValueError('invalid query SMILES')
generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
qfp = generator.GetFingerprint(qmol)
valid = [(i, fp) for i, fp in enumerate(library_fps) if fp is not None]
sims = DataStructs.BulkTanimotoSimilarity(qfp, [fp for _, fp in valid])
return [(source_i, sim) for (source_i, _), sim in zip(valid, sims)
if sim >= threshold]Goal: Rank a library by how much each compound "contains" the features of the query (asymmetric).
Approach: Tversky with alpha=1, beta=0 rewards compounds containing query bits (substructure-like) while ignoring extra bits in the compound.
from rdkit import DataStructs
def tversky_substructure_like(qfp, lib_fps, alpha=1.0, beta=0.0):
return [DataStructs.TverskySimilarity(qfp, f, alpha, beta) for f in lib_fps if f]Use case: identifying analogs that extend a pharmacophore vs. exact-similarity ranking.
Goal: Group a library around Taylor-Butina centroids whose assigned neighbors are within the selected distance cutoff.
Approach: Compute upper-triangle distance matrix, apply Taylor-Butina with chosen distance cutoff.
from rdkit.ML.Cluster import Butina
def cluster(mols, cutoff=0.4):
generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
fps = [generator.GetFingerprint(m) for m in mols]
n = len(fps)
dists = []
for i in range(1, n):
sims = DataStructs.BulkTanimotoSimilarity(fps[i], fps[:i])
dists.extend([1 - s for s in sims])
return Butina.ClusterData(dists, n, cutoff, isDistData=True)cutoff=0.4 means each assigned member was a neighbor of its selected centroid at Tanimoto >= 0.6. It does not guarantee that every pair of non-centroid members has Tanimoto >= 0.6. The first molecule in each returned cluster is the cluster centroid.
Trade-off: Butina materializes O(N^2) pairwise distances. Benchmark memory and runtime on the actual library; for much larger collections, use an approximate method such as an MHFP6 LSH forest.
Goal: Select N diverse compounds from a library by maximizing the minimum pairwise distance.
from rdkit.SimDivFilters import rdSimDivPickers
picker = rdSimDivPickers.MaxMinPicker()
n_pick = 100
n_lib = len(fps)
selected = picker.LazyBitVectorPick(fps, n_lib, n_pick, seed=42)LazyBitVectorPick is memory-efficient (does not materialize full distance matrix).
Goal: Find the largest substructure shared across a set of molecules.
Approach: rdFMCS.FindMCS with parameters controlling atom/bond equivalence.
from rdkit.Chem import rdFMCS
def mcs_smarts(mols, timeout=60, ring_match='strict', atom_match='elements'):
params = rdFMCS.MCSParameters()
params.Timeout = timeout
if ring_match == 'strict':
params.BondCompareParameters.MatchFusedRings = True
params.BondCompareParameters.MatchFusedRingsStrict = True
params.BondCompareParameters.RingMatchesRingOnly = True
if atom_match == 'elements':
params.AtomCompareParameters.MatchValences = False
result = rdFMCS.FindMCS(mols, params)
return result.smartsString, result.numAtoms, result.numBondsUse cases: identify scaffold across a series, build scaffold hopping queries, generate consensus pharmacophore.
Limit: MCS search can become combinatorial as input count, size, and structural divergence grow. Set a finite timeout, inspect result.canceled, and consider pre-clustering or reducing the comparison set; no molecule-count or atom-count boundary guarantees tractability.
Goal: Detect pairs of similar molecules with dissimilar activities (cliffs).
Approach: Compute pairwise ECFP4 Tanimoto + pIC50 delta. Flag pairs with high similarity and large activity gap.
def activity_cliffs(df, sim_threshold=0.85, activity_gap=2.0, activity_col='pIC50'):
generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
mols = [Chem.MolFromSmiles(s) for s in df['smiles']]
if any(mol is None for mol in mols):
raise ValueError('activity-cliff input contains invalid SMILES')
fps = [generator.GetFingerprint(mol) for mol in mols]
cliffs = []
for i in range(len(fps)):
sims = DataStructs.BulkTanimotoSimilarity(fps[i], fps[i+1:])
for j_off, sim in enumerate(sims):
j = i + 1 + j_off
if sim >= sim_threshold:
gap = abs(df[activity_col].iloc[i] - df[activity_col].iloc[j])
if gap >= activity_gap:
cliffs.append((i, j, sim, gap))
return cliffsActivity cliffs flag (a) measurement noise, (b) cryptic SAR (e.g. ring-flip changing dihedral), (c) protein conformational selection, or (d) actually informative SAR. Cliffs are an opportunity for medchem investigation, not necessarily an error.
For large libraries, direct all-pairs comparison becomes expensive. The mhfp package provides an LSH-forest helper for approximate nearest-neighbor retrieval over MHFP6 fingerprints. Measure recall and latency against an exact subset for the project dataset.
from mhfp.encoder import MHFPEncoder
from mhfp.lsh_forest import LSHForestHelper
encoder = MHFPEncoder(2048)
def build_index(smiles_list):
forest = LSHForestHelper()
fingerprints = []
for i, smiles in enumerate(smiles_list):
fp = encoder.encode(smiles, radius=3)
fingerprints.append(fp)
forest.add(i, fp)
forest.index()
return forest, fingerprints
def query_index(forest, qmol, fingerprints, k=10):
qfp = encoder.encode_mol(qmol, radius=3)
return forest.query(qfp, k=k, data=fingerprints)The returned neighbors are approximate in MHFP6 space. Benchmark them against an exact MHFP-distance search on a representative subset before choosing LSH parameters.
Trigger: Library spans drug-like + natural products + peptides + metabolites.
Mechanism: A local circular fingerprint may not preserve the distinctions needed for a particular mixed-modality retrieval task.
Symptom: Known relevant neighbors are not enriched above background, or retrieval metrics and neighborhood stability are poor on target-relevant controls. A low mean pairwise similarity alone is not a universal saturation test.
Fix: Benchmark ECFP4 against alternatives such as MHFP6 or MAP4 using held-out analog recovery, scaffold-aware retrieval, or another task-aligned metric. Do not transfer raw-score thresholds between fingerprint families.
Trigger: Library >100k molecules.
Mechanism: Butina requires upper-triangle distance matrix, ~5e9 floats for 100k compounds.
Symptom: OOM error or hours of CPU.
Fix: Use approximate clustering (HDBSCAN on UMAP-reduced fingerprints) or LSH-based clustering on MHFP6.
Trigger: Mol set with low overlap, large molecules, or many input mols.
Mechanism: MCS search is NP-hard; algorithm tries every atom-mapping permutation within timeout.
Symptom: Returns small partial MCS or empty result.
Fix: Raise timeout; reduce input mol count; pre-cluster by Tanimoto first then MCS within clusters.
Trigger: Comparing fingerprints between two molecules that hash to the same bits.
Mechanism: A folded hashed fingerprint can map distinct atom environments to the same bits; collision frequency depends on molecule size, radius, and fingerprint length.
Symptom: Two structurally different molecules report Tanimoto 1.0.
Fix: For exact identity, compare canonical SMILES or InChIKey, not fingerprint. Use unhashed sparse fingerprint to disambiguate.
Trigger: Threshold tuned on ECFP4 applied to RDKit FP, AtomPair, or MACCS.
Mechanism: Bit-density and fragment-resolution differ; Tanimoto distributions shift.
Symptom: "Similar" set is much larger or smaller than expected.
Fix: Re-tune the threshold per fingerprint and dataset. AtomPair ~0.55, MACCS ~0.85, ECFP4 ~0.7, and FCFP4 ~0.6 are repository starting heuristics, not universal equivalents.
If a pair flags as an activity cliff under one representation but not another, treat that as representation sensitivity. Inspect atom mappings, fingerprint environments, assay uncertainty, and the exact structural change; disagreement alone does not establish which substituent caused the activity difference.
| Symptom | Cause | Fix |
|---|---|---|
BulkTanimotoSimilarity output is treated as bit counts | The API returns similarity values, normally floats | Keep the returned values as similarities; inspect input vector types if the output is unexpected |
| Reported similarity > 1 | Custom formula, malformed data, negative features, or an incorrectly normalized external implementation | Verify the coefficient definition and inputs; standard nonnegative RDKit Tanimoto and Tversky similarities are bounded by 1 |
| Cluster centroids change when input order changes | Taylor-Butina tie handling and assignment are order-sensitive | Standardize and sort inputs by a stable identifier before clustering; record reordering and the input order |
| MaxMinPicker returns first N inputs | All-zero initial similarity matrix | Seed picker explicitly: picker.LazyBitVectorPick(fps, n_lib, n_pick, seed=42) |
| Activity cliff "false positives" | Bit-collisions inflate similarity | Use sparse Morgan or compare canonical SMILES for exact ID |
| Diverse subset has duplicates | Standardization not applied | Canonicalize via chemoinformatics/molecular-standardization first |
| Tanimoto incompatible with neural fingerprint | Continuous-valued fingerprint | Use cosine or sklearn cdist with 'cosine' metric |
rdFMCS API documentation -- MCS ring-comparison parameter names and semantics. https://www.rdkit.org/docs/source/rdkit.Chem.rdFMCS.html© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in chemoinformatics/similarity-searching of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Similarity Searching next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Similarity Searching this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.6k | Automated safety check: Pass | MIT | |
| Hypothesis Generationspacering-net/codeg | 3.9k | 14 repos | ~3.6k | Automated safety check: Notes | MIT | |
| GitHub Deep Researchbytedance/deer-flow | 84k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Nature Paper CardYuan1z0825/nature-skills | 47k | 2 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Content Research Writerweapp-tailwindcss/weapp-tailwindcss | 1.9k | 25 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Last30daysmvanhorn/last30days-skill | 64k | — | ~7.9k | Automated safety check: Notes | MIT |
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
bytedance/deer-flow
Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.
Yuan1z0825/nature-skills
Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.
weapp-tailwindcss/weapp-tailwindcss
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.
mvanhorn/last30days-skill
Research what people actually say about any topic in the last 30 days.
spacering-net/codeg
Structured manuscript/grant review with checklist-based evaluation.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures…. Bio Similarity Searching is an agent skill from GPTomics/bioSkills. Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures, scaffold-hopping vs lead-optimization regimes, activity-cliff diagnosis, and large-library nearest-neighbor methods (BulkTanimoto, MHFP6 LSH forest, USRCAT).
Bio Similarity Searching fits situations like: ranking compounds by structural resemblance to a query; clustering libraries; finding analogs; diagnosing activity cliffs.
Run `npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a claude-code`. Or copy the skill folder (chemoinformatics/similarity-searching in GPTomics/bioSkills) into .claude/skills/bio-similarity-searching in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a codex`. Or copy the skill folder (chemoinformatics/similarity-searching in GPTomics/bioSkills) into .agents/skills/bio-similarity-searching in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-similarity-searching, .gemini/skills/bio-similarity-searching, .github/skills/bio-similarity-searching and .opencode/skills/bio-similarity-searching in your project.
Going by SKILL.md and its folder, Bio Similarity Searching needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: doi.org and rdkit.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Similarity Searching is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Similarity Searching: Hypothesis Generation (spacering-net/codeg, 3.9k stars), GitHub Deep Research (bytedance/deer-flow, 84k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars) and Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.