Agent skill

Bio Similarity Searching

by GPTomics in GPTomics/bioSkills

Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures…

MITAuto-check passedResearch & Science

Install Bio Similarity Searching

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-similarity-searching --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/chemoinformatics/similarity-searching .claude/skills/bio-similarity-searching && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-similarity-searching
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.6k tokens
SKILL.md length
1,623 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures…

  • Ranking compounds by structural resemblance to a query
  • SKILL.md covers Version Compatibility, Similarity Coefficient Taxonomy, When to Use Which Coefficient and Tanimoto Thresholds…, plus 13 more sections
  • Runs Python scripts from its folder; calls pip
  • Clustering libraries

What it does

Bio Similarity Searching is an agent skill from GPTomics/bioSkills. Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures, scaffold-hopping vs lead-optimization regimes, activity-cliff diagnosis, and large-library nearest-neighbor methods (BulkTanimoto, MHFP6 LSH forest, USRCAT). Use when ranking compounds by structural resemblance to a query, clustering libraries, finding analogs, or diagnosing activity cliffs.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/similarity_search.py` and `usage-guide.md`).

It sits in Research & Science. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Ranking compounds by structural resemblance to a query
  • Clustering libraries
  • Finding analogs
  • Diagnosing activity cliffs

Example prompts

  • “Use the bio-similarity-searching skill to perform molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count…”
  • “/bio-similarity-searching”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org
    • rdkit.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Similarity Searching loads about 4.6k tokens when it runs. Until then it costs about 126 tokens; SKILL.md has 1,623 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,623 words, ~4,582 tokens.

Download SKILL.mdSave it as .claude/skills/bio-similarity-searching/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-similarity-searching
description
Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures, scaffold-hopping vs lead-optimization regimes, activity-cliff diagnosis, and large-library nearest-neighbor methods (BulkTanimoto, MHFP6 LSH forest, USRCAT). Use when ranking compounds by structural resemblance to a query, clustering libraries, finding analogs, or diagnosing activity cliffs.
tool_type
python
primary_tool
RDKit

Version Compatibility

Reference examples tested with: RDKit 2024.09+, scikit-learn 1.4+, mhfp 1.9+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Similarity Searching

Find structurally similar compounds and cluster libraries by similarity. The choice of similarity coefficient and fingerprint is task-aware: Tanimoto for symmetric similarity in lead optimization, Tversky for asymmetric "substructure-like" queries, Dice for higher sensitivity in low-similarity regimes, and MaxCommon Substructure (MCS) for scaffold-hopping. Tanimoto similarity above 0.7 is not a guarantee of activity preservation; activity cliffs (similar molecules with dissimilar activities) are common (Maggiora 2014).

For fingerprint choice, see chemoinformatics/molecular-descriptors. For 3D shape similarity, see chemoinformatics/shape-similarity.

Similarity Coefficient Taxonomy

CoefficientFormulaRangeSymmetricUse caseFails when
Tanimotoc / (a + b - c)0-1YesDefault for ECFP4 similarity, ranking analogsSaturates at low similarity (drug vs natural product)
Dice2c / (a + b)0-1YesBit or nonnegative sparse-count vectors when Dice semantics are intendedThresholds depend on vector type; analog choice subjective
Cosine (Ochiai)c / sqrt(a*b)0-1YesCount vectors, weighted similarityNot standard for bit vectors
Tversky alpha,betac / (alpha*(a-c) + beta*(b-c) + c)0-1No when alpha != betaAsymmetric "is A a substructure of B" queriesParameter choice subjective; alpha=1,beta=0 = substructure-like
Hamming(a + b - 2c) / nBits0-1YesBinary bit vectors when bit-wise disagreement mattersDoes not preserve count magnitude
Russell-Raoc / nBits0-1YesSparse fingerprintsBiased by fingerprint density
Kulczynski(c/a + c/b) / 20-1YesWhen fingerprints have very different bit-densityLess standard

Where a = set bits in fp1, b = set bits in fp2, c = bits in common.

When to Use Which Coefficient

ScenarioCoefficientWhy
Standard analog search (drug-like, ECFP4)Tanimoto, start near 0.7Repository starting heuristic; calibrate against project actives and analog judgments
Sensitive search at lower similarityDice, threshold 0.45Dice is roughly 2*Tanimoto/(1+Tanimoto); more sensitive in middle range
Substructure-like rankingTversky alpha=1, beta=0Asymmetric: rewards compounds containing query features
Count fingerprints (neural, atom-environment)CosineBit-vector Tanimoto loses information
Activity-cliff diagnosisTanimoto + property differenceDetect ECFP4>=0.85 but
Cross-target / scaffold-hoppingFCFP4 Tanimoto OR AtomPair TanimotoPharmacophore-equivalent matches different scaffolds
Metabolomics / natural productsMHFP6 JaccardECFP4 saturates near 0.2 across diverse classes
3D shapeTanimoto on shape volume overlapSee shape-similarity skill

Tanimoto Thresholds (Repository Starting Heuristics)

ThresholdInterpretationCaveat
>=0.85Likely same scaffold + close analogActivity cliffs still possible
0.70-0.85Same series, R-group variationStandard "similar" threshold
0.55-0.70Related chemotype, different decorationUseful for series expansion
0.35-0.55Distant analog, possible scaffold hopMany false positives
<0.35Mostly noise; use 3D shape or pharmacophore insteadECFP4 not informative

These bands are working defaults for ECFP4-like fingerprints, not transferable calibration. Inspect the target dataset's similarity distribution and known series before setting a cutoff. Maggiora's similarity principle states "similar molecules tend to have similar activity" -- but activity cliffs (Stumpfe & Bajorath 2012) violate this. Treat high ECFP4 similarity as a prioritization signal, not evidence that activity will be preserved.

Decision Tree by Scenario

GoalWorkflowTools
Find analogs of a hit (lead opt)ECFP4 Tanimoto >=0.7 searchRDKit BulkTanimotoSimilarity
Find scaffold hopsFCFP4 OR AtomPair Tanimoto >=0.5 + filter MCSRDKit + rdFMCS
Cluster library by chemotypeButina clustering at Tanimoto 0.6 cutoffRDKit Butina.ClusterData
Diversity samplingMaxMin selection on TanimotoRDKit rdSimDivPickers.MaxMinPicker
Nearest neighbors in >1M libraryLSH forest with MHFP6mhfp.lsh_forest.LSHForestHelper
Activity cliff diagnosisTanimoto + pIC50 delta scatterCustom analysis
3D similarity (shape)USRCAT / Open3DAlign / ROCSshape-similarity skill

Tanimoto Similarity (single query, large library)

Goal: Rank a library by ECFP4 Tanimoto similarity to a query molecule, returning hits above a threshold.

Approach: Generate ECFP4 fingerprints for all molecules once, then use BulkTanimotoSimilarity for O(N) lookup.

python
from rdkit import Chem, DataStructs
from rdkit.Chem import rdFingerprintGenerator

def precompute_fps(smiles_list, radius=2, nBits=2048):
    generator = rdFingerprintGenerator.GetMorganGenerator(
        radius=radius, fpSize=nBits)
    fps = []
    for smi in smiles_list:
        mol = Chem.MolFromSmiles(smi)
        if mol is None:
            fps.append(None)
        else:
            fps.append(generator.GetFingerprint(mol))
    return fps

def search(query_smi, library_fps, threshold=0.7):
    qmol = Chem.MolFromSmiles(query_smi)
    if qmol is None:
        raise ValueError('invalid query SMILES')
    generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
    qfp = generator.GetFingerprint(qmol)
    valid = [(i, fp) for i, fp in enumerate(library_fps) if fp is not None]
    sims = DataStructs.BulkTanimotoSimilarity(qfp, [fp for _, fp in valid])
    return [(source_i, sim) for (source_i, _), sim in zip(valid, sims)
            if sim >= threshold]

Goal: Rank a library by how much each compound "contains" the features of the query (asymmetric).

Approach: Tversky with alpha=1, beta=0 rewards compounds containing query bits (substructure-like) while ignoring extra bits in the compound.

python
from rdkit import DataStructs

def tversky_substructure_like(qfp, lib_fps, alpha=1.0, beta=0.0):
    return [DataStructs.TverskySimilarity(qfp, f, alpha, beta) for f in lib_fps if f]

Use case: identifying analogs that extend a pharmacophore vs. exact-similarity ranking.

Butina Clustering

Goal: Group a library around Taylor-Butina centroids whose assigned neighbors are within the selected distance cutoff.

Approach: Compute upper-triangle distance matrix, apply Taylor-Butina with chosen distance cutoff.

python
from rdkit.ML.Cluster import Butina

def cluster(mols, cutoff=0.4):
    generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
    fps = [generator.GetFingerprint(m) for m in mols]
    n = len(fps)
    dists = []
    for i in range(1, n):
        sims = DataStructs.BulkTanimotoSimilarity(fps[i], fps[:i])
        dists.extend([1 - s for s in sims])
    return Butina.ClusterData(dists, n, cutoff, isDistData=True)

cutoff=0.4 means each assigned member was a neighbor of its selected centroid at Tanimoto >= 0.6. It does not guarantee that every pair of non-centroid members has Tanimoto >= 0.6. The first molecule in each returned cluster is the cluster centroid.

Trade-off: Butina materializes O(N^2) pairwise distances. Benchmark memory and runtime on the actual library; for much larger collections, use an approximate method such as an MHFP6 LSH forest.

Diversity Selection (MaxMin)

Goal: Select N diverse compounds from a library by maximizing the minimum pairwise distance.

python
from rdkit.SimDivFilters import rdSimDivPickers

picker = rdSimDivPickers.MaxMinPicker()
n_pick = 100
n_lib = len(fps)
selected = picker.LazyBitVectorPick(fps, n_lib, n_pick, seed=42)

LazyBitVectorPick is memory-efficient (does not materialize full distance matrix).

Maximum Common Substructure

Goal: Find the largest substructure shared across a set of molecules.

Approach: rdFMCS.FindMCS with parameters controlling atom/bond equivalence.

python
from rdkit.Chem import rdFMCS

def mcs_smarts(mols, timeout=60, ring_match='strict', atom_match='elements'):
    params = rdFMCS.MCSParameters()
    params.Timeout = timeout
    if ring_match == 'strict':
        params.BondCompareParameters.MatchFusedRings = True
        params.BondCompareParameters.MatchFusedRingsStrict = True
        params.BondCompareParameters.RingMatchesRingOnly = True
    if atom_match == 'elements':
        params.AtomCompareParameters.MatchValences = False
    result = rdFMCS.FindMCS(mols, params)
    return result.smartsString, result.numAtoms, result.numBonds

Use cases: identify scaffold across a series, build scaffold hopping queries, generate consensus pharmacophore.

Limit: MCS search can become combinatorial as input count, size, and structural divergence grow. Set a finite timeout, inspect result.canceled, and consider pre-clustering or reducing the comparison set; no molecule-count or atom-count boundary guarantees tractability.

Activity Cliff Diagnosis

Goal: Detect pairs of similar molecules with dissimilar activities (cliffs).

Approach: Compute pairwise ECFP4 Tanimoto + pIC50 delta. Flag pairs with high similarity and large activity gap.

python
def activity_cliffs(df, sim_threshold=0.85, activity_gap=2.0, activity_col='pIC50'):
    generator = rdFingerprintGenerator.GetMorganGenerator(radius=2, fpSize=2048)
    mols = [Chem.MolFromSmiles(s) for s in df['smiles']]
    if any(mol is None for mol in mols):
        raise ValueError('activity-cliff input contains invalid SMILES')
    fps = [generator.GetFingerprint(mol) for mol in mols]
    cliffs = []
    for i in range(len(fps)):
        sims = DataStructs.BulkTanimotoSimilarity(fps[i], fps[i+1:])
        for j_off, sim in enumerate(sims):
            j = i + 1 + j_off
            if sim >= sim_threshold:
                gap = abs(df[activity_col].iloc[i] - df[activity_col].iloc[j])
                if gap >= activity_gap:
                    cliffs.append((i, j, sim, gap))
    return cliffs

Activity cliffs flag (a) measurement noise, (b) cryptic SAR (e.g. ring-flip changing dihedral), (c) protein conformational selection, or (d) actually informative SAR. Cliffs are an opportunity for medchem investigation, not necessarily an error.

Large-Library Nearest Neighbor (MHFP6 + LSH Forest)

For large libraries, direct all-pairs comparison becomes expensive. The mhfp package provides an LSH-forest helper for approximate nearest-neighbor retrieval over MHFP6 fingerprints. Measure recall and latency against an exact subset for the project dataset.

python
from mhfp.encoder import MHFPEncoder
from mhfp.lsh_forest import LSHForestHelper

encoder = MHFPEncoder(2048)

def build_index(smiles_list):
    forest = LSHForestHelper()
    fingerprints = []
    for i, smiles in enumerate(smiles_list):
        fp = encoder.encode(smiles, radius=3)
        fingerprints.append(fp)
        forest.add(i, fp)
    forest.index()
    return forest, fingerprints

def query_index(forest, qmol, fingerprints, k=10):
    qfp = encoder.encode_mol(qmol, radius=3)
    return forest.query(qfp, k=k, data=fingerprints)

The returned neighbors are approximate in MHFP6 space. Benchmark them against an exact MHFP-distance search on a representative subset before choosing LSH parameters.

Show full SKILL.md (631 more words)Show less

Per-Tool Failure Modes

ECFP4 Tanimoto -- saturation in diverse libraries

Trigger: Library spans drug-like + natural products + peptides + metabolites.

Mechanism: A local circular fingerprint may not preserve the distinctions needed for a particular mixed-modality retrieval task.

Symptom: Known relevant neighbors are not enriched above background, or retrieval metrics and neighborhood stability are poor on target-relevant controls. A low mean pairwise similarity alone is not a universal saturation test.

Fix: Benchmark ECFP4 against alternatives such as MHFP6 or MAP4 using held-out analog recovery, scaffold-aware retrieval, or another task-aligned metric. Do not transfer raw-score thresholds between fingerprint families.

Butina clustering -- O(N^2) memory blowup

Trigger: Library >100k molecules.

Mechanism: Butina requires upper-triangle distance matrix, ~5e9 floats for 100k compounds.

Symptom: OOM error or hours of CPU.

Fix: Use approximate clustering (HDBSCAN on UMAP-reduced fingerprints) or LSH-based clustering on MHFP6.

MCS -- exponential timeout

Trigger: Mol set with low overlap, large molecules, or many input mols.

Mechanism: MCS search is NP-hard; algorithm tries every atom-mapping permutation within timeout.

Symptom: Returns small partial MCS or empty result.

Fix: Raise timeout; reduce input mol count; pre-cluster by Tanimoto first then MCS within clusters.

Tanimoto = 1.0 != same molecule

Trigger: Comparing fingerprints between two molecules that hash to the same bits.

Mechanism: A folded hashed fingerprint can map distinct atom environments to the same bits; collision frequency depends on molecule size, radius, and fingerprint length.

Symptom: Two structurally different molecules report Tanimoto 1.0.

Fix: For exact identity, compare canonical SMILES or InChIKey, not fingerprint. Use unhashed sparse fingerprint to disambiguate.

Similarity threshold transfer fails

Trigger: Threshold tuned on ECFP4 applied to RDKit FP, AtomPair, or MACCS.

Mechanism: Bit-density and fragment-resolution differ; Tanimoto distributions shift.

Symptom: "Similar" set is much larger or smaller than expected.

Fix: Re-tune the threshold per fingerprint and dataset. AtomPair ~0.55, MACCS ~0.85, ECFP4 ~0.7, and FCFP4 ~0.6 are repository starting heuristics, not universal equivalents.

Reconciliation: Cliffs Across Methods

If a pair flags as an activity cliff under one representation but not another, treat that as representation sensitivity. Inspect atom mappings, fingerprint environments, assay uncertainty, and the exact structural change; disagreement alone does not establish which substituent caused the activity difference.

Common Errors

SymptomCauseFix
BulkTanimotoSimilarity output is treated as bit countsThe API returns similarity values, normally floatsKeep the returned values as similarities; inspect input vector types if the output is unexpected
Reported similarity > 1Custom formula, malformed data, negative features, or an incorrectly normalized external implementationVerify the coefficient definition and inputs; standard nonnegative RDKit Tanimoto and Tversky similarities are bounded by 1
Cluster centroids change when input order changesTaylor-Butina tie handling and assignment are order-sensitiveStandardize and sort inputs by a stable identifier before clustering; record reordering and the input order
MaxMinPicker returns first N inputsAll-zero initial similarity matrixSeed picker explicitly: picker.LazyBitVectorPick(fps, n_lib, n_pick, seed=42)
Activity cliff "false positives"Bit-collisions inflate similarityUse sparse Morgan or compare canonical SMILES for exact ID
Diverse subset has duplicatesStandardization not appliedCanonicalize via chemoinformatics/molecular-standardization first
Tanimoto incompatible with neural fingerprintContinuous-valued fingerprintUse cosine or sklearn cdist with 'cosine' metric

References

  • chemoinformatics/molecular-descriptors - Generate fingerprints for similarity
  • chemoinformatics/molecular-standardization - Canonicalize before comparing
  • chemoinformatics/substructure-search - SMARTS pattern-based searching
  • chemoinformatics/scaffold-analysis - Scaffold-based similarity (Bemis-Murcko, MMPA)
  • chemoinformatics/shape-similarity - 3D shape similarity (USRCAT, ROCS)
  • machine-learning/biomarker-discovery - ML on similarity features

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in chemoinformatics/similarity-searching of GPTomics/bioSkills.

  • SKILL.md
  • examples/similarity_search.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Similarity Searching next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Similarity Searching compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Similarity Searching this skillGPTomics/bioSkills1.2k1 repos~4.6kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow84k4 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills47k2 repos~2.1kAutomated safety check: PassApache-2.0
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT
Last30daysmvanhorn/last30days-skill64k—~7.9kAutomated safety check: NotesMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    47k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Last30days

    mvanhorn/last30days-skill

    Research what people actually say about any topic in the last 30 days.

    64k GitHub stars~7.9k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.9k GitHub starsUsed in 17 repos~5.9k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Similarity Searching

What does Bio Similarity Searching do?

Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures…. Bio Similarity Searching is an agent skill from GPTomics/bioSkills. Performs molecular similarity searching using Tanimoto, Tversky, Dice, and cosine coefficients on bit/count fingerprints with explicit choice rules for symmetric vs asymmetric measures, scaffold-hopping vs lead-optimization regimes, activity-cliff diagnosis, and large-library nearest-neighbor methods (BulkTanimoto, MHFP6 LSH forest, USRCAT).

When should I use Bio Similarity Searching?

Bio Similarity Searching fits situations like: ranking compounds by structural resemblance to a query; clustering libraries; finding analogs; diagnosing activity cliffs.

How do I install Bio Similarity Searching in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a claude-code`. Or copy the skill folder (chemoinformatics/similarity-searching in GPTomics/bioSkills) into .claude/skills/bio-similarity-searching in your project. Claude Code loads it when a task matches its description.

How do I install Bio Similarity Searching in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a codex`. Or copy the skill folder (chemoinformatics/similarity-searching in GPTomics/bioSkills) into .agents/skills/bio-similarity-searching in your project. Codex loads it when a task matches its description.

Can I use Bio Similarity Searching in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-similarity-searching -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-similarity-searching, .gemini/skills/bio-similarity-searching, .github/skills/bio-similarity-searching and .opencode/skills/bio-similarity-searching in your project.

What does Bio Similarity Searching need to run?

Going by SKILL.md and its folder, Bio Similarity Searching needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Similarity Searching access the network?

SKILL.md names 2 domains. As links in the text: doi.org and rdkit.org. This is read from the text; nothing was executed.

Is Bio Similarity Searching safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Similarity Searching use?

Bio Similarity Searching is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Similarity Searching use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Similarity Searching?

Skills that share tags, products or a category with Bio Similarity Searching: Hypothesis Generation (spacering-net/codeg, 3.9k stars), GitHub Deep Research (bytedance/deer-flow, 84k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars) and Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Similarity Searching?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.