Agent skill

Claw Semantic Sim

by ClawBio in ClawBio/ClawBio

Semantic Similarity Index for disease research literature using PubMedBERT embeddings

MITAuto-check passedAI & LLM Engineering

Install Claw Semantic Sim

skills CLI
$ npx skills add ClawBio/ClawBio --skill claw-semantic-sim -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio claw-semantic-sim --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/claw-semantic-sim .claude/skills/claw-semantic-sim && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
claw-semantic-sim
GitHub stars
1.2k
Used in
3 other repos
Token cost
~1.4k tokens
SKILL.md length
487 words
Files
1
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Semantic Similarity Index for disease research literature using PubMedBERT embeddings

  • Works in 6 steps: Takes a disease list (GBD taxonomy) as… → Retrieves PubMed abstracts (2000-2025)… → Generates 768-dimensional PubMedBERT… → …
  • Tasks that involve Embeddings
  • SKILL.md covers What it does, Why this exists, Key Finding and Pipeline, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Claw Semantic Sim is an agent skill from ClawBio/ClawBio. Semantic Similarity Index for disease research literature using PubMedBERT embeddings

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Embeddings. It works with PubMed. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Tasks that involve Embeddings

Example prompts

  • “/claw-semantic-sim”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Takes a disease list (GBD taxonomy) as input
  2. Retrieves PubMed abstracts (2000-2025) for each disease with quality filtering
  3. Generates 768-dimensional PubMedBERT embeddings for every abstract
  4. Computes four semantic equity metrics per disease
  5. Generates publication-quality multi-panel figures
  6. Produces a markdown report with all metrics, rankings, and reproducibility bundle

What it can do on your machine

Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Claw Semantic Sim loads about 1.4k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 487 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 487 words, ~1,410 tokens.

Download SKILL.mdSave it as .claude/skills/claw-semantic-sim/SKILL.md (or your agent's skills folder).
name
claw-semantic-sim
description
Semantic Similarity Index for disease research literature using PubMedBERT embeddings
license
MIT
metadata.version
0.1.0
metadata.author
Manuel Corpas
metadata.tags
health-equity, semantic-analysis, NLP, PubMedBERT, disease-neglect

🦖 Semantic Similarity Index

Measure how isolated or connected disease research is across the global biomedical literature, using PubMedBERT embeddings on PubMed abstracts spanning 175 GBD diseases.

What it does

  1. Takes a disease list (GBD taxonomy) as input
  2. Retrieves PubMed abstracts (2000-2025) for each disease with quality filtering
  3. Generates 768-dimensional PubMedBERT embeddings for every abstract
  4. Computes four semantic equity metrics per disease:
    • Semantic Isolation Index (SII): average cosine distance to k-nearest disease neighbours; higher = more isolated, less connected research
    • Knowledge Transfer Potential (KTP): cross-disease centroid similarity; higher = more potential for research spillover
    • Research Clustering Coefficient (RCC): within-disease embedding variance; higher = more diverse research approaches
    • Temporal Semantic Drift: cosine distance between yearly centroids; measures how research focus evolves
  5. Generates publication-quality multi-panel figures:
    • Panel A: Semantic isolation by disease category (boxplot)
    • Panel B: Top 20 most semantically isolated diseases (bar chart, NTD/Global South colour-coded)
    • Panel C: Semantic isolation vs research volume (scatter with regression)
    • Panel D: NTD vs non-NTD significance test (Welch's t-test, Cohen's d)
  6. Produces a markdown report with all metrics, rankings, and reproducibility bundle

Why this exists

If you ask ChatGPT to "measure research neglect for diseases," it will:

  • Not know which embedding model to use for biomedical text
  • Hallucinate metrics that sound plausible but have no methodological grounding
  • Skip quality filtering (year coverage, abstract coverage, minimum papers)
  • Not handle MPS acceleration or checkpointed batch processing
  • Produce a single scatter plot with no disease classification

This skill encodes the correct methodological decisions:

  • Uses PubMedBERT (the gold-standard biomedical language model)
  • Fetches from PubMed with exponential backoff and NCBI rate limiting
  • Quality filters: year coverage >= 70%, abstract coverage >= 95%, minimum 50 papers
  • Batch embedding with Apple MPS acceleration and CPU fallback
  • Checkpointed processing (resume after interruption)
  • HDF5 storage with gzip compression and SHA-256 checksums
  • Classification against WHO NTD list and Global South priority diseases
  • Statistical significance testing (Welch's t-test, Cohen's d)
Show full SKILL.md (173 more words)Show less

Key Finding

Neglected tropical diseases (NTDs) are significantly more semantically isolated than other conditions (P < 0.001, Cohen's d = 0.8+). They exist in knowledge silos with limited cross-disciplinary research bridges. The 25 most isolated diseases are disproportionately Global South priority conditions.

Pipeline

05-00-heim-sem-setup.py     # Validate environment, create directories
05-01-heim-sem-fetch.py     # Retrieve PubMed abstracts (checkpointed)
05-02-heim-sem-embed.py     # Generate PubMedBERT embeddings (MPS/CPU)
05-03-heim-sem-compute.py   # Compute SII, KTP, RCC, temporal drift
05-04-heim-sem-figures.py   # Generate publication figures
05-05-heim-sem-integrate.py # Merge with biobank + clinical trial dimensions

Status

Not yet implemented. semantic_sim.py and the 05-00..05-05 pipeline scripts described above are not present in this repository yet; there is no runnable demo.

Example Output

Semantic Similarity Index
=========================
Diseases analysed: 175
Total PubMed abstracts: 13,100,000
Embedding model: PubMedBERT (768-dim)

Metric Ranges:
  SII: 0.0412 - 0.1893
  KTP: 0.6234 - 0.9187
  RCC: 0.0891 - 0.3421

Key Finding:
  NTDs show +38% higher semantic isolation
  P < 0.0001, Cohen's d = 0.84
  14/25 most isolated diseases are Global South priority

Figures saved to: demo_report/
  Fig5_Semantic_Structure.png (300 dpi)
  Fig5_Semantic_Structure.pdf (vector)

Reproducibility:
  commands.sh | environment.yml | checksums.sha256

Interpretation Guide

  • High SII: Disease research exists in a knowledge silo; limited cross-disciplinary bridges
  • Low KTP: Research on this disease has few methodological overlaps with others
  • High RCC: Diverse research approaches within the disease (many subtopics)
  • High Temporal Drift: Research focus has shifted significantly over time
  • NTDs shown in red, Global South diseases in orange, others in grey
  • The scatter plot (Panel C) reveals the inverse relationship between research volume and isolation

Citation

If you use this skill in a publication, please cite:

  • Corpas, M. et al. (2026). HEIM: Health Equity Index for Measuring structural bias in biomedical research. Under review.
  • Corpas, M. (2026). ClawBio. https://github.com/ClawBio/ClawBio

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/claw-semantic-sim of ClawBio/ClawBio.

Open the folder on GitHubat commit 5e045e3

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Claw Semantic Sim next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Claw Semantic Sim compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Claw Semantic Sim this skillClawBio/ClawBio1.2k3 repos~1.4kAutomated safety check: PassMIT
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Unimoljinzhezenggroup/computational-chemistry-agent-skills1481 repos~1.5kAutomated safety check: PassLGPL-3.0-or-later
ExploreZimoLiao/scholaraio576—~755Automated safety check: PassMIT
Esmdavila7/claude-code-templates32k10 repos~2.6kAutomated safety check: WarnMIT
Kermt EmbedNVIDIA-BioNeMo/bionemo-agent-toolkit478—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Unimol

    jinzhezenggroup/computational-chemistry-agent-skills

    A standardized CLI wrapper for Uni-Mol molecular ML workflows that handles representation extraction (embeddings), model training (regression/classification), and property prediction with built-in…

    148 GitHub starsUsed in 1 repo~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Explore

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to survey a journal or field, fetch papers from OpenAlex, cluster topics, build exploration embeddings, or search named explore libraries under…

    576 GitHub stars~755 tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Esm

    davila7/claude-code-templates

    Comprehensive toolkit for protein language models including ESM3 (generative multimodal protein design across sequence, structure, and function) and ESM C (efficient protein embeddings and…

    32k GitHub starsUsed in 10 repos~2.6k tokens
    AI & LLM EngineeringAuto-check: warnings
  • Kermt Embed

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (groverbase / cmim / hybrid / finetuned).

    478 GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Sc Clustering

    TianGzlab/OmicsClaw

    Load when building the neighbour graph, embedding (UMAP/t-SNE/diffmap/PHATE), and clustering (Leiden/Louvain) on a normalised single-cell AnnData.

    161 GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Works with

Questions about Claw Semantic Sim

What does Claw Semantic Sim do?

Semantic Similarity Index for disease research literature using PubMedBERT embeddings. Claw Semantic Sim is an agent skill from ClawBio/ClawBio.

When should I use Claw Semantic Sim?

Claw Semantic Sim fits situations like: tasks that involve Embeddings.

How do I install Claw Semantic Sim in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill claw-semantic-sim -a claude-code`. Or copy the skill folder (skills/claw-semantic-sim in ClawBio/ClawBio) into .claude/skills/claw-semantic-sim in your project. Claude Code loads it when a task matches its description.

How do I install Claw Semantic Sim in Codex?

Run `npx skills add ClawBio/ClawBio --skill claw-semantic-sim -a codex`. Or copy the skill folder (skills/claw-semantic-sim in ClawBio/ClawBio) into .agents/skills/claw-semantic-sim in your project. Codex loads it when a task matches its description.

Can I use Claw Semantic Sim in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill claw-semantic-sim -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/claw-semantic-sim, .gemini/skills/claw-semantic-sim, .github/skills/claw-semantic-sim and .opencode/skills/claw-semantic-sim in your project.

What does Claw Semantic Sim need to run?

SKILL.md names no scripts, command-line tools or credentials: Claw Semantic Sim is instructions for the agent only. Our summary lists: Python 3.

Does Claw Semantic Sim access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Claw Semantic Sim safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Claw Semantic Sim use?

Claw Semantic Sim is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Claw Semantic Sim use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Claw Semantic Sim?

Skills that share tags, products or a category with Claw Semantic Sim: Esmfold2 (JimLiu/science-skills, 227 stars), Unimol (jinzhezenggroup/computational-chemistry-agent-skills, 148 stars), Explore (ZimoLiao/scholaraio, 576 stars) and Esm (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Claw Semantic Sim?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.