Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (groverbase / cmim / hybrid / finetuned).

Apache-2.0Auto-check passedAI & LLM Engineering

Install Kermt Embed

skills CLI
$ npx skills add NVIDIA-BioNeMo/bionemo-agent-toolkit --skill kermt-embed -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA-BioNeMo/bionemo-agent-toolkit kermt-embed --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA-BioNeMo/bionemo-agent-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/open-models-skills/kermt/kermt-embed .claude/skills/kermt-embed && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kermt-embed
GitHub stars
478
Token cost
~1.7k tokens
SKILL.md length
564 words
Files
2
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (groverbase / cmim / hybrid / finetuned).

  • Works in 7 steps: Pre-flight: container + system probe. → Compute run directory. → Resolve & validate the checkpoint. → …
  • Tasks that involve Drug discovery and cheminformatics
  • SKILL.md covers Hardware requirements, Inputs, Workflow and Hard rules, plus 2 more sections
  • Calls jq and git

What it does

Kermt Embed is an agent skill from NVIDIA-BioNeMo/bionemo-agent-toolkit. Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (groverbase / cmim / hybrid / finetuned). Writes one .npy per readout type (atomfromatom, bondfromatom, atomfrombond, bondfrombond) plus canonicalsmiles.npy and validity.npy. Calls task/extractembeddings.py (which featurizes SMILES on the fly — no pre-computed features needed).

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`). Compatibility notes: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.

It sits in AI & LLM Engineering, covering Drug discovery and cheminformatics and Embeddings. The repository describes itself as: Turn any agent into a life science expert with NVIDIA BioNeMo skills. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Drug discovery and cheminformatics
  • Tasks that involve Embeddings

Example prompts

  • “/kermt-embed”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Pre-flight: container + system probe.
  2. Compute run directory.
  3. Resolve & validate the checkpoint.
  4. Validate the data.
  5. Prepare the data (clean-only — no features step).
  6. Launch the runner (blocking).
  7. Report to the user.

What it can do on your machine

Read from SKILL.md and the folder at commit 2113472. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.

    From compatibility in the SKILL.md frontmatter.

Context cost

Kermt Embed loads about 1.7k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 564 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA-BioNeMo/bionemo-agent-toolkit at commit 2113472, republished under its Apache-2.0 licence (© NVIDIA-BioNeMo). 564 words, ~1,680 tokens.

Download SKILL.mdSave it as .claude/skills/kermt-embed/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
kermt-embed
description
Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (grover_base / cmim / hybrid / finetuned). Writes one .npy per readout type (atom_from_atom, bond_from_atom, atom_from_bond, bond_from_bond) plus canonical_smiles.npy and validity.npy. Calls task/extract_embeddings.py (which featurizes SMILES on the fly — no pre-computed features needed).
compatibility
Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
license
Apache-2.0
metadata.owner
evax@nvidia.com
metadata.classification
workflow-skill
metadata.risk_tier
skill

kermt-embed

Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint. The skill is the workflow orchestrator: validate ckpt, validate CSV, clean SMILES, launch the runner blocking, return the per-readout .npy files.

Hardware requirements

  • GPUs: 1 (single-GPU).
  • VRAM: ≥ 4 GB for the default batch_size 64.
  • Disk: depends on output size — roughly a few MB per 1k molecules at hidden 800 per readout, so ~10–20 MB per 1k molecules across the 4 readouts. Plus a small canonical_smiles.npy + validity.npy per run.
  • Driver / CUDA: any host supporting CUDA 12.6.

Inputs

Required:

  • --csv <path> — SMILES CSV. First column is smiles; other columns are ignored (no targets needed).

Checkpoint (optional — defaults to the released model if omitted):

  • --ckpt <path> — any encoder-bearing checkpoint. Grover_base, cmim, hybrid, and finetuned ckpts are all accepted. The validator only refuses ckpts with no encoder. If omitted, the skill offers to download the released pretrained hybrid model nvidia/NV-KERMT-70M-v2 and embed with it — see "Resolve & validate the checkpoint" (workflow step 3).
  • --pretrained-release — explicit opt-in to use the released model without the interactive prompt (for non-interactive / agent runs). Mutually exclusive with --ckpt.
  • --model-dir <dir> — where to save the downloaded bundle (default $KERMT_REPO/models/NV-KERMT-70M-v2/). An already-complete bundle there is reused, not re-downloaded.

Optional:

  • --batch-size N — override the configured default (64).
  • --gpus 0 — single GPU id (default 0).
  • --from-prepare <dir> — skip the prepare step and reuse an existing prepare_data.json in <dir>.

Workflow

Let $KERMT_REPO be the path to your kermt repo checkout.

  1. Pre-flight: container + system probe.

    $KERMT_REPO/agent/scripts/kermt_container.sh check_system
  2. Compute run directory.

    RUN_DIR=$KERMT_REPO/runs/embed_$(date -u +%Y-%m-%dT%H-%M-%SZ)
  3. Resolve & validate the checkpoint.

    Resolve — only if --ckpt was omitted. Default to the released pretrained hybrid model nvidia/NV-KERMT-70M-v2:

    • Consent gate. Unless --pretrained-release was passed, ask the user: "No checkpoint given — download the released model nvidia/NV-KERMT-70M-v2 (NVIDIA Open Model License, https://huggingface.co/nvidia/NV-KERMT-70M-v2) and embed with it? [y/N]". Never download without an explicit yes (or --pretrained-release). If both --ckpt and --pretrained-release are given, abort — they conflict.
    • Save location. Default $KERMT_REPO/models/NV-KERMT-70M-v2/; honor --model-dir <dir> if given. An already-complete bundle is reused.
    • Download (foreground; ~282 MB on first fetch):
      $KERMT_REPO/agent/scripts/kermt_container.sh run --model-dir <save-dir> -- \
          "python agent/scripts/fetch_released_model.py --out /model"
      Parse the JSON; abort on ok: false (surface errors). On success set <user-ckpt> = <save-dir>/kermt_contrastive_v2.0.pt.

    Validate the resolved (or user-provided) ckpt:

    $KERMT_REPO/agent/scripts/kermt_container.sh run --ckpt <user-ckpt> -- \
        "python agent/scripts/check_checkpoint.py --mode embed --ckpt /ckpt"

    Parse JSON. Abort on ok: false. The validator only refuses encoder-less ckpts (rare).

  4. Validate the data.

    $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> -- \
        "python agent/scripts/check_data.py --mode embed --csv /data/<basename>"
  5. Prepare the data (clean-only — no features step).

    $KERMT_REPO/agent/scripts/kermt_container.sh run --data <user-csv> --run-dir $RUN_DIR -- \
        "python agent/scripts/prepare_data.py --mode embed \\
             --csv /data/<basename> --out /runs/data"

    Outputs land at $RUN_DIR/data/prepare_data.json with a single clean_csv path. task/extract_embeddings.py featurizes from SMILES on the fly.

  6. Launch the runner (blocking).

    $KERMT_REPO/agent/scripts/kermt_container.sh run \\
        --ckpt <user-ckpt> --run-dir $RUN_DIR -- \\
        "python agent/scripts/run_extract_embeddings.py \\
             --ckpt /ckpt \\
             --prepare-manifest /runs/data/prepare_data.json \\
             --out /runs \\
             [--gpus 0 --batch-size N]"
  7. Report to the user.

    • Embeddings directory: $RUN_DIR/out/
      • atom_from_atom.npy, bond_from_atom.npy, atom_from_bond.npy, bond_from_bond.npy (the 4 standard readouts; each shape (N_rows, hidden_size))
      • metadata.pkl — pickle of a dict containing canonical_smiles (RDKit-canonicalized SMILES per row), valid (boolean per-row: did RDKit parse it), plus other run metadata.
    • Manifest: $RUN_DIR/run.json
    • Log: $RUN_DIR/logs/embed.log
Show full SKILL.md (116 more words)Show less

Hard rules

  • Never download the released model without consent. When --ckpt is omitted, download nvidia/NV-KERMT-70M-v2 only after an explicit user "yes" or an explicit --pretrained-release flag. --ckpt and --pretrained-release are mutually exclusive.
  • Never modify the user's ckpt. The runner reads-only via task/extract_embeddings.py's --checkpoint <path> flag.
  • Arch comes from the ckpt. No --hidden-size flag etc. on this runner; task/extract_embeddings.py reads arch from the ckpt's saved_args.

Common errors

  • prepare_data manifest is missing required output 'clean_csv' → prepare ran with --skip-clean but no source CSV given. Re-run prepare without it.
  • --gpus '0,1' is single-GPU only → pass a single id.

Replayability

bash
$(jq -r .cmd_replay $RUN_DIR/run.json)

If ok_to_replay: false (dirty kermt repo worktree at launch time), pin the commit via repo.commit and git checkout it first.

© NVIDIA-BioNeMo, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in open-models-skills/kermt/kermt-embed of NVIDIA-BioNeMo/bionemo-agent-toolkit.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit 2113472

Compare with similar skills

Kermt Embed next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kermt Embed compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kermt Embed this skillNVIDIA-BioNeMo/bionemo-agent-toolkit478—~1.7kAutomated safety check: PassApache-2.0
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Unimoljinzhezenggroup/computational-chemistry-agent-skills1481 repos~1.5kAutomated safety check: PassLGPL-3.0-or-later
Disease Reversal PredictionInternScience/scp1691 repos~931Automated safety check: PassMIT
MolfeatK-Dense-AI/scientific-agent-skills48k1 repos~2.4kAutomated safety check: NotesApache-2.0
Kermt EmbedNVIDIA/skills3.5k1 repos~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Unimol

    jinzhezenggroup/computational-chemistry-agent-skills

    A standardized CLI wrapper for Uni-Mol molecular ML workflows that handles representation extraction (embeddings), model training (regression/classification), and property prediction with built-in…

    148 GitHub starsUsed in 1 repo~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Predict a molecule's ability to reverse disease states using DLEPS (Disease-Ligand Embedding Projection Score) for drug repositioning and discovery.

    169 GitHub starsUsed in 1 repo~931 tokens
    AI & LLM EngineeringAuto-check passed
  • Molfeat

    K-Dense-AI/scientific-agent-skills

    Featurizes small molecules with Molfeat for QSAR/QSPR, chemical similarity, virtual screening, and molecular ML.

    48k GitHub starsUsed in 1 repo~2.4k tokens
    Research & ScienceAuto-check: notes
  • Kermt Embed

    NVIDIA/skills

    Official

    Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint.

    3.5k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Molfeat Molecular Featurization

    jaechang-hits/SciAgent-Skills

    Molecular featurization hub (100+ featurizers) for ML. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check passed

More from NVIDIA-BioNeMo/bionemo-agent-toolkit

All 14 skills in this repo
  • Complexa Binder Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided…

    478 GitHub stars~3.1k tokensUpdated today
    Auto-check: notes
  • Parabricks

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Route NVIDIA Parabricks pbrun tools, assess GPU/runtime readiness, and provide version-aware command guidance for FASTQ/BAM processing, RNA-seq, variant calling, BAM QC, and GVCF workflows.

    478 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Protein Binder Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Orchestrate an end-to-end de novo protein binder design campaign against a protein target by composing BioNeMo NIM skills.

    478 GitHub stars~1.4k tokensUpdated today
    Auto-check: notes
  • Cuequivariance

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Build and debug cuEquivariance irreps, custom Irrep subclasses, Clebsch-Gordan tensor products, and equivariant or segmented polynomials.

    478 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Genomics Workflow Acceleration

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    A skill your agent uses when accelerating existing genomics workflows with NVIDIA Parabricks, improving runtime or price/performance, converting pipeline steps to GPUs, or comparing CPU and GPU…

    478 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Complexa Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    End-to-end Proteina-Complexa design pipeline driver. An agent skill from NVIDIA-BioNeMo/bionemo-agent-toolkit.

    478 GitHub stars~4.1k tokensUpdated today
    Auto-check: notes

Questions about Kermt Embed

What does Kermt Embed do?

Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (groverbase / cmim / hybrid / finetuned). Kermt Embed is an agent skill from NVIDIA-BioNeMo/bionemo-agent-toolkit. Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint (groverbase / cmim / hybrid / finetuned).

When should I use Kermt Embed?

Kermt Embed fits situations like: tasks that involve Drug discovery and cheminformatics; tasks that involve Embeddings.

How do I install Kermt Embed in Claude Code?

Run `npx skills add NVIDIA-BioNeMo/bionemo-agent-toolkit --skill kermt-embed -a claude-code`. Or copy the skill folder (open-models-skills/kermt/kermt-embed in NVIDIA-BioNeMo/bionemo-agent-toolkit) into .claude/skills/kermt-embed in your project. Claude Code loads it when a task matches its description.

How do I install Kermt Embed in Codex?

Run `npx skills add NVIDIA-BioNeMo/bionemo-agent-toolkit --skill kermt-embed -a codex`. Or copy the skill folder (open-models-skills/kermt/kermt-embed in NVIDIA-BioNeMo/bionemo-agent-toolkit) into .agents/skills/kermt-embed in your project. Codex loads it when a task matches its description.

Can I use Kermt Embed in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-BioNeMo/bionemo-agent-toolkit --skill kermt-embed -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kermt-embed, .gemini/skills/kermt-embed, .github/skills/kermt-embed and .opencode/skills/kermt-embed in your project.

What does Kermt Embed need to run?

Going by SKILL.md and its folder, Kermt Embed needs the command-line tools its instructions call (jq and git). Our summary lists: Python 3; Docker. Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron..

Does Kermt Embed access the network?

SKILL.md names 1 domain. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is Kermt Embed safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kermt Embed use?

Kermt Embed is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kermt Embed use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kermt Embed?

Skills that share tags, products or a category with Kermt Embed: Esmfold2 (JimLiu/science-skills, 227 stars), Unimol (jinzhezenggroup/computational-chemistry-agent-skills, 148 stars), Disease Reversal Prediction (InternScience/scp, 169 stars) and Molfeat (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kermt Embed?

NVIDIA-BioNeMo (a GitHub organization) maintains it in NVIDIA-BioNeMo/bionemo-agent-toolkit, which has 478 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA-BioNeMo/bionemo-agent-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.