Nvmolkit Usage
NVIDIA-BioNeMo/bionemo-agent-toolkit
Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF…
Molecular featurization hub (100+ featurizers) for ML. An agent skill from jaechang-hits/SciAgent-Skills.
$ npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills molfeat-molecular-featurization --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/molfeat-molecular-featurization .claude/skills/molfeat-molecular-featurization && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "molfeat-molecular-featurization" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/molfeat-molecular-featurization into .claude/skills/molfeat-molecular-featurization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "molfeat-molecular-featurization", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/molfeat-molecular-featurizationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills molfeat-molecular-featurization --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/structural-biology-drug-discovery/molfeat-molecular-featurization .agents/skills/molfeat-molecular-featurization && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "molfeat-molecular-featurization" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/molfeat-molecular-featurization into .agents/skills/molfeat-molecular-featurization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "molfeat-molecular-featurization", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills molfeat-molecular-featurization --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/structural-biology-drug-discovery/molfeat-molecular-featurization .cursor/skills/molfeat-molecular-featurization && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "molfeat-molecular-featurization" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/molfeat-molecular-featurization into .cursor/skills/molfeat-molecular-featurization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "molfeat-molecular-featurization", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/structural-biology-drug-discovery/molfeat-molecular-featurization--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills molfeat-molecular-featurization --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/structural-biology-drug-discovery/molfeat-molecular-featurization .gemini/skills/molfeat-molecular-featurization && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "molfeat-molecular-featurization" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/molfeat-molecular-featurization into .gemini/skills/molfeat-molecular-featurization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "molfeat-molecular-featurization", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills molfeat-molecular-featurizationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/structural-biology-drug-discovery/molfeat-molecular-featurization .github/skills/molfeat-molecular-featurization && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "molfeat-molecular-featurization" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/molfeat-molecular-featurization into .github/skills/molfeat-molecular-featurization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "molfeat-molecular-featurization", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills molfeat-molecular-featurization --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/structural-biology-drug-discovery/molfeat-molecular-featurization .opencode/skills/molfeat-molecular-featurization && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "molfeat-molecular-featurization" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/molfeat-molecular-featurization into .opencode/skills/molfeat-molecular-featurization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "molfeat-molecular-featurization", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
molfeat-molecular-featurizationMolecular featurization hub (100+ featurizers) for ML. An agent skill from jaechang-hits/SciAgent-Skills.
Molfeat Molecular Featurization is an agent skill from jaechang-hits/SciAgent-Skills. Molecular featurization hub (100+ featurizers) for ML. SMILES to fingerprints (ECFP, MACCS, MAP4), descriptors (RDKit 2D, Mordred), pretrained embeddings (ChemBERTa, GIN, Graphormer), pharmacophores. Scikit-learn compatible with parallelization/caching. For QSAR, virtual screening, similarity, and molecular DL.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/api_reference.md` and `references/available_featurizers.md`).
It sits in Research & Science, covering Drug discovery and cheminformatics and Embeddings. It works with scikit-learn and RDKit. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvpipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
molfeat-docs.datamol.iogithub.compypi.orgportal.valencelabs.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Molfeat Molecular Featurization loads about 4.3k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 950 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 950 words, ~4,325 tokens.
.claude/skills/molfeat-molecular-featurization/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Molfeat is a comprehensive Python library for molecular featurization that unifies 100+ pre-trained embeddings and hand-crafted featurizers under a scikit-learn compatible API. Convert SMILES strings into numerical representations (fingerprints, descriptors, deep learning embeddings) for QSAR modeling, virtual screening, similarity searching, and chemical space analysis.
uv pip install molfeat
# Optional extras for specific featurizer types
uv pip install "molfeat[transformer]" # ChemBERTa, ChemGPT, MolT5
uv pip install "molfeat[dgl]" # GIN graph neural networks
uv pip install "molfeat[graphormer]" # Graphormer models
uv pip install "molfeat[fcd]" # FCD descriptors
uv pip install "molfeat[map4]" # MAP4 fingerprints
uv pip install "molfeat[all]" # All dependenciesfrom molfeat.calc import FPCalculator
from molfeat.trans import MoleculeTransformer
smiles = ["CCO", "CC(=O)O", "c1ccccc1", "CC(C)O"]
# Create fingerprint calculator + transformer
calc = FPCalculator("ecfp", radius=3, fpSize=2048)
transformer = MoleculeTransformer(calc, n_jobs=-1)
# Featurize batch in parallel
features = transformer(smiles)
print(f"Shape: {features.shape}") # (4, 2048)
# Save configuration for reproducibility
transformer.to_state_yaml_file("featurizer_config.yml")Molfeat organizes featurization into three layers:
| Layer | Class | Purpose | Use When |
|---|---|---|---|
| Calculator | molfeat.calc.* | Single molecule → feature vector | Custom loops, single molecules |
| Transformer | molfeat.trans.MoleculeTransformer | Batch processing with parallelization | Datasets, scikit-learn pipelines |
| Store | molfeat.store.ModelStore | Discovery and loading of pretrained models | Finding available featurizers |
Calculators are callable: calc("CCO") returns a numpy array. Transformers wrap calculators for batch processing: transformer(smiles_list) returns a 2D array. Pretrained transformers (PretrainedMolTransformer) add batched GPU inference and caching.
| Task | Recommended | Dimensions | Speed |
|---|---|---|---|
| General QSAR | ecfp (radius=3) | 2048 | Fast |
| Scaffold similarity | maccs | 167 | Very fast |
| Large-scale screening | map4 | 1024 | Fast |
| Interpretable models | desc2D (RDKitDescriptors2D) | 200+ | Fast |
| Comprehensive descriptors | mordred | 1800+ | Medium |
| Transfer learning | ChemBERTa-77M-MLM | 768 | Slow* |
| Graph-based DL | gin-supervised-masking | Variable | Slow* |
| Pharmacophore | fcfp or cats2D | 2048 / 21 | Fast |
| 3D shape | usr / usrcat | 12 / 60 | Fast |
*First run slow; subsequent runs cached.
Save and reload exact featurizer configuration for reproducibility:
# Save
transformer.to_state_yaml_file("config.yml")
transformer.to_state_json_file("config.json")
# Reload
loaded = MoleculeTransformer.from_state_yaml_file("config.yml")from molfeat.calc import FPCalculator
# ECFP — most popular, general-purpose
ecfp = FPCalculator("ecfp", radius=3, fpSize=2048)
fp = ecfp("CCO")
print(f"ECFP shape: {fp.shape}") # (2048,)
# MACCS keys — 167-bit structural keys, fast scaffold similarity
maccs = FPCalculator("maccs")
fp = maccs("c1ccccc1")
print(f"MACCS shape: {fp.shape}") # (167,)
# Count-based fingerprints (non-binary)
ecfp_count = FPCalculator("ecfp-count", radius=3, fpSize=2048)
# MAP4 — MinHashed atom-pair, efficient for large databases
map4 = FPCalculator("map4")
print(f"MAP4 shape: {map4('CCO').shape}") # (1024,)Available fingerprint types: ecfp, fcfp, maccs, rdkit, avalon, pattern, layered, atompair, topological, map4, secfp, erg, estate (and count variants with -count suffix).
from molfeat.calc import RDKitDescriptors2D, MordredDescriptors
# RDKit 2D — 200+ named properties (MW, logP, TPSA, etc.)
desc2d = RDKitDescriptors2D()
descriptors = desc2d("CCO")
print(f"2D descriptors: {len(descriptors)}") # 200+
print(f"Feature names: {desc2d.columns[:5]}")
# Mordred — 1800+ comprehensive descriptors
mordred = MordredDescriptors()
descriptors = mordred("c1ccccc1O")
print(f"Mordred descriptors: {len(descriptors)}") # 1800+from molfeat.calc import CATSCalculator, USRDescriptors
# CATS — pharmacophore point pair distributions
cats = CATSCalculator(mode="2D", scale="raw")
descriptors = cats("CC(C)Cc1ccc(C)cc1C")
print(f"CATS shape: {descriptors.shape}") # (21,)
# USR — ultrafast shape recognition
usr = USRDescriptors()
shape = usr("CC(=O)Oc1ccccc1C(=O)O")
print(f"USR shape: {shape.shape}") # (12,)from molfeat.trans import MoleculeTransformer, FeatConcat
from molfeat.calc import FPCalculator
smiles = ["CCO", "CC(=O)O", "c1ccccc1", "CC(C)O", "CCCC"]
# Parallel batch processing
transformer = MoleculeTransformer(FPCalculator("ecfp"), n_jobs=-1)
features = transformer(smiles)
print(f"Batch shape: {features.shape}") # (5, 2048)
# Concatenate multiple featurizers
concat = FeatConcat([
FPCalculator("maccs"), # 167 dims
FPCalculator("ecfp") # 2048 dims
])
combo_transformer = MoleculeTransformer(concat, n_jobs=-1)
combo_features = combo_transformer(smiles)
print(f"Combined shape: {combo_features.shape}") # (5, 2215)
# Error-tolerant processing
safe_transformer = MoleculeTransformer(
FPCalculator("ecfp"), n_jobs=-1,
ignore_errors=True, verbose=True
)
features = safe_transformer(["CCO", "invalid", "c1ccccc1"])
# Returns None for failed moleculesfrom molfeat.trans.pretrained import PretrainedMolTransformer
# ChemBERTa — RoBERTa trained on 77M PubChem compounds
chemberta = PretrainedMolTransformer("ChemBERTa-77M-MLM", n_jobs=-1)
embeddings = chemberta(["CCO", "CC(=O)O", "c1ccccc1"])
print(f"ChemBERTa shape: {embeddings.shape}") # (3, 768)
# GIN — graph neural network pretrained on ChEMBL
gin = PretrainedMolTransformer("gin-supervised-masking", n_jobs=-1)
graph_emb = gin(["CCO", "CC(=O)O"])
print(f"GIN shape: {graph_emb.shape}")from molfeat.store.modelstore import ModelStore
store = ModelStore()
print(f"Total available: {len(store.available_models)}")
# Search for specific model
results = store.search(name="ChemBERTa")
for model in results:
print(f" {model.name}: {model.description}")
# View usage and load
card = store.search(name="ChemBERTa-77M-MLM")[0]
card.usage()
transformer = store.load("ChemBERTa-77M-MLM")from molfeat.calc import FPCalculator
from molfeat.trans import MoleculeTransformer
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import cross_val_score
# Featurize molecules
transformer = MoleculeTransformer(FPCalculator("ecfp", radius=3), n_jobs=-1)
X = transformer(smiles_train)
print(f"Features shape: {X.shape}")
# Train and evaluate
model = RandomForestRegressor(n_estimators=100)
scores = cross_val_score(model, X, y_train, cv=5, scoring='r2')
print(f"R² = {scores.mean():.3f} ± {scores.std():.3f}")
# Save for deployment
transformer.to_state_yaml_file("production_featurizer.yml")from sklearn.ensemble import RandomForestClassifier
# Step 1: Featurize known actives/inactives
transformer = MoleculeTransformer(FPCalculator("ecfp"), n_jobs=-1)
X_train = transformer(train_smiles)
# Step 2: Train classifier
clf = RandomForestClassifier(n_estimators=500, n_jobs=-1)
clf.fit(X_train, train_labels)
# Step 3: Screen library (e.g., 1M compounds)
X_screen = transformer(screening_smiles)
predictions = clf.predict_proba(X_screen)[:, 1]
# Step 4: Rank and select top hits
top_indices = predictions.argsort()[::-1][:1000]
top_hits = [screening_smiles[i] for i in top_indices]
print(f"Top 1000 hits selected from {len(screening_smiles)} compounds")from molfeat.calc import FPCalculator, RDKitDescriptors2D
from sklearn.metrics import roc_auc_score
featurizers = {
'ECFP': FPCalculator("ecfp"),
'MACCS': FPCalculator("maccs"),
'Descriptors': RDKitDescriptors2D(),
}
for name, calc in featurizers.items():
transformer = MoleculeTransformer(calc, n_jobs=-1)
X_train = transformer(smiles_train)
X_test = transformer(smiles_test)
clf = RandomForestClassifier(n_estimators=100)
clf.fit(X_train, y_train)
auc = roc_auc_score(y_test, clf.predict_proba(X_test)[:, 1])
print(f"{name}: AUC = {auc:.3f}")from sklearn.pipeline import Pipeline
from sklearn.ensemble import RandomForestClassifier
pipeline = Pipeline([
('featurizer', MoleculeTransformer(FPCalculator("ecfp"), n_jobs=-1)),
('classifier', RandomForestClassifier(n_estimators=100))
])
pipeline.fit(smiles_train, y_train)
predictions = pipeline.predict(smiles_test)from sklearn.metrics.pairwise import cosine_similarity
calc = FPCalculator("ecfp")
query_fp = calc("CC(=O)Oc1ccccc1C(=O)O").reshape(1, -1) # Aspirin
transformer = MoleculeTransformer(calc, n_jobs=-1)
db_fps = transformer(database_smiles)
similarities = cosine_similarity(query_fp, db_fps)[0]
top_k = similarities.argsort()[-10:][::-1]
for i in top_k:
print(f" {database_smiles[i]}: {similarities[i]:.3f}")import numpy as np
def featurize_chunks(smiles_list, transformer, chunk_size=10000):
all_features = []
for i in range(0, len(smiles_list), chunk_size):
chunk = smiles_list[i:i+chunk_size]
features = transformer(chunk)
all_features.append(features)
print(f"Processed {min(i+chunk_size, len(smiles_list))}/{len(smiles_list)}")
return np.vstack(all_features)| Parameter | Module | Default | Description |
|---|---|---|---|
method | FPCalculator | — | Fingerprint type: ecfp, maccs, map4, etc. |
radius | FPCalculator | 3 | Circular fingerprint radius |
fpSize | FPCalculator | 2048 | Fingerprint bit length |
counting | FPCalculator | False | Count vector instead of binary |
n_jobs | MoleculeTransformer | 1 | Parallel workers (-1 = all cores) |
ignore_errors | MoleculeTransformer | False | Skip invalid molecules (returns None) |
verbose | MoleculeTransformer | False | Log processing details |
dtype | MoleculeTransformer | float64 | Output type (float32 for memory) |
mode | CATSCalculator | "2D" | Distance calculation mode |
scale | CATSCalculator | "raw" | Scaling: raw, num, count |
n_jobs=-1 for parallel processing on all CPU cores — significant speedup for batch featurizationignore_errors=True for large datasets — invalid SMILES won't crash the pipelineto_state_yaml_file() for reproducibility — recreate exact featurizer laterMoleculeTransformer(calc, dtype=np.float32)FeatConcat to capture complementary molecular information| Problem | Cause | Solution |
|---|---|---|
ValueError: unsupported featurizer | Unknown method name | Check FPCalculator supported types or use ModelStore.search() |
ImportError for pretrained model | Missing optional dependency | Install extras: pip install "molfeat[transformer]" or "molfeat[dgl]" |
None in output array | Invalid SMILES with ignore_errors=True | Filter results: [f for f in features if f is not None] |
| Memory error on large dataset | Too many molecules at once | Process in chunks of 10K-50K (see Recipes) |
| Slow pretrained model inference | First run downloads model weights | Normal — subsequent runs use cache |
| Shape mismatch in pipeline | Mixed valid/invalid molecules | Ensure ignore_errors=True and filter None before ML model |
| Reproducibility issues | Different molfeat versions | Pin version and save config: transformer.to_state_yaml_file() |
datamol-cheminformatics — High-level molecular manipulation (standardization, I/O, conformers)rdkit-cheminformatics — Low-level cheminformatics (substructure, reactions, 3D)scikit-learn — ML models consuming molfeat featuresMain SKILL.md + 2 reference files. Original total: 1,273 lines (SKILL.md 510 + api_reference.md 429 + available_featurizers.md 334). Scripts: none. Examples: 724 lines (examples.md).
references/available_featurizers.md: Complete catalog of all 100+ featurizers organized by category — transformer models, GNNs, descriptors, fingerprints, pharmacophore, shape, scaffold, graph featurizers. Includes dimensions, dependencies, and selection guidance per category. Purely lookup-oriented content preserved as reference.
references/api_reference.md: Detailed API reference for molfeat.calc, molfeat.trans, and molfeat.store modules. Covers SerializableCalculator base class, all calculator subclasses with parameters, MoleculeTransformer methods, PretrainedMolTransformer, FeatConcat, ModelStore/ModelCard API, data type control, and PyTorch integration patterns.
Original file disposition:
SKILL.md (510 lines) → Core API modules 1-6, Key Concepts (architecture, selection guide), Quick Start, Workflows 1-3. "Choosing the Right Featurizer" → Key Concepts selection guide table. "Advanced Features" (custom preprocessing, batch processing, caching) → Recipes + Best Practices. "Common Featurizers Reference" table → Key Concepts selection guide. "Performance Tips" → Best Practices. Per-use-case disposition: QSAR Modeling → Workflow 1, Virtual Screening → Workflow 2, Similarity Search → Recipe, Chemical Space → When to Use bullet, scikit-learn Pipeline → Recipe, Featurizer Comparison → Workflow 3references/api_reference.md (429 lines) → Migrated to new references/api_reference.md. Core patterns (FPCalculator, MoleculeTransformer, basic ModelStore) relocated to SKILL.md Core API modules 1-6. Detailed class methods, SerializableCalculator base class, PrecomputedMolTransformer, and PyTorch integration retained in referencereferences/available_featurizers.md (334 lines) → Migrated to new references/available_featurizers.md. Top-level summary → Key Concepts selection guide table. Full categorized catalog retained in referencereferences/examples.md (724 lines) → Fully consolidated inline: installation → Prerequisites; calculator examples → Core API 1-3; transformer examples → Core API 4; pretrained examples → Core API 5; ML integration → Workflows 1-3 + Recipes; advanced patterns (custom preprocessing, caching, chunk processing) → Recipes + Best Practices; troubleshooting → Troubleshooting table. No separate reference file needed — all content absorbed into SKILL.md sectionsRetention: ~490 lines (SKILL.md) + ~170 lines (available_featurizers) + ~190 lines (api_reference) = ~850 / 1,273 original (excl. examples.md treated as consolidated) = ~67%. Including examples.md in denominator: ~850 / 1,997 = ~43%.
© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/structural-biology-drug-discovery/molfeat-molecular-featurization of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Molfeat Molecular Featurization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Molfeat Molecular Featurization this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~4.3k | Automated safety check: Pass | Apache-2.0 | |
| Nvmolkit UsageNVIDIA-BioNeMo/bionemo-agent-toolkit | 479 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| MolfeatK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~2.4k | Automated safety check: Notes | Apache-2.0 | |
| Unimoljinzhezenggroup/computational-chemistry-agent-skills | 148 | — | ~1.5k | Automated safety check: Pass | LGPL-3.0-or-later | |
| Nvmolkit UsageNVIDIA/skills | 3.6k | 1 repos | ~4.8k | Automated safety check: Pass | Apache-2.0 | |
| DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3k | Automated safety check: Notes | MIT |
NVIDIA-BioNeMo/bionemo-agent-toolkit
Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF…
K-Dense-AI/scientific-agent-skills
Featurizes small molecules with Molfeat for QSAR/QSPR, chemical similarity, virtual screening, and molecular ML.
jinzhezenggroup/computational-chemistry-agent-skills
A standardized CLI wrapper for Uni-Mol molecular ML workflows that handles representation extraction (embeddings), model training (regression/classification), and property prediction with built-in…
NVIDIA/skills
A skill your agent uses when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches.
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
locbp-uzh/biopipelines
Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Molecular featurization hub (100+ featurizers) for ML. An agent skill from jaechang-hits/SciAgent-Skills. Molfeat Molecular Featurization is an agent skill from jaechang-hits/SciAgent-Skills. Molecular featurization hub (100+ featurizers) for ML.
Molfeat Molecular Featurization fits situations like: tasks that involve Drug discovery and cheminformatics; tasks that involve Embeddings.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/molfeat-molecular-featurization in jaechang-hits/SciAgent-Skills) into .claude/skills/molfeat-molecular-featurization in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/molfeat-molecular-featurization in jaechang-hits/SciAgent-Skills) into .agents/skills/molfeat-molecular-featurization in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/molfeat-molecular-featurization, .gemini/skills/molfeat-molecular-featurization, .github/skills/molfeat-molecular-featurization and .opencode/skills/molfeat-molecular-featurization in your project.
Going by SKILL.md and its folder, Molfeat Molecular Featurization needs the command-line tools its instructions call (uv and pip). Our summary lists: Python 3.
SKILL.md names 4 domains. As links in the text: molfeat-docs.datamol.io, github.com, pypi.org and portal.valencelabs.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Molfeat Molecular Featurization is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Molfeat Molecular Featurization: Nvmolkit Usage (NVIDIA-BioNeMo/bionemo-agent-toolkit, 479 stars), Molfeat (K-Dense-AI/scientific-agent-skills, 48k stars), Unimol (jinzhezenggroup/computational-chemistry-agent-skills, 148 stars) and Nvmolkit Usage (NVIDIA/skills, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.