Agent skill

Molfeat Molecular Featurization

by jaechang-hits in jaechang-hits/SciAgent-Skills

Molecular featurization hub (100+ featurizers) for ML. An agent skill from jaechang-hits/SciAgent-Skills.

Apache-2.0Auto-check passedResearch & Science

Install Molfeat Molecular Featurization

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills molfeat-molecular-featurization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/molfeat-molecular-featurization .claude/skills/molfeat-molecular-featurization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
molfeat-molecular-featurization
GitHub stars
374
Used in
1 other repo
Token cost
~4.3k tokens
SKILL.md length
950 words
Files
3 (incl. references)
Skills in repo
169
Repo updated
First seen
Licence
Apache-2.0

At a glance

Molecular featurization hub (100+ featurizers) for ML. An agent skill from jaechang-hits/SciAgent-Skills.

  • Works in 6 steps: Fingerprint Calculators → Descriptor Calculators → Pharmacophore & Shape Calculators → …
  • Tasks that involve Drug discovery and cheminformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 10 more sections
  • Calls uv and pip

What it does

Molfeat Molecular Featurization is an agent skill from jaechang-hits/SciAgent-Skills. Molecular featurization hub (100+ featurizers) for ML. SMILES to fingerprints (ECFP, MACCS, MAP4), descriptors (RDKit 2D, Mordred), pretrained embeddings (ChemBERTa, GIN, Graphormer), pharmacophores. Scikit-learn compatible with parallelization/caching. For QSAR, virtual screening, similarity, and molecular DL.

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/api_reference.md` and `references/available_featurizers.md`).

It sits in Research & Science, covering Drug discovery and cheminformatics and Embeddings. It works with scikit-learn and RDKit. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Drug discovery and cheminformatics
  • Tasks that involve Embeddings

Example prompts

  • “/molfeat-molecular-featurization”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Fingerprint Calculators
  2. Descriptor Calculators
  3. Pharmacophore & Shape Calculators
  4. Batch Processing with Transformers
  5. Pretrained Model Embeddings
  6. ModelStore — Discovering Featurizers

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • molfeat-docs.datamol.io
    • github.com
    • pypi.org
    • portal.valencelabs.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Molfeat Molecular Featurization loads about 4.3k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 950 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 950 words, ~4,325 tokens.

Download SKILL.mdSave it as .claude/skills/molfeat-molecular-featurization/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
molfeat-molecular-featurization
description
Molecular featurization hub (100+ featurizers) for ML. SMILES to fingerprints (ECFP, MACCS, MAP4), descriptors (RDKit 2D, Mordred), pretrained embeddings (ChemBERTa, GIN, Graphormer), pharmacophores. Scikit-learn compatible with parallelization/caching. For QSAR, virtual screening, similarity, and molecular DL.
license
Apache-2.0

Molfeat — Molecular Featurization Hub

Overview

Molfeat is a comprehensive Python library for molecular featurization that unifies 100+ pre-trained embeddings and hand-crafted featurizers under a scikit-learn compatible API. Convert SMILES strings into numerical representations (fingerprints, descriptors, deep learning embeddings) for QSAR modeling, virtual screening, similarity searching, and chemical space analysis.

When to Use

  • Building QSAR/QSPR models requiring molecular features as input
  • Virtual screening — ranking compound libraries by predicted activity
  • Similarity searching against molecular databases
  • Chemical space analysis — clustering, visualization, dimensionality reduction
  • Deep learning on molecules using pretrained embeddings (ChemBERTa, GIN)
  • Featurization pipelines integrating with scikit-learn or PyTorch
  • Comparing multiple molecular representations for benchmarking
  • For molecular manipulation and filtering use datamol instead; for substructure-based molecular operations use rdkit-cheminformatics

Prerequisites

bash
uv pip install molfeat

# Optional extras for specific featurizer types
uv pip install "molfeat[transformer]"   # ChemBERTa, ChemGPT, MolT5
uv pip install "molfeat[dgl]"           # GIN graph neural networks
uv pip install "molfeat[graphormer]"    # Graphormer models
uv pip install "molfeat[fcd]"           # FCD descriptors
uv pip install "molfeat[map4]"          # MAP4 fingerprints
uv pip install "molfeat[all]"           # All dependencies

Quick Start

python
from molfeat.calc import FPCalculator
from molfeat.trans import MoleculeTransformer

smiles = ["CCO", "CC(=O)O", "c1ccccc1", "CC(C)O"]

# Create fingerprint calculator + transformer
calc = FPCalculator("ecfp", radius=3, fpSize=2048)
transformer = MoleculeTransformer(calc, n_jobs=-1)

# Featurize batch in parallel
features = transformer(smiles)
print(f"Shape: {features.shape}")  # (4, 2048)

# Save configuration for reproducibility
transformer.to_state_yaml_file("featurizer_config.yml")

Key Concepts

Architecture: Calculator → Transformer → Store

Molfeat organizes featurization into three layers:

LayerClassPurposeUse When
Calculatormolfeat.calc.*Single molecule → feature vectorCustom loops, single molecules
Transformermolfeat.trans.MoleculeTransformerBatch processing with parallelizationDatasets, scikit-learn pipelines
Storemolfeat.store.ModelStoreDiscovery and loading of pretrained modelsFinding available featurizers

Calculators are callable: calc("CCO") returns a numpy array. Transformers wrap calculators for batch processing: transformer(smiles_list) returns a 2D array. Pretrained transformers (PretrainedMolTransformer) add batched GPU inference and caching.

Featurizer Selection Guide
TaskRecommendedDimensionsSpeed
General QSARecfp (radius=3)2048Fast
Scaffold similaritymaccs167Very fast
Large-scale screeningmap41024Fast
Interpretable modelsdesc2D (RDKitDescriptors2D)200+Fast
Comprehensive descriptorsmordred1800+Medium
Transfer learningChemBERTa-77M-MLM768Slow*
Graph-based DLgin-supervised-maskingVariableSlow*
Pharmacophorefcfp or cats2D2048 / 21Fast
3D shapeusr / usrcat12 / 60Fast

*First run slow; subsequent runs cached.

State Persistence

Save and reload exact featurizer configuration for reproducibility:

python
# Save
transformer.to_state_yaml_file("config.yml")
transformer.to_state_json_file("config.json")

# Reload
loaded = MoleculeTransformer.from_state_yaml_file("config.yml")

Core API

1. Fingerprint Calculators
python
from molfeat.calc import FPCalculator

# ECFP — most popular, general-purpose
ecfp = FPCalculator("ecfp", radius=3, fpSize=2048)
fp = ecfp("CCO")
print(f"ECFP shape: {fp.shape}")  # (2048,)

# MACCS keys — 167-bit structural keys, fast scaffold similarity
maccs = FPCalculator("maccs")
fp = maccs("c1ccccc1")
print(f"MACCS shape: {fp.shape}")  # (167,)

# Count-based fingerprints (non-binary)
ecfp_count = FPCalculator("ecfp-count", radius=3, fpSize=2048)

# MAP4 — MinHashed atom-pair, efficient for large databases
map4 = FPCalculator("map4")
print(f"MAP4 shape: {map4('CCO').shape}")  # (1024,)

Available fingerprint types: ecfp, fcfp, maccs, rdkit, avalon, pattern, layered, atompair, topological, map4, secfp, erg, estate (and count variants with -count suffix).

2. Descriptor Calculators
python
from molfeat.calc import RDKitDescriptors2D, MordredDescriptors

# RDKit 2D — 200+ named properties (MW, logP, TPSA, etc.)
desc2d = RDKitDescriptors2D()
descriptors = desc2d("CCO")
print(f"2D descriptors: {len(descriptors)}")  # 200+
print(f"Feature names: {desc2d.columns[:5]}")

# Mordred — 1800+ comprehensive descriptors
mordred = MordredDescriptors()
descriptors = mordred("c1ccccc1O")
print(f"Mordred descriptors: {len(descriptors)}")  # 1800+
3. Pharmacophore & Shape Calculators
python
from molfeat.calc import CATSCalculator, USRDescriptors

# CATS — pharmacophore point pair distributions
cats = CATSCalculator(mode="2D", scale="raw")
descriptors = cats("CC(C)Cc1ccc(C)cc1C")
print(f"CATS shape: {descriptors.shape}")  # (21,)

# USR — ultrafast shape recognition
usr = USRDescriptors()
shape = usr("CC(=O)Oc1ccccc1C(=O)O")
print(f"USR shape: {shape.shape}")  # (12,)
4. Batch Processing with Transformers
python
from molfeat.trans import MoleculeTransformer, FeatConcat
from molfeat.calc import FPCalculator

smiles = ["CCO", "CC(=O)O", "c1ccccc1", "CC(C)O", "CCCC"]

# Parallel batch processing
transformer = MoleculeTransformer(FPCalculator("ecfp"), n_jobs=-1)
features = transformer(smiles)
print(f"Batch shape: {features.shape}")  # (5, 2048)

# Concatenate multiple featurizers
concat = FeatConcat([
    FPCalculator("maccs"),      # 167 dims
    FPCalculator("ecfp")        # 2048 dims
])
combo_transformer = MoleculeTransformer(concat, n_jobs=-1)
combo_features = combo_transformer(smiles)
print(f"Combined shape: {combo_features.shape}")  # (5, 2215)

# Error-tolerant processing
safe_transformer = MoleculeTransformer(
    FPCalculator("ecfp"), n_jobs=-1,
    ignore_errors=True, verbose=True
)
features = safe_transformer(["CCO", "invalid", "c1ccccc1"])
# Returns None for failed molecules
5. Pretrained Model Embeddings
python
from molfeat.trans.pretrained import PretrainedMolTransformer

# ChemBERTa — RoBERTa trained on 77M PubChem compounds
chemberta = PretrainedMolTransformer("ChemBERTa-77M-MLM", n_jobs=-1)
embeddings = chemberta(["CCO", "CC(=O)O", "c1ccccc1"])
print(f"ChemBERTa shape: {embeddings.shape}")  # (3, 768)

# GIN — graph neural network pretrained on ChEMBL
gin = PretrainedMolTransformer("gin-supervised-masking", n_jobs=-1)
graph_emb = gin(["CCO", "CC(=O)O"])
print(f"GIN shape: {graph_emb.shape}")
6. ModelStore — Discovering Featurizers
python
from molfeat.store.modelstore import ModelStore

store = ModelStore()
print(f"Total available: {len(store.available_models)}")

# Search for specific model
results = store.search(name="ChemBERTa")
for model in results:
    print(f"  {model.name}: {model.description}")

# View usage and load
card = store.search(name="ChemBERTa-77M-MLM")[0]
card.usage()
transformer = store.load("ChemBERTa-77M-MLM")

Common Workflows

Workflow 1: QSAR Model Building
python
from molfeat.calc import FPCalculator
from molfeat.trans import MoleculeTransformer
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import cross_val_score

# Featurize molecules
transformer = MoleculeTransformer(FPCalculator("ecfp", radius=3), n_jobs=-1)
X = transformer(smiles_train)
print(f"Features shape: {X.shape}")

# Train and evaluate
model = RandomForestRegressor(n_estimators=100)
scores = cross_val_score(model, X, y_train, cv=5, scoring='r2')
print(f"R² = {scores.mean():.3f} ± {scores.std():.3f}")

# Save for deployment
transformer.to_state_yaml_file("production_featurizer.yml")
Workflow 2: Virtual Screening Pipeline
python
from sklearn.ensemble import RandomForestClassifier

# Step 1: Featurize known actives/inactives
transformer = MoleculeTransformer(FPCalculator("ecfp"), n_jobs=-1)
X_train = transformer(train_smiles)

# Step 2: Train classifier
clf = RandomForestClassifier(n_estimators=500, n_jobs=-1)
clf.fit(X_train, train_labels)

# Step 3: Screen library (e.g., 1M compounds)
X_screen = transformer(screening_smiles)
predictions = clf.predict_proba(X_screen)[:, 1]

# Step 4: Rank and select top hits
top_indices = predictions.argsort()[::-1][:1000]
top_hits = [screening_smiles[i] for i in top_indices]
print(f"Top 1000 hits selected from {len(screening_smiles)} compounds")
Workflow 3: Featurizer Benchmarking
python
from molfeat.calc import FPCalculator, RDKitDescriptors2D
from sklearn.metrics import roc_auc_score

featurizers = {
    'ECFP': FPCalculator("ecfp"),
    'MACCS': FPCalculator("maccs"),
    'Descriptors': RDKitDescriptors2D(),
}

for name, calc in featurizers.items():
    transformer = MoleculeTransformer(calc, n_jobs=-1)
    X_train = transformer(smiles_train)
    X_test = transformer(smiles_test)
    clf = RandomForestClassifier(n_estimators=100)
    clf.fit(X_train, y_train)
    auc = roc_auc_score(y_test, clf.predict_proba(X_test)[:, 1])
    print(f"{name}: AUC = {auc:.3f}")

Common Recipes

Recipe: Scikit-learn Pipeline Integration
python
from sklearn.pipeline import Pipeline
from sklearn.ensemble import RandomForestClassifier

pipeline = Pipeline([
    ('featurizer', MoleculeTransformer(FPCalculator("ecfp"), n_jobs=-1)),
    ('classifier', RandomForestClassifier(n_estimators=100))
])
pipeline.fit(smiles_train, y_train)
predictions = pipeline.predict(smiles_test)
python
from sklearn.metrics.pairwise import cosine_similarity

calc = FPCalculator("ecfp")
query_fp = calc("CC(=O)Oc1ccccc1C(=O)O").reshape(1, -1)  # Aspirin

transformer = MoleculeTransformer(calc, n_jobs=-1)
db_fps = transformer(database_smiles)

similarities = cosine_similarity(query_fp, db_fps)[0]
top_k = similarities.argsort()[-10:][::-1]
for i in top_k:
    print(f"  {database_smiles[i]}: {similarities[i]:.3f}")
Recipe: Chunk Processing for Large Datasets
python
import numpy as np

def featurize_chunks(smiles_list, transformer, chunk_size=10000):
    all_features = []
    for i in range(0, len(smiles_list), chunk_size):
        chunk = smiles_list[i:i+chunk_size]
        features = transformer(chunk)
        all_features.append(features)
        print(f"Processed {min(i+chunk_size, len(smiles_list))}/{len(smiles_list)}")
    return np.vstack(all_features)

Key Parameters

ParameterModuleDefaultDescription
methodFPCalculator—Fingerprint type: ecfp, maccs, map4, etc.
radiusFPCalculator3Circular fingerprint radius
fpSizeFPCalculator2048Fingerprint bit length
countingFPCalculatorFalseCount vector instead of binary
n_jobsMoleculeTransformer1Parallel workers (-1 = all cores)
ignore_errorsMoleculeTransformerFalseSkip invalid molecules (returns None)
verboseMoleculeTransformerFalseLog processing details
dtypeMoleculeTransformerfloat64Output type (float32 for memory)
modeCATSCalculator"2D"Distance calculation mode
scaleCATSCalculator"raw"Scaling: raw, num, count

Best Practices

  • Use n_jobs=-1 for parallel processing on all CPU cores — significant speedup for batch featurization
  • Start with ECFP for initial baselines — best general-purpose fingerprint before trying deep learning
  • Use ignore_errors=True for large datasets — invalid SMILES won't crash the pipeline
  • Save configurations with to_state_yaml_file() for reproducibility — recreate exact featurizer later
  • Use float32 when memory matters: MoleculeTransformer(calc, dtype=np.float32)
  • Cache pretrained embeddings — first ChemBERTa/GIN inference is slow, subsequent runs use cache
  • Process in chunks for datasets >100K — prevents memory exhaustion (see Recipes)
  • Combine fingerprints with FeatConcat to capture complementary molecular information

Troubleshooting

ProblemCauseSolution
ValueError: unsupported featurizerUnknown method nameCheck FPCalculator supported types or use ModelStore.search()
ImportError for pretrained modelMissing optional dependencyInstall extras: pip install "molfeat[transformer]" or "molfeat[dgl]"
None in output arrayInvalid SMILES with ignore_errors=TrueFilter results: [f for f in features if f is not None]
Memory error on large datasetToo many molecules at onceProcess in chunks of 10K-50K (see Recipes)
Slow pretrained model inferenceFirst run downloads model weightsNormal — subsequent runs use cache
Shape mismatch in pipelineMixed valid/invalid moleculesEnsure ignore_errors=True and filter None before ML model
Reproducibility issuesDifferent molfeat versionsPin version and save config: transformer.to_state_yaml_file()
Show full SKILL.md (329 more words)Show less
  • datamol-cheminformatics — High-level molecular manipulation (standardization, I/O, conformers)
  • rdkit-cheminformatics — Low-level cheminformatics (substructure, reactions, 3D)
  • scikit-learn — ML models consuming molfeat features

References

Bundled Resources

Main SKILL.md + 2 reference files. Original total: 1,273 lines (SKILL.md 510 + api_reference.md 429 + available_featurizers.md 334). Scripts: none. Examples: 724 lines (examples.md).

references/available_featurizers.md: Complete catalog of all 100+ featurizers organized by category — transformer models, GNNs, descriptors, fingerprints, pharmacophore, shape, scaffold, graph featurizers. Includes dimensions, dependencies, and selection guidance per category. Purely lookup-oriented content preserved as reference.

references/api_reference.md: Detailed API reference for molfeat.calc, molfeat.trans, and molfeat.store modules. Covers SerializableCalculator base class, all calculator subclasses with parameters, MoleculeTransformer methods, PretrainedMolTransformer, FeatConcat, ModelStore/ModelCard API, data type control, and PyTorch integration patterns.

Original file disposition:

  • SKILL.md (510 lines) → Core API modules 1-6, Key Concepts (architecture, selection guide), Quick Start, Workflows 1-3. "Choosing the Right Featurizer" → Key Concepts selection guide table. "Advanced Features" (custom preprocessing, batch processing, caching) → Recipes + Best Practices. "Common Featurizers Reference" table → Key Concepts selection guide. "Performance Tips" → Best Practices. Per-use-case disposition: QSAR Modeling → Workflow 1, Virtual Screening → Workflow 2, Similarity Search → Recipe, Chemical Space → When to Use bullet, scikit-learn Pipeline → Recipe, Featurizer Comparison → Workflow 3
  • references/api_reference.md (429 lines) → Migrated to new references/api_reference.md. Core patterns (FPCalculator, MoleculeTransformer, basic ModelStore) relocated to SKILL.md Core API modules 1-6. Detailed class methods, SerializableCalculator base class, PrecomputedMolTransformer, and PyTorch integration retained in reference
  • references/available_featurizers.md (334 lines) → Migrated to new references/available_featurizers.md. Top-level summary → Key Concepts selection guide table. Full categorized catalog retained in reference
  • references/examples.md (724 lines) → Fully consolidated inline: installation → Prerequisites; calculator examples → Core API 1-3; transformer examples → Core API 4; pretrained examples → Core API 5; ML integration → Workflows 1-3 + Recipes; advanced patterns (custom preprocessing, caching, chunk processing) → Recipes + Best Practices; troubleshooting → Troubleshooting table. No separate reference file needed — all content absorbed into SKILL.md sections

Retention: ~490 lines (SKILL.md) + ~170 lines (available_featurizers) + ~190 lines (api_reference) = ~850 / 1,273 original (excl. examples.md treated as consolidated) = ~67%. Including examples.md in denominator: ~850 / 1,997 = ~43%.

© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/structural-biology-drug-discovery/molfeat-molecular-featurization of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/api_reference.md
  • references/available_featurizers.md

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Molfeat Molecular Featurization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Molfeat Molecular Featurization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Molfeat Molecular Featurization this skilljaechang-hits/SciAgent-Skills3741 repos~4.3kAutomated safety check: PassApache-2.0
Nvmolkit UsageNVIDIA-BioNeMo/bionemo-agent-toolkit479—~4.4kAutomated safety check: PassApache-2.0
MolfeatK-Dense-AI/scientific-agent-skills48k1 repos~2.4kAutomated safety check: NotesApache-2.0
Unimoljinzhezenggroup/computational-chemistry-agent-skills148—~1.5kAutomated safety check: PassLGPL-3.0-or-later
Nvmolkit UsageNVIDIA/skills3.6k1 repos~4.8kAutomated safety check: PassApache-2.0
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT

Similar skills

  • Nvmolkit Usage

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF…

    479 GitHub stars~4.4k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Molfeat

    K-Dense-AI/scientific-agent-skills

    Featurizes small molecules with Molfeat for QSAR/QSPR, chemical similarity, virtual screening, and molecular ML.

    48k GitHub starsUsed in 1 repo~2.4k tokens
    Research & ScienceAuto-check: notes
  • Unimol

    jinzhezenggroup/computational-chemistry-agent-skills

    A standardized CLI wrapper for Uni-Mol molecular ML workflows that handles representation extraction (embeddings), model training (regression/classification), and property prediction with built-in…

    148 GitHub stars~1.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Nvmolkit Usage

    NVIDIA/skills

    Official

    A skill your agent uses when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches.

    3.6k GitHub starsUsed in 1 repo~4.8k tokens
    Research & ScienceAuto-check passed
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 10 days ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Molfeat Molecular Featurization

What does Molfeat Molecular Featurization do?

Molecular featurization hub (100+ featurizers) for ML. An agent skill from jaechang-hits/SciAgent-Skills. Molfeat Molecular Featurization is an agent skill from jaechang-hits/SciAgent-Skills. Molecular featurization hub (100+ featurizers) for ML.

When should I use Molfeat Molecular Featurization?

Molfeat Molecular Featurization fits situations like: tasks that involve Drug discovery and cheminformatics; tasks that involve Embeddings.

How do I install Molfeat Molecular Featurization in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/molfeat-molecular-featurization in jaechang-hits/SciAgent-Skills) into .claude/skills/molfeat-molecular-featurization in your project. Claude Code loads it when a task matches its description.

How do I install Molfeat Molecular Featurization in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/molfeat-molecular-featurization in jaechang-hits/SciAgent-Skills) into .agents/skills/molfeat-molecular-featurization in your project. Codex loads it when a task matches its description.

Can I use Molfeat Molecular Featurization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill molfeat-molecular-featurization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/molfeat-molecular-featurization, .gemini/skills/molfeat-molecular-featurization, .github/skills/molfeat-molecular-featurization and .opencode/skills/molfeat-molecular-featurization in your project.

What does Molfeat Molecular Featurization need to run?

Going by SKILL.md and its folder, Molfeat Molecular Featurization needs the command-line tools its instructions call (uv and pip). Our summary lists: Python 3.

Does Molfeat Molecular Featurization access the network?

SKILL.md names 4 domains. As links in the text: molfeat-docs.datamol.io, github.com, pypi.org and portal.valencelabs.com. This is read from the text; nothing was executed.

Is Molfeat Molecular Featurization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Molfeat Molecular Featurization use?

Molfeat Molecular Featurization is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Molfeat Molecular Featurization use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.

What are the alternatives to Molfeat Molecular Featurization?

Skills that share tags, products or a category with Molfeat Molecular Featurization: Nvmolkit Usage (NVIDIA-BioNeMo/bionemo-agent-toolkit, 479 stars), Molfeat (K-Dense-AI/scientific-agent-skills, 48k stars), Unimol (jinzhezenggroup/computational-chemistry-agent-skills, 148 stars) and Nvmolkit Usage (NVIDIA/skills, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Molfeat Molecular Featurization?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.