Molecode
AtomFlow-AI/MoleCode
A skill your agent uses for deterministic molecule understanding, graph-level editing, generation, and validation with MoleCode — an explicit Mermaid graph in which every atom and bond is a typed…
Therapeutics Data Commons (TDC) AI-ready drug discovery datasets.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pytdc-therapeutics-data-commons --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons .claude/skills/pytdc-therapeutics-data-commons && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pytdc-therapeutics-data-commons" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons into .claude/skills/pytdc-therapeutics-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytdc-therapeutics-data-commons", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commonsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pytdc-therapeutics-data-commons --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons .agents/skills/pytdc-therapeutics-data-commons && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pytdc-therapeutics-data-commons" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons into .agents/skills/pytdc-therapeutics-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytdc-therapeutics-data-commons", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pytdc-therapeutics-data-commons --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons .cursor/skills/pytdc-therapeutics-data-commons && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pytdc-therapeutics-data-commons" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons into .cursor/skills/pytdc-therapeutics-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytdc-therapeutics-data-commons", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pytdc-therapeutics-data-commons --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons .gemini/skills/pytdc-therapeutics-data-commons && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pytdc-therapeutics-data-commons" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons into .gemini/skills/pytdc-therapeutics-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytdc-therapeutics-data-commons", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills pytdc-therapeutics-data-commonsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons .github/skills/pytdc-therapeutics-data-commons && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pytdc-therapeutics-data-commons" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons into .github/skills/pytdc-therapeutics-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytdc-therapeutics-data-commons", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pytdc-therapeutics-data-commons --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons .opencode/skills/pytdc-therapeutics-data-commons && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pytdc-therapeutics-data-commons" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons into .opencode/skills/pytdc-therapeutics-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytdc-therapeutics-data-commons", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pytdc-therapeutics-data-commonsTherapeutics Data Commons (TDC) AI-ready drug discovery datasets.
Pytdc Therapeutics Data Commons is an agent skill from jaechang-hits/SciAgent-Skills. Therapeutics Data Commons (TDC) AI-ready drug discovery datasets. Curated ADME, toxicity, DTI, DDI with scaffold/cold splits, standardized metrics, molecular oracles, and ADMET benchmarks for therapeutic ML and property prediction. For chemical database queries use chembl-database-bioactivity; for featurization use molfeat.
Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/datasets_catalog.md` and `references/oracles_utilities.md`).
It sits in Research & Science, covering Drug discovery and cheminformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
tdcommons.aitdc.readthedocs.iogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pytdc Therapeutics Data Commons loads about 4.5k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 1,073 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its MIT licence (© jaechang-hits). 1,073 words, ~4,458 tokens.
.claude/skills/pytdc-therapeutics-data-commons/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.PyTDC is an open-science platform providing AI-ready datasets and benchmarks for drug discovery. It organizes therapeutics data into three categories: single-instance prediction (molecular/protein properties), multi-instance prediction (drug-target interactions), and generation (molecule design, retrosynthesis). All datasets come with standardized splits, evaluation metrics, and molecular oracles.
chembl-database-bioactivity insteadmolfeat insteaduv pip install PyTDC
# Core deps: numpy, pandas, scikit-learn, tqdm, fuzzywuzzy
# Optional: rdkit (scaffold splits), torch-geometric (PyG conversion)API Note: TDC downloads datasets on first access (~10-500 MB per dataset). Specify path='data/' to control download location. No API key required.
from tdc.single_pred import ADME
from tdc import Evaluator
# Load dataset with scaffold split
data = ADME(name='Caco2_Wang')
split = data.get_split(method='scaffold', seed=42, frac=[0.7, 0.1, 0.2])
train, valid, test = split['train'], split['valid'], split['test']
print(f"Train: {len(train)}, Valid: {len(valid)}, Test: {len(test)}")
# Train: ~640, Valid: ~91, Test: ~182
# Evaluate predictions
evaluator = Evaluator(name='MAE')
# score = evaluator(test['Y'].values, predictions)Load datasets for predicting properties of individual molecules or proteins.
from tdc.single_pred import ADME, Tox, HTS, QM
# ADME — pharmacokinetic properties
data = ADME(name='Caco2_Wang') # Intestinal permeability (regression)
data = ADME(name='BBB_Martins') # Blood-brain barrier (binary)
data = ADME(name='Lipophilicity_AstraZeneca') # LogD (regression)
data = ADME(name='Solubility_AqSolDB') # Aqueous solubility
# Toxicity — adverse effects
data = Tox(name='hERG') # Cardiotoxicity (binary)
data = Tox(name='AMES') # Mutagenicity (binary)
data = Tox(name='DILI') # Drug-induced liver injury
data = Tox(name='ClinTox') # Clinical trial toxicity
# Access data as DataFrame
df = data.get_data(format='df')
print(df.columns.tolist())
# ['Drug_ID', 'Drug', 'Y'] — Drug is SMILES, Y is target label
print(f"Dataset size: {len(df)}, Label range: [{df['Y'].min():.2f}, {df['Y'].max():.2f}]")Other single-prediction tasks: HTS (screening), QM (quantum mechanics), Yields, Epitope, Develop, CRISPROutcome.
Load datasets for predicting interactions between pairs of biomedical entities.
from tdc.multi_pred import DTI, DDI, PPI
# Drug-Target Interaction — binding affinity
data = DTI(name='BindingDB_Kd') # 52,284 pairs, Kd values
data = DTI(name='DAVIS') # 30,056 pairs, kinase binding
data = DTI(name='KIBA') # 118,254 pairs, kinase bioactivity
# Drug-Drug Interaction — interaction type prediction
data = DDI(name='DrugBank') # 191,808 pairs, 86 interaction types
# Protein-Protein Interaction
data = PPI(name='HuRI')
# Multi-instance data format
df = data.get_data(format='df')
print(df.columns.tolist())
# ['Drug_ID', 'Drug', 'Target_ID', 'Target', 'Y']
# Drug=SMILES, Target=protein sequence, Y=binding affinity or classOther multi-instance tasks: GDA, DrugRes, DrugSyn, PeptideMHC, AntibodyAff, MTI, Catalyst, TrialOutcome.
Load training sets and oracles for molecule generation and retrosynthesis.
from tdc.generation import MolGen, RetroSyn, PairMolGen
from tdc import Oracle
# Molecule generation — training data
data = MolGen(name='ChEMBL_V29') # 1.6M drug-like SMILES
split = data.get_split()
train_smiles = split['train']['Drug'].tolist()
# Oracle scoring — evaluate generated molecules
oracle = Oracle(name='GSK3B') # GSK3B inhibition predictor (0-1)
score = oracle('CC(C)Cc1ccc(cc1)C(C)C(O)=O')
print(f"GSK3B score: {score:.4f}")
# Batch evaluation
scores = oracle(['CCO', 'c1ccccc1', 'CC(=O)O'])
print(f"Batch scores: {scores}")
# Retrosynthesis — reaction prediction
data = RetroSyn(name='USPTO') # 1.9M reactions
split = data.get_split()
# Paired generation — prodrug design
data = PairMolGen(name='Prodrug')Apply meaningful data splits and standardized evaluation metrics.
from tdc.single_pred import ADME
from tdc.multi_pred import DTI
from tdc import Evaluator
# Scaffold split — ensures chemical diversity between sets
data = ADME(name='Caco2_Wang')
split = data.get_split(method='scaffold', seed=42, frac=[0.7, 0.1, 0.2])
# Cold splits — for DTI (unseen drugs/targets in test set)
data = DTI(name='BindingDB_Kd')
cold_drug = data.get_split(method='cold_drug', seed=1)
cold_target = data.get_split(method='cold_target', seed=1)
# Verify no overlap in cold split
train_drugs = set(cold_drug['train']['Drug_ID'])
test_drugs = set(cold_drug['test']['Drug_ID'])
print(f"Drug overlap: {len(train_drugs & test_drugs)}") # 0
# Evaluation metrics
eval_mae = Evaluator(name='MAE')
eval_auc = Evaluator(name='ROC-AUC')
eval_spearman = Evaluator(name='Spearman')
# score = eval_mae(y_true, y_pred)Available split methods: random, scaffold (Bemis-Murcko), cold_drug, cold_target, cold_drug_target, temporal.
Available metrics: Classification — ROC-AUC, PR-AUC, F1, Accuracy, Kappa. Regression — RMSE, MAE, R2, MSE. Ranking — Spearman, Pearson. Multi-label — Micro-F1, Macro-F1.
Run standardized multi-seed evaluation protocols for model comparison.
from tdc.benchmark_group import admet_group
# Load ADMET benchmark (22 datasets)
group = admet_group(path='data/')
# Standard 5-seed evaluation protocol
benchmark = group.get('Caco2_Wang')
predictions = {}
for seed in [1, 2, 3, 4, 5]:
train_df = benchmark['train']
valid_df = benchmark['valid']
test_df = benchmark['test']
# Train your model on train_df, tune on valid_df
# predictions[seed] = model.predict(test_df['Drug'])
predictions[seed] = test_df['Y'].values # placeholder
# Get benchmark results
results = group.evaluate(predictions)
print(f"Mean MAE: {results['Caco2_Wang'][0]:.4f} ± {results['Caco2_Wang'][1]:.4f}")| Category | Import Path | Task Examples | Data Format |
|---|---|---|---|
| Single-Instance | tdc.single_pred | ADME, Tox, HTS, QM | Drug (SMILES) + Y (label) |
| Multi-Instance | tdc.multi_pred | DTI, DDI, PPI, DrugSyn | Drug + Target + Y |
| Generation | tdc.generation | MolGen, RetroSyn | SMILES collections |
| Benchmark | tdc.benchmark_group | admet_group | Curated splits |
| Category | Examples | Speed | Output Range |
|---|---|---|---|
| Biochemical | DRD2, GSK3B, JNK3, 5HT2A | Medium (ML) | 0-1 probability |
| Physicochemical | QED, SA, LogP, MW | Fast (rule-based) | Varies by metric |
| Composite | Isomer_Meta, Median1/2, Rediscovery | Medium | 0-1 combined |
| Specialized | ASKCOS, Docking, Vina | Slow (external) | Varies |
| Utility | Function | Example |
|---|---|---|
| Format conversion | MolConvert(src, dst) | SMILES → PyG, ECFP, SELFIES, DGL |
| Molecule filters | MolFilter(filters) | PAINS, BMS, Glaxo, drug-likeness |
| Label binarization | label_transform() | Continuous → binary at threshold |
| Unit conversion | label_transform(from_unit, to_unit) | nM → pIC50 |
| ID resolution | cid2smiles(), uniprot2seq() | PubChem CID → SMILES |
| Dataset listing | retrieve_dataset_names(task) | List all ADME datasets |
from tdc.single_pred import ADME
from tdc import Evaluator
import numpy as np
data = ADME(name='Caco2_Wang')
evaluator = Evaluator(name='MAE')
results = []
for seed in [1, 2, 3, 4, 5]:
split = data.get_split(method='scaffold', seed=seed)
train, valid, test = split['train'], split['valid'], split['test']
# model.fit(train['Drug'], train['Y'])
# preds = model.predict(test['Drug'])
preds = test['Y'].values + np.random.normal(0, 0.1, len(test)) # placeholder
score = evaluator(test['Y'].values, preds)
results.append(score)
print(f"Seed {seed}: MAE = {score:.4f}")
print(f"Mean MAE: {np.mean(results):.4f} ± {np.std(results):.4f}")from tdc import Oracle
import numpy as np
# Define multi-objective scoring
oracles = {
'QED': (Oracle(name='QED'), 0.3), # drug-likeness
'SA': (Oracle(name='SA'), 0.3), # synthetic accessibility
'GSK3B': (Oracle(name='GSK3B'), 0.4), # target activity
}
test_smiles = ['CC(C)Cc1ccc(cc1)C(C)C(O)=O', 'c1ccc2c(c1)cc1ccc3cccc4ccc2c1c34']
for smi in test_smiles:
scores = {}
weighted_sum = 0
for name, (oracle, weight) in oracles.items():
score = oracle(smi)
scores[name] = score
weighted_sum += score * weight
print(f"SMILES: {smi[:30]}...")
print(f" Scores: {scores}")
print(f" Weighted: {weighted_sum:.4f}")DTI(name='BindingDB_Kd') (Core API Module 2)data.get_split(method='cold_drug', seed=seed) (Core API Module 4)Evaluator(name='Spearman') (Module 4)| Parameter | Function/Module | Default | Range/Options | Effect |
|---|---|---|---|---|
method | get_split() | 'scaffold' | random, scaffold, cold_drug, cold_target, temporal | Split strategy for train/test |
seed | get_split() | 42 | 1-5 for benchmarks | Reproducibility; use 5 seeds for benchmarks |
frac | get_split() | [0.7, 0.1, 0.2] | Sum must equal 1.0 | Train/valid/test proportions |
name | Evaluator() | — | MAE, RMSE, ROC-AUC, Spearman, etc. | Evaluation metric |
name | Oracle() | — | QED, SA, GSK3B, DRD2, etc. | Scoring function for molecules |
src/dst | MolConvert() | — | SMILES, SELFIES, PyG, DGL, ECFP4 | Molecular representation formats |
path | admet_group() | 'data/' | Any directory | Dataset download/cache location |
format | get_data() | 'df' | 'df', 'dict' | Output data format |
cold_drug tests generalization to unseen drugs, cold_target to unseen targetsMolFilter (PAINS, drug-likeness) before training to remove problematic compoundspath — avoids re-downloading large datasets across sessionsfrom tdc.single_pred import ADME
from tdc.utils import retrieve_dataset_names
# List all ADME datasets
datasets = retrieve_dataset_names('ADME')
print(f"Available ADME datasets: {datasets}")
# Load and inspect
data = ADME(name='Caco2_Wang')
df = data.get_data(format='df')
print(f"Size: {len(df)}")
print(f"Label stats: mean={df['Y'].mean():.2f}, std={df['Y'].std():.2f}")
print(f"SMILES example: {df['Drug'].iloc[0]}")from tdc.chem_utils import MolConvert
# SMILES to multiple representations
smiles = 'CC(C)Cc1ccc(cc1)C(C)C(O)=O'
converter_ecfp = MolConvert(src='SMILES', dst='ECFP4')
converter_selfies = MolConvert(src='SMILES', dst='SELFIES')
ecfp = converter_ecfp(smiles)
selfies = converter_selfies(smiles)
print(f"ECFP4 shape: {ecfp.shape}") # (1024,) binary fingerprint
print(f"SELFIES: {selfies}")from tdc import Oracle
# Define property constraints
constraints = {
'QED': (Oracle(name='QED'), 0.5, 1.0), # min, max
'SA': (Oracle(name='SA'), 1.0, 4.0),
'LogP': (Oracle(name='LogP'), -0.5, 5.0),
}
def check_constraints(smiles):
"""Check if molecule satisfies all property constraints."""
results = {}
all_pass = True
for name, (oracle, lo, hi) in constraints.items():
score = oracle(smiles)
passed = lo <= score <= hi
results[name] = {'score': score, 'passed': passed}
all_pass = all_pass and passed
return all_pass, results
passed, details = check_constraints('CC(C)Cc1ccc(cc1)C(C)C(O)=O')
for name, info in details.items():
status = "PASS" if info['passed'] else "FAIL"
print(f"{name}: {info['score']:.3f} [{status}]")| Problem | Cause | Solution |
|---|---|---|
ModuleNotFoundError: tdc | Package not installed | uv pip install PyTDC |
| Scaffold split fails | Missing RDKit dependency | uv pip install rdkit for scaffold decomposition |
| Dataset download timeout | Large dataset or slow connection | Set path='data/' for persistent cache; retry |
KeyError on dataset name | Wrong name or task category | Use retrieve_dataset_names('ADME') to list valid names |
| Oracle returns NaN | Invalid SMILES or RDKit parse failure | Validate SMILES with RDKit MolFromSmiles() first |
| Cold split empty test set | Too few unique entities | Use frac=[0.7, 0.1, 0.2] with larger datasets |
| Benchmark evaluation error | Wrong prediction format | Pass dict with seeds as keys: {1: preds1, 2: preds2, ...} |
| Memory error on large dataset | Full dataset loaded to memory | Process in chunks or use smaller split fractions |
| PyG conversion fails | torch-geometric not installed | uv pip install torch-geometric for graph conversion |
Covers: complete catalog of all TDC datasets organized by task category (single-instance, multi-instance, generation) with dataset names, sizes, label types, and data sources. Relocated inline: top ADME/Tox/DTI datasets with code examples consolidated into Core API Modules 1-2. Omitted: None — all dataset entries preserved in catalog format.
Covers: detailed oracle documentation (all 17+ oracles with parameters, speed tiers, output ranges, custom oracle template) and data processing utilities (format conversion targets, molecule filter types, label transformation, entity resolution). Consolidated from original oracles.md + utilities.md. Relocated inline: Quick Start oracle usage pattern, top evaluation metrics, split methods, MolConvert pattern → Core API Modules 3-4. Omitted: distribution learning KS-test example — niche statistical comparison; leaderboard submission guide — platform-specific.
load_and_split_data.py (215 lines): scaffold/cold split patterns → Core API Module 4; custom split fractions → Key Parameters; evaluation examples → Workflow 1. Thin wrappers around get_split() and Evaluator().benchmark_evaluation.py (328 lines): 5-seed protocol → Core API Module 5 + Workflow 1; multi-dataset evaluation → Workflow 1 pattern; leaderboard guide → omitted (platform-specific).molecular_generation.py (405 lines): single/batch oracle usage → Core API Module 3; multi-objective scoring → Workflow 2; constraint satisfaction → Recipe 3; distribution learning → omitted (niche).© jaechang-hits, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Pytdc Therapeutics Data Commons next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pytdc Therapeutics Data Commons this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~4.5k | Automated safety check: Pass | MIT | |
| MolecodeAtomFlow-AI/MoleCode | 306 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Drug DiscoveryTommy-yw/RunbookHermes | 546 | 1 repos | ~2.3k | Automated safety check: Pass | MIT | |
| DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3k | Automated safety check: Notes | MIT | |
| Biomedical Analysis Dispatchxjtulyc/MedgeClaw | 617 | 1 repos | ~2k | Automated safety check: Pass | None | |
| Biopipelineslocbp-uzh/biopipelines | 109 | — | ~2.4k | Automated safety check: Pass | MIT |
AtomFlow-AI/MoleCode
A skill your agent uses for deterministic molecule understanding, graph-level editing, generation, and validation with MoleCode — an explicit Mermaid graph in which every atom and bond is a typed…
Tommy-yw/RunbookHermes
Pharmaceutical research assistant for drug discovery workflows.
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
xjtulyc/MedgeClaw
Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.
locbp-uzh/biopipelines
Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…
wu-yc/LabClaw
Retrieves chemical compound information from PubChem and ChEMBL with disambiguation, cross-referencing, and quality assessment.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
Therapeutics Data Commons (TDC) AI-ready drug discovery datasets. Pytdc Therapeutics Data Commons is an agent skill from jaechang-hits/SciAgent-Skills. Therapeutics Data Commons (TDC) AI-ready drug discovery datasets.
Pytdc Therapeutics Data Commons fits situations like: tasks that involve Drug discovery and cheminformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons in jaechang-hits/SciAgent-Skills) into .claude/skills/pytdc-therapeutics-data-commons in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons in jaechang-hits/SciAgent-Skills) into .agents/skills/pytdc-therapeutics-data-commons in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pytdc-therapeutics-data-commons, .gemini/skills/pytdc-therapeutics-data-commons, .github/skills/pytdc-therapeutics-data-commons and .opencode/skills/pytdc-therapeutics-data-commons in your project.
Going by SKILL.md and its folder, Pytdc Therapeutics Data Commons needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md names 3 domains. As links in the text: tdcommons.ai, tdc.readthedocs.io and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Pytdc Therapeutics Data Commons is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Pytdc Therapeutics Data Commons: Molecode (AtomFlow-AI/MoleCode, 306 stars), Drug Discovery (Tommy-yw/RunbookHermes, 546 stars), DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars) and Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.