Agent skill

Pytdc Therapeutics Data Commons

by jaechang-hits in jaechang-hits/SciAgent-Skills

Therapeutics Data Commons (TDC) AI-ready drug discovery datasets.

MITAuto-check passedResearch & Science

Install Pytdc Therapeutics Data Commons

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills pytdc-therapeutics-data-commons --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons .claude/skills/pytdc-therapeutics-data-commons && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pytdc-therapeutics-data-commons
GitHub stars
374
Used in
1 other repo
Token cost
~4.5k tokens
SKILL.md length
1,073 words
Files
3 (incl. references)
Skills in repo
169
Repo updated
First seen
Licence
MIT

At a glance

Therapeutics Data Commons (TDC) AI-ready drug discovery datasets.

  • Works in 6 steps: Load DTI dataset —… → Apply cold_drug split —… → Verify zero drug overlap between train… → …
  • Tasks that involve Drug discovery and cheminformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 8 more sections
  • Calls uv

What it does

Pytdc Therapeutics Data Commons is an agent skill from jaechang-hits/SciAgent-Skills. Therapeutics Data Commons (TDC) AI-ready drug discovery datasets. Curated ADME, toxicity, DTI, DDI with scaffold/cold splits, standardized metrics, molecular oracles, and ADMET benchmarks for therapeutic ML and property prediction. For chemical database queries use chembl-database-bioactivity; for featurization use molfeat.

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/datasets_catalog.md` and `references/oracles_utilities.md`).

It sits in Research & Science, covering Drug discovery and cheminformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is MIT.

When your agent uses it

  • Tasks that involve Drug discovery and cheminformatics

Example prompts

  • “Use the pytdc-therapeutics-data-commons skill to therapeutic Data Commons (TDC) AI-ready drug discovery datasets”
  • “/pytdc-therapeutics-data-commons”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Load DTI dataset — DTI(name='BindingDB_Kd') (Core API Module 2)
  2. Apply cold_drug split — data.get_split(method='cold_drug', seed=seed) (Core API Module 4)
  3. Verify zero drug overlap between train and test sets (Module 4 overlap check)
  4. Train model on train set, predict on test set
  5. Evaluate with Spearman correlation — Evaluator(name='Spearman') (Module 4)
  6. Repeat for 5 seeds and report mean ± std (same pattern as Workflow 1)

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • tdcommons.ai
    • tdc.readthedocs.io
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pytdc Therapeutics Data Commons loads about 4.5k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 1,073 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its MIT licence (© jaechang-hits). 1,073 words, ~4,458 tokens.

Download SKILL.mdSave it as .claude/skills/pytdc-therapeutics-data-commons/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
pytdc-therapeutics-data-commons
description
Therapeutics Data Commons (TDC) AI-ready drug discovery datasets. Curated ADME, toxicity, DTI, DDI with scaffold/cold splits, standardized metrics, molecular oracles, and ADMET benchmarks for therapeutic ML and property prediction. For chemical database queries use chembl-database-bioactivity; for featurization use molfeat.
license
MIT

PyTDC (Therapeutics Data Commons)

Overview

PyTDC is an open-science platform providing AI-ready datasets and benchmarks for drug discovery. It organizes therapeutics data into three categories: single-instance prediction (molecular/protein properties), multi-instance prediction (drug-target interactions), and generation (molecule design, retrosynthesis). All datasets come with standardized splits, evaluation metrics, and molecular oracles.

When to Use

  • Loading curated ADME, toxicity, or bioactivity datasets for ML model training
  • Benchmarking drug discovery models with standardized 5-seed evaluation protocols
  • Predicting drug-target or drug-drug interactions with proper cold-split evaluation
  • Generating novel molecules and scoring them with molecular oracles (QED, SA, DRD2, GSK3B)
  • Accessing scaffold-based or temporal train/test splits for pharmaceutical ML
  • Converting molecular representations (SMILES to PyG graphs, ECFP fingerprints, SELFIES)
  • For chemical database queries (compound search, bioactivity), use chembl-database-bioactivity instead
  • For molecular featurization beyond format conversion, use molfeat instead

Prerequisites

bash
uv pip install PyTDC
# Core deps: numpy, pandas, scikit-learn, tqdm, fuzzywuzzy
# Optional: rdkit (scaffold splits), torch-geometric (PyG conversion)

API Note: TDC downloads datasets on first access (~10-500 MB per dataset). Specify path='data/' to control download location. No API key required.

Quick Start

python
from tdc.single_pred import ADME
from tdc import Evaluator

# Load dataset with scaffold split
data = ADME(name='Caco2_Wang')
split = data.get_split(method='scaffold', seed=42, frac=[0.7, 0.1, 0.2])
train, valid, test = split['train'], split['valid'], split['test']
print(f"Train: {len(train)}, Valid: {len(valid)}, Test: {len(test)}")
# Train: ~640, Valid: ~91, Test: ~182

# Evaluate predictions
evaluator = Evaluator(name='MAE')
# score = evaluator(test['Y'].values, predictions)

Core API

Module 1: Single-Instance Prediction — Dataset Access

Load datasets for predicting properties of individual molecules or proteins.

python
from tdc.single_pred import ADME, Tox, HTS, QM

# ADME — pharmacokinetic properties
data = ADME(name='Caco2_Wang')       # Intestinal permeability (regression)
data = ADME(name='BBB_Martins')       # Blood-brain barrier (binary)
data = ADME(name='Lipophilicity_AstraZeneca')  # LogD (regression)
data = ADME(name='Solubility_AqSolDB')         # Aqueous solubility

# Toxicity — adverse effects
data = Tox(name='hERG')              # Cardiotoxicity (binary)
data = Tox(name='AMES')              # Mutagenicity (binary)
data = Tox(name='DILI')              # Drug-induced liver injury
data = Tox(name='ClinTox')           # Clinical trial toxicity

# Access data as DataFrame
df = data.get_data(format='df')
print(df.columns.tolist())
# ['Drug_ID', 'Drug', 'Y'] — Drug is SMILES, Y is target label
print(f"Dataset size: {len(df)}, Label range: [{df['Y'].min():.2f}, {df['Y'].max():.2f}]")

Other single-prediction tasks: HTS (screening), QM (quantum mechanics), Yields, Epitope, Develop, CRISPROutcome.

Module 2: Multi-Instance Prediction — Interaction Datasets

Load datasets for predicting interactions between pairs of biomedical entities.

python
from tdc.multi_pred import DTI, DDI, PPI

# Drug-Target Interaction — binding affinity
data = DTI(name='BindingDB_Kd')      # 52,284 pairs, Kd values
data = DTI(name='DAVIS')             # 30,056 pairs, kinase binding
data = DTI(name='KIBA')              # 118,254 pairs, kinase bioactivity

# Drug-Drug Interaction — interaction type prediction
data = DDI(name='DrugBank')           # 191,808 pairs, 86 interaction types

# Protein-Protein Interaction
data = PPI(name='HuRI')

# Multi-instance data format
df = data.get_data(format='df')
print(df.columns.tolist())
# ['Drug_ID', 'Drug', 'Target_ID', 'Target', 'Y']
# Drug=SMILES, Target=protein sequence, Y=binding affinity or class

Other multi-instance tasks: GDA, DrugRes, DrugSyn, PeptideMHC, AntibodyAff, MTI, Catalyst, TrialOutcome.

Module 3: Generation Tasks — Molecular Design

Load training sets and oracles for molecule generation and retrosynthesis.

python
from tdc.generation import MolGen, RetroSyn, PairMolGen
from tdc import Oracle

# Molecule generation — training data
data = MolGen(name='ChEMBL_V29')     # 1.6M drug-like SMILES
split = data.get_split()
train_smiles = split['train']['Drug'].tolist()

# Oracle scoring — evaluate generated molecules
oracle = Oracle(name='GSK3B')         # GSK3B inhibition predictor (0-1)
score = oracle('CC(C)Cc1ccc(cc1)C(C)C(O)=O')
print(f"GSK3B score: {score:.4f}")

# Batch evaluation
scores = oracle(['CCO', 'c1ccccc1', 'CC(=O)O'])
print(f"Batch scores: {scores}")

# Retrosynthesis — reaction prediction
data = RetroSyn(name='USPTO')         # 1.9M reactions
split = data.get_split()

# Paired generation — prodrug design
data = PairMolGen(name='Prodrug')
Module 4: Data Splits and Evaluation

Apply meaningful data splits and standardized evaluation metrics.

python
from tdc.single_pred import ADME
from tdc.multi_pred import DTI
from tdc import Evaluator

# Scaffold split — ensures chemical diversity between sets
data = ADME(name='Caco2_Wang')
split = data.get_split(method='scaffold', seed=42, frac=[0.7, 0.1, 0.2])

# Cold splits — for DTI (unseen drugs/targets in test set)
data = DTI(name='BindingDB_Kd')
cold_drug = data.get_split(method='cold_drug', seed=1)
cold_target = data.get_split(method='cold_target', seed=1)

# Verify no overlap in cold split
train_drugs = set(cold_drug['train']['Drug_ID'])
test_drugs = set(cold_drug['test']['Drug_ID'])
print(f"Drug overlap: {len(train_drugs & test_drugs)}")  # 0

# Evaluation metrics
eval_mae = Evaluator(name='MAE')
eval_auc = Evaluator(name='ROC-AUC')
eval_spearman = Evaluator(name='Spearman')
# score = eval_mae(y_true, y_pred)

Available split methods: random, scaffold (Bemis-Murcko), cold_drug, cold_target, cold_drug_target, temporal.

Available metrics: Classification — ROC-AUC, PR-AUC, F1, Accuracy, Kappa. Regression — RMSE, MAE, R2, MSE. Ranking — Spearman, Pearson. Multi-label — Micro-F1, Macro-F1.

Module 5: Benchmark Groups

Run standardized multi-seed evaluation protocols for model comparison.

python
from tdc.benchmark_group import admet_group

# Load ADMET benchmark (22 datasets)
group = admet_group(path='data/')

# Standard 5-seed evaluation protocol
benchmark = group.get('Caco2_Wang')
predictions = {}

for seed in [1, 2, 3, 4, 5]:
    train_df = benchmark['train']
    valid_df = benchmark['valid']
    test_df = benchmark['test']
    # Train your model on train_df, tune on valid_df
    # predictions[seed] = model.predict(test_df['Drug'])
    predictions[seed] = test_df['Y'].values  # placeholder

# Get benchmark results
results = group.evaluate(predictions)
print(f"Mean MAE: {results['Caco2_Wang'][0]:.4f} ± {results['Caco2_Wang'][1]:.4f}")

Key Concepts

Dataset Organization
CategoryImport PathTask ExamplesData Format
Single-Instancetdc.single_predADME, Tox, HTS, QMDrug (SMILES) + Y (label)
Multi-Instancetdc.multi_predDTI, DDI, PPI, DrugSynDrug + Target + Y
Generationtdc.generationMolGen, RetroSynSMILES collections
Benchmarktdc.benchmark_groupadmet_groupCurated splits
Oracle Categories
CategoryExamplesSpeedOutput Range
BiochemicalDRD2, GSK3B, JNK3, 5HT2AMedium (ML)0-1 probability
PhysicochemicalQED, SA, LogP, MWFast (rule-based)Varies by metric
CompositeIsomer_Meta, Median1/2, RediscoveryMedium0-1 combined
SpecializedASKCOS, Docking, VinaSlow (external)Varies
Data Processing Utilities
UtilityFunctionExample
Format conversionMolConvert(src, dst)SMILES → PyG, ECFP, SELFIES, DGL
Molecule filtersMolFilter(filters)PAINS, BMS, Glaxo, drug-likeness
Label binarizationlabel_transform()Continuous → binary at threshold
Unit conversionlabel_transform(from_unit, to_unit)nM → pIC50
ID resolutioncid2smiles(), uniprot2seq()PubChem CID → SMILES
Dataset listingretrieve_dataset_names(task)List all ADME datasets

Common Workflows

Workflow 1: Multi-Seed ADME Model Evaluation
python
from tdc.single_pred import ADME
from tdc import Evaluator
import numpy as np

data = ADME(name='Caco2_Wang')
evaluator = Evaluator(name='MAE')

results = []
for seed in [1, 2, 3, 4, 5]:
    split = data.get_split(method='scaffold', seed=seed)
    train, valid, test = split['train'], split['valid'], split['test']
    # model.fit(train['Drug'], train['Y'])
    # preds = model.predict(test['Drug'])
    preds = test['Y'].values + np.random.normal(0, 0.1, len(test))  # placeholder
    score = evaluator(test['Y'].values, preds)
    results.append(score)
    print(f"Seed {seed}: MAE = {score:.4f}")

print(f"Mean MAE: {np.mean(results):.4f} ± {np.std(results):.4f}")
Workflow 2: Multi-Objective Molecular Scoring
python
from tdc import Oracle
import numpy as np

# Define multi-objective scoring
oracles = {
    'QED': (Oracle(name='QED'), 0.3),      # drug-likeness
    'SA': (Oracle(name='SA'), 0.3),         # synthetic accessibility
    'GSK3B': (Oracle(name='GSK3B'), 0.4),   # target activity
}

test_smiles = ['CC(C)Cc1ccc(cc1)C(C)C(O)=O', 'c1ccc2c(c1)cc1ccc3cccc4ccc2c1c34']

for smi in test_smiles:
    scores = {}
    weighted_sum = 0
    for name, (oracle, weight) in oracles.items():
        score = oracle(smi)
        scores[name] = score
        weighted_sum += score * weight
    print(f"SMILES: {smi[:30]}...")
    print(f"  Scores: {scores}")
    print(f"  Weighted: {weighted_sum:.4f}")
Workflow 3: Cold-Split DTI Evaluation (text-only)
  1. Load DTI dataset — DTI(name='BindingDB_Kd') (Core API Module 2)
  2. Apply cold_drug split — data.get_split(method='cold_drug', seed=seed) (Core API Module 4)
  3. Verify zero drug overlap between train and test sets (Module 4 overlap check)
  4. Train model on train set, predict on test set
  5. Evaluate with Spearman correlation — Evaluator(name='Spearman') (Module 4)
  6. Repeat for 5 seeds and report mean ± std (same pattern as Workflow 1)

Key Parameters

ParameterFunction/ModuleDefaultRange/OptionsEffect
methodget_split()'scaffold'random, scaffold, cold_drug, cold_target, temporalSplit strategy for train/test
seedget_split()421-5 for benchmarksReproducibility; use 5 seeds for benchmarks
fracget_split()[0.7, 0.1, 0.2]Sum must equal 1.0Train/valid/test proportions
nameEvaluator()—MAE, RMSE, ROC-AUC, Spearman, etc.Evaluation metric
nameOracle()—QED, SA, GSK3B, DRD2, etc.Scoring function for molecules
src/dstMolConvert()—SMILES, SELFIES, PyG, DGL, ECFP4Molecular representation formats
pathadmet_group()'data/'Any directoryDataset download/cache location
formatget_data()'df''df', 'dict'Output data format

Best Practices

  1. Always use scaffold splits for molecular property prediction — random splits leak structural information and inflate performance metrics
  2. Report 5-seed evaluations with mean ± std — single-seed results are unreliable for method comparison
  3. Use cold splits for interaction prediction — cold_drug tests generalization to unseen drugs, cold_target to unseen targets
  4. Filter molecules early with MolFilter (PAINS, drug-likeness) before training to remove problematic compounds
  5. Normalize oracles appropriately — QED returns 0-1, SA returns 1-10 (lower is better), binding scores vary. Check oracle documentation before combining
  6. Cache datasets by specifying a persistent path — avoids re-downloading large datasets across sessions
Show full SKILL.md (386 more words)Show less

Common Recipes

Recipe 1: Dataset Exploration and Statistics
python
from tdc.single_pred import ADME
from tdc.utils import retrieve_dataset_names

# List all ADME datasets
datasets = retrieve_dataset_names('ADME')
print(f"Available ADME datasets: {datasets}")

# Load and inspect
data = ADME(name='Caco2_Wang')
df = data.get_data(format='df')
print(f"Size: {len(df)}")
print(f"Label stats: mean={df['Y'].mean():.2f}, std={df['Y'].std():.2f}")
print(f"SMILES example: {df['Drug'].iloc[0]}")
Recipe 2: Molecule Format Conversion Pipeline
python
from tdc.chem_utils import MolConvert

# SMILES to multiple representations
smiles = 'CC(C)Cc1ccc(cc1)C(C)C(O)=O'

converter_ecfp = MolConvert(src='SMILES', dst='ECFP4')
converter_selfies = MolConvert(src='SMILES', dst='SELFIES')

ecfp = converter_ecfp(smiles)
selfies = converter_selfies(smiles)
print(f"ECFP4 shape: {ecfp.shape}")    # (1024,) binary fingerprint
print(f"SELFIES: {selfies}")
Recipe 3: Custom Oracle with Constraint Satisfaction
python
from tdc import Oracle

# Define property constraints
constraints = {
    'QED': (Oracle(name='QED'), 0.5, 1.0),       # min, max
    'SA': (Oracle(name='SA'), 1.0, 4.0),
    'LogP': (Oracle(name='LogP'), -0.5, 5.0),
}

def check_constraints(smiles):
    """Check if molecule satisfies all property constraints."""
    results = {}
    all_pass = True
    for name, (oracle, lo, hi) in constraints.items():
        score = oracle(smiles)
        passed = lo <= score <= hi
        results[name] = {'score': score, 'passed': passed}
        all_pass = all_pass and passed
    return all_pass, results

passed, details = check_constraints('CC(C)Cc1ccc(cc1)C(C)C(O)=O')
for name, info in details.items():
    status = "PASS" if info['passed'] else "FAIL"
    print(f"{name}: {info['score']:.3f} [{status}]")

Troubleshooting

ProblemCauseSolution
ModuleNotFoundError: tdcPackage not installeduv pip install PyTDC
Scaffold split failsMissing RDKit dependencyuv pip install rdkit for scaffold decomposition
Dataset download timeoutLarge dataset or slow connectionSet path='data/' for persistent cache; retry
KeyError on dataset nameWrong name or task categoryUse retrieve_dataset_names('ADME') to list valid names
Oracle returns NaNInvalid SMILES or RDKit parse failureValidate SMILES with RDKit MolFromSmiles() first
Cold split empty test setToo few unique entitiesUse frac=[0.7, 0.1, 0.2] with larger datasets
Benchmark evaluation errorWrong prediction formatPass dict with seeds as keys: {1: preds1, 2: preds2, ...}
Memory error on large datasetFull dataset loaded to memoryProcess in chunks or use smaller split fractions
PyG conversion failstorch-geometric not installeduv pip install torch-geometric for graph conversion

Bundled Resources

references/datasets_catalog.md

Covers: complete catalog of all TDC datasets organized by task category (single-instance, multi-instance, generation) with dataset names, sizes, label types, and data sources. Relocated inline: top ADME/Tox/DTI datasets with code examples consolidated into Core API Modules 1-2. Omitted: None — all dataset entries preserved in catalog format.

references/oracles_utilities.md

Covers: detailed oracle documentation (all 17+ oracles with parameters, speed tiers, output ranges, custom oracle template) and data processing utilities (format conversion targets, molecule filter types, label transformation, entity resolution). Consolidated from original oracles.md + utilities.md. Relocated inline: Quick Start oracle usage pattern, top evaluation metrics, split methods, MolConvert pattern → Core API Modules 3-4. Omitted: distribution learning KS-test example — niche statistical comparison; leaderboard submission guide — platform-specific.

Script Disposition
  • load_and_split_data.py (215 lines): scaffold/cold split patterns → Core API Module 4; custom split fractions → Key Parameters; evaluation examples → Workflow 1. Thin wrappers around get_split() and Evaluator().
  • benchmark_evaluation.py (328 lines): 5-seed protocol → Core API Module 5 + Workflow 1; multi-dataset evaluation → Workflow 1 pattern; leaderboard guide → omitted (platform-specific).
  • molecular_generation.py (405 lines): single/batch oracle usage → Core API Module 3; multi-objective scoring → Workflow 2; constraint satisfaction → Recipe 3; distribution learning → omitted (niche).
  • chembl-database-bioactivity — for querying ChEMBL compound/target/activity data directly
  • molfeat — for advanced molecular featurization beyond TDC's built-in MolConvert
  • rdkit-cheminformatics — for molecular manipulation, substructure search, descriptor calculation

References

© jaechang-hits, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/datasets_catalog.md
  • references/oracles_utilities.md

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pytdc Therapeutics Data Commons next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pytdc Therapeutics Data Commons compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pytdc Therapeutics Data Commons this skilljaechang-hits/SciAgent-Skills3741 repos~4.5kAutomated safety check: PassMIT
MolecodeAtomFlow-AI/MoleCode306—~1.9kAutomated safety check: PassMIT
Drug DiscoveryTommy-yw/RunbookHermes5461 repos~2.3kAutomated safety check: PassMIT
DiffDock Molecular DockingK-Dense-AI/scientific-agent-skills48k1 repos~3kAutomated safety check: NotesMIT
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone
Biopipelineslocbp-uzh/biopipelines109—~2.4kAutomated safety check: PassMIT

Similar skills

  • Molecode

    AtomFlow-AI/MoleCode

    A skill your agent uses for deterministic molecule understanding, graph-level editing, generation, and validation with MoleCode — an explicit Mermaid graph in which every atom and bond is a typed…

    306 GitHub stars~1.9k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Drug Discovery

    Tommy-yw/RunbookHermes

    Pharmaceutical research assistant for drug discovery workflows.

    546 GitHub starsUsed in 1 repo~2.3k tokens
    Research & ScienceAuto-check passed
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Research & ScienceAuto-check: notes
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Biopipelines

    locbp-uzh/biopipelines

    Design and run computational protein and ligand workflows on a GPU: binder and enzyme design, de novo backbone generation, inverse folding and sequence redesign, structure prediction, protein-ligand…

    109 GitHub stars~2.4k tokensUpdated 11 days ago
    Research & ScienceAuto-check passed
  • Retrieves chemical compound information from PubChem and ChEMBL with disambiguation, cross-referencing, and quality assessment.

    1.1k GitHub starsUsed in 2 repos~2.3k tokens
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Pytdc Therapeutics Data Commons

What does Pytdc Therapeutics Data Commons do?

Therapeutics Data Commons (TDC) AI-ready drug discovery datasets. Pytdc Therapeutics Data Commons is an agent skill from jaechang-hits/SciAgent-Skills. Therapeutics Data Commons (TDC) AI-ready drug discovery datasets.

When should I use Pytdc Therapeutics Data Commons?

Pytdc Therapeutics Data Commons fits situations like: tasks that involve Drug discovery and cheminformatics.

How do I install Pytdc Therapeutics Data Commons in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a claude-code`. Or copy the skill folder (skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons in jaechang-hits/SciAgent-Skills) into .claude/skills/pytdc-therapeutics-data-commons in your project. Claude Code loads it when a task matches its description.

How do I install Pytdc Therapeutics Data Commons in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a codex`. Or copy the skill folder (skills/structural-biology-drug-discovery/pytdc-therapeutics-data-commons in jaechang-hits/SciAgent-Skills) into .agents/skills/pytdc-therapeutics-data-commons in your project. Codex loads it when a task matches its description.

Can I use Pytdc Therapeutics Data Commons in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill pytdc-therapeutics-data-commons -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pytdc-therapeutics-data-commons, .gemini/skills/pytdc-therapeutics-data-commons, .github/skills/pytdc-therapeutics-data-commons and .opencode/skills/pytdc-therapeutics-data-commons in your project.

What does Pytdc Therapeutics Data Commons need to run?

Going by SKILL.md and its folder, Pytdc Therapeutics Data Commons needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Pytdc Therapeutics Data Commons access the network?

SKILL.md names 3 domains. As links in the text: tdcommons.ai, tdc.readthedocs.io and github.com. This is read from the text; nothing was executed.

Is Pytdc Therapeutics Data Commons safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pytdc Therapeutics Data Commons use?

Pytdc Therapeutics Data Commons is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pytdc Therapeutics Data Commons use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.3k tokens, read only when the agent opens those files.

What are the alternatives to Pytdc Therapeutics Data Commons?

Skills that share tags, products or a category with Pytdc Therapeutics Data Commons: Molecode (AtomFlow-AI/MoleCode, 306 stars), Drug Discovery (Tommy-yw/RunbookHermes, 546 stars), DiffDock Molecular Docking (K-Dense-AI/scientific-agent-skills, 48k stars) and Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pytdc Therapeutics Data Commons?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.