Therapeutics Data Commons (PyTDC) for AI-ready therapeutic ML datasets and benchmarks; use it when you need standardized dataset loading, meaningful splits (e.g., scaffold/cold-start), and…

MITAuto-check passedData & Analytics

Install Pytdc

skills CLI
$ npx skills add aipoch/medical-research-skills --skill pytdc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills pytdc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Evidence Insight/pytdc' .claude/skills/pytdc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pytdc
GitHub stars
2k
Token cost
~1.7k tokens
SKILL.md length
506 words
Files
8 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Therapeutics Data Commons (PyTDC) for AI-ready therapeutic ML datasets and benchmarks; use it when you need standardized dataset loading, meaningful splits (e.g., scaffold/cold-start), and…

  • Works in 5 steps: Dataset Access Pattern → Splitting Strategies (Key Parameters) → Standardized Evaluation → …
  • You need standardized dataset loading
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 1 more section
  • Runs Python scripts from its folder; calls uv

What it does

Pytdc is an agent skill from aipoch/medical-research-skills. Therapeutics Data Commons (PyTDC) for AI-ready therapeutic ML datasets and benchmarks; use it when you need standardized dataset loading, meaningful splits (e.g., scaffold/cold-start), and consistent evaluation for ADME/Toxicity/DTI/DDI or molecular optimization.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `pytdc_audit_result_v1.json`, `references/datasets.md` and `references/oracles.md`).

It sits in Data & Analytics. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • You need standardized dataset loading
  • Meaningful splits (e.g.
  • Scaffold/cold-start)
  • Consistent evaluation for ADME/Toxicity/DTI/DDI

Example prompts

  • “Use the pytdc skill to therapeutic Data Commons (PyTDC) for AI-ready therapeutic ML datasets and benchmarks; use it when you need standardized…”
  • “/pytdc”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Dataset Access Pattern
  2. Splitting Strategies (Key Parameters)
  3. Standardized Evaluation
  4. Data Schemas (What Columns to Expect)
  5. Oracles for Molecular Optimization

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pytdc loads about 1.7k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 506 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 506 words, ~1,749 tokens.

Download SKILL.mdSave it as .claude/skills/pytdc/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
pytdc
description
Therapeutics Data Commons (PyTDC) for AI-ready therapeutic ML datasets and benchmarks; use it when you need standardized dataset loading, meaningful splits (e.g., scaffold/cold-start), and consistent evaluation for ADME/Toxicity/DTI/DDI or molecular optimization.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

  • You need curated, AI-ready datasets for drug discovery tasks (e.g., ADME, toxicity, bioactivity, DTI/DDI).
  • You want standardized benchmarks with consistent evaluation protocols (including multi-seed benchmark groups).
  • You require meaningful data splits such as scaffold splits (chemical diversity) or cold-start splits (unseen drugs/targets).
  • You are building models for single-instance prediction (molecular/protein properties) or multi-instance prediction (interactions).
  • You are doing goal-directed molecular generation and need property Oracles for scoring/optimization.

Key Features

  • Dataset access by task type
    • Single-instance prediction: ADME, Toxicity (Tox), HTS, QM, and more.
    • Multi-instance prediction: DTI, DDI, PPI, and more.
    • Generation: MolGen, RetroSyn, PairMolGen.
  • Standardized splitting utilities
    • random, scaffold, and cold-start variants such as cold_drug, cold_target, cold_drug_target, plus temporal where applicable.
  • Unified evaluation
    • Built-in Evaluator with common metrics (ROC-AUC, PR-AUC, RMSE, MAE, Spearman, etc.).
  • Benchmark groups
    • Curated collections (e.g., ADMET group) with recommended evaluation protocols (commonly 5 seeds).
  • Chem/data utilities
    • Molecular format conversion (e.g., SMILES → PyG), filtering, balancing, negative sampling, entity retrieval (CID→SMILES, UniProt→sequence).
  • Molecular Oracles
    • Property scoring functions usable for goal-directed generation workflows (see references/oracles.md).

Dependencies

Install (recommended):

bash
uv pip install PyTDC

Upgrade:

bash
uv pip install PyTDC --upgrade

Core runtime dependencies (installed automatically; versions depend on the PyTDC release you install):

  • PyTDC (latest from PyPI)
  • numpy
  • pandas
  • scikit-learn
  • tqdm
  • seaborn
  • fuzzywuzzy

Optional dependencies may be pulled in automatically depending on which submodules you use (e.g., graph backends or chemistry toolchains).

Example Usage

A complete runnable example that:

  1. loads an ADME dataset,
  2. performs a scaffold split,
  3. trains a simple baseline model,
  4. evaluates with a standard metric,
  5. queries an Oracle score for a SMILES.
python
# pip install PyTDC scikit-learn

from tdc.single_pred import ADME
from tdc import Evaluator, Oracle

from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import Ridge

def main():
    # 1) Load a single-instance prediction dataset (ADME)
    data = ADME(name="Caco2_Wang")

    # 2) Create a scaffold split (train/valid/test)
    split = data.get_split(method="scaffold", seed=42, frac=[0.7, 0.1, 0.2])
    train, valid, test = split["train"], split["valid"], split["test"]

    # 3) Train a simple baseline model on SMILES strings
    #    (character n-gram + ridge regression; replace with your own model)
    model = Pipeline(
        steps=[
            ("featurizer", CountVectorizer(analyzer="char", ngram_range=(2, 5))),
            ("regressor", Ridge(alpha=1.0)),
        ]
    )
    model.fit(train["Drug"], train["Y"])

    # 4) Evaluate on the test set using a TDC Evaluator
    y_pred = model.predict(test["Drug"])
    evaluator = Evaluator(name="MAE")
    mae = evaluator(test["Y"], y_pred)
    print(f"Test MAE: {mae:.4f}")

    # 5) Oracle scoring example (property scoring for a SMILES)
    oracle = Oracle(name="DRD2")
    score = oracle("CC(C)Cc1ccc(cc1)C(C)C(O)=O")
    print(f"DRD2 Oracle score: {score}")

if __name__ == "__main__":
    main()

Related references and templates (if present in this skill package):

  • Oracle catalog and usage: references/oracles.md
  • Utility functions (splits, processing, retrieval): references/utilities.md
  • Dataset catalog: references/datasets.md
  • Workflow templates: scripts/load_and_split_data.py, scripts/benchmark_evaluation.py, scripts/molecular_generation.py

Implementation Details

1) Dataset Access Pattern

PyTDC datasets follow a consistent interface:

python
from tdc.<problem> import <Task>

data = <Task>(name="<DatasetName>")
df = data.get_data(format="df")
split = data.get_split(method="scaffold", seed=1, frac=[0.7, 0.1, 0.2])
  • <problem> is typically one of:
    • single_pred (single-entity property prediction)
    • multi_pred (pairwise/multi-entity interaction prediction)
    • generation (molecule/reaction generation tasks)
Show full SKILL.md (194 more words)Show less
2) Splitting Strategies (Key Parameters)

Use get_split(...) to obtain {"train": ..., "valid": ..., "test": ...}.

Common parameters:

  • method: split strategy
  • seed: random seed for reproducibility
  • frac: [train, valid, test] fractions (when supported)

Typical methods:

  • random: random shuffling split
  • scaffold: Bemis–Murcko scaffold-based split to reduce scaffold leakage and improve chemical generalization
  • Cold-start (commonly for interaction tasks like DTI/DDI):
    • cold_drug: test contains unseen drugs
    • cold_target: test contains unseen targets
    • cold_drug_target: test contains unseen drugs and targets
  • temporal: time-based split for datasets with timestamps (when available)

Example:

python
split = data.get_split(method="cold_target", seed=1)
3) Standardized Evaluation

TDC provides a unified evaluator:

python
from tdc import Evaluator

evaluator = Evaluator(name="ROC-AUC")  # classification
score = evaluator(y_true, y_pred)

Choose metrics appropriate to the task type:

  • Classification: ROC-AUC, PR-AUC, F1, Accuracy, etc.
  • Regression: RMSE, MAE, R2, Spearman, Pearson, etc.
4) Data Schemas (What Columns to Expect)

While schemas vary by task, common conventions include:

  • Single-instance prediction (e.g., ADME/Tox):

    • Drug (often SMILES) and label Y
    • sometimes Drug_ID / Compound_ID
  • Multi-instance prediction (e.g., DTI):

    • Drug (SMILES), Target (protein sequence), label Y
    • plus identifiers such as Drug_ID, Target_ID
5) Oracles for Molecular Optimization

Oracles provide a callable scoring interface:

python
from tdc import Oracle

oracle = Oracle(name="GSK3B")
score = oracle("CCO...")
scores = oracle(["SMILES1", "SMILES2"])

Use Oracles to:

  • score candidate molecules during generation,
  • define optimization objectives,
  • compare molecules under consistent property predictors.

For the full list of Oracles and their expected inputs/outputs, see references/oracles.md.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in scientific-skills/Evidence Insight/pytdc of aipoch/medical-research-skills.

  • SKILL.md
  • pytdc_audit_result_v1.json
  • references/datasets.md
  • references/oracles.md
  • references/utilities.md
  • scripts/benchmark_evaluation.py
  • scripts/load_and_split_data.py
  • scripts/molecular_generation.py

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Pytdc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pytdc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pytdc this skillaipoch/medical-research-skills2k—~1.7kAutomated safety check: PassMIT
Exploratory Data Analysisspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: PassMIT
Scientific Figure MakingChenLiu-1996/figures4papers8.3k—~557Automated safety check: PassCustom licence
Academic Figure SkillTingxiYu/academic-figure-skill4801 repos~7kAutomated safety check: PassApache-2.0
Statistical Powerspacering-net/codeg3.9k1 repos~3.6kAutomated safety check: NotesMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0

Similar skills

  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Figure Making

    ChenLiu-1996/figures4papers

    Covers publication-ready matplotlib figures for academic papers, slides, and reports—bars, trends, scatter, heatmaps, and multi-panel layouts—with this…

    8.3k GitHub stars~557 tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed
  • Academic Figure Skill

    TingxiYu/academic-figure-skill

    Academic-grade scientific figure creation for Nature/Cell/Science journals.

    480 GitHub starsUsed in 1 repo~7k tokens
    Data & AnalyticsAuto-check passed
  • Statistical Power

    spacering-net/codeg

    Sample-size and statistical power calculations for planning studies.

    3.9k GitHub starsUsed in 1 repo~3.6k tokens
    Data & AnalyticsAuto-check: notes
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 11 days ago
    Data & AnalyticsAuto-check: notes
  • Mathodology Figure Presets

    sweetcornna/mathodology

    A skill your agent uses when selecting, designing, generating or reviewing scientific figures, complex modeling charts, paper illustrations or image2-assisted visuals.

    303 GitHub stars~1k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 22 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 22 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 22 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 22 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 22 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed

Questions about Pytdc

What does Pytdc do?

Therapeutics Data Commons (PyTDC) for AI-ready therapeutic ML datasets and benchmarks; use it when you need standardized dataset loading, meaningful splits (e.g., scaffold/cold-start), and…. Pytdc is an agent skill from aipoch/medical-research-skills., scaffold/cold-start), and consistent evaluation for ADME/Toxicity/DTI/DDI or molecular optimization.

When should I use Pytdc?

Pytdc fits situations like: you need standardized dataset loading; meaningful splits (e.g; scaffold/cold-start); consistent evaluation for ADME/Toxicity/DTI/DDI.

How do I install Pytdc in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill pytdc -a claude-code`. Or copy the skill folder (scientific-skills/Evidence Insight/pytdc in aipoch/medical-research-skills) into .claude/skills/pytdc in your project. Claude Code loads it when a task matches its description.

How do I install Pytdc in Codex?

Run `npx skills add aipoch/medical-research-skills --skill pytdc -a codex`. Or copy the skill folder (scientific-skills/Evidence Insight/pytdc in aipoch/medical-research-skills) into .agents/skills/pytdc in your project. Codex loads it when a task matches its description.

Can I use Pytdc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill pytdc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pytdc, .gemini/skills/pytdc, .github/skills/pytdc and .opencode/skills/pytdc in your project.

What does Pytdc need to run?

Going by SKILL.md and its folder, Pytdc needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Pytdc access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Pytdc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pytdc use?

Pytdc is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pytdc use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.3k tokens, read only when the agent opens those files.

What are the alternatives to Pytdc?

Skills that share tags, products or a category with Pytdc: Exploratory Data Analysis (spacering-net/codeg, 3.9k stars), Scientific Figure Making (ChenLiu-1996/figures4papers, 8.3k stars), Academic Figure Skill (TingxiYu/academic-figure-skill, 480 stars) and Statistical Power (spacering-net/codeg, 3.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pytdc?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,978 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.