Senior ML Engineer
davila7/claude-code-templates
World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems.
Agent skill
by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs
Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-train --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/08-distributed-training/ray-train .claude/skills/ray-train && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ray-train" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/ray-train into .claude/skills/ray-train/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-train", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/ray-trainType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-train --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/08-distributed-training/ray-train .agents/skills/ray-train && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ray-train" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/ray-train into .agents/skills/ray-train/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-train", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-train --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/08-distributed-training/ray-train .cursor/skills/ray-train && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ray-train" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/ray-train into .cursor/skills/ray-train/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-train", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 08-distributed-training/ray-train--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-train --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/08-distributed-training/ray-train .gemini/skills/ray-train && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ray-train" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/ray-train into .gemini/skills/ray-train/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-train", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-trainInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/08-distributed-training/ray-train .github/skills/ray-train && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ray-train" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/ray-train into .github/skills/ray-train/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-train", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-train --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/08-distributed-training/ray-train .opencode/skills/ray-train && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ray-train" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/08-distributed-training/ray-train into .opencode/skills/ray-train/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-train", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ray-trainScales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.
This skill walks through Ray Train in five workflows. You can wrap existing single-GPU PyTorch code in a TorchTrainer with a ScalingConfig, train a Hugging Face model through TransformersTrainer, run a distributed hyperparameter search with Ray Tune and the ASHAScheduler, add checkpointing so failed workers restart, and move to multiple machines. For the multi-node case it shows connecting to a Ray cluster and starting the head node with ray start.
The text says Ray takes care of distributed coordination, GPU allocation, fault tolerance, checkpointing and metric aggregation, and that the same training code runs on one GPU or many. It also lists situations that favor Ray Train, such as elastic scaling while a job runs. A multi-node reference file adds detail. The excerpt is cut off before the alternatives section.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.ray.iogithub.comforms.gleFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ray Train Distributed Training loads about 2.7k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 332 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 332 words, ~2,673 tokens.
.claude/skills/ray-train/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Ray Train scales machine learning training from single GPU to multi-node clusters with minimal code changes.
Installation:
pip install -U "ray[train]"Basic PyTorch training (single node):
import ray
from ray import train
from ray.train import ScalingConfig
from ray.train.torch import TorchTrainer
import torch
import torch.nn as nn
# Define training function
def train_func(config):
# Your normal PyTorch code
model = nn.Linear(10, 1)
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)
# Prepare for distributed (Ray handles device placement)
model = train.torch.prepare_model(model)
for epoch in range(10):
# Your training loop
output = model(torch.randn(32, 10))
loss = output.sum()
loss.backward()
optimizer.step()
optimizer.zero_grad()
# Report metrics (logged automatically)
train.report({"loss": loss.item(), "epoch": epoch})
# Run distributed training
trainer = TorchTrainer(
train_func,
scaling_config=ScalingConfig(
num_workers=4, # 4 GPUs/workers
use_gpu=True
)
)
result = trainer.fit()
print(f"Final loss: {result.metrics['loss']}")That's it! Ray handles:
Original single-GPU code:
model = MyModel().cuda()
optimizer = torch.optim.Adam(model.parameters())
for epoch in range(epochs):
for batch in dataloader:
loss = model(batch)
loss.backward()
optimizer.step()Ray Train version (scales to multi-GPU/multi-node):
from ray.train.torch import TorchTrainer
from ray import train
def train_func(config):
model = MyModel()
optimizer = torch.optim.Adam(model.parameters())
# Prepare for distributed (automatic device placement)
model = train.torch.prepare_model(model)
dataloader = train.torch.prepare_data_loader(dataloader)
for epoch in range(epochs):
for batch in dataloader:
loss = model(batch)
loss.backward()
optimizer.step()
# Report metrics
train.report({"loss": loss.item()})
# Scale to 8 GPUs
trainer = TorchTrainer(
train_func,
scaling_config=ScalingConfig(num_workers=8, use_gpu=True)
)
trainer.fit()Benefits: Same code runs on 1 GPU or 1000 GPUs
from ray.train.huggingface import TransformersTrainer
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments
def train_func(config):
# Load model and tokenizer
model = AutoModelForCausalLM.from_pretrained("gpt2")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
# Training arguments (HuggingFace API)
training_args = TrainingArguments(
output_dir="./output",
num_train_epochs=3,
per_device_train_batch_size=8,
learning_rate=2e-5,
)
# Ray automatically handles distributed training
from transformers import Trainer
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
)
trainer.train()
# Scale to multi-node (2 nodes × 8 GPUs = 16 workers)
trainer = TransformersTrainer(
train_func,
scaling_config=ScalingConfig(
num_workers=16,
use_gpu=True,
resources_per_worker={"GPU": 1}
)
)
result = trainer.fit()from ray import tune
from ray.train.torch import TorchTrainer
from ray.tune.schedulers import ASHAScheduler
def train_func(config):
# Use hyperparameters from config
lr = config["lr"]
batch_size = config["batch_size"]
model = MyModel()
optimizer = torch.optim.Adam(model.parameters(), lr=lr)
model = train.torch.prepare_model(model)
for epoch in range(10):
# Training loop
loss = train_epoch(model, optimizer, batch_size)
train.report({"loss": loss, "epoch": epoch})
# Define search space
param_space = {
"lr": tune.loguniform(1e-5, 1e-2),
"batch_size": tune.choice([16, 32, 64, 128])
}
# Run 20 trials with early stopping
tuner = tune.Tuner(
TorchTrainer(
train_func,
scaling_config=ScalingConfig(num_workers=4, use_gpu=True)
),
param_space=param_space,
tune_config=tune.TuneConfig(
num_samples=20,
scheduler=ASHAScheduler(metric="loss", mode="min")
)
)
results = tuner.fit()
best = results.get_best_result(metric="loss", mode="min")
print(f"Best hyperparameters: {best.config}")Result: Distributed hyperparameter search across cluster
from ray import train
from ray.train import Checkpoint
def train_func(config):
model = MyModel()
optimizer = torch.optim.Adam(model.parameters())
# Try to resume from checkpoint
checkpoint = train.get_checkpoint()
if checkpoint:
with checkpoint.as_directory() as checkpoint_dir:
state = torch.load(f"{checkpoint_dir}/model.pt")
model.load_state_dict(state["model"])
optimizer.load_state_dict(state["optimizer"])
start_epoch = state["epoch"]
else:
start_epoch = 0
model = train.torch.prepare_model(model)
for epoch in range(start_epoch, 100):
loss = train_epoch(model, optimizer)
# Save checkpoint every 10 epochs
if epoch % 10 == 0:
checkpoint = Checkpoint.from_directory(
train.get_context().get_trial_dir()
)
torch.save({
"model": model.state_dict(),
"optimizer": optimizer.state_dict(),
"epoch": epoch
}, checkpoint.path / "model.pt")
train.report({"loss": loss}, checkpoint=checkpoint)
trainer = TorchTrainer(
train_func,
scaling_config=ScalingConfig(num_workers=8, use_gpu=True)
)
# Automatically resumes from checkpoint if training fails
result = trainer.fit()from ray.train import ScalingConfig
# Connect to Ray cluster
ray.init(address="auto") # Or ray.init("ray://head-node:10001")
# Train across 4 nodes × 8 GPUs = 32 workers
trainer = TorchTrainer(
train_func,
scaling_config=ScalingConfig(
num_workers=32,
use_gpu=True,
resources_per_worker={"GPU": 1, "CPU": 4},
placement_strategy="SPREAD" # Spread across nodes
)
)
result = trainer.fit()Launch Ray cluster:
# On head node
ray start --head --port=6379
# On worker nodes
ray start --address=<head-node-ip>:6379Use Ray Train when:
Key advantages:
Use alternatives instead:
Issue: Ray cluster not connecting
Check ray status:
ray status
# Should show:
# - Nodes: 4
# - GPUs: 32
# - Workers: ReadyIf not connected:
# Restart head node
ray stop
ray start --head --port=6379 --dashboard-host=0.0.0.0
# Restart worker nodes
ray stop
ray start --address=<head-ip>:6379Issue: Out of memory
Reduce workers or use gradient accumulation:
scaling_config=ScalingConfig(
num_workers=4, # Reduce from 8
use_gpu=True
)
# In train_func, accumulate gradients
for i, batch in enumerate(dataloader):
loss = model(batch) / accumulation_steps
loss.backward()
if (i + 1) % accumulation_steps == 0:
optimizer.step()
optimizer.zero_grad()Issue: Slow training
Check if data loading is bottleneck:
import time
def train_func(config):
for epoch in range(epochs):
start = time.time()
for batch in dataloader:
data_time = time.time() - start
# Train...
start = time.time()
print(f"Data loading: {data_time:.3f}s")If data loading is slow, increase workers:
dataloader = DataLoader(dataset, num_workers=8)Multi-node setup: See references/multi-node.md for Ray cluster deployment on AWS, GCP, Kubernetes, and SLURM.
Hyperparameter tuning: See references/hyperparameter-tuning.md for Ray Tune integration, search algorithms (Optuna, HyperOpt), and population-based training.
Custom training loops: See references/custom-loops.md for advanced Ray Train usage, custom backends, and integration with other frameworks.
Supported accelerators:
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in 08-distributed-training/ray-train of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
Ray Train Distributed Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ray Train Distributed Training this skillOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.7k | Automated safety check: Pass | MIT | |
| Senior ML Engineerdavila7/claude-code-templates | 32k | 2 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Technology Selectiondotnet/skills | 5.6k | 1 repos | ~2.1k | Automated safety check: Pass | MIT | |
| GPU OptimizerMathews-Tom/armory | 328 | — | ~3.5k | Automated safety check: Notes | MIT | |
| Formattingbrendanhasz/probflow | 175 | — | ~381 | Automated safety check: Pass | MIT | |
| ML Engineerdavila7/claude-code-templates | 32k | 9 repos | ~2.3k | Automated safety check: Pass | MIT |
davila7/claude-code-templates
World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems.
dotnet/skills
Guides technology selection and implementation of AI and ML features in .NET 8+ applications using ML.NET, Microsoft.Extensions.AI (MEAI), Microsoft Agent Framework (MAF), GitHub Copilot SDK, ONNX…
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
brendanhasz/probflow
Ensure consistent code formatting using the uv package manager and pre-commit.
davila7/claude-code-templates
Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
Works with
Categories
Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery. This skill walks through Ray Train in five workflows. You can wrap existing single-GPU PyTorch code in a TorchTrainer with a ScalingConfig, train a Hugging Face model through TransformersTrainer, run a distributed hyperparameter search with Ray Tune and the ASHAScheduler, add checkpointing so failed workers restart, and move to multiple machines.
Ray Train Distributed Training fits situations like: training a model across several machines instead of one; running a hyperparameter sweep spread over a cluster; making a long training job survive worker failures; porting a single-GPU PyTorch training loop to multiple GPUs.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a claude-code`. Or copy the skill folder (08-distributed-training/ray-train in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/ray-train in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a codex`. Or copy the skill folder (08-distributed-training/ray-train in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/ray-train in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-train -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ray-train, .gemini/skills/ray-train, .github/skills/ray-train and .opencode/skills/ray-train in your project.
Going by SKILL.md and its folder, Ray Train Distributed Training needs the command-line tools its instructions call (pip). Our summary lists: Python with `ray[train]`; GPUs, and a Ray cluster for multi-node runs.
SKILL.md names 3 domains. As links in the text: docs.ray.io, github.com and forms.gle. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ray Train Distributed Training is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ray Train Distributed Training: Senior ML Engineer (davila7/claude-code-templates, 32k stars), Technology Selection (dotnet/skills, 5.6k stars), GPU Optimizer (Mathews-Tom/armory, 328 stars) and Formatting (brendanhasz/probflow, 175 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.