SHAP Model Explainability
davila7/claude-code-templates
Explains machine learning predictions with SHAP: picking the right explainer, computing Shapley values and drawing waterfall, beeswarm, bar and force plots.
Uses Ray Data to read, transform and write large datasets across a cluster for ML training and batch inference, with streaming execution and optional GPU steps.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/05-data-processing/ray-data .claude/skills/ray-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ray-data" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/05-data-processing/ray-data into .claude/skills/ray-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/05-data-processing/ray-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/05-data-processing/ray-data .agents/skills/ray-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ray-data" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/05-data-processing/ray-data into .agents/skills/ray-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/05-data-processing/ray-data .cursor/skills/ray-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ray-data" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/05-data-processing/ray-data into .cursor/skills/ray-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 05-data-processing/ray-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/05-data-processing/ray-data .gemini/skills/ray-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ray-data" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/05-data-processing/ray-data into .gemini/skills/ray-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/05-data-processing/ray-data .github/skills/ray-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ray-data" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/05-data-processing/ray-data into .github/skills/ray-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ray-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/05-data-processing/ray-data .opencode/skills/ray-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ray-data" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/05-data-processing/ray-data into .opencode/skills/ray-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ray-dataUses Ray Data to read, transform and write large datasets across a cluster for ML training and batch inference, with streaming execution and optional GPU steps.
Ray Data is a distributed data library for machine learning workloads, and this skill shows how to use it after installing the ray[data] extra. It covers reading Parquet, CSV, JSON, image and other data from cloud storage or from Python objects, then transforming it with vectorized batch maps, row-by-row maps, filters and group-by aggregations. Streaming execution lets it process data larger than memory.
Further sections cover GPU-accelerated preprocessing, writing results back out as Parquet, repartitioning to control parallelism, and tuning batch size. A worked example passes a dataset into Ray Train's TorchTrainer, and reference files go deeper on integration and transformations. The skill points to Pandas for small single-machine data, Dask for tabular and SQL-like work, and Spark for enterprise ETL.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.ray.iogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ray Data for ML Pipelines loads about 1.8k tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 268 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 268 words, ~1,826 tokens.
.claude/skills/ray-data/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Distributed data processing library for ML and AI workloads.
Use Ray Data when:
Key features:
Use alternatives instead:
pip install -U 'ray[data]'import ray
# Read Parquet files
ds = ray.data.read_parquet("s3://bucket/data/*.parquet")
# Transform data (lazy execution)
ds = ds.map_batches(lambda batch: {"processed": batch["text"].str.lower()})
# Consume data
for batch in ds.iter_batches(batch_size=100):
print(batch)import ray
from ray.train import ScalingConfig
from ray.train.torch import TorchTrainer
# Create dataset
train_ds = ray.data.read_parquet("s3://bucket/train/*.parquet")
def train_func(config):
# Access dataset in training
train_ds = ray.train.get_dataset_shard("train")
for epoch in range(10):
for batch in train_ds.iter_batches(batch_size=32):
# Train on batch
pass
# Train with Ray
trainer = TorchTrainer(
train_func,
datasets={"train": train_ds},
scaling_config=ScalingConfig(num_workers=4, use_gpu=True)
)
trainer.fit()import ray
# Parquet (recommended for ML)
ds = ray.data.read_parquet("s3://bucket/data/*.parquet")
# CSV
ds = ray.data.read_csv("s3://bucket/data/*.csv")
# JSON
ds = ray.data.read_json("gs://bucket/data/*.json")
# Images
ds = ray.data.read_images("s3://bucket/images/")# From list
ds = ray.data.from_items([{"id": i, "value": i * 2} for i in range(1000)])
# From range
ds = ray.data.range(1000000) # Synthetic data
# From pandas
import pandas as pd
df = pd.DataFrame({"col1": [1, 2, 3], "col2": [4, 5, 6]})
ds = ray.data.from_pandas(df)# Batch transformation (fast)
def process_batch(batch):
batch["doubled"] = batch["value"] * 2
return batch
ds = ds.map_batches(process_batch, batch_size=1000)# Row-by-row (slower)
def process_row(row):
row["squared"] = row["value"] ** 2
return row
ds = ds.map(process_row)# Filter rows
ds = ds.filter(lambda row: row["value"] > 100)# Group by column
ds = ds.groupby("category").count()
# Custom aggregation
ds = ds.groupby("category").map_groups(lambda group: {"sum": group["value"].sum()})# Use GPU for preprocessing
def preprocess_images_gpu(batch):
import torch
images = torch.tensor(batch["image"]).cuda()
# GPU preprocessing
processed = images * 255
return {"processed": processed.cpu().numpy()}
ds = ds.map_batches(
preprocess_images_gpu,
batch_size=64,
num_gpus=1 # Request GPU
)# Write to Parquet
ds.write_parquet("s3://bucket/output/")
# Write to CSV
ds.write_csv("output/")
# Write to JSON
ds.write_json("output/")# Control parallelism
ds = ds.repartition(100) # 100 blocks for 100-core cluster# Larger batches = faster vectorized ops
ds.map_batches(process_fn, batch_size=10000) # vs batch_size=100# Process data larger than memory
ds = ray.data.read_parquet("s3://huge-dataset/")
for batch in ds.iter_batches(batch_size=1000):
process(batch) # Streamed, not loaded to memoryimport ray
# Load model
def load_model():
# Load once per worker
return MyModel()
# Inference function
class BatchInference:
def __init__(self):
self.model = load_model()
def __call__(self, batch):
predictions = self.model(batch["input"])
return {"prediction": predictions}
# Run distributed inference
ds = ray.data.read_parquet("s3://data/")
predictions = ds.map_batches(BatchInference, batch_size=32, num_gpus=1)
predictions.write_parquet("s3://output/")# Multi-step pipeline
ds = (
ray.data.read_parquet("s3://raw/")
.map_batches(clean_data)
.map_batches(tokenize)
.map_batches(augment)
.write_parquet("s3://processed/")
)# Convert to PyTorch
torch_ds = ds.to_torch(label_column="label", batch_size=32)
for batch in torch_ds:
# batch is dict with tensors
inputs, labels = batch["features"], batch["label"]# Convert to TensorFlow
tf_ds = ds.to_tf(feature_columns=["image"], label_column="label", batch_size=32)
for features, labels in tf_ds:
# Train model
pass| Format | Read | Write | Use Case |
|---|---|---|---|
| Parquet | ✅ | ✅ | ML data (recommended) |
| CSV | ✅ | ✅ | Tabular data |
| JSON | ✅ | ✅ | Semi-structured |
| Images | ✅ | ❌ | Computer vision |
| NumPy | ✅ | ✅ | Arrays |
| Pandas | ✅ | ❌ | DataFrames |
Scaling (processing 100GB data):
GPU acceleration (image preprocessing):
Production deployments:
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in 05-data-processing/ray-data of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.
Ray Data for ML Pipelines next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ray Data for ML Pipelines this skillOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~1.8k | Automated safety check: Pass | MIT | |
| SHAP Model Explainabilitydavila7/claude-code-templates | 32k | 12 repos | ~4.6k | Automated safety check: Pass | MIT | |
| Technology Selectiondotnet/skills | 5.6k | 2 repos | ~2.1k | Automated safety check: Pass | MIT | |
| Optimize For GPUK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.4k | Automated safety check: Pass | MIT | |
| ML Model Trainingsecondsky/claude-skills | 227 | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Ingesting Dataancoleman/ai-design-components | 526 | — | ~1.9k | Automated safety check: Pass | MIT |
davila7/claude-code-templates
Explains machine learning predictions with SHAP: picking the right explainer, computing Shapley values and drawing waterfall, beeswarm, bar and force plots.
dotnet/skills
Guides technology selection and implementation of AI and ML features in .NET 8+ applications using ML.NET, Microsoft.Extensions.AI (MEAI), Microsoft Agent Framework (MAF), GitHub Copilot SDK, ONNX…
K-Dense-AI/scientific-agent-skills
GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.
secondsky/claude-skills
Train ML models with scikit-learn, PyTorch, TensorFlow. An agent skill from secondsky/claude-skills.
ancoleman/ai-design-components
Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.
ancoleman/ai-design-components
Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Works with
Categories
Uses Ray Data to read, transform and write large datasets across a cluster for ML training and batch inference, with streaming execution and optional GPU steps. Ray Data is a distributed data library for machine learning workloads, and this skill shows how to use it after installing the ray[data] extra. It covers reading Parquet, CSV, JSON, image and other data from cloud storage or from Python objects, then transforming it with vectorized batch maps, row-by-row maps, filters and group-by aggregations.
Ray Data for ML Pipelines fits situations like: preprocessing a dataset too big for one machine before model training; building a batch inference pipeline over images, audio or video; moving a data-prep script from a laptop to a cluster; feeding a Ray Train job from Parquet files in cloud storage.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a claude-code`. Or copy the skill folder (05-data-processing/ray-data in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/ray-data in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a codex`. Or copy the skill folder (05-data-processing/ray-data in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/ray-data in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ray-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ray-data, .gemini/skills/ray-data, .github/skills/ray-data and .opencode/skills/ray-data in your project.
Going by SKILL.md and its folder, Ray Data for ML Pipelines needs the command-line tools its instructions call (pip). Our summary lists: Python with `ray[data]`; A Ray cluster for distributed runs.
SKILL.md names 2 domains. As links in the text: docs.ray.io and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ray Data for ML Pipelines is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 879 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ray Data for ML Pipelines: SHAP Model Explainability (davila7/claude-code-templates, 32k stars), Technology Selection (dotnet/skills, 5.6k stars), Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars) and ML Model Training (secondsky/claude-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,313 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.