Agent skill

Coreweave Core Workflow B

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Run distributed GPU training jobs on CoreWeave with multi-node PyTorch.

MITAuto-check passedAI & LLM Engineering

Install Coreweave Core Workflow B

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-core-workflow-b -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-core-workflow-b --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/coreweave-core-workflow-b .claude/skills/coreweave-core-workflow-b && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coreweave-core-workflow-b
GitHub stars
2.8k
Token cost
~1.2k tokens
SKILL.md length
264 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Run distributed GPU training jobs on CoreWeave with multi-node PyTorch.

  • Works in 3 steps: Single-Node Multi-GPU Training → Persistent Storage for Training Data → Monitor Training Progress
  • Training models across multiple GPUs
  • SKILL.md covers Overview, Prerequisites, Instructions and Error Handling, plus 4 more sections
  • Calls kubectl

What it does

Coreweave Core Workflow B is an agent skill from jeremylongshore/tons-of-skills-marketplace. Run distributed GPU training jobs on CoreWeave with multi-node PyTorch. Use when training models across multiple GPUs, setting up distributed training, or running fine-tuning jobs on CoreWeave H100 clusters. Trigger with phrases like "coreweave training", "coreweave multi-gpu", "distributed training coreweave", "fine-tune on coreweave".

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Deep learning, GPU and accelerator computing and Fine-tuning. It works with PyTorch. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Training models across multiple GPUs
  • Setting up distributed training
  • Running fine-tuning jobs on CoreWeave H100 clusters
  • With phrases like coreweave training

Example prompts

  • “coreweave training”
  • “coreweave multi-gpu”
  • “distributed training coreweave”
  • “/coreweave-core-workflow-b”

Requirements

  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(kubectl:*), Grep

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Single-Node Multi-GPU Training
  2. Persistent Storage for Training Data
  3. Monitor Training Progress

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(kubectl:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.coreweave.com
    • pytorch.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Coreweave Core Workflow B loads about 1.2k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 264 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 264 words, ~1,237 tokens.

Download SKILL.mdSave it as .claude/skills/coreweave-core-workflow-b/SKILL.md (or your agent's skills folder).
name
coreweave-core-workflow-b
description
Run distributed GPU training jobs on CoreWeave with multi-node PyTorch. Use when training models across multiple GPUs, setting up distributed training, or running fine-tuning jobs on CoreWeave H100 clusters. Trigger with phrases like "coreweave training", "coreweave multi-gpu", "distributed training coreweave", "fine-tune on coreweave".
allowed-tools
Read, Write, Edit, Bash(kubectl:*), Grep
compatibility
Designed for Claude Code
version
1.11.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, gpu-cloud, kubernetes, inference, coreweave

CoreWeave Core Workflow: GPU Training

Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.

Overview

Run distributed GPU training on CoreWeave: single-node multi-GPU and multi-node training with PyTorch DDP, Slurm-on-Kubernetes, and shared storage.

Prerequisites

  • CKS cluster with multi-GPU node pools (8xA100 or 8xH100)
  • Shared storage (CoreWeave PVC or NFS)
  • Training container with PyTorch and NCCL

Instructions

Step 1: Single-Node Multi-GPU Training
yaml
# training-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: llm-finetune
spec:
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: trainer
          image: ghcr.io/myorg/trainer:latest
          command: ["torchrun"]
          args:
            - "--nproc_per_node=8"
            - "train.py"
            - "--model_name=meta-llama/Llama-3.1-8B"
            - "--batch_size=4"
            - "--epochs=3"
          resources:
            limits:
              nvidia.com/gpu: "8"
              memory: 512Gi
              cpu: "64"
          volumeMounts:
            - name: data
              mountPath: /data
            - name: checkpoints
              mountPath: /checkpoints
      volumes:
        - name: data
          persistentVolumeClaim:
            claimName: training-data
        - name: checkpoints
          persistentVolumeClaim:
            claimName: model-checkpoints
      affinity:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
              - matchExpressions:
                  - key: gpu.nvidia.com/class
                    operator: In
                    values: ["A100_NVLINK_A100_SXM4_80GB"]
Step 2: Persistent Storage for Training Data
yaml
# storage.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: training-data
spec:
  accessModes: ["ReadWriteMany"]
  resources:
    requests:
      storage: 500Gi
  storageClassName: shared-hdd-ord1
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: model-checkpoints
spec:
  accessModes: ["ReadWriteMany"]
  resources:
    requests:
      storage: 200Gi
  storageClassName: shared-ssd-ord1
Step 3: Monitor Training Progress
bash
# Watch training logs
kubectl logs -f job/llm-finetune

# Check GPU utilization
kubectl exec -it $(kubectl get pod -l job-name=llm-finetune -o name) -- nvidia-smi

# Check training metrics
kubectl exec -it $(kubectl get pod -l job-name=llm-finetune -o name) -- \
  cat /checkpoints/training_log.json | tail -5

Error Handling

ErrorCauseSolution
NCCL timeoutNetwork issue between GPUsUse NVLink nodes (SXM4/SXM5)
OOMKilledBatch size too largeReduce batch size or use gradient accumulation
Checkpoint save failedPVC fullIncrease storage or prune old checkpoints
Job evictedPreemptionUse on-demand nodes for training

Output

  • A GPU training Job bound to explicitly selected node, storage, and checkpoint resources.
  • A repeatable monitoring trail: pod state, GPU utilization, and training metrics are available to the authorized operator without exposing model inputs or credentials.
  • Durable checkpoints on the approved PVC so a failed or preempted job can resume from a known state rather than silently restarting training.

Examples

Before scheduling a costly multi-GPU run, submit a small trusted smoke job to the same namespace and inspect its scheduling event and GPU allocation:

bash
kubectl apply -f training-job.yaml
kubectl get job llm-finetune --watch
kubectl get pods -l job-name=llm-finetune -o wide
kubectl logs job/llm-finetune --tail=100

If the job cannot schedule, stop before increasing quota or changing node selectors. Confirm the namespace quota, approved GPU class, and PVC binding with the platform owner; preserve the failed event output with secrets and customer data redacted.

Resources

Next Steps

For troubleshooting, see coreweave-common-errors.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/coreweave-core-workflow-b of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Coreweave Core Workflow B next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coreweave Core Workflow B compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coreweave Core Workflow B this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.2kAutomated safety check: PassMIT
ML Training RecipesOrchestra-Research/AI-Research-SKILLs13k1 repos~2.8kAutomated safety check: PassMIT
OpenVLA-OFT Fine-TuningOrchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT
OpenPI Fine-Tuning and ServingOrchestra-Research/AI-Research-SKILLs13k—~3.6kAutomated safety check: PassMIT
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • ML Training Recipes

    Orchestra-Research/AI-Research-SKILLs

    PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM.

    13k GitHub starsUsed in 1 repo~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • OpenVLA-OFT Fine-Tuning

    Orchestra-Research/AI-Research-SKILLs

    Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

    13k GitHub stars~3.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • OpenPI Fine-Tuning and Serving

    Orchestra-Research/AI-Research-SKILLs

    Fine-tunes and serves Physical Intelligence's pi0, pi0-fast and pi0.5 robot policies with JAX or PyTorch, including checkpoint conversion and policy servers.

    13k GitHub stars~3.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Accelerate

    Orchestra-Research/AI-Research-SKILLs

    Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups.

    13k GitHub starsUsed in 5 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated yesterday
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Coreweave Core Workflow B

What does Coreweave Core Workflow B do?

Run distributed GPU training jobs on CoreWeave with multi-node PyTorch. Coreweave Core Workflow B is an agent skill from jeremylongshore/tons-of-skills-marketplace. Run distributed GPU training jobs on CoreWeave with multi-node PyTorch.

When should I use Coreweave Core Workflow B?

Coreweave Core Workflow B fits situations like: training models across multiple GPUs; setting up distributed training; running fine-tuning jobs on CoreWeave H100 clusters; with phrases like coreweave training.

How do I install Coreweave Core Workflow B in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-core-workflow-b -a claude-code`. Or copy the skill folder (skills/.curated/coreweave-core-workflow-b in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/coreweave-core-workflow-b in your project. Claude Code loads it when a task matches its description.

How do I install Coreweave Core Workflow B in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-core-workflow-b -a codex`. Or copy the skill folder (skills/.curated/coreweave-core-workflow-b in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/coreweave-core-workflow-b in your project. Codex loads it when a task matches its description.

Can I use Coreweave Core Workflow B in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-core-workflow-b -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coreweave-core-workflow-b, .gemini/skills/coreweave-core-workflow-b, .github/skills/coreweave-core-workflow-b and .opencode/skills/coreweave-core-workflow-b in your project.

What does Coreweave Core Workflow B need to run?

Going by SKILL.md and its folder, Coreweave Core Workflow B needs the command-line tools its instructions call (kubectl). Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(kubectl:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Coreweave Core Workflow B access the network?

SKILL.md names 2 domains. As links in the text: docs.coreweave.com and pytorch.org. This is read from the text; nothing was executed.

Is Coreweave Core Workflow B safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Coreweave Core Workflow B use?

Coreweave Core Workflow B is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coreweave Core Workflow B use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Coreweave Core Workflow B?

Skills that share tags, products or a category with Coreweave Core Workflow B: ML Training Recipes (Orchestra-Research/AI-Research-SKILLs, 13k stars), OpenVLA-OFT Fine-Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars), OpenPI Fine-Tuning and Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coreweave Core Workflow B?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.