Agent skill

Coreweave Common Errors

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors.

MITAuto-check passedAI & LLM Engineering

Install Coreweave Common Errors

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-common-errors -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-common-errors --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/coreweave-common-errors .claude/skills/coreweave-common-errors && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coreweave-common-errors
GitHub stars
2.8k
Token cost
~1.1k tokens
SKILL.md length
393 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors.

  • Works in 7 steps: Pod Stuck Pending -- No GPU Available → CUDA Out of Memory → Image Pull BackOff → …
  • Pods are stuck Pending
  • SKILL.md covers Overview, Prerequisites, Instructions and Error Reference, plus 5 more sections
  • Calls kubectl and jq; needs GH_TOKEN

What it does

Coreweave Common Errors is an agent skill from jeremylongshore/tons-of-skills-marketplace. Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors. Use when pods are stuck Pending, GPUs are not allocated, or experiencing CUDA and NCCL errors. Trigger with phrases like "coreweave error", "coreweave pod pending", "coreweave gpu not found", "coreweave debug", "fix coreweave".

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA and Kubernetes. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Pods are stuck Pending
  • GPUs are not allocated
  • Experiencing CUDA and NCCL errors
  • With phrases like coreweave error

Example prompts

  • “coreweave error”
  • “coreweave pod pending”
  • “coreweave gpu not found”
  • “/coreweave-common-errors”

Requirements

  • Docker
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Bash(kubectl:*), Grep

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Pod Stuck Pending -- No GPU Available
  2. CUDA Out of Memory
  3. Image Pull BackOff
  4. NCCL Timeout (Multi-GPU)
  5. PVC Not Mounting
  6. Node Affinity Mismatch
  7. Service Not Reachable

What it can do on your machine

Read from SKILL.md and the folder at commit 80f86df. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash(kubectl:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.coreweave.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GH_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Coreweave Common Errors loads about 1.1k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 393 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit 80f86df, republished under its MIT licence (© jeremylongshore). 393 words, ~1,062 tokens.

Download SKILL.mdSave it as .claude/skills/coreweave-common-errors/SKILL.md (or your agent's skills folder).
name
coreweave-common-errors
description
Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors. Use when pods are stuck Pending, GPUs are not allocated, or experiencing CUDA and NCCL errors. Trigger with phrases like "coreweave error", "coreweave pod pending", "coreweave gpu not found", "coreweave debug", "fix coreweave".
allowed-tools
Read, Bash(kubectl:*), Grep
compatibility
Designed for Claude Code
version
1.11.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, gpu-cloud, kubernetes, inference, coreweave

CoreWeave Common Errors

Overview

Use this triage guide to classify common GPU, Kubernetes, storage, and connectivity failures before changing capacity or credentials. Capture the smallest redacted evidence set and use a reversible fix in the affected namespace.

Prerequisites

  • Read-only access to the affected namespace, pod events, quota, and node labels.
  • The workload name, expected GPU class, and a named service or platform owner.

Instructions

  1. Identify the pod, Job, or Service and collect its status plus recent events.
  2. Match the symptom to the table below, then validate the proposed cause with the listed read-only command before applying a fix.
  3. Make the smallest namespace-scoped change, verify recovery, and record the redacted event/command outcome in the incident or change record.

Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.

Error Reference

1. Pod Stuck Pending -- No GPU Available
text
kubectl describe pod <pod-name> | grep -A5 Events
# "0/N nodes are available: insufficient nvidia.com/gpu"

Fix: Check GPU availability: kubectl get nodes -l gpu.nvidia.com/class=A100_PCIE_80GB. Try a different GPU type or region.

2. CUDA Out of Memory
torch.cuda.OutOfMemoryError: CUDA out of memory

Fix: Reduce batch size, enable gradient checkpointing, or use a larger GPU (A100-80GB instead of 40GB).

3. Image Pull BackOff

Fix: Create an imagePullSecret:

bash
kubectl create secret docker-registry regcred \
  --docker-server=ghcr.io \
  --docker-username=$GH_USER \
  --docker-password=$GH_TOKEN
4. NCCL Timeout (Multi-GPU)
NCCL error: unhandled system error

Fix: Ensure all GPUs are on the same node (NVLink). For multi-node, use InfiniBand-connected nodes.

5. PVC Not Mounting

Fix: Check storage class availability: kubectl get sc. Use CoreWeave storage classes like shared-hdd-ord1 or shared-ssd-ord1.

Show full SKILL.md (161 more words)Show less
6. Node Affinity Mismatch

Fix: List valid GPU class labels:

bash
kubectl get nodes -o json | jq -r '.items[].metadata.labels["gpu.nvidia.com/class"]' | sort -u
7. Service Not Reachable

Fix: Check Service and Endpoints:

text
kubectl get svc,endpoints <service-name>

Output

  • A classified failure with supporting redacted events and a bounded recovery action.
  • A verified recovery result or a clear escalation to the platform owner.

Error Handling

Triage failureSafe response
Root cause remains unclearStop speculative changes and collect a redacted debug bundle.
Quota or capacity change is requiredObtain the namespace owner approval; do not alter cluster-wide quotas.
Credential failure is suspectedRotate/revoke through the approved secret manager; never print or paste the credential.
Data or model artifact may be corruptQuarantine it and verify checksum before retrying.

Examples

For a Pending GPU pod, collect events and quota before selecting another GPU class:

bash
kubectl -n research describe pod trainer-0
kubectl -n research describe resourcequota
kubectl get nodes -l gpu.nvidia.com/class=A100_PCIE_80GB

If capacity is unavailable, leave the Job unchanged and escalate with the redacted events. Do not remove affinity or quota controls merely to force scheduling.

Resources

Next Steps

For diagnostics, see coreweave-debug-bundle.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/coreweave-common-errors of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit 80f86df

Compare with similar skills

Coreweave Common Errors next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coreweave Common Errors compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coreweave Common Errors this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMIT
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone
Metal Kernelpytorch/pytorch104k—~4.9kAutomated safety check: PassCustom licence

Similar skills

  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Coreweave Common Errors

What does Coreweave Common Errors do?

Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors. Coreweave Common Errors is an agent skill from jeremylongshore/tons-of-skills-marketplace. Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors.

When should I use Coreweave Common Errors?

Coreweave Common Errors fits situations like: pods are stuck Pending; GPUs are not allocated; experiencing CUDA and NCCL errors; with phrases like coreweave error.

How do I install Coreweave Common Errors in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-common-errors -a claude-code`. Or copy the skill folder (skills/.curated/coreweave-common-errors in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/coreweave-common-errors in your project. Claude Code loads it when a task matches its description.

How do I install Coreweave Common Errors in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-common-errors -a codex`. Or copy the skill folder (skills/.curated/coreweave-common-errors in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/coreweave-common-errors in your project. Codex loads it when a task matches its description.

Can I use Coreweave Common Errors in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-common-errors -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coreweave-common-errors, .gemini/skills/coreweave-common-errors, .github/skills/coreweave-common-errors and .opencode/skills/coreweave-common-errors in your project.

What does Coreweave Common Errors need to run?

Going by SKILL.md and its folder, Coreweave Common Errors needs the command-line tools its instructions call (kubectl and jq) and credentials named GH_TOKEN. Our summary lists: Docker. Its frontmatter pre-approves these tools: Read, Bash(kubectl:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Coreweave Common Errors access the network?

SKILL.md names 1 domain. As links in the text: docs.coreweave.com. This is read from the text; nothing was executed.

Is Coreweave Common Errors safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Coreweave Common Errors use?

Coreweave Common Errors is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coreweave Common Errors use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Coreweave Common Errors?

Skills that share tags, products or a category with Coreweave Common Errors: Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Cuda Index Width (pytorch/pytorch, 104k stars), Graphsignal (graphsignal/graphsignal, 257 stars) and LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coreweave Common Errors?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,825 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 9, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.