Agent skill

Coreweave Incident Runbook

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Incident response runbook for CoreWeave GPU workload failures.

MITAuto-check passedDevOps & Cloud

Install Coreweave Incident Runbook

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-incident-runbook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-incident-runbook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/coreweave-incident-runbook .claude/skills/coreweave-incident-runbook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coreweave-incident-runbook
GitHub stars
2.8k
Token cost
~1k tokens
SKILL.md length
373 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Incident response runbook for CoreWeave GPU workload failures.

  • Works in 4 steps: Declare severity and scope, then capture… → Stabilize impact with the documented… → Collect only redacted, bounded… → …
  • Inference services are down
  • SKILL.md covers Overview, Prerequisites, Instructions and Triage Steps, plus 7 more sections
  • Calls kubectl

What it does

Coreweave Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Incident response runbook for CoreWeave GPU workload failures. Use when inference services are down, GPUs are unavailable, or responding to production incidents on CoreWeave. Trigger with phrases like "coreweave incident", "coreweave outage", "coreweave runbook", "coreweave service down".

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering Runbooks and postmortems and Incident response. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Inference services are down
  • GPUs are unavailable
  • Responding to production incidents on CoreWeave
  • With phrases like coreweave incident

Example prompts

  • “coreweave incident”
  • “coreweave outage”
  • “coreweave runbook”
  • “/coreweave-incident-runbook”

Requirements

  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Bash(kubectl:*), Grep

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Declare severity and scope, then capture pod, event, node, and service status.
  2. Stabilize impact with the documented scale, failover, or rollback action before root-cause work.
  3. Collect only redacted, bounded diagnostics and escalate hardware or capacity issues through CoreWeave support.
  4. Verify recovery against the service SLO, update stakeholders, and create follow-up work for root cause and prevention.

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash(kubectl:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • status.coreweave.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Coreweave Incident Runbook loads about 1k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 373 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 373 words, ~1,016 tokens.

Download SKILL.mdSave it as .claude/skills/coreweave-incident-runbook/SKILL.md (or your agent's skills folder).
name
coreweave-incident-runbook
description
Incident response runbook for CoreWeave GPU workload failures. Use when inference services are down, GPUs are unavailable, or responding to production incidents on CoreWeave. Trigger with phrases like "coreweave incident", "coreweave outage", "coreweave runbook", "coreweave service down".
allowed-tools
Read, Bash(kubectl:*), Grep
compatibility
Designed for Claude Code
version
1.11.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, gpu-cloud, kubernetes, inference, coreweave

CoreWeave Incident Runbook

Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.

Overview

Respond to GPU workload incidents by stabilizing customer impact, preserving redacted evidence, and restoring a known-good state. The incident commander owns communications and escalation; responders use only the access required for triage.

Prerequisites

  • An incident ID, named commander, affected namespace/service, and on-call route.
  • Authorized read-only cluster access plus a documented production rollback revision.
  • A redaction policy for logs, prompts, model artifacts, and credentials.

Instructions

  1. Declare severity and scope, then capture pod, event, node, and service status.
  2. Stabilize impact with the documented scale, failover, or rollback action before root-cause work.
  3. Collect only redacted, bounded diagnostics and escalate hardware or capacity issues through CoreWeave support.
  4. Verify recovery against the service SLO, update stakeholders, and create follow-up work for root cause and prevention.

Triage Steps

bash
# 1. Check pod status
kubectl get pods -l app=inference -o wide

# 2. Check recent events
kubectl get events --sort-by=.lastTimestamp | tail -20

# 3. Check node status
kubectl get nodes -l gpu.nvidia.com/class -o wide

# 4. Check GPU health
kubectl exec -it $(kubectl get pod -l app=inference -o name | head -1) -- nvidia-smi

Common Incidents

Inference Service Down
  1. Check pod status and events
  2. If OOMKilled: reduce batch size or upgrade GPU
  3. If ImagePullBackOff: check registry credentials
  4. If Pending: check GPU quota and availability
GPU Node Failure
  1. Pods will be rescheduled automatically
  2. If no capacity: scale down non-critical workloads
  3. Contact CoreWeave support for extended outages
Model Loading Failure
  1. Check HuggingFace token secret exists
  2. Verify model name spelling
  3. Check PVC has sufficient storage
  4. Review container logs for download errors
Show full SKILL.md (139 more words)Show less

Rollback

bash
kubectl rollout undo deployment/inference

Output

  • A time-stamped incident record with scope, owner, stabilization action, and redacted evidence.
  • A verified recovery or an explicit escalation with a safe customer-impact mitigation.

Error Handling

Incident complicationRequired response
Rollback failsStop repeated deploy attempts, escalate to the platform owner, and preserve events.
Diagnostics include a secretRestrict distribution, rotate the secret, and recollect redacted evidence.
No GPU capacity is availablePrioritize critical services under approved policy; do not remove tenant quotas.
Recovery cannot meet SLODeclare the continuing impact and use the approved fallback or communication path.

Examples

For a failing production rollout, stabilize first and capture only the relevant evidence:

bash
kubectl -n inference-prod rollout undo deployment/inference
kubectl -n inference-prod rollout status deployment/inference --timeout=10m
kubectl -n inference-prod get events --sort-by=.lastTimestamp | tail -30

Record the revision, outcome, and redacted events in the incident. Do not retry a broken image or expose a troubleshooting endpoint publicly.

Resources

Next Steps

For data handling, see coreweave-data-handling.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/coreweave-incident-runbook of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Coreweave Incident Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coreweave Incident Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coreweave Incident Runbook this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0
Activation Governance Chaos RolloutAli-Marandi/DataSense107—~1.9kAutomated safety check: PassMIT
Incident Response686f6c61/alfred-dev117—~1.1kAutomated safety check: PassMIT
Superset Incident Triagesuperset-sh/superset15k—~1kAutomated safety check: PassCustom licence
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT

Similar skills

  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Incident Response

    686f6c61/alfred-dev

    Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.

    117 GitHub stars~1.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Coreweave Incident Runbook

What does Coreweave Incident Runbook do?

Incident response runbook for CoreWeave GPU workload failures. Coreweave Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Incident response runbook for CoreWeave GPU workload failures.

When should I use Coreweave Incident Runbook?

Coreweave Incident Runbook fits situations like: inference services are down; GPUs are unavailable; responding to production incidents on CoreWeave; with phrases like coreweave incident.

How do I install Coreweave Incident Runbook in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-incident-runbook -a claude-code`. Or copy the skill folder (skills/.curated/coreweave-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/coreweave-incident-runbook in your project. Claude Code loads it when a task matches its description.

How do I install Coreweave Incident Runbook in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-incident-runbook -a codex`. Or copy the skill folder (skills/.curated/coreweave-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/coreweave-incident-runbook in your project. Codex loads it when a task matches its description.

Can I use Coreweave Incident Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-incident-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coreweave-incident-runbook, .gemini/skills/coreweave-incident-runbook, .github/skills/coreweave-incident-runbook and .opencode/skills/coreweave-incident-runbook in your project.

What does Coreweave Incident Runbook need to run?

Going by SKILL.md and its folder, Coreweave Incident Runbook needs the command-line tools its instructions call (kubectl). Its frontmatter pre-approves these tools: Read, Bash(kubectl:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Coreweave Incident Runbook access the network?

SKILL.md names 1 domain. As links in the text: status.coreweave.com. This is read from the text; nothing was executed.

Is Coreweave Incident Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Coreweave Incident Runbook use?

Coreweave Incident Runbook is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coreweave Incident Runbook use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Coreweave Incident Runbook?

Skills that share tags, products or a category with Coreweave Incident Runbook: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coreweave Incident Runbook?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.