Agent skill

Coreweave Cost Tuning

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Optimize CoreWeave GPU cloud costs with right-sizing and scheduling.

MITAuto-check passedAI & LLM Engineering

Install Coreweave Cost Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-cost-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-cost-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/coreweave-cost-tuning .claude/skills/coreweave-cost-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coreweave-cost-tuning
GitHub stars
2.8k
Token cost
~1.3k tokens
SKILL.md length
446 words
Files
1
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Optimize CoreWeave GPU cloud costs with right-sizing and scheduling.

  • Works in 4 steps: Baseline seven days of GPU utilization,… → Propose one reversible change—right-size… → Compare the canary against its SLO and… → …
  • Reducing GPU spend
  • SKILL.md covers Overview, Prerequisites, GPU Pricing Reference… and Cost Optimization Strategies, plus 6 more sections
  • Calls kubectl

What it does

Coreweave Cost Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize CoreWeave GPU cloud costs with right-sizing and scheduling. Use when reducing GPU spend, selecting cost-effective instances, or implementing scale-to-zero for dev workloads. Trigger with phrases like "coreweave cost", "coreweave pricing", "reduce coreweave spend", "coreweave budget".

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering GPU and accelerator computing. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Reducing GPU spend
  • Selecting cost-effective instances
  • Implementing scale-to-zero for dev workloads
  • With phrases like coreweave cost

Example prompts

  • “coreweave cost”
  • “coreweave pricing”
  • “reduce coreweave spend”
  • “/coreweave-cost-tuning”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(kubectl:*), Grep

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Baseline seven days of GPU utilization, queue time, request latency, error rate,
  2. Propose one reversible change—right-size a non-production workload, set a
  3. Compare the canary against its SLO and budget for an agreed observation window.
  4. Record the instance choice, owner, forecast, and rollback trigger in the

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(kubectl:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • coreweave.com
    • docs.coreweave.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Coreweave Cost Tuning loads about 1.3k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 446 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 446 words, ~1,251 tokens.

Download SKILL.mdSave it as .claude/skills/coreweave-cost-tuning/SKILL.md (or your agent's skills folder).
name
coreweave-cost-tuning
description
Optimize CoreWeave GPU cloud costs with right-sizing and scheduling. Use when reducing GPU spend, selecting cost-effective instances, or implementing scale-to-zero for dev workloads. Trigger with phrases like "coreweave cost", "coreweave pricing", "reduce coreweave spend", "coreweave budget".
allowed-tools
Read, Write, Edit, Bash(kubectl:*), Grep
compatibility
Designed for Claude Code
version
1.11.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, gpu-cloud, kubernetes, inference, coreweave

CoreWeave Cost Tuning

Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.

Overview

Reduce GPU spend by matching a workload's memory, throughput, availability, and latency requirements to the smallest approved capacity. Cost changes must preserve the service SLO and retain a measured rollback path; approximate public prices are planning inputs, not a billing source of truth.

Prerequisites

  • Read-only access to workload utilization, namespace quota, and billing allocation data.
  • A named service owner and target SLO for the workload being resized.
  • Approval for any change that can reduce production capacity or alter availability.

GPU Pricing Reference (approximate)

GPUPer GPU/hourBest For
A100 40GB PCIe~$1.50Development, smaller models
A100 80GB PCIe~$2.21Production inference
H100 80GB PCIe~$4.76High-throughput inference
H100 SXM5 (8x)~$6.15/GPUTraining, multi-GPU
L40~$1.10Image generation, light inference

Cost Optimization Strategies

Scale-to-Zero for Dev/Staging
yaml
autoscaling.knative.dev/minScale: "0"
autoscaling.knative.dev/scaleDownDelay: "5m"
Right-Size GPU Selection
python
def recommend_gpu(model_size_b: float, inference_only: bool = True) -> str:
    if model_size_b <= 7:
        return "L40" if inference_only else "A100_PCIE_80GB"
    elif model_size_b <= 13:
        return "A100_PCIE_80GB"
    elif model_size_b <= 70:
        return "A100_PCIE_80GB (4x tensor parallel)"
    else:
        return "H100_SXM5 (8x tensor parallel)"
Quantization to Use Smaller GPUs

Use AWQ or GPTQ quantization to fit larger models on smaller GPUs:

bash
# 70B model at 4-bit fits on single A100-80GB instead of 4x
vllm serve meta-llama/Llama-3.1-70B-Instruct-AWQ --quantization awq

Instructions

  1. Baseline seven days of GPU utilization, queue time, request latency, error rate, and allocated cost by namespace; do not decide from a single peak.
  2. Propose one reversible change—right-size a non-production workload, set a conservative scale-down delay, or run a quantized canary.
  3. Compare the canary against its SLO and budget for an agreed observation window. Promote only if quality, latency, and error rate remain within the published limit.
  4. Record the instance choice, owner, forecast, and rollback trigger in the change record. Revert capacity immediately if the workload breaches its SLO.
Show full SKILL.md (184 more words)Show less

Output

  • A documented utilization baseline and cost allocation for the selected workload.
  • A right-sizing or scale-to-zero recommendation with its performance guardrails.
  • A reversible change record containing owner, observation window, and rollback threshold.

Error Handling

ConditionLikely causeSafe response
Latency rises after downsizingInsufficient GPU capacity or queueingRestore the previous resource request and investigate with the service owner.
Cost data is incompleteLabels or billing export are missingStop the optimization; repair allocation labels before making a pricing decision.
Quantized canary loses qualityModel or quantization setting is unsuitableRoute traffic back to the baseline model and retain the evaluation result.
Scale-to-zero causes cold-start failuresDelay or startup budget is too smallRestore minimum replicas for the affected service and tune in a non-production lane.

Examples

Run an approved staging canary with an explicit resource limit, then compare it to the baseline before touching production:

bash
kubectl -n inference-staging patch deployment summarizer \
  --type merge -p '{"spec":{"template":{"spec":{"containers":[{"name":"server","resources":{"limits":{"nvidia.com/gpu":1}}}]}}}}'
kubectl -n inference-staging rollout status deployment/summarizer --timeout=10m

If p95 latency, error rate, or evaluation quality crosses the signed threshold, restore the prior manifest and attach the redacted measurements to the change record.

Resources

Next Steps

For architecture patterns, see coreweave-reference-architecture.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/.curated/coreweave-cost-tuning of jeremylongshore/tons-of-skills-marketplace.

Open the folder on GitHubat commit cfae287

Compare with similar skills

Coreweave Cost Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coreweave Cost Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coreweave Cost Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.3kAutomated safety check: PassMIT
Areno Debug RuntimeinclusionAI/AReno323—~486Automated safety check: PassApache-2.0
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Dynamo Interconnect CheckNVIDIA/skills3.6k1 repos~1.6kAutomated safety check: PassApache-2.0
Lambda LabsLuciole-Studio/Misaka-Agent1711 repos~3kAutomated safety check: WarnMIT
DGX Spark Memory and Thermal Opswshobson/agents40k—~2kAutomated safety check: PassMIT

Similar skills

  • Areno Debug Runtime

    inclusionAI/AReno

    Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

    323 GitHub stars~486 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink.

    3.6k GitHub starsUsed in 1 repo~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Lambda Labs

    Luciole-Studio/Misaka-Agent

    On-demand GPU cloud instances for ML training. An agent skill from Luciole-Studio/Misaka-Agent.

    171 GitHub starsUsed in 1 repo~3k tokens
    AI & LLM EngineeringAuto-check: warnings
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • SkyPilot Multi-Cloud Orchestration

    Orchestra-Research/AI-Research-SKILLs

    Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

    13k GitHub starsUsed in 4 repos~2.4k tokens
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Coreweave Cost Tuning

What does Coreweave Cost Tuning do?

Optimize CoreWeave GPU cloud costs with right-sizing and scheduling. Coreweave Cost Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize CoreWeave GPU cloud costs with right-sizing and scheduling.

When should I use Coreweave Cost Tuning?

Coreweave Cost Tuning fits situations like: reducing GPU spend; selecting cost-effective instances; implementing scale-to-zero for dev workloads; with phrases like coreweave cost.

How do I install Coreweave Cost Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-cost-tuning -a claude-code`. Or copy the skill folder (skills/.curated/coreweave-cost-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/coreweave-cost-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Coreweave Cost Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-cost-tuning -a codex`. Or copy the skill folder (skills/.curated/coreweave-cost-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/coreweave-cost-tuning in your project. Codex loads it when a task matches its description.

Can I use Coreweave Cost Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-cost-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coreweave-cost-tuning, .gemini/skills/coreweave-cost-tuning, .github/skills/coreweave-cost-tuning and .opencode/skills/coreweave-cost-tuning in your project.

What does Coreweave Cost Tuning need to run?

Going by SKILL.md and its folder, Coreweave Cost Tuning needs the command-line tools its instructions call (kubectl). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(kubectl:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Coreweave Cost Tuning access the network?

SKILL.md names 2 domains. As links in the text: coreweave.com and docs.coreweave.com. This is read from the text; nothing was executed.

Is Coreweave Cost Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Coreweave Cost Tuning use?

Coreweave Cost Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coreweave Cost Tuning use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Coreweave Cost Tuning?

Skills that share tags, products or a category with Coreweave Cost Tuning: Areno Debug Runtime (inclusionAI/AReno, 323 stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars), Dynamo Interconnect Check (NVIDIA/skills, 3.6k stars) and Lambda Labs (Luciole-Studio/Misaka-Agent, 171 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coreweave Cost Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.