Happy Infra Metrics and Grafana
slopus/happy
Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.
Agent skill
by jeremylongshore in jeremylongshore/tons-of-skills-marketplace
Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a…
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-gpu-cost-leak-hunter --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/coreweave-gpu-cost-leak-hunter .claude/skills/coreweave-gpu-cost-leak-hunter && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "coreweave-gpu-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-cost-leak-hunter into .claude/skills/coreweave-gpu-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-cost-leak-hunter", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-cost-leak-hunterType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-gpu-cost-leak-hunter --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.curated/coreweave-gpu-cost-leak-hunter .agents/skills/coreweave-gpu-cost-leak-hunter && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "coreweave-gpu-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-cost-leak-hunter into .agents/skills/coreweave-gpu-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-cost-leak-hunter", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-gpu-cost-leak-hunter --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.curated/coreweave-gpu-cost-leak-hunter .cursor/skills/coreweave-gpu-cost-leak-hunter && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "coreweave-gpu-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-cost-leak-hunter into .cursor/skills/coreweave-gpu-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-cost-leak-hunter", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jeremylongshore/tons-of-skills-marketplace.git --path skills/.curated/coreweave-gpu-cost-leak-hunter--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-gpu-cost-leak-hunter --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.curated/coreweave-gpu-cost-leak-hunter .gemini/skills/coreweave-gpu-cost-leak-hunter && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "coreweave-gpu-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-cost-leak-hunter into .gemini/skills/coreweave-gpu-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-cost-leak-hunter", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-gpu-cost-leak-hunterInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.curated/coreweave-gpu-cost-leak-hunter .github/skills/coreweave-gpu-cost-leak-hunter && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "coreweave-gpu-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-cost-leak-hunter into .github/skills/coreweave-gpu-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-cost-leak-hunter", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace coreweave-gpu-cost-leak-hunter --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.curated/coreweave-gpu-cost-leak-hunter .opencode/skills/coreweave-gpu-cost-leak-hunter && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "coreweave-gpu-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/coreweave-gpu-cost-leak-hunter into .opencode/skills/coreweave-gpu-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "coreweave-gpu-cost-leak-hunter", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
coreweave-gpu-cost-leak-hunterHunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a…
Coreweave GPU Cost Leak Hunter is an agent skill from jeremylongshore/tons-of-skills-marketplace. Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a CFO-grokkable, dollar-ranked FinOps report. CoreWeave ships no cost dashboard and no billing API, so the spend view is built from PromQL against its managed Grafana. Use when a user asks why their CoreWeave GPU bill is high, wants to find wasted GPU spend or idle reservations, or needs a GPU FinOps cost report. Trigger with…
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `ARD.md`, `PRD.md` and `eval-spec.yaml`). Compatibility notes: Designed for Claude Code
It sits in DevOps & Cloud, covering Monitoring and alerting, Cloud cost optimization and GPU and accelerator computing. It works with Prometheus and Grafana. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditGlobBash(curl:*)Bash(jq:*)Bash(kubectl get:*)Bash(python3:*)From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
curlkubectljqpython3From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.coreweave.comcoreweave.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
CW_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Coreweave GPU Cost Leak Hunter loads about 3.5k tokens when it runs, and up to ~8.7k if it reads all its reference files. Until then it costs about 163 tokens; SKILL.md has 1,286 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,286 words, ~3,503 tokens.
.claude/skills/coreweave-gpu-cost-leak-hunter/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Community-contributed. Not affiliated with, endorsed by, or sponsored by CoreWeave, Inc. CoreWeave is a registered trademark of CoreWeave, Inc.
Audits a CoreWeave GPU cluster for real-dollar cost leaks — idle reserved capacity, GPUs on the wrong SKU, allocated-but-idle instances, and steady on-demand spend that should be committed — then emits a CFO-grokkable, dollar-ranked FinOps report.
CoreWeave ships no cost dashboard and no billing API (usage-monitoring docs). There is also no single "dollars" metric — spend is reconstructed by querying usage from CoreWeave's managed Grafana in PromQL and multiplying each resource's usage by its rate-card price. This skill does exactly that, then ranks the leaks by monthly dollar impact.
The math is deterministic: PromQL returns usage counts, and the bundled
scripts/rank-and-report.py does every multiplication, sum, and ranking — the agent
never eyeballs a number. Two of the four categories are billed waste (Confirmed);
the other two are a right-sizing model (Estimated) and a commitment decision
(At-risk), labeled so a CFO never reads a modeled number as recoverable cash.
Deep domain knowledge lives in references/, loaded only when a leak needs it.
admin, metrics, or write group in the CoreWeave Cloud
Console (usage-monitoring docs). This is the hard dependency; Step 1 probes it
and fails fast if the group is missing.$CW_PROM_URL (the Grafana data-source
proxy, e.g. https://grafana.ORG.coreweave.com/api/datasources/proxy/uid/UID)
and a bearer token in $CW_TOKEN for curl.kubeconfig for the cluster (CoreWeave-issued) so kubectl get can corroborate
live GPU allocation and node labels.references/gpu-right-sizing.md (dated
snapshot of coreweave.com/pricing) or the customer's contract.jq and python3 for parsing query JSON and running the ranker.Authentication. All auth comes from the environment ($CW_PROM_URL, $CW_TOKEN,
$KUBECONFIG) — no secrets are hardcoded. Grafana enforces the group membership above
on every query.
The pipeline is detect → price → rank → report. PromQL returns usage; the dollar
arithmetic runs in scripts/; deep knowledge loads from references/ on demand:
Probe billing:instance:total before anything else. An HTTP 401/403 or empty result
means the token's principal is not in admin/metrics/write — STOP and report it;
do not continue into the scans.
curl -sS -H "Authorization: Bearer $CW_TOKEN" \
--data-urlencode 'query=count(billing:instance:total)' \
"$CW_PROM_URL/api/v1/query" | jq -r '.status, (.data.result | length)'If status is not success with a non-empty result, load
references/promql-billing-setup.md and report
the missing group access verbatim. Stop here.
Reconstruct 30-day GPU node-hours per instance type. CoreWeave has no dollars metric, so this returns usage — the ranker multiplies by the rate card. Write the JSON to the working dir for the ranker.
curl -sS -H "Authorization: Bearer $CW_TOKEN" \
--data-urlencode 'query=sum by (instance_type) (sum_over_time(billing:instance:total[30d:1h]))' \
"$CW_PROM_URL/api/v1/query" > "$OUT/baseline.json"The rate card and the per-category PromQL live in
references/gpu-cost-leak-categories.md.
Load it now — the four scans below reference its recording-rule notes.
Reserved GPUs bill at the committed rate whether used or not. A reserved GPU
sitting below a utilization floor is confirmed waste — you paid for it and it did no
work. Cross reserved allocation (billing_gpu, filtered by the reservation label)
against SM-active from DCGM.
curl -sS -H "Authorization: Bearer $CW_TOKEN" --data-urlencode \
'query=sum by (instance_type,node) (avg_over_time(billing_gpu{reservation!=""}[30d:1h]))
and on(node) (avg by (node) (avg_over_time(DCGM_FI_PROF_SM_ACTIVE[30d:1h])) < 0.05)' \
"$CW_PROM_URL/api/v1/query" > "$OUT/leak1-idle-reserved.json"The reservation label key is provider-specific — confirm yours with kubectl get nodes --show-labels. Waste = idle reserved GPU-hours × committed rate (ranker input).
H100/H200 running small-model (~7B–30B) inference is over-paying: for that regime
L40S is cheaper per token (directional — see gpu-right-sizing.md). Flag those
instance-hours; the ranker re-prices them at the L40S rate.
curl -sS -H "Authorization: Bearer $CW_TOKEN" --data-urlencode \
'query=sum by (instance_type) (sum_over_time(billing:instance:total{instance_type=~".*(h100|h200).*"}[30d:1h]))' \
"$CW_PROM_URL/api/v1/query" > "$OUT/leak2-wrong-gpu.json"This is Estimated: the rate delta is exact rate-card math, but throughput
equivalence on L40S is a model. Confirm the served model size with the cluster owner
before acting; FP8 serving needs Hopper/Ada, not Ampere (see gpu-right-sizing.md).
On-demand GPUs that are allocated (billing) but running at low SM-utilization / low MFU bill the full on-demand rate for no work — confirmed billed waste, the GPU twin of an idle cluster.
curl -sS -H "Authorization: Bearer $CW_TOKEN" --data-urlencode \
'query=(avg by (node,instance_type) (avg_over_time(DCGM_FI_PROF_SM_ACTIVE[30d:1h])) < 0.05)
and on(node) (sum by (node) (avg_over_time(billing_gpu{reservation=""}[30d:1h])) > 0)' \
"$CW_PROM_URL/api/v1/query" > "$OUT/leak3-idle-ondemand.json"Corroborate with kubectl get pods -A --field-selector=status.phase=Running to
confirm nothing is actually scheduled on the flagged node. Waste = idle on-demand
GPU-hours × on-demand rate.
A stable on-demand floor — GPUs of one type always running across the window — is
paying on-demand for capacity a commitment discounts up to 60% (pricing).
Measure the always-on floor with min_over_time.
curl -sS -H "Authorization: Bearer $CW_TOKEN" --data-urlencode \
'query=min_over_time(sum by (instance_type) (billing:instance:total{reservation=""})[30d:1h])' \
"$CW_PROM_URL/api/v1/query" > "$OUT/leak4-commit-gap.json"This is At-risk: the up-to-60% saving is pending a commitment decision, and a
commitment is itself a paid obligation — see the over-reservation caution in
gpu-cost-leak-categories.md. Savings = floor GPU-hours × on-demand rate × discount.
Assemble one leak object per category from the PromQL usage results plus the rate
card, then pipe them to the deterministic ranker — the LLM does NOT do the arithmetic.
Because CoreWeave exposes no dollars metric, each object carries usage_gpu_hours and
its rate-card rate; the ranker multiplies usage × rate itself (and applies the
re-price or discount factor). Each object's kind (confirmed / estimated /
at-risk) tells the renderer to split the headline confirmed-vs-pending, rank
descending by monthly dollars, and stamp a Confidence column.
OUT="${OUT:-$(pwd)/cost-leak-out}" && mkdir -p "$OUT"
# Each Step wrote a leak-N.json {category, root_cause, fix, kind, usage_gpu_hours,
# rate_usd_per_gpu_hour, ...}; the ranker does usage × rate deterministically.
jq -s '.' "$OUT"/leak-*.json | \
python3 scripts/rank-and-report.py \
--monthly-spend 180000 --window-end "$WINDOW_END" \
--out "$OUT/cost-leak-report.md"Render the output using the verbatim template in
references/cfo-output-format.md. Use Glob to
collect the per-leak JSON, Write the report, and Edit it to rescale the headline
spend on request.
A $180K/month CoreWeave GPU cluster is burning ~$40K/month (confirmed), plus up to ~$29K/month pending review —
each with a /year companion.# | Where it's leaking | $/month | Confidence | The fix), one row per category, highest dollar impact first, each fix a single change.$/GPU-hour rates for the cluster engineer.| Error | Cause | Solution |
|---|---|---|
HTTP 401/403 on /api/v1/query | Token principal not in admin/metrics/write | Run Step 1; report the group requirement from promql-billing-setup.md. Stop. |
Empty result for billing:instance:total | Wrong data-source proxy UID, or org has no billing metrics enabled | Verify $CW_PROM_URL points at the Grafana Prometheus proxy; confirm in Grafana Explore. |
DCGM_* series absent | DCGM exporter not scraped on the node pool | Skip Leaks 1/3 utilization filter for that pool; note "utilization unavailable" rather than reporting $0. |
reservation label missing | Provider label key differs per org | Confirm the reservation/committed label with kubectl get nodes --show-labels; substitute it in the query. |
Ranker prints ~$0/month confirmed | A kind value was mis-cased and dropped from the sum | The ranker normalizes case; verify each leak object's kind is one of the three tiers. |
Runs the full pipeline. The access probe passes, the four scans return rows, and the ranker emits a split, confidence-stamped report:
### A $180K/month CoreWeave GPU cluster is burning **~$44,986/month** (confirmed), plus up to **~$29,110/month** pending review
Trailing 30 days ending 2026-06-22. Confirmed **~$540K/year**; up to **~$349K/year** more pending review. Spend is reconstructed from PromQL against CoreWeave's managed Grafana (no billing API). Every line below is one change.
| # | Where it's leaking | $/month | Confidence | The fix |
|---|---|--:|---|---|
| 1 | **Idle reserved GPUs** — reserved capacity billing around the clock below a utilization floor | **$26,280** | Confirmed | Right-size or release the reservation |
| 2 | **Allocated-but-idle on-demand GPUs** — nodes up at <5% SM-active, paying full rate for no work | **$18,706** | Confirmed | Scale-to-zero / deschedule the idle nodes |
| 3 | **H100/H200 on small-model inference** — L40S is cheaper per token in the 7B–30B regime | **$16,629** | Estimated | Move small inference to L40S |
| 4 | **Steady on-demand that should be committed** — an always-on floor paying on-demand | **$12,481** | At-risk | Commit the stable floor (up to 60% off) |
**The #1 line alone — idle reserved gpus (confirmed) — is ~$315K/year, fixed in one setting.**User asks "are we paying for idle reserved GPUs?" The skill runs Step 3 only, crosses
billing_gpu{reservation!=""} against DCGM_FI_PROF_SM_ACTIVE, and reports each
reserved node below the floor with its 30-day committed spend.
references/gpu-cost-leak-categories.md — the four leak categories: definition, PromQL, root cause, the one fix.references/cfo-output-format.md — verbatim CFO report template + the never-sum invariant.references/gpu-right-sizing.md — L40S/H100/H200/A100/L40 decision table + the FP8 rule (figures flagged directional).references/promql-billing-setup.md — the billing metrics + group access, cited.coreweave-cost-tuning authors cost-control config; this skill detects leaks and dollarizes them.© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references) in skills/.curated/coreweave-gpu-cost-leak-hunter of jeremylongshore/tons-of-skills-marketplace.
Open the folder on GitHubat commit cfae287
Coreweave GPU Cost Leak Hunter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Coreweave GPU Cost Leak Hunter this skilljeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~3.5k | Automated safety check: Pass | MIT | |
| Happy Infra Metrics and Grafanaslopus/happy | 24k | — | ~2k | Automated safety check: Notes | MIT | |
| Syncmetapawurb/hotpath-rs | 1.9k | — | ~1.2k | Automated safety check: Notes | MIT | |
| Optimize Slurm TopologyNVlabs/alpasim | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Dashboard Previewm4r1k/Eneru | 149 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Alerting Irmgrafana/skills | 282 | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 |
slopus/happy
Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.
pawurb/hotpath-rs
Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).
NVlabs/alpasim
Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.
m4r1k/Eneru
Visually verify Eneru browser-dashboard changes against a live daemon or audit an exact deployment.
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
grafana/skills
Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units…
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
jeremylongshore/tons-of-skills-marketplace
Execute proactive auto-loading: automatically detects and loads agents.md files.
jeremylongshore/tons-of-skills-marketplace
Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.
jeremylongshore/tons-of-skills-marketplace
Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Works with
Categories
Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a…. Coreweave GPU Cost Leak Hunter is an agent skill from jeremylongshore/tons-of-skills-marketplace. Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a CFO-grokkable, dollar-ranked FinOps report.
Coreweave GPU Cost Leak Hunter fits situations like: A user asks why their CoreWeave GPU bill is high; wants to find wasted GPU spend; idle reservations; needs a GPU FinOps cost report.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a claude-code`. Or copy the skill folder (skills/.curated/coreweave-gpu-cost-leak-hunter in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/coreweave-gpu-cost-leak-hunter in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a codex`. Or copy the skill folder (skills/.curated/coreweave-gpu-cost-leak-hunter in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/coreweave-gpu-cost-leak-hunter in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill coreweave-gpu-cost-leak-hunter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coreweave-gpu-cost-leak-hunter, .gemini/skills/coreweave-gpu-cost-leak-hunter, .github/skills/coreweave-gpu-cost-leak-hunter and .opencode/skills/coreweave-gpu-cost-leak-hunter in your project.
Going by SKILL.md and its folder, Coreweave GPU Cost Leak Hunter needs Python for the scripts in its folder, the command-line tools its instructions call (curl, kubectl, jq and python3) and credentials named CW_TOKEN. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Glob, Bash(curl:*), Bash(jq:*), Bash(kubectl get:*), Bash(python3:*). Compatibility (from SKILL.md): Designed for Claude Code.
SKILL.md names 2 domains. As links in the text: docs.coreweave.com and coreweave.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Coreweave GPU Cost Leak Hunter is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Coreweave GPU Cost Leak Hunter: Happy Infra Metrics and Grafana (slopus/happy, 24k stars), Syncmeta (pawurb/hotpath-rs, 1.9k stars), Optimize Slurm Topology (NVlabs/alpasim, 1.3k stars) and Dashboard Preview (m4r1k/Eneru, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.
Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.