Greptimedb Perses Dashboard
GreptimeTeam/dashboard
Generate Perses dashboards or single panels for GreptimeDB. An agent skill from GreptimeTeam/dashboard.
Agent skill
by jeremylongshore in jeremylongshore/tons-of-skills-marketplace
Hunt down Databricks cost leaks — wasted DBUs, idle clusters, oversized SQL warehouses, and untagged runaway spend — and produce a FinOps cost report.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-cost-leak-hunter --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/databricks-cost-leak-hunter .claude/skills/databricks-cost-leak-hunter && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "databricks-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-cost-leak-hunter into .claude/skills/databricks-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-cost-leak-hunter", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-cost-leak-hunterType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-cost-leak-hunter --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.curated/databricks-cost-leak-hunter .agents/skills/databricks-cost-leak-hunter && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "databricks-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-cost-leak-hunter into .agents/skills/databricks-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-cost-leak-hunter", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-cost-leak-hunter --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.curated/databricks-cost-leak-hunter .cursor/skills/databricks-cost-leak-hunter && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "databricks-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-cost-leak-hunter into .cursor/skills/databricks-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-cost-leak-hunter", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jeremylongshore/tons-of-skills-marketplace.git --path skills/.curated/databricks-cost-leak-hunter--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-cost-leak-hunter --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.curated/databricks-cost-leak-hunter .gemini/skills/databricks-cost-leak-hunter && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "databricks-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-cost-leak-hunter into .gemini/skills/databricks-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-cost-leak-hunter", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-cost-leak-hunterInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.curated/databricks-cost-leak-hunter .github/skills/databricks-cost-leak-hunter && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "databricks-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-cost-leak-hunter into .github/skills/databricks-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-cost-leak-hunter", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-cost-leak-hunter --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.curated/databricks-cost-leak-hunter .opencode/skills/databricks-cost-leak-hunter && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "databricks-cost-leak-hunter" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-cost-leak-hunter into .opencode/skills/databricks-cost-leak-hunter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-cost-leak-hunter", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
databricks-cost-leak-hunterHunt down Databricks cost leaks — wasted DBUs, idle clusters, oversized SQL warehouses, and untagged runaway spend — and produce a FinOps cost report.
Databricks Cost Leak Hunter is an agent skill from jeremylongshore/tons-of-skills-marketplace. Hunt down Databricks cost leaks — wasted DBUs, idle clusters, oversized SQL warehouses, and untagged runaway spend — and produce a FinOps cost report. Use when a user asks why their Databricks bill is high, wants to find cost leaks / wasted DBUs / idle clusters, or needs a FinOps cost report. Trigger with "databricks cost", "why is my databricks bill", "find wasted spend", "cost leak".
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts and reference files (for example `docs/ADR.md`, `docs/CFO-ONE-PAGER.md` and `docs/ONE-PAGER.md`). Compatibility notes: Designed for Claude Code
It sits in DevOps & Cloud, covering Cloud cost optimization and SQL. It works with Databricks, SQL and Model Context Protocol. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBash(databricks:*)Bash(jq:*)Globmcp__databricks-workspace-mcp__clusters_getmcp__databricks-workspace-mcp__clusters_eventsmcp__databricks-workspace-mcp__clusters_listmcp__databricks-workspace-mcp__instance_pools_list…and 1 more on the same allowed-tools line.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
databricksjqpython3From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.databricks.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DATABRICKS_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Databricks Cost Leak Hunter loads about 4.6k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 104 tokens; SKILL.md has 1,669 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,669 words, ~4,644 tokens.
.claude/skills/databricks-cost-leak-hunter/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Audits a Databricks workspace for real-dollar cost leaks — idle compute, jobs on the wrong SKU, overprovisioned clusters, and the Photon premium paid without the speedup — then emits a CFO-grokkable, dollar-ranked FinOps remediation report.
This skill finds where Databricks money is leaking and how much, in dollars, per
month. Confirmed-spend figures come from the customer's own system.billing.usage
joined to system.billing.list_prices — never an estimate. Two of the four
categories (overprovisioning, Photon premium) are explicitly modeled/at-risk
amounts, labeled as such so a CFO never confuses them with recoverable spend. The
skill surfaces four named leak categories, ranks them by monthly dollar impact, and
explains each root cause in FinOps language a CFO can act on without an engineer to
translate.
It is architecturally distinct from the v1 databricks-cost-tuning skill: that one
AUTHORS policy (creates cluster policies, spot configs). This one DETECTS leaks and
reports them, dollarized and ranked. The math is deterministic — bundled scripts in
scripts/ do the arithmetic so the agent never eyeballs numbers — and deep domain
knowledge lives in references/ loaded only when a leak needs it.
The skill uses two data planes. Dollar figures come from the Databricks CLI
Statement Execution API (databricks api post /api/2.0/sql/statements) reading
system.* — authenticated by the CLI's own DATABRICKS_HOST+DATABRICKS_TOKEN /
databricks auth login. The live config/event evidence that explains why a leak
exists (auto-termination setting, node type, autoscale floor, pool min_idle) comes
from the custom databricks-workspace-mcp control-plane tools — the one MCP
dependency. The SQL produces the number; the workspace MCP turns it into a verified,
single-config-change fix.
Read access to the billing system tables is the hard dependency and the most common
failure. system.billing.usage requires a metastore-admin grant chain — the
skill detects a missing grant upfront and reports it, rather than failing mid-flow.
GRANT USE CATALOG ON CATALOG system TO <principal>GRANT USE SCHEMA ON SCHEMA system.billing TO <principal>GRANT SELECT ON TABLE system.billing.usage TO <principal>GRANT SELECT ON TABLE system.billing.list_prices TO <principal>system.compute): system.compute.clusters, system.compute.node_timeline.databricks auth login, or DATABRICKS_HOST + DATABRICKS_TOKEN)
and jq for parsing JSON tool output. The dollar queries
run through the CLI Statement Execution API — UC enforces the grant chain above.DATABRICKS_WAREHOUSE_ID env var set to a running SQL warehouse — every
statement-execution call (Step 1's probe included) requires it.databricks-workspace-mcp registered (its own PAT/U2M/M2M auth; PAT
unsupported in Databricks-App deployment mode). It reads the live REST API, not
system.*, so it needs no system-table grants. If it is absent the skill still
produces dollar figures but cannot corroborate live config — it then accepts
pasted config input.Authentication. The CLI Statement Execution API uses DATABRICKS_HOST +
DATABRICKS_TOKEN or databricks auth login; UC enforces the metastore grant chain
on every system.* read. The custom databricks-workspace-mcp authenticates
separately via its own PAT / U2M / M2M token. No secrets are hardcoded — all auth
comes from the environment or the registered MCP server.
Run the upfront grant check before any analysis — see Step 1.
The pipeline is detect → compute → rank → report. SQL detection runs through the
CLI Statement Execution API; config corroboration runs on databricks-workspace-mcp;
the dollar arithmetic runs in scripts/; deep knowledge loads from references/ on
demand.
Probe the billing tables before anything else. If the probe errors with a
permission message, STOP and report the exact missing grant — do not continue into
the leak scans. Requires DATABRICKS_WAREHOUSE_ID (a running SQL warehouse).
databricks api post /api/2.0/sql/statements --json '{
"warehouse_id": "'"$DATABRICKS_WAREHOUSE_ID"'",
"statement": "SELECT 1 FROM system.billing.usage LIMIT 1",
"wait_timeout": "30s"
}' | jq -r '.status.state, .status.error.message // "ok"'If state is not SUCCEEDED, load
${CLAUDE_SKILL_DIR}/references/system-tables-setup.md
and report the missing grant chain to the user verbatim. Stop here.
Establish the trailing-30-day total spend so every leak can be expressed as a share
of a real number, and capture the window's MAX(usage_date) to stamp into the
report. The price-window join (usage × list_prices.pricing.default, matched on
sku_name AND usage_unit within the price-effective window, currency_code='USD')
is the dollar primitive reused by every category query.
# The CLI does NOT expand ${VARS} inside a --json @file, so inject the warehouse
# id with jq at call time (the static template carries only wait_timeout + statement).
databricks api post /api/2.0/sql/statements --json "$(
jq --arg wh "$DATABRICKS_WAREHOUSE_ID" '. + {warehouse_id: $wh}' \
"${CLAUDE_SKILL_DIR}/scripts/sql/spend-baseline.sql.json"
)"The canonical CTE and full per-category SQL live in
${CLAUDE_SKILL_DIR}/references/cost-leak-categories.md.
Load it now — the four detection queries below all reference its priced CTE.
Join priced All-Purpose usage to system.compute.clusters; flag clusters whose
latest-change auto_termination_minutes = 0. Rank by 30-day idle spend. This is
confirmed spend — money actually billed for idle compute.
SELECT p.usage_metadata.cluster_id AS cluster_id,
COALESCE(c.cluster_name, 'unknown') AS cluster_name,
c.auto_termination_minutes,
ROUND(SUM(p.usd), 2) AS spend_30d_usd
FROM priced p
JOIN cluster_cfg c ON p.usage_metadata.cluster_id = c.cluster_id
WHERE p.billing_origin_product = 'ALL_PURPOSE'
AND c.auto_termination_minutes = 0
GROUP BY p.usage_metadata.cluster_id, c.cluster_name, c.auto_termination_minutes
HAVING SUM(p.usd) > 0
ORDER BY spend_30d_usd DESC;Corroborate each flagged cluster's live config with databricks-workspace-mcp
clusters_get (confirm autotermination_minutes = 0 right now) and clusters_events
(measure the idle gap between RUNNING and TERMINATING).
The signature leak: a usage row with a job_id in usage_metadata AND
billing_origin_product = 'ALL_PURPOSE' ($0.55/DBU) instead of $0.15/DBU). Re-price the same DBUs at the current Jobs rate to compute savings.
This is confirmed savings — a deterministic re-pricing delta. The JOBS_COMPUTE
(jobs_rate
CTE is deduped to one USD rate per usage_unit so the join cannot fan out (see
cost-leak-categories.md).
SELECT p.usage_metadata.job_id AS job_id,
ROUND(SUM(p.usd), 2) AS spend_on_all_purpose_30d_usd,
ROUND(SUM(p.usd) - SUM(p.usage_quantity * jr.jobs_unit_price), 2)
AS potential_savings_30d_usd
FROM priced p
JOIN jobs_rate jr ON p.usage_unit = jr.usage_unit
WHERE p.billing_origin_product = 'ALL_PURPOSE'
AND p.usage_metadata.job_id IS NOT NULL
GROUP BY p.usage_metadata.job_id
HAVING SUM(p.usd) - SUM(p.usage_quantity * jr.jobs_unit_price) > 0
ORDER BY potential_savings_30d_usd DESC;Confirm the live compute type is All-Purpose (not Jobs) with clusters_list /
clusters_get (cluster_source) before recommending the move. The job_id +
ALL_PURPOSE billing signal is itself dollar-accurate; the REST check is
belt-and-suspenders.
Aggregate mean CPU from system.compute.node_timeline, join to 30-day spend, flag
clusters burning real dollars at chronically low utilization (< 25%). This figure is
an estimate (est_overprovision = spend × (1 − CPU%)), not billed waste — it is
the one modeled number in the pipeline and is labeled est_* everywhere.
SELECT s.cluster_id,
ROUND(u.avg_cpu_pct, 1) AS avg_cpu_pct,
ROUND(s.spend_30d_usd, 2) AS spend_30d_usd,
ROUND(s.spend_30d_usd * (1 - LEAST(u.avg_cpu_pct,100)/100.0), 2)
AS est_overprovision_30d_usd
FROM spend s
JOIN util u ON s.cluster_id = u.cluster_id
WHERE u.avg_cpu_pct < 25 AND s.spend_30d_usd > 0
ORDER BY est_overprovision_30d_usd DESC;Corroborate the configured floor with clusters_get (REST nested
autoscale.min_workers / autoscale.max_workers); for the idle-pool variant use
instance_pools_list (min_idle_instances + stats.idle_count) — pool waste is NOT
a billing row.
Photon is not a column on system.compute.clusters; it is billing-visible via
the SKU. Isolate usage whose sku_name ILIKE '%PHOTON%' and surface the ~2× premium
portion as the at-risk amount — money for review against actual runtime gain, not
confirmed waste.
SELECT p.usage_metadata.cluster_id AS cluster_id,
ROUND(SUM(p.usd), 2) AS photon_spend_30d_usd,
ROUND(SUM(p.usd) / 2.0, 2) AS photon_premium_at_risk_30d_usd
FROM priced p
WHERE p.sku_name ILIKE '%PHOTON%'
AND p.billing_origin_product IN ('ALL_PURPOSE','JOBS_COMPUTE')
AND p.usage_metadata.cluster_id IS NOT NULL
GROUP BY p.usage_metadata.cluster_id
HAVING SUM(p.usd) > 0
ORDER BY photon_premium_at_risk_30d_usd DESC;Confirm Photon is live and worth keeping with databricks-workspace-mcp
clusters_get (REST runtime_engine — a config-plane field, not a system column);
for DLT pipelines use pipelines_get (spec.photon / serverless / edition). See
${CLAUDE_SKILL_DIR}/references/dlt-tier-cost-tradeoffs.md
when the leak touches DLT/serverless tiers.
Pass each category's query result to the deterministic ranker — the LLM does NOT do
the arithmetic. Each leak object carries a kind field
(confirmed / estimated / at-risk) so the renderer can split the headline into
confirmed-recoverable vs estimated/at-risk-pending-review and stamp a Confidence
column. The script sums per-category figures by kind, ranks descending by monthly
dollar impact, annualizes the headline and #1 line, stamps the trailing-30-day window
end date, and renders the CFO-grokkable report.
# Per-category results and the rendered report are RUNTIME outputs — they go to
# a working dir ($OUT), never the skill package. Steps 3–6 wrote leak-*.json here.
OUT="${OUT:-$(pwd)/cost-leak-out}" && mkdir -p "$OUT"
jq -s '.' "$OUT"/leak-*.json | \
python3 "${CLAUDE_SKILL_DIR}/scripts/rank-and-report.py" \
--monthly-spend 100000 \
--window-end "$WINDOW_END_DATE" \
--out "$OUT/cost-leak-report.md"Use Glob to collect the per-category leak-*.json results, Write the rendered
report, and Edit it if the user wants the headline spend rescaled. Render the
output using the verbatim template in
${CLAUDE_SKILL_DIR}/references/cfo-output-format.md.
$OUT/cost-leak-report.md in the working dir) leading
with a split headline that never sums confirmed and unconfirmed dollars under
one verb — ### A $<spend>/month workspace is burning **~$<confirmed>/month** (confirmed), plus up to **~$<at-risk>/month** pending review — each with its
~$<annualized>/year companion.Trailing 30 days ending <window-end>) so every figure has an explicit calendar
window, not just a /month cadence label.Confidence column
(# | Where it's leaking | $/month | Confidence | The fix), one row per category,
ranked highest dollar impact first, $/month right-aligned, each fix a single
config change. Root-cause cells use plain-business language — no raw DBU unit in
the CFO-visible text (DBU detail stays in the per-leak detail artifacts).system.billing.usage, while
overprovision (estimated) and Photon premium (at-risk) are labeled as modeled.databricks-workspace-mcp
and the underlying $/DBU rates for engineers.| Error | Cause | Solution |
|---|---|---|
PERMISSION_DENIED on system.billing.usage | Metastore-admin grant chain missing | Run Step 1; report the exact GRANT USE CATALOG / USE SCHEMA / SELECT chain from system-tables-setup.md. Stop, do not continue. |
| CLI not authenticated / token expired | No valid DATABRICKS_HOST + token | Re-run databricks auth login; verify with databricks current-user me. |
Empty / unset DATABRICKS_WAREHOUSE_ID | Required warehouse for statement execution not set | Set DATABRICKS_WAREHOUSE_ID to a running SQL warehouse before Step 1. |
list_prices join returns NULL usd | Custom/negotiated pricing not in list_prices, or usage_unit mismatch | Join on sku_name AND usage_unit within the price window with currency_code='USD'; if still NULL, use the customer's contracted rate card from references/cost-leak-categories.md. |
| Workspace MCP missing | Server not registered | Degrade gracefully: report it absent, run the dollar half, accept pasted config for corroboration. Never fail silently mid-flow. |
node_timeline empty for a cluster | Serverless/short-lived compute, or monitoring lag | Skip Leak 3 for that cluster; note "utilization unavailable" rather than reporting $0 overprovision. |
Untagged spend / no cluster_name | Clusters lack CostCenter/Team tags | Attribute by cluster_id; flag attribution as incomplete in the report footer. |
Runs the full pipeline. The grant check passes, the four scans return rows, and the ranker emits the CFO report with a split, confidence-stamped headline:
### A $100K/month Databricks workspace is burning **~$19,000/month** (confirmed), plus up to **~$8,000/month** pending review
Trailing 30 days ending 2026-06-22. Confirmed ~$228K/year; up to ~$96K/year more pending review. Every line below is one config change.
| # | Where it's leaking | $/month | Confidence | The fix |
|---|---|--:|---|---|
| 1 | Clusters that never shut themselves off — paying around the clock for compute nobody is using | **$12,000** | Confirmed | Set auto-shutoff (e.g. 30 min) |
| 2 | Scheduled batch jobs running on the premium notebook tier — ~3.6× the batch rate for identical work | **$7,000** | Confirmed | Move job clusters to the batch tier |
| 3 | Clusters sized for peak, idling most of the time — typically 30–50% oversized | **$5,000** | Estimated | Turn on autoscaling, drop the floor |
| 4 | Paying a ~2× speed-engine premium on jobs that don't run faster | **$3,000** | At-risk | Turn off the speed engine where it adds no gain |
**The #1 line alone — idle clusters (confirmed) — is ~$144K/year, fixed in one setting.**User asks "find idle clusters wasting money." The skill runs Step 3 only, joins the
spend to clusters_get, and reports each auto_termination_minutes = 0 cluster with
its 30-day idle spend and the live idle gap from clusters_events.
User asks "are any jobs on the wrong compute?" Step 4 returns each job_id running
on All-Purpose with potential_savings_30d_usd, corroborated by clusters_get
confirming cluster_source is not JOB — the single fix is "move to Jobs Compute."
${CLAUDE_SKILL_DIR}/references/cost-leak-categories.md — the four leak categories: definition, real detection SQL, FinOps root cause, remediation.${CLAUDE_SKILL_DIR}/references/cfo-output-format.md — verbatim CFO report template + the 90-second-skim rules.${CLAUDE_SKILL_DIR}/references/system-tables-setup.md — metastore-admin grant chain + access verification.${CLAUDE_SKILL_DIR}/references/dlt-tier-cost-tradeoffs.md — DLT / serverless / Photon cost-tier encyclopedia, loaded on demand.© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (scripts, references) in skills/.curated/databricks-cost-leak-hunter of jeremylongshore/tons-of-skills-marketplace.
Open the folder on GitHubat commit cfae287
Databricks Cost Leak Hunter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Databricks Cost Leak Hunter this skilljeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~4.6k | Automated safety check: Pass | MIT | |
| Greptimedb Perses DashboardGreptimeTeam/dashboard | 111 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| Databricks Data Discoverydatabricks/databricks-agent-skills | 345 | — | ~2.1k | Automated safety check: Pass | Custom licence | |
| TMA1 Observability Querytma1-ai/tma1 | 119 | — | ~5.1k | Automated safety check: Notes | Apache-2.0 | |
| Azure Synapse AnalyticsMicrosoftDocs/Agent-Skills | 776 | 1 repos | ~13k | Automated safety check: Pass | CC-BY-4.0 | |
| Trigger.dev Cost Savings Auditpapermark/papermark | 9.2k | — | ~1.3k | Automated safety check: Pass | Custom licence |
GreptimeTeam/dashboard
Generate Perses dashboards or single panels for GreptimeDB. An agent skill from GreptimeTeam/dashboard.
databricks/databricks-agent-skills
Discover, explore, and query Databricks data via Genie — the CLI equivalent of the Genie One MCP.
tma1-ai/tma1
Answers questions about agent spend, token use, traces, events, errors and tool usage by running read-only SQL against a local TMA1 observability store.
MicrosoftDocs/Agent-Skills
Expert knowledge for Azure Synapse Analytics development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration…
papermark/papermark
Audits Trigger.dev tasks, schedules and run history for wasteful machine sizes, retries, polling and cron frequency to cut spend.
StarRocks/starrocks
Proposes verified fixes for failing SQL examples in StarRocks docs across three languages and several versions, then opens draft pull requests and never merges.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
jeremylongshore/tons-of-skills-marketplace
Execute proactive auto-loading: automatically detects and loads agents.md files.
jeremylongshore/tons-of-skills-marketplace
Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.
jeremylongshore/tons-of-skills-marketplace
Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Works with
Categories
Hunt down Databricks cost leaks — wasted DBUs, idle clusters, oversized SQL warehouses, and untagged runaway spend — and produce a FinOps cost report. Databricks Cost Leak Hunter is an agent skill from jeremylongshore/tons-of-skills-marketplace. Hunt down Databricks cost leaks — wasted DBUs, idle clusters, oversized SQL warehouses, and untagged runaway spend — and produce a FinOps cost report.
Databricks Cost Leak Hunter fits situations like: A user asks why their Databricks bill is high; wants to find cost leaks / wasted DBUs / idle clusters; needs a FinOps cost report; with databricks cost.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a claude-code`. Or copy the skill folder (skills/.curated/databricks-cost-leak-hunter in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/databricks-cost-leak-hunter in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a codex`. Or copy the skill folder (skills/.curated/databricks-cost-leak-hunter in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/databricks-cost-leak-hunter in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-cost-leak-hunter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/databricks-cost-leak-hunter, .gemini/skills/databricks-cost-leak-hunter, .github/skills/databricks-cost-leak-hunter and .opencode/skills/databricks-cost-leak-hunter in your project.
Going by SKILL.md and its folder, Databricks Cost Leak Hunter needs Python for the scripts in its folder, the command-line tools its instructions call (databricks, jq and python3) and credentials named DATABRICKS_TOKEN. Our summary lists: Python 3; A credential in DATABRICKS_TOKEN. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(databricks:*), Bash(jq:*), Glob, mcp__databricks-workspace-mcp__clusters_get, mcp__databricks-workspace-mcp__clusters_events, mcp__databricks-workspace-mcp__clusters_list, mcp__databricks-workspace-mcp__instance_pools_list, mcp__databricks-workspace-mcp__pipelines_get. Compatibility (from SKILL.md): Designed for Claude Code.
SKILL.md names 1 domain. As links in the text: docs.databricks.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Databricks Cost Leak Hunter is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Databricks Cost Leak Hunter: Greptimedb Perses Dashboard (GreptimeTeam/dashboard, 111 stars), Databricks Data Discovery (databricks/databricks-agent-skills, 345 stars), TMA1 Observability Query (tma1-ai/tma1, 119 stars) and Azure Synapse Analytics (MicrosoftDocs/Agent-Skills, 776 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.
Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.