Happy Infra Metrics and Grafana
slopus/happy
Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.
Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid…
$ npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install grafana/skills prometheus-cardinality-troubleshooter --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/grafana-cloud/prometheus-cardinality-troubleshooter .claude/skills/prometheus-cardinality-troubleshooter && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "prometheus-cardinality-troubleshooter" agent skill from https://github.com/grafana/skills/tree/main/skills/grafana-cloud/prometheus-cardinality-troubleshooter into .claude/skills/prometheus-cardinality-troubleshooter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prometheus-cardinality-troubleshooter", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/grafana/skills/tree/main/skills/grafana-cloud/prometheus-cardinality-troubleshooterType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install grafana/skills prometheus-cardinality-troubleshooter --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/grafana-cloud/prometheus-cardinality-troubleshooter .agents/skills/prometheus-cardinality-troubleshooter && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "prometheus-cardinality-troubleshooter" agent skill from https://github.com/grafana/skills/tree/main/skills/grafana-cloud/prometheus-cardinality-troubleshooter into .agents/skills/prometheus-cardinality-troubleshooter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prometheus-cardinality-troubleshooter", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install grafana/skills prometheus-cardinality-troubleshooter --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/grafana-cloud/prometheus-cardinality-troubleshooter .cursor/skills/prometheus-cardinality-troubleshooter && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "prometheus-cardinality-troubleshooter" agent skill from https://github.com/grafana/skills/tree/main/skills/grafana-cloud/prometheus-cardinality-troubleshooter into .cursor/skills/prometheus-cardinality-troubleshooter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prometheus-cardinality-troubleshooter", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/grafana/skills.git --path skills/grafana-cloud/prometheus-cardinality-troubleshooter--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install grafana/skills prometheus-cardinality-troubleshooter --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/grafana-cloud/prometheus-cardinality-troubleshooter .gemini/skills/prometheus-cardinality-troubleshooter && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "prometheus-cardinality-troubleshooter" agent skill from https://github.com/grafana/skills/tree/main/skills/grafana-cloud/prometheus-cardinality-troubleshooter into .gemini/skills/prometheus-cardinality-troubleshooter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prometheus-cardinality-troubleshooter", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install grafana/skills prometheus-cardinality-troubleshooterInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/grafana-cloud/prometheus-cardinality-troubleshooter .github/skills/prometheus-cardinality-troubleshooter && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "prometheus-cardinality-troubleshooter" agent skill from https://github.com/grafana/skills/tree/main/skills/grafana-cloud/prometheus-cardinality-troubleshooter into .github/skills/prometheus-cardinality-troubleshooter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prometheus-cardinality-troubleshooter", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install grafana/skills prometheus-cardinality-troubleshooter --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/grafana-cloud/prometheus-cardinality-troubleshooter .opencode/skills/prometheus-cardinality-troubleshooter && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "prometheus-cardinality-troubleshooter" agent skill from https://github.com/grafana/skills/tree/main/skills/grafana-cloud/prometheus-cardinality-troubleshooter into .opencode/skills/prometheus-cardinality-troubleshooter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "prometheus-cardinality-troubleshooter", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
prometheus-cardinality-troubleshooterDiagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid…
Prometheus Cardinality Troubleshooter is an agent skill from grafana/skills, published by the product's own GitHub organization. Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid memory growth. Walks through tsdb status endpoints, per-metric and per-label drill-downs, common-culprit galleries, and remediation paths. Use when the user is currently experiencing a cardinality fire. For preventing cardinality issues at the source, route to prometheus-label-strategy. For post-ingest aggregation…
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Monitoring and alerting. It works with Prometheus and Grafana. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1ccacf2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curljqFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
prometheus-prod-xx.grafana.netFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Prometheus Cardinality Troubleshooter loads about 4.6k tokens when it runs. Until then it costs about 158 tokens; SKILL.md has 1,680 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from grafana/skills at commit 1ccacf2, republished under its Apache-2.0 licence (© grafana). 1,680 words, ~4,559 tokens.
.claude/skills/prometheus-cardinality-troubleshooter/SKILL.md (or your agent's skills folder).You are an expert in diagnosing live Prometheus cardinality problems. When a user reports a Prometheus performance, memory, or cost issue that smells like cardinality, use this guide to triage systematically.
This skill is diagnostic and operational. For schema design and prevention, route to prometheus-label-strategy.
Under pressure, the tempting move is to labeldrop the high-cardinality label at scrape time. Do not. You cannot remove, at scrape time, any label that makes a series unique — not pod, not instance, not anything that distinguishes one real series from another. It looks like it stops the bleeding; it actually breaks the data:
rate() and increase() return garbage, often absurdly high values.rate() absurd?" and there's nothing to point to.The only safe remediations are:
action: drop on __name__) — you're discarding the whole metric, not merging distinct series.path, user_id, etc.).pod. Route to adaptive-metrics.Everywhere below that says "drop a label," read it through this rule: drop whole metrics, fix the source, or use Adaptive Metrics — never labeldrop a distinguishing label.
| Symptom | Likely Cause | First Action |
|---|---|---|
| Prometheus OOMKilled or memory growing linearly | Active series growth (often from a new bad metric or label) | Active Series triage |
| Single PromQL query slow or OOMs the querier | One or more metrics in the query have high cardinality | Per-query drill-down |
| Remote write lagging, WAL growing | Sample throughput spike — series count OR scrape interval changed | Active Series triage + check scrape intervals |
429 Too Many Samples / out of bounds errors | Hitting Mimir/Cortex ingester per-tenant series limit | Per-metric drill-down, find the new offender |
| Grafana Cloud Active Series bill spiked | New metric, new label, or rollout creating churn | Per-metric drill-down + churn check |
| Grafana Cloud DPM bill spiked but Active Series flat | Scrape interval shortened, OR remote_write sending duplicates | DPM-side issue — route to dpm-finder |
series_limit_per_user errors after a deploy | Application change introduced a new bad label | Recent change diff |
| Series count grows then resets every restart | Series churn from ephemeral label values | Churn diagnosis |
# Total active series in the local Prometheus
prometheus_tsdb_head_series
# Or for Mimir / Grafana Cloud Metrics (per tenant)
cortex_ingester_memory_series{user="<tenant>"}Compare to recent history:
# Growth over the last 7 days
deriv(prometheus_tsdb_head_series[7d]) * 86400A growth rate > a few % per day on a stable application set is a red flag.
Prometheus exposes a built-in cardinality breakdown:
curl -s http://prometheus:9090/api/v1/status/tsdb | jqReturns:
seriesCountByMetricName — top metrics by series countlabelValueCountByLabelName — top labels by unique value countmemoryInBytesByLabelName — top labels by memory footprintseriesCountByLabelValuePair — top label-value pairs by series countThis is usually the fastest path to "which metric / which label is the problem."
For Grafana Cloud:
# Same endpoint, authenticated against the per-tenant Mimir
curl -s -u "<user>:<token>" \
"https://prometheus-prod-XX.grafana.net/api/prom/api/v1/status/tsdb" | jq"seriesCountByMetricName": [
{ "name": "http_request_duration_seconds_bucket", "value": 184320 },
{ "name": "go_gc_duration_seconds", "value": 80 },
...
]Heuristics:
_bucket) at the top is almost always the answer — those have a 14× multiplier (bucket count + 3). The fix is usually reducing the labels on the underlying histogram at the source (in instrumentation code), not stripping them at scrape and not touching the buckets themselves._total, _count, _sum) — that's a histogram or summary, count all variants together for the true impact"labelValueCountByLabelName": [
{ "name": "url", "value": 84210 },
{ "name": "trace_id", "value": 41000 },
{ "name": "pod", "value": 1820 }
]Red flags:
trace_id, request_id, session_id, query, email, path, url — these should never be labels. They belong in exemplars, logs, or traces.pod with thousands of values — see Churn diagnosis; recent churn often inflates this numberOnce you've identified a suspect metric, find which label is responsible.
# How many unique values does each label have on this metric?
count by (__name__) (
count by (__name__, label_name_here) (
http_request_duration_seconds_bucket
)
)Repeat per label, or use the helper:
# Via the Prometheus HTTP API
curl -s "http://prometheus:9090/api/v1/labels?match[]=http_request_duration_seconds_bucket" | jq -r '.data[]' | \
while read label; do
count=$(curl -s "http://prometheus:9090/api/v1/label/${label}/values?match[]=http_request_duration_seconds_bucket" | jq '.data | length')
echo "${count} ${label}"
done | sort -rn | head -20# Top 20 path values for http_requests_total
topk(20,
count by (path) (http_requests_total)
)If you see UUIDs, hashes, timestamps, or numeric IDs in the top values → that label has unbounded values from the source.
# Series-per-instance breakdown — if uneven, one instance is misbehaving
sum by (job, instance) ({__name__=~"my_metric.*"})If the cardinality fire started recently, the cause is almost always a recent change. Diff what's there now against what was there before.
Via Grafana Cloud cardinality dashboard, or:
# Current metrics
group by (__name__) ({__name__!=""})
# Compare to last week (offset)
group by (__name__) ({__name__!=""} offset 7d)Diff externally. A new metric near the top of seriesCountByMetricName that wasn't there a week ago → that's your offender.
# Active series correlated with build_info
prometheus_tsdb_head_series
# Overlay with:
changes(app_build_info[1d])A vertical step in series count aligned with a deploy is conclusive.
High churn means series are being created and abandoned faster than they age out. Symptoms: series count keeps climbing, then drops sharply on Prometheus restart.
# Series created vs. removed per second
rate(prometheus_tsdb_head_series_created_total[5m])
rate(prometheus_tsdb_head_series_removed_total[5m])
# Ratio of churned to live
prometheus_tsdb_head_series_created_total / prometheus_tsdb_head_seriesA creation rate that materially exceeds the removal rate, sustained, means cardinality is on a one-way trip up. Common causes:
| Cause | Tell |
|---|---|
Pod rollouts emitting pod label | Churn spike aligns with deploy timing; affects pod-discovered scrapes |
version / git_sha / image_tag label on every metric | Churn spike on every deploy across many metrics |
Ephemeral hostnames in instance | Cloud autoscaling event timing |
| Bug: dynamic label names | Churn climbs forever, never plateaus |
| Application bug emitting fresh UUIDs as labels | Linear unbounded growth, no deploy correlation |
# A churn-driven head block carries old series until tsdb compaction
prometheus_tsdb_head_chunks
go_memstats_heap_inuse_bytes{job="prometheus"}Restarting Prometheus drops churned series but is not a fix. The fix is at the source.
Tell: *_bucket metric at the top of seriesCountByMetricName. Multiplier ≈ 14×.
Fix:
path, method, or status_code in the instrumentation code (don't labeldrop them at scrape — that merges distinct histograms and corrupts the buckets). For series already in Grafana Cloud you can't change, aggregate them with Adaptive Metrics.Tell: kube_pod_labels or kube_pod_annotations at the top, with label_* or annotation_* labels driving cardinality.
Fix: configure kube-state-metrics with --metric-labels-allowlist and --metric-annotations-allowlist. By default it emits all labels and annotations as series.
# kube-state-metrics flags
--metric-labels-allowlist=pods=[app,team,version]
--metric-annotations-allowlist=pods=[checksum/config]Tell: http_requests_total (or framework equivalent) grew 10×+ overnight. topk(20, count by (path) (http_requests_total)) shows hundreds of /users/123456-style values.
Fix: the real fix is to template the path in application code (/users/:id) — route the user to prometheus-label-strategy. For series already in Grafana Cloud, Adaptive Metrics can aggregate path away correctly — route to adaptive-metrics.
Do not "normalize" path with a relabel replacement rule — collapsing /users/123, /users/456, … into one /users/:id value at scrape merges distinct series and produces duplicate-sample errors and broken rate(). The merge has to happen at the source (templating) or post-ingest (Adaptive Metrics), never at scrape.
If you must stop a production fire right now and templating isn't deployable yet, the only safe scrape-time action is to drop the entire offending metric (you lose it completely until the code fix lands — a deliberate trade, not a silent corruption):
# Emergency: drop the whole metric until the source is templated
metric_relabel_configs:
- source_labels: [__name__]
regex: http_requests_total
action: dropTell: A metric you don't recognize in the top 10. Grep the source — often a _details or _per_request debug metric the developer forgot to gate.
Fix: drop entirely at scrape:
metric_relabel_configs:
- source_labels: [__name__]
regex: my_app_request_details
action: dropOpen a ticket against the team to remove it from the code.
Tell: Series count for one job is several × what it should be. Looking at one series, you see both an app-emitted instance=... AND the target instance=... collided into something weird (Prometheus renames the conflicting one to exported_instance).
Fix: the right fix is in the application — stop emitting instance/node/host from code; they belong to the scrape target. Confirm honor_labels is false (the default) so the target labels win.
If you need a scrape-time stopgap, you may remove a label only where it exactly duplicates a target label — that's the one safe labeldrop, because the target label still provides uniqueness. Scope it tightly to the duplicated names and never include pod (or any other label that is the source of uniqueness):
# Stopgap ONLY for app-emitted duplicates of target labels.
# Drops the `exported_*` collisions — NOT pod, which makes K8s series unique.
metric_relabel_configs:
- regex: exported_(instance|node|host)
action: labeldropThen the target labels from relabel_configs apply cleanly. Prefer fixing the app.
Tell: A federated Prometheus or Mimir global view has way more series than expected. Each source has its own cluster / region label, multiplying.
Fix: this is usually expected — federation by design preserves source labels. If the series count is too high, federate only aggregated recording rules, not raw metrics:
- job_name: federate
honor_labels: true
metrics_path: /federate
params:
'match[]':
- '{__name__=~".*:.*"}' # Recording-rule naming convention onlyCardinality fire confirmed
│
├── Need to stop the bleeding NOW (production OOM, ingest 429s)
│ └── Drop the ENTIRE offending metric via metric_relabel_configs (action: drop on __name__)
│ (also applies to Alloy/Agent — same syntax)
│ Do NOT labeldrop a distinguishing label — it breaks the data, see "The One Rule".
│ Then schedule the proper fix.
│
├── It's a Grafana Cloud Active Series bill issue, not a perf issue
│ ├── Cardinality is structural and you can't fix the app
│ │ └── Route to `adaptive-metrics` skill (post-ingest aggregation rules — the safe way)
│ └── You want metric-by-metric DPM breakdown
│ └── Route to `dpm-finder` skill
│
├── It's a fixable application bug (unbounded label, debug metric in prod)
│ ├── Short-term: drop the whole metric at scrape, OR aggregate via Adaptive Metrics
│ └── Long-term: fix in code; route to `prometheus-label-strategy` for design guidance
│
├── It's histogram cardinality
│ ├── Reduce labels on the underlying histogram AT THE SOURCE (14× win per label)
│ ├── Reduce bucket count if appropriate
│ └── Consider native histograms for high-resolution latency
│
└── It's churn (deploy-driven)
├── Stop EMITTING `version`/`git_sha`/`instance` from app code (use info-metric for version)
├── Keep `pod` — never drop it; if pod-level series are too costly, use Adaptive Metrics
└── Verify K8s SD relabel rules aren't mapping in `uid` or other ephemeral fieldsThese are the safe scrape-time emergency actions: dropping an entire unwanted metric. They do not merge distinct series, so they don't corrupt the data.
⚠️ There is intentionally no
labeldropof a distinguishing label and no value-normalizing relabel here. Both merge distinct series and breakrate()/DPM (see The One Rule). To reduce cardinality without dropping the whole metric, fix the source or use Adaptive Metrics (route toadaptive-metrics). The only safelabeldropis removing a label that exactly duplicates a target label (e.g.exported_instance) — see App-emitted labels colliding with target labels.
For Prometheus scrape_configs:
metric_relabel_configs:
# Drop a specific bad metric entirely
- source_labels: [__name__]
regex: bad_metric_name
action: drop
# Drop a set of debug/temporary metrics by name prefix
- source_labels: [__name__]
regex: debug_.*
action: dropFor Grafana Alloy (prometheus.relabel component):
prometheus.relabel "drop_bad_metric" {
forward_to = [prometheus.remote_write.default.receiver]
rule {
source_labels = ["__name__"]
regex = "bad_metric_name"
action = "drop"
}
}Always test in staging first, and prefer fixing the source or using Adaptive Metrics over any scrape-time drop.
prometheus-label-strategyadaptive-metricsdpm-finderpromqlalloyloki-label-analyzer (different system, same family of problems)This skill's lane is diagnosis under pressure. Prevention, design, and post-ingest cost optimization live elsewhere.
© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/grafana-cloud/prometheus-cardinality-troubleshooter of grafana/skills.
Open the folder on GitHubat commit 1ccacf2
Prometheus Cardinality Troubleshooter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Prometheus Cardinality Troubleshooter this skillgrafana/skills | 282 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Happy Infra Metrics and Grafanaslopus/happy | 24k | — | ~2k | Automated safety check: Notes | MIT | |
| Syncmetapawurb/hotpath-rs | 1.9k | — | ~1.2k | Automated safety check: Notes | MIT | |
| Optimize Slurm TopologyNVlabs/alpasim | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Dashboard Previewm4r1k/Eneru | 149 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Graftm4r1k/Eneru | 149 | 1 repos | ~2.3k | Automated safety check: Pass | MIT |
slopus/happy
Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.
pawurb/hotpath-rs
Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).
NVlabs/alpasim
Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.
m4r1k/Eneru
Visually verify Eneru browser-dashboard changes against a live daemon or audit an exact deployment.
m4r1k/Eneru
This repo is indexed by graft/. An agent skill from m4r1k/Eneru.
m4r1k/Eneru
Mandatory pre-release deep review for minor/major releases (X.Y.0 / X.0.0).
grafana/skills
Write or review k6 documentation across the three k6 repositories - k6-DefinitelyTyped (TypeScript types), k6-docs (user documentation), and k6 (release notes / changelog).
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
grafana/skills
Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units…
grafana/skills
A skill your agent uses when the user wants to performance-test, load-test, or stress-test a public website end-to-end with k6.
grafana/skills
Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.
grafana/skills
Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config…
Works with
Categories
Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid…. Prometheus Cardinality Troubleshooter is an agent skill from grafana/skills, published by the product's own GitHub organization. Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid memory growth.
Prometheus Cardinality Troubleshooter fits situations like: the user is currently experiencing a cardinality fire; tasks that involve Monitoring and alerting.
Run `npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a claude-code`. Or copy the skill folder (skills/grafana-cloud/prometheus-cardinality-troubleshooter in grafana/skills) into .claude/skills/prometheus-cardinality-troubleshooter in your project. Claude Code loads it when a task matches its description.
Run `npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a codex`. Or copy the skill folder (skills/grafana-cloud/prometheus-cardinality-troubleshooter in grafana/skills) into .agents/skills/prometheus-cardinality-troubleshooter in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/skills --skill prometheus-cardinality-troubleshooter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prometheus-cardinality-troubleshooter, .gemini/skills/prometheus-cardinality-troubleshooter, .github/skills/prometheus-cardinality-troubleshooter and .opencode/skills/prometheus-cardinality-troubleshooter in your project.
Going by SKILL.md and its folder, Prometheus Cardinality Troubleshooter needs the command-line tools its instructions call (curl and jq).
SKILL.md names 1 domain. In commands or code: prometheus-prod-xx.grafana.net; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Prometheus Cardinality Troubleshooter is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Prometheus Cardinality Troubleshooter: Happy Infra Metrics and Grafana (slopus/happy, 24k stars), Syncmeta (pawurb/hotpath-rs, 1.9k stars), Optimize Slurm Topology (NVlabs/alpasim, 1.3k stars) and Dashboard Preview (m4r1k/Eneru, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
grafana (a GitHub organization, an official publisher) maintains it in grafana/skills, which has 282 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on October 8, 2026.
Source: grafana/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.