Official agent skill

Prometheus Label Strategy

by grafana in grafana/skills

Expert evaluator for Prometheus label strategy on Grafana Cloud.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Prometheus Label Strategy

skills CLI
$ npx skills add grafana/skills --skill prometheus-label-strategy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install grafana/skills prometheus-label-strategy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/grafana-cloud/prometheus-label-strategy .claude/skills/prometheus-label-strategy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prometheus-label-strategy
GitHub stars
282
Token cost
~6.4k tokens
SKILL.md length
2,931 words
Files
1
Skills in repo
51
Repo updated
First seen
Licence
Apache-2.0

At a glance

Expert evaluator for Prometheus label strategy on Grafana Cloud.

  • Works in 5 steps: Fix in the Application (best) → relabel_configs (target-time relabeling) → Adaptive Metrics (Grafana Cloud —… → …
  • The user asks to evaluate
  • SKILL.md covers The One Rule: Never Drop a…, Core Concepts, Label Evaluation Framework and Evaluation Output Format, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Prometheus Label Strategy is an agent skill from grafana/skills, published by the product's own GitHub organization. Expert evaluator for Prometheus label strategy on Grafana Cloud. Audits, designs, and improves label schemas using cardinality scoring, access-pattern alignment, static vs. dynamic label rules, histogram bucket discipline, and instrumentation hygiene. Prevents high cardinality at the source — in application code and scrape target labels — without dropping labels that make series unique (which breaks the data). For reducing the cost of series already in Grafana Cloud, routes to the adaptive-metrics skill. Use when…

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Monitoring and alerting. It works with Prometheus and Grafana. The licence is Apache-2.0.

When your agent uses it

  • The user asks to evaluate
  • Improve Prometheus labels —
  • Asks how to prevent high cardinality at the source

Example prompts

  • “why is my Prometheus slow / expensive right now”
  • “/prometheus-label-strategy”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Fix in the Application (best)
  2. relabel_configs (target-time relabeling)
  3. Adaptive Metrics (Grafana Cloud — post-ingest, the safe way to reduce cardinality)
  4. metric_relabel_configs (narrow, safe uses only)
  5. Recording Rules (query-time cardinality reduction)

What it can do on your machine

Read from SKILL.md and the folder at commit 1ccacf2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and promql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prometheus Label Strategy loads about 6.4k tokens when it runs. Until then it costs about 194 tokens; SKILL.md has 2,931 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~194
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from grafana/skills at commit 1ccacf2, republished under its Apache-2.0 licence (© grafana). 2,931 words, ~6,375 tokens.

Download SKILL.mdSave it as .claude/skills/prometheus-label-strategy/SKILL.md (or your agent's skills folder).
name
prometheus-label-strategy
description
Expert evaluator for Prometheus label strategy on Grafana Cloud. Audits, designs, and improves label schemas using cardinality scoring, access-pattern alignment, static vs. dynamic label rules, histogram bucket discipline, and instrumentation hygiene. Prevents high cardinality at the source — in application code and scrape target labels — without dropping labels that make series unique (which breaks the data). For reducing the cost of series already in Grafana Cloud, routes to the adaptive-metrics skill. Use when the user asks to evaluate, audit, design, or improve Prometheus labels — or asks how to prevent high cardinality at the source. For "why is my Prometheus slow / expensive right now" triage, see prometheus-cardinality-troubleshooter.
license
Apache-2.0

Prometheus Label Strategy Evaluator

You are an expert in Prometheus label strategy. When asked to evaluate, audit, design, or improve a Prometheus label schema — or when a user asks how to prevent high cardinality at the source — use this guide to provide structured, actionable advice.

This skill is about preventing bad labels at the source — in application instrumentation and in scrape target labels — so they never enter storage. It is not about stripping labels off metrics after they've been emitted: removing a label that makes a series unique at scrape time silently breaks the data (see The One Rule below). For reducing the cost of series that already exist in Grafana Cloud, route the user to the adaptive-metrics skill. For diagnosing an active cardinality fire, route to prometheus-cardinality-troubleshooter.


The One Rule: Never Drop a Label That Makes a Series Unique

You cannot remove, at scrape time, any label that makes a series unique. Not pod, not instance, not anything that distinguishes one real series from another. This includes metric_relabel_configs with action: labeldrop and the equivalent prometheus.relabel rules in Alloy.

It looks like a cardinality win. It is not — it breaks the data, silently and permanently:

  • Counter resets get mixed together. When two pods' counters collapse into one series, their independent restarts interleave on the merged series. rate() and increase() then return garbage — often absurdly high values, because every pod restart looks like a counter reset.
  • DPM inflates instead of dropping. Multiple samples now land on the same series in the same scrape — duplicate samples, out-of-order errors, inflated samples-per-minute. People come back weeks later asking "why is my DPM so high?" or "why is rate() returning absurd numbers?" — and there is no evidence left in the data of where it broke.
  • The aggregation is wrong, not just coarse. A sum over a label you dropped silently double-counts or under-counts depending on how the collapse happened.

The trap is that none of this errors at config time. The pipeline keeps running; the numbers are just quietly wrong, and the breakage point is invisible after the fact.

The right tools, in order:

  1. Don't emit the bad label in the first place — fix the application code. This is the only place a label can be removed without consequence, because the series was never unique on it to begin with.
  2. For series already flowing into Grafana Cloud that you can't fix at the source → Adaptive Metrics. This is exactly what it is for: it aggregates series correctly — counter-reset-aware, with a recorded audit trail, and reversible — instead of blindly stripping labels. Route the user to the adaptive-metrics skill.

metric_relabel_configs has a couple of narrow, safe uses (dropping an entire unwanted metric; removing a label that exactly duplicates a target label) — covered in Source-Side Prevention — but reducing cardinality by dropping a distinguishing label is never one of them.


Core Concepts

Series are the fundamental unit in Prometheus. Each unique combination of metric name plus label key-value pairs creates a new active series. Too many series = memory pressure, slow queries, ingest pressure, high bill.

Cardinality = the number of unique values a label can have. Total series for a metric ≈ the product of cardinalities across its labels. A metric with path (100 values), status_code (10 values), method (5 values), and instance (50 values) = 250,000 series per metric. Adding one more high-cardinality label often 10–100×s the count.

The dual impact rule: High-cardinality labels hurt on both paths:

  • Ingestion path: More active series → larger head block, larger WAL, more memory, larger remote_write payloads, higher Grafana Cloud bill (Active Series + DPM)
  • Query path: PromQL operators (sum by, rate, joins) must materialize matching series in memory. High cardinality balloons query memory and latency

Series churn is the silent killer. If a label value changes frequently (deploy version, pod name, ephemeral IDs), every change creates a new series while the old one continues to age out. Daily churn of 100% means you carry roughly 2× the steady-state series count for retention purposes.

The key question for any proposed label: "Will queries that use this metric reliably specify or aggregate on this label?" If no → it should NOT be a label.


Label Evaluation Framework

When auditing a label set, assess each label against these criteria.

Cardinality Scoring
Label ExampleCardinalityVerdict
env (prod/staging/dev)2–5 values✅ Good
job (Prometheus scrape job)5–50 values✅ Good
cluster, regionTens✅ Good
namespace (K8s)Tens–low hundreds✅ Acceptable
service, workload, containerTens–hundreds✅ Acceptable
instance (host:port)Hundreds–low thousands⚠️ Evaluate — fine on per-instance metrics, risky on aggregated ones
pod (K8s)Thousands + transient = high churn⚠️ Required for K8s monitoring and series uniqueness — keep it. If pod-level series are too expensive, reduce them with Adaptive Metrics; never drop at scrape
path / route (HTTP)Bounded if templated; unbounded if raw URLs⚠️ Only with templated values (/users/:id)
version, image_tag, git_shaGrows on every deploy → churn⚠️ Use sparingly; consider info-metric pattern
user_id, request_id, trace_idUnbounded❌ Never as label — use exemplars
customer_id, tenant_idOften unbounded❌ Only acceptable for small fixed tenant counts
error_message, query, sqlUnbounded text❌ Never
Access Pattern Alignment

For each label, ask:

  • Do queries on this metric reliably aggregate by or filter on this label?
  • Does this label logically segment the metric the way users think about it?
  • Would removing this label force users to use exemplars, logs, or traces instead — and would that be acceptable for the rare lookup case?
Static vs. Dynamic Label Values
  • Static / target labels (set once per scrape target via relabel_configs, e.g., env=prod, cluster=us-east, team=payments) add cardinality proportional to targets, not requests. Cheap and high-value. Use freely.
  • Dynamic / sample labels (emitted by the application per measurement, e.g., status_code, method, cache_hit) multiply cardinality by value count. Keep possible values in the single digits or low tens. The application code is the source of truth — fix it there, not in Prometheus.
Consistency Check
  • Label names consistent across services? (status vs status_code vs http_status produces three separate label families — joins break)
  • Label values normalized? (200 vs "200", GET vs get, Error vs error)
  • Naming convention consistent? Prometheus convention is snake_case for both metric and label names
  • Same concept, same name across services? (service vs svc vs app_name)
Histogram Bucket Discipline (critical, often missed)

Every histogram metric multiplies its base cardinality by (bucket count + 3) — buckets via _bucket{le="..."} plus _sum, _count, and _created (Prometheus 2.39+).

  • Default prometheus.DefBuckets has 11 buckets → 14× multiplier
  • A histogram with method, path, status already at 1,000 series becomes 14,000 series after adding histogram cardinality
  • Always trim histogram label cardinality first — labels matter 14× more on histograms than on counters/gauges
  • Consider native histograms (Prometheus 2.40+) which use a single sparse series instead of one-per-bucket — major cardinality reduction for high-resolution latency tracking
Info-Metric Pattern (for high-churn metadata)

When you want to know about a label (e.g., version, git_sha, image_tag) without paying for it on every metric, use an info metric:

# A single low-cardinality counter/gauge of value 1, with the metadata attached
app_build_info{app="payment-api", version="2.4.1", git_sha="a1b2c3"} 1

Then join at query time. The classic approach is a vector match with group_left:

promql
sum by (version) (
  rate(http_requests_total{app="payment-api"}[5m])
  * on (app) group_left (version) app_build_info
)

The version label lives on exactly one series per build, not on every metric.

The info() function (simpler join)

PromQL's info() function (experimental, Prometheus 3.0+; enable with --enable-feature=promql-experimental-functions) automates the info-metric join so you don't have to hand-write the * on (...) group_left (...) match:

promql
info(
  rate(http_requests_total{app="payment-api"}[5m]),
  {version=~".+"}
)

info(v, [labelselector]) takes a range/instant vector v and, for each series, finds matching info metrics and adds their labels. The optional second argument is a label-matcher restricting which info labels are attached (here, only version). By default info() joins against the conventional target_info metric and matches on identifying labels (e.g. instance, job), so it's especially ergonomic for OpenTelemetry-style target_info. For custom info metrics like app_build_info the explicit group_left form above is still the most portable.

Prefer info() when you're on Prometheus 3.x and joining against target_info; fall back to the explicit group_left match for older versions, custom info metrics, or when the experimental feature flag isn't enabled.


Evaluation Output Format

When auditing a label set, produce a report in this structure:

## Prometheus Label Strategy Audit

### Summary
[1-2 sentence overall assessment — total estimated active series, biggest risks]

### Per-Label Analysis
| Metric Family | Label | Cardinality | Used in Queries? | Verdict | Action |
|---|---|---|---|---|---|
| http_requests_total | path | Unbounded (raw URLs) | Sometimes | ❌ Remove | Template in code: `/users/:id` not `/users/12345` |
| http_requests_total | pod | High + churn | Rarely | ⚠️ Keep — makes the series unique | If too expensive, aggregate away with Adaptive Metrics; query by `workload` for the common case |

### Histogram-Specific Findings
[Highlight any histograms with high label cardinality — these are 14×+ amplified]

### Estimated Impact
- Active series reduction: [X series → Y series]
- DPM reduction: [X DPM → Y DPM]  (samples-per-minute = series × ~6 at 10s scrape)
- Memory impact: [if measurable]

### Recommended Label Set
[Final recommended labels per metric family]

### Implementation Plan
1. [Code changes — instrumentation hygiene: stop emitting bad labels at the source]
2. [Scrape target labels — relabel_configs (additive: env, cluster, team, workload)]
3. [Post-ingest cost reduction on series you can't fix at the source — Adaptive Metrics]
4. [Recording rules to materialize useful aggregates]

These should be set as target labels (via relabel_configs on the scrape job, NOT emitted by the app) — they're per-target, low cardinality, high query value:

LabelPurposeNotes
jobPrometheus scrape job nameSet automatically by Prometheus
instanceTarget endpoint (host:port)Set automatically; rename via relabel_configs to a friendlier value if needed
envEnvironment (prod, staging, dev)Set via static_configs labels or service discovery
clusterMulti-cluster differentiationCritical for federation/Mimir multi-tenant
regionGeographic region
team / squadOwnership — also useful for access control
serviceLogical service identityOne service may span multiple jobs

These should NOT be re-emitted by the application. If the app emits a cluster label, it duplicates the target label and creates collisions / honor_labels decisions you don't want to make.


Kubernetes Patterns

LabelSourceNotes
namespacePod metadataAlways keep
containerPod specLow cardinality, useful for multi-container pods
workloadDerived: {controller_kind}/{controller_name}Add as a stable aggregation key alongside pod — static, predictable. It's an addition, not a replacement: don't use it as an excuse to drop pod
serviceK8s ServiceIf scraping via Service
Handling the pod Label

pod is high-cardinality and transient — it rolls on every deploy and restart, so it dominates churn and series count. But it is also a label that makes K8s series unique, and Kubernetes monitoring (per-pod resource attribution, kube-state-metrics joins) depends on it. The One Rule applies: do not drop pod at scrape time. Collapsing pods into one series mixes their counter resets and breaks rate().

Instead:

  • Add workload ({controller_kind}/{controller_name}) as a target label via relabel_configs, so dashboards and alerts can aggregate on the stable workload identity (sum by (workload)) without touching pod. This is additive — it removes nothing.
  • Don't emit pod from application code — let it come from Kubernetes service discovery, so there is exactly one source of truth (see below).
  • If pod-level series are genuinely too expensive in Grafana Cloud, reduce them with Adaptive Metrics, which aggregates pod away correctly (post-ingest, counter-reset-aware, reversible) rather than corrupting the raw data at scrape. Route to the adaptive-metrics skill.
Don't Map Ephemeral Fields into Labels in the First Place

uid regenerates on every pod recreation and has no legitimate query use. The fix is to never map it into a label — leave it out of your relabel_configs. (It isn't in default kubernetes_sd_configs output unless you explicitly target it.) Don't try to labeldrop it after the fact — by then it's already distinguishing series, and removing it breaks the data exactly like dropping any other unique label.

One Source of Truth for Target Identity

instance, pod, node, and host should come from scrape target labels, not from application code. If the app also emits its own instance/node, you get duplicates and honor_labels collisions. The fix is in the application — stop emitting them — not a scrape-time labeldrop. (Removing a label that exactly duplicates a target label is the one narrow exception; see metric_relabel_configs.)

Show full SKILL.md (1,147 more words)Show less
kube-state-metrics label propagation ⚠️
  • kube_pod_labels{label_app_kubernetes_io_*=...} can carry dozens of metadata labels
  • Each unique pod label combination is a new series
  • Restrict at the source with kube-state-metrics' --metric-labels-allowlist — this controls what is ever emitted, so it's prevention, not destructive after-the-fact dropping

Source-Side Prevention: Where to Fix What

There are five levers, in order of preference:

1. Fix in the Application (best)

Bad labels emitted by the app are the root cause. Examples:

  • HTTP paths: use templated routes (/users/:id) not raw paths
  • Error metrics: use a small enum (error_type="timeout") not the error message string
  • User-scoped metrics: don't include user_id — use exemplars to point to logs/traces
  • Free-form input: never emit user-supplied strings as label values

If you control the code, this is always the right fix. It saves cost on every downstream system (Prometheus, remote_write, Mimir, Grafana Cloud).

2. relabel_configs (target-time relabeling)

Runs before the scrape. Used to:

  • Set target labels (env, cluster, team) on discovered targets
  • Drop entire targets you don't want to scrape
  • Rewrite instance to a friendly value
  • Add identity from service discovery metadata
yaml
scrape_configs:
  - job_name: my-app
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      # Set workload from controller metadata
      - source_labels: [__meta_kubernetes_pod_controller_kind, __meta_kubernetes_pod_controller_name]
        target_label: workload
        separator: /
      # Set env from a pod label
      - source_labels: [__meta_kubernetes_pod_label_env]
        target_label: env
      # Only scrape pods explicitly opted in
      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
        regex: "true"
        action: keep
3. Adaptive Metrics (Grafana Cloud — post-ingest, the safe way to reduce cardinality)

When the cardinality is structural and you can't fix it at the source — the label legitimately exists and makes series unique, you just don't need every value at full resolution — Adaptive Metrics is the correct tool, and the only safe way to reduce the cost of series that already exist.

It works after ingest, as aggregation rules applied in Grafana Cloud. Crucially, it aggregates series correctly:

  • It handles counter resets properly, so rate() and increase() stay accurate.
  • It records what was aggregated, so there's an audit trail — you can answer "why did this change?" later.
  • It's reversible: drop a rule and the full-resolution series come back.

This is the difference between "the data is now cheaper" (Adaptive Metrics) and "the data is now wrong" (labeldrop at scrape). Route the user to the adaptive-metrics skill for rule design.

4. metric_relabel_configs (narrow, safe uses only)

Runs after the scrape, before storage.

⚠️ Do not use metric_relabel_configs (or Alloy prometheus.relabel) to drop a label that distinguishes series — pod, instance, user_id, path, anything. See The One Rule. It looks like a cardinality fix and silently breaks rate(), inflates DPM, and corrupts aggregations. Use the application code (lever 1) or Adaptive Metrics (lever 3) instead. The same caution applies to normalizing a label value (e.g. collapsing status_code to 2xx) at scrape — it merges distinct series and produces duplicate-sample errors; do that in code or via Adaptive Metrics, never here.

The genuinely safe uses are:

  • Drop an entire metric you never want stored — you're discarding the whole metric, not collapsing distinct series into one:
    yaml
    metric_relabel_configs:
      - source_labels: [__name__]
        regex: my_app_request_details
        action: drop
  • Remove a label that exactly duplicates a target label. If the app emits its own cluster/instance that already comes from the scrape target, the target label still provides uniqueness, so removing the duplicate breaks nothing. Prefer fixing the app, but this is a safe stopgap.

That's the whole list. If you're reaching for metric_relabel_configs to bring down a series count, you almost certainly want Adaptive Metrics instead.

5. Recording Rules (query-time cardinality reduction)

Pre-aggregate expensive series into a lower-cardinality recorded series. Stored at the same data point density but with far fewer series.

yaml
groups:
  - name: http-requests-aggregates
    interval: 30s
    rules:
      # Drop pod/instance dimension; keep only service-level rollup
      - record: service:http_requests:rate5m
        expr: sum by (service, env, cluster, status_code) (rate(http_requests_total[5m]))

Queries that target the rollup are dramatically cheaper. The raw series still exist — recording rules don't reduce ingest cost (use Adaptive Metrics for that — not a scrape-time labeldrop). They reduce query cost.


Instrumentation Hygiene (for app developers)

If the user is writing instrumentation code, these are the rules:

RuleWhy
Never use unbounded user input as a label valueemail, user_id, query string, error message — they're the #1 cardinality bug
Template HTTP paths before recording/users/{id} not /users/12345. Most frameworks do this via routing metadata
Bound error labels via small enumserror_type="timeout" not error="connection to db-shard-7 timed out at 14:32:09"
Don't put version / git_sha / build_id on every metricUse an info metric and join at query time
Don't emit pod / node / host from codeComes from scrape targets — duplicating creates collisions
Avoid dynamically constructed label names (keys)metric{[user]=1} cannot be bounded — use a fixed key
Use histograms sparingly and trim labels first14× cardinality amplification
Prefer exemplars over labels for trace correlationExemplars carry trace_id without inflating cardinality
Exemplars (the escape hatch)

Exemplars attach a trace_id (or any key-value pair) to specific samples without making it a label dimension. The ideal home for high-cardinality correlation data.

Requires OpenMetrics format, Prometheus 2.26+, scrape config:

yaml
scrape_configs:
  - job_name: my-app
    enable_protobuf_negotiation: true
    # Or for text-format:
    follow_redirects: true

And on the Prometheus server:

yaml
storage:
  exemplars:
    max_exemplars: 100000

Use exemplars for:

  • trace_id correlation (Tempo, Jaeger)
  • request_id for specific debug lookups
  • Any sparse "useful when you need it" key

Query exemplars via Grafana's exemplars-on-graph feature, not via PromQL aggregation.


The 80/20 Rule

The most impactful improvements almost always come from these five changes:

  1. Drop unbounded labels at the app layer — path (untemplated), user_id, error_message. Single biggest win.
  2. Trim histogram label cardinality before anything else — 14× amplification on every histogram.
  3. Don't emit pod/instance/node from application code — let them come from scrape targets, and add a stable workload target label to aggregate on. (Never drop the real pod at scrape to cut cardinality — if pod-level series are too expensive, use Adaptive Metrics.)
  4. Use info metrics for version / git_sha / image_tag — eliminates deploy-driven churn.
  5. Set target labels via relabel_configs, not app code — env, cluster, team, service should never be emitted by the application.

Focus on these before anything else.


Labels to Avoid — Quick Reference

LabelWhyAlternative
user_id, customer_id (large tenant base)UnboundedExemplars; aggregate by tenant_tier
request_id, trace_idUnboundedExemplars
path / route (raw URLs)UnboundedTemplate in code: /users/:id
error_message, query, sqlUnbounded textBounded error_type enum
version, git_sha, image_tag (on every metric)Churn on every deployInfo metric pattern
App-emitted pod (duplicating SD)Should come from K8s service discovery, not codeStop emitting it in code; keep the discovered pod. Never drop the real pod to cut cardinality — use Adaptive Metrics
uid (K8s)Unbounded; regenerates on restartNever map it into a label in the first place (leave it out of relabel_configs)
Application-emitted instance, node, hostShould come from scrape targetStop emitting in code (removing an exact target-label duplicate at scrape is the only safe drop)
Dynamically-named label keysCannot be boundedUse fixed keys with bounded values
Raw status_code on histograms14× amplificationBucket to status_class (2xx, 4xx, 5xx)

When to Route Elsewhere

  • "Reduce my Grafana Cloud bill" / "reduce cardinality on series already ingested" → engage adaptive-metrics skill (post-ingest aggregation rules — the safe, counter-reset-aware way; never labeldrop distinguishing labels at scrape)
  • "Which metrics are driving my DPM?" → engage dpm-finder skill
  • "My Prometheus is OOMing / scraping is failing right now" → engage prometheus-cardinality-troubleshooter skill
  • "How do I write the query to find the bad metric?" → engage promql skill
  • "How do I configure relabel rules in Alloy?" → engage alloy skill

This skill's lane is strategy and design. Other skills own diagnosis and operational remediation.

© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/grafana-cloud/prometheus-label-strategy of grafana/skills.

Open the folder on GitHubat commit 1ccacf2

Compare with similar skills

Prometheus Label Strategy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prometheus Label Strategy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prometheus Label Strategy this skillgrafana/skills282—~6.4kAutomated safety check: PassApache-2.0
Happy Infra Metrics and Grafanaslopus/happy24k—~2kAutomated safety check: NotesMIT
Syncmetapawurb/hotpath-rs1.9k—~1.2kAutomated safety check: NotesMIT
Optimize Slurm TopologyNVlabs/alpasim1.3k—~1.6kAutomated safety check: PassApache-2.0
Dashboard Previewm4r1k/Eneru149—~1.4kAutomated safety check: PassMIT
Graftm4r1k/Eneru1491 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Syncmeta

    pawurb/hotpath-rs

    Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).

    1.9k GitHub stars~1.2k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.

    1.3k GitHub stars~1.6k tokensUpdated 22 days ago
    DevOps & CloudAuto-check passed
  • Visually verify Eneru browser-dashboard changes against a live daemon or audit an exact deployment.

    149 GitHub stars~1.4k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Graft

    m4r1k/Eneru

    This repo is indexed by graft/. An agent skill from m4r1k/Eneru.

    149 GitHub starsUsed in 1 repo~2.3k tokens
    DevOps & CloudAuto-check passed
  • Release Review

    m4r1k/Eneru

    Mandatory pre-release deep review for minor/major releases (X.Y.0 / X.0.0).

    149 GitHub stars~1.9k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed

More from grafana/skills

All 51 skills in this repo
  • K6 Docs

    grafana/skills

    Official

    Write or review k6 documentation across the three k6 repositories - k6-DefinitelyTyped (TypeScript types), k6-docs (user documentation), and k6 (release notes / changelog).

    282 GitHub stars~678 tokensUpdated 2 days ago
    Auto-check passed
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Dashboarding

    grafana/skills

    Official

    Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units…

    282 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • K6 Perf Test Website

    grafana/skills

    Official

    A skill your agent uses when the user wants to performance-test, load-test, or stress-test a public website end-to-end with k6.

    282 GitHub stars~3.3k tokensUpdated 2 days ago
    Auto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    282 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Adaptive Metrics

    grafana/skills

    Official

    Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config…

    282 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Prometheus Label Strategy

What does Prometheus Label Strategy do?

Expert evaluator for Prometheus label strategy on Grafana Cloud. Prometheus Label Strategy is an agent skill from grafana/skills, published by the product's own GitHub organization. Expert evaluator for Prometheus label strategy on Grafana Cloud.

When should I use Prometheus Label Strategy?

Prometheus Label Strategy fits situations like: the user asks to evaluate; improve Prometheus labels —; asks how to prevent high cardinality at the source.

How do I install Prometheus Label Strategy in Claude Code?

Run `npx skills add grafana/skills --skill prometheus-label-strategy -a claude-code`. Or copy the skill folder (skills/grafana-cloud/prometheus-label-strategy in grafana/skills) into .claude/skills/prometheus-label-strategy in your project. Claude Code loads it when a task matches its description.

How do I install Prometheus Label Strategy in Codex?

Run `npx skills add grafana/skills --skill prometheus-label-strategy -a codex`. Or copy the skill folder (skills/grafana-cloud/prometheus-label-strategy in grafana/skills) into .agents/skills/prometheus-label-strategy in your project. Codex loads it when a task matches its description.

Can I use Prometheus Label Strategy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/skills --skill prometheus-label-strategy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prometheus-label-strategy, .gemini/skills/prometheus-label-strategy, .github/skills/prometheus-label-strategy and .opencode/skills/prometheus-label-strategy in your project.

What does Prometheus Label Strategy need to run?

SKILL.md names no scripts, command-line tools or credentials: Prometheus Label Strategy is instructions for the agent only.

Does Prometheus Label Strategy access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prometheus Label Strategy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prometheus Label Strategy use?

Prometheus Label Strategy is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prometheus Label Strategy use?

About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prometheus Label Strategy?

Skills that share tags, products or a category with Prometheus Label Strategy: Happy Infra Metrics and Grafana (slopus/happy, 24k stars), Syncmeta (pawurb/hotpath-rs, 1.9k stars), Optimize Slurm Topology (NVlabs/alpasim, 1.3k stars) and Dashboard Preview (m4r1k/Eneru, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prometheus Label Strategy?

grafana (a GitHub organization, an official publisher) maintains it in grafana/skills, which has 282 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on October 8, 2026.

Source: grafana/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.