Agent skill

Prometheus Grafana

by BagelHole in BagelHole/DevOps-Security-Agent-Skills

Set up metrics collection and visualization with Prometheus and Grafana.

MITAuto-check passedDevOps & Cloud

Install Prometheus Grafana

skills CLI
$ npx skills add BagelHole/DevOps-Security-Agent-Skills --skill prometheus-grafana -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BagelHole/DevOps-Security-Agent-Skills prometheus-grafana --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BagelHole/DevOps-Security-Agent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/devops/observability/prometheus-grafana .claude/skills/prometheus-grafana && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prometheus-grafana
GitHub stars
1.2k
Token cost
~2.5k tokens
SKILL.md length
221 words
Files
7 (incl. scripts, references, assets)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Set up metrics collection and visualization with Prometheus and Grafana.

  • Implementing monitoring
  • SKILL.md covers When to Use This Skill, Prerequisites, Prometheus Setup and Kubernetes Deployment, plus 9 more sections
  • Runs Shell scripts from its folder; calls helm; reaches prometheus-community.github.io and hooks.slack.com; needs GF_SECURITY_ADMIN_PASSWORD
  • Metrics collection

What it does

Prometheus Grafana is an agent skill from BagelHole/DevOps-Security-Agent-Skills. Set up metrics collection and visualization with Prometheus and Grafana. Configure scrape targets, create PromQL queries, build dashboards, and implement alerting. Use when implementing monitoring, metrics collection, or visualization for applications and infrastructure.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts, reference files and assets (for example `assets/grafana-dashboard-template.json`, `assets/prometheus-config.yaml` and `references/alerting-rules.md`).

It sits in DevOps & Cloud, covering Monitoring and alerting. It works with Prometheus, Grafana and Kubernetes. The repository describes itself as: Agent-ready DevOps, security, infrastructure, and compliance knowledge base with 80+ skills across Kubernetes, Terraform, AWS/Azure/GCP, AI platform operations, container… The licence is MIT.

When your agent uses it

  • Implementing monitoring
  • Metrics collection
  • Visualization for applications and infrastructure

Example prompts

  • “/prometheus-grafana”

Requirements

  • Node.js
  • A Bash shell
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 0365f57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • helm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • prometheus-community.github.io
    • hooks.slack.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GF_SECURITY_ADMIN_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prometheus Grafana loads about 2.5k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 73 tokens; SKILL.md has 221 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from BagelHole/DevOps-Security-Agent-Skills at commit 0365f57, republished under its MIT licence (© BagelHole). 221 words, ~2,516 tokens.

Download SKILL.mdSave it as .claude/skills/prometheus-grafana/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
prometheus-grafana
description
Set up metrics collection and visualization with Prometheus and Grafana. Configure scrape targets, create PromQL queries, build dashboards, and implement alerting. Use when implementing monitoring, metrics collection, or visualization for applications and infrastructure.
license
MIT
metadata.author
devops-skills
metadata.version
1.0

Prometheus & Grafana

Collect metrics and visualize system performance with the Prometheus-Grafana stack.

When to Use This Skill

Use this skill when:

  • Setting up metrics collection infrastructure
  • Creating monitoring dashboards
  • Writing PromQL queries for analysis
  • Configuring alerting rules
  • Monitoring Kubernetes clusters

Prerequisites

  • Docker or Kubernetes for deployment
  • Network access to monitored targets
  • Basic understanding of metrics concepts

Prometheus Setup

Docker Deployment
yaml
# docker-compose.yml
version: '3.8'

services:
  prometheus:
    image: prom/prometheus:v2.48.0
    ports:
      - "9090:9090"
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml
      - ./rules:/etc/prometheus/rules
      - prometheus-data:/prometheus
    command:
      - '--config.file=/etc/prometheus/prometheus.yml'
      - '--storage.tsdb.path=/prometheus'
      - '--storage.tsdb.retention.time=15d'

  grafana:
    image: grafana/grafana:10.2.0
    ports:
      - "3000:3000"
    volumes:
      - grafana-data:/var/lib/grafana
    environment:
      - GF_SECURITY_ADMIN_PASSWORD=admin

volumes:
  prometheus-data:
  grafana-data:
Configuration
yaml
# prometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

alerting:
  alertmanagers:
    - static_configs:
        - targets:
            - alertmanager:9093

rule_files:
  - /etc/prometheus/rules/*.yml

scrape_configs:
  - job_name: 'prometheus'
    static_configs:
      - targets: ['localhost:9090']

  - job_name: 'node'
    static_configs:
      - targets:
          - 'node-exporter:9100'

  - job_name: 'applications'
    static_configs:
      - targets:
          - 'app1:8080'
          - 'app2:8080'
    metrics_path: /metrics

Kubernetes Deployment

Using Helm
bash
# Add Prometheus community Helm repo
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts

# Install kube-prometheus-stack
helm install prometheus prometheus-community/kube-prometheus-stack \
  --namespace monitoring \
  --create-namespace \
  --set grafana.adminPassword=admin
ServiceMonitor
yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: myapp
  namespace: monitoring
spec:
  selector:
    matchLabels:
      app: myapp
  endpoints:
    - port: metrics
      interval: 30s
      path: /metrics
  namespaceSelector:
    matchNames:
      - default

PromQL Queries

Basic Queries
promql
# Current CPU usage
node_cpu_seconds_total{mode="idle"}

# Rate of HTTP requests per second
rate(http_requests_total[5m])

# Average response time
avg(http_request_duration_seconds_sum / http_request_duration_seconds_count)

# Memory usage percentage
(1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100
Aggregations
promql
# Sum requests by status code
sum by (status_code) (rate(http_requests_total[5m]))

# Average CPU by instance
avg by (instance) (rate(node_cpu_seconds_total{mode!="idle"}[5m]))

# Top 5 endpoints by request count
topk(5, sum by (endpoint) (rate(http_requests_total[5m])))

# 95th percentile latency
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
Time-Based Queries
promql
# Compare to 1 hour ago
http_requests_total - http_requests_total offset 1h

# Predict disk space in 4 hours
predict_linear(node_filesystem_avail_bytes[1h], 4 * 3600)

# Changes in last 5 minutes
changes(up[5m])

# Average over 24 hours
avg_over_time(http_requests_total[24h])

Alerting Rules

yaml
# rules/alerts.yml
groups:
  - name: application
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m]))
          / sum(rate(http_requests_total[5m])) > 0.05
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "High error rate detected"
          description: "Error rate is {{ $value | humanizePercentage }}"

      - alert: ServiceDown
        expr: up == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "Service {{ $labels.instance }} is down"

      - alert: HighMemoryUsage
        expr: |
          (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) > 0.9
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "High memory usage on {{ $labels.instance }}"
          description: "Memory usage is {{ $value | humanizePercentage }}"

      - alert: DiskSpaceLow
        expr: |
          (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) < 0.1
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Disk space low on {{ $labels.instance }}"

Alertmanager

yaml
# alertmanager.yml
global:
  resolve_timeout: 5m
  slack_api_url: 'https://hooks.slack.com/services/xxx'

route:
  receiver: 'slack-notifications'
  group_by: ['alertname', 'severity']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  routes:
    - match:
        severity: critical
      receiver: 'pagerduty'

receivers:
  - name: 'slack-notifications'
    slack_configs:
      - channel: '#alerts'
        send_resolved: true
        title: '{{ .Status | toUpper }}: {{ .CommonAnnotations.summary }}'
        text: '{{ .CommonAnnotations.description }}'

  - name: 'pagerduty'
    pagerduty_configs:
      - service_key: 'xxx'
        severity: critical

Grafana Dashboards

Dashboard JSON Structure
json
{
  "dashboard": {
    "title": "Application Metrics",
    "panels": [
      {
        "title": "Request Rate",
        "type": "graph",
        "targets": [
          {
            "expr": "sum(rate(http_requests_total[5m])) by (status_code)",
            "legendFormat": "{{ status_code }}"
          }
        ],
        "gridPos": {"x": 0, "y": 0, "w": 12, "h": 8}
      },
      {
        "title": "Latency P95",
        "type": "gauge",
        "targets": [
          {
            "expr": "histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))"
          }
        ],
        "gridPos": {"x": 12, "y": 0, "w": 6, "h": 8}
      }
    ]
  }
}
Provisioning Dashboards
yaml
# grafana/provisioning/dashboards/dashboards.yml
apiVersion: 1

providers:
  - name: 'default'
    orgId: 1
    folder: ''
    type: file
    disableDeletion: false
    updateIntervalSeconds: 30
    options:
      path: /var/lib/grafana/dashboards
Data Source Provisioning
yaml
# grafana/provisioning/datasources/prometheus.yml
apiVersion: 1

datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true
    editable: false

Recording Rules

yaml
# rules/recording.yml
groups:
  - name: aggregations
    interval: 30s
    rules:
      - record: job:http_requests:rate5m
        expr: sum by (job) (rate(http_requests_total[5m]))

      - record: instance:node_cpu:avg_rate5m
        expr: |
          avg by (instance) (
            rate(node_cpu_seconds_total{mode!="idle"}[5m])
          )

      - record: job:http_latency:p95
        expr: |
          histogram_quantile(0.95,
            sum by (job, le) (rate(http_request_duration_seconds_bucket[5m]))
          )

Application Instrumentation

Go Application
go
import (
    "github.com/prometheus/client_golang/prometheus"
    "github.com/prometheus/client_golang/prometheus/promhttp"
)

var httpRequests = prometheus.NewCounterVec(
    prometheus.CounterOpts{
        Name: "http_requests_total",
        Help: "Total HTTP requests",
    },
    []string{"method", "endpoint", "status"},
)

func init() {
    prometheus.MustRegister(httpRequests)
}

// Expose metrics endpoint
http.Handle("/metrics", promhttp.Handler())
Node.js Application
javascript
const client = require('prom-client');

const httpRequests = new client.Counter({
  name: 'http_requests_total',
  help: 'Total HTTP requests',
  labelNames: ['method', 'endpoint', 'status']
});

// Middleware
app.use((req, res, next) => {
  res.on('finish', () => {
    httpRequests.inc({
      method: req.method,
      endpoint: req.path,
      status: res.statusCode
    });
  });
  next();
});

// Expose metrics
app.get('/metrics', async (req, res) => {
  res.set('Content-Type', client.register.contentType);
  res.end(await client.register.metrics());
});

Common Issues

Issue: Targets Not Discovered

Problem: Prometheus not scraping targets Solution: Check network connectivity, verify target labels

Issue: High Memory Usage

Problem: Prometheus using excessive memory Solution: Reduce retention, use recording rules, limit cardinality

Issue: Slow Queries

Problem: PromQL queries timing out Solution: Use recording rules, limit time ranges, optimize queries

Issue: Missing Data Points

Problem: Gaps in metrics data Solution: Check scrape interval, verify target availability

Best Practices

  • Use recording rules for frequently-used queries
  • Limit label cardinality to prevent memory issues
  • Set appropriate retention based on storage capacity
  • Use histogram metrics for latency measurement
  • Implement proper alerting thresholds
  • Version control dashboards as code
  • Use federation for large-scale deployments
  • Regularly review and prune unused metrics

© BagelHole, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references, assets) in devops/observability/prometheus-grafana of BagelHole/DevOps-Security-Agent-Skills.

  • SKILL.md
  • assets/grafana-dashboard-template.json
  • assets/prometheus-config.yaml
  • references/alerting-rules.md
  • references/promql-cheatsheet.md
  • scripts/backup-grafana.sh
  • scripts/prometheus-health-check.sh

Open the folder on GitHubat commit 0365f57

Compare with similar skills

Prometheus Grafana next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prometheus Grafana compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prometheus Grafana this skillBagelHole/DevOps-Security-Agent-Skills1.2k—~2.5kAutomated safety check: PassMIT
Grafana Dashboardpando85/kaniop132—~987Automated safety check: PassAGPL-3.0
Alloygrafana/skills282—~1.3kAutomated safety check: PassApache-2.0
Qdrant Monitoring Setupqdrant/skills2542 repos~874Automated safety check: PassApache-2.0
Prometheus Grafanasickn33/agentic-awesome-skills47k1 repos~2.7kAutomated safety check: PassMIT
Beylagrafana/skills282—~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Grafana Dashboard

    pando85/kaniop

    Improve and validate the Kaniop Grafana dashboard against repository metrics and the grigri live cluster.

    132 GitHub stars~987 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Alloy

    grafana/skills

    Official

    Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…

    282 GitHub stars~1.3k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Official

    Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization.

    254 GitHub starsUsed in 2 repos~874 tokens
    DevOps & CloudAuto-check passed
  • Prometheus Grafana

    sickn33/agentic-awesome-skills

    Set up metrics collection and visualization with Prometheus and Grafana.

    47k GitHub starsUsed in 1 repo~2.7k tokens
    DevOps & CloudAuto-check passed
  • Beyla

    grafana/skills

    Official

    Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart.

    282 GitHub stars~1.1k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Fleet Management

    grafana/skills

    Official

    Manage a fleet of Grafana Alloy collectors with Fleet Management — author Alloy pipelines once, target them via attribute matchers (env="production", regex region=~"us-."), push remotely via OpAMP…

    282 GitHub stars~1.3k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed

More from BagelHole/DevOps-Security-Agent-Skills

All 44 skills in this repo
  • Hashicorp Vault

    BagelHole/DevOps-Security-Agent-Skills

    Manage secrets and PKI with HashiCorp Vault. An agent skill from BagelHole/DevOps-Security-Agent-Skills.

    1.2k GitHub stars~2k tokensUpdated 4 mo ago
    Auto-check passed
  • Incident Response

    BagelHole/DevOps-Security-Agent-Skills

    Handle security incidents with IR playbooks and procedures. An agent skill from BagelHole/DevOps-Security-Agent-Skills.

    1.2k GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Kubernetes Ops

    BagelHole/DevOps-Security-Agent-Skills

    Deploy, scale, and manage Kubernetes workloads. An agent skill from BagelHole/DevOps-Security-Agent-Skills.

    1.2k GitHub stars~2.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Linux Hardening

    BagelHole/DevOps-Security-Agent-Skills

    Apply CIS benchmarks and secure Linux servers. An agent skill from BagelHole/DevOps-Security-Agent-Skills.

    1.2k GitHub stars~662 tokensUpdated 4 mo ago
    Auto-check: notes
  • Vulnerability Scanning

    BagelHole/DevOps-Security-Agent-Skills

    Scan systems and dependencies for CVEs and security vulnerabilities.

    1.2k GitHub stars~2.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Argocd Gitops

    BagelHole/DevOps-Security-Agent-Skills

    Implement GitOps with ArgoCD for declarative Kubernetes deployments.

    1.2k GitHub stars~2.4k tokensUpdated 4 mo ago
    Auto-check passed

Categories

Questions about Prometheus Grafana

What does Prometheus Grafana do?

Set up metrics collection and visualization with Prometheus and Grafana. Prometheus Grafana is an agent skill from BagelHole/DevOps-Security-Agent-Skills. Set up metrics collection and visualization with Prometheus and Grafana.

When should I use Prometheus Grafana?

Prometheus Grafana fits situations like: implementing monitoring; metrics collection; visualization for applications and infrastructure.

How do I install Prometheus Grafana in Claude Code?

Run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill prometheus-grafana -a claude-code`. Or copy the skill folder (devops/observability/prometheus-grafana in BagelHole/DevOps-Security-Agent-Skills) into .claude/skills/prometheus-grafana in your project. Claude Code loads it when a task matches its description.

How do I install Prometheus Grafana in Codex?

Run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill prometheus-grafana -a codex`. Or copy the skill folder (devops/observability/prometheus-grafana in BagelHole/DevOps-Security-Agent-Skills) into .agents/skills/prometheus-grafana in your project. Codex loads it when a task matches its description.

Can I use Prometheus Grafana in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill prometheus-grafana -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prometheus-grafana, .gemini/skills/prometheus-grafana, .github/skills/prometheus-grafana and .opencode/skills/prometheus-grafana in your project.

What does Prometheus Grafana need to run?

Going by SKILL.md and its folder, Prometheus Grafana needs a shell for the scripts in its folder, the command-line tools its instructions call (helm) and credentials named GF_SECURITY_ADMIN_PASSWORD. Our summary lists: Node.js; A Bash shell; Docker.

Does Prometheus Grafana access the network?

SKILL.md names 2 domains. In commands or code: prometheus-community.github.io and hooks.slack.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Prometheus Grafana safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Prometheus Grafana use?

Prometheus Grafana is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prometheus Grafana use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Prometheus Grafana?

Skills that share tags, products or a category with Prometheus Grafana: Grafana Dashboard (pando85/kaniop, 132 stars), Alloy (grafana/skills, 282 stars), Qdrant Monitoring Setup (qdrant/skills, 254 stars) and Prometheus Grafana (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prometheus Grafana?

BagelHole (a GitHub user) maintains it in BagelHole/DevOps-Security-Agent-Skills, which has 1,152 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on May 22, 2026.

Source: BagelHole/DevOps-Security-Agent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.