Official agent skill

Infrastructure

by grafana in grafana/skills

Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — k8s-monitoring Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy…

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Infrastructure

skills CLI
$ npx skills add grafana/skills --skill infrastructure -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install grafana/skills infrastructure --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/grafana/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/grafana-cloud/infrastructure .claude/skills/infrastructure && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
infrastructure
GitHub stars
281
Token cost
~1.2k tokens
SKILL.md length
157 words
Files
3 (incl. references)
Skills in repo
51
Repo updated
First seen
Licence
Apache-2.0

At a glance

Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — k8s-monitoring Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy…

  • Works in 3 steps: Onboard a Kubernetes cluster… → Monitor a Linux host → Pull AWS / Azure / GCP metrics
  • Onboarding a new cluster
  • SKILL.md covers Prerequisites, Common Workflows, Troubleshooting and Resources
  • Calls kubectl, helm and curl; reaches grafana.github.io

What it does

Infrastructure is an agent skill from grafana/skills, published by the product's own GitHub organization. Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — k8s-monitoring Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy prometheus.exporter.unix for Linux hosts, cAdvisor + Docker discovery for containers, and CloudWatch / Azure Monitor / Google Cloud Monitoring datasource setup. Use when onboarding a new cluster or VM fleet to Grafana Cloud, picking the right Helm values for K8s scraping, wiring kube-state-metrics + node-exporter + cAdvisor, alerting on…

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/clouds-and-hosts.md` and `references/k8s-monitoring-values.md`).

It sits in DevOps & Cloud, covering Container orchestration, Monitoring and alerting and Web scraping. It works with Kubernetes, Grafana, Google Cloud and Helm. The licence is Apache-2.0.

When your agent uses it

  • Onboarding a new cluster
  • VM fleet to Grafana Cloud
  • Picking the right Helm values for K8s scraping
  • Wiring kube-state-metrics + node-exporter + cAdvisor

Example prompts

  • “monitor my cluster”
  • “send K8s metrics to Grafana”
  • “scrape EC2 metrics”
  • “/infrastructure”

Requirements

  • Docker

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Onboard a Kubernetes cluster (k8s-monitoring chart)
  2. Monitor a Linux host
  3. Pull AWS / Azure / GCP metrics

What it can do on your machine

Read from SKILL.md and the folder at commit 1ccacf2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • helm
    • curl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • grafana.github.io

    Also links to:

    • grafana.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Infrastructure loads about 1.2k tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 207 tokens; SKILL.md has 157 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~207
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from grafana/skills at commit 1ccacf2, republished under its Apache-2.0 licence (© grafana). 157 words, ~1,188 tokens.

Download SKILL.mdSave it as .claude/skills/infrastructure/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
infrastructure
description
Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — `k8s-monitoring` Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy `prometheus.exporter.unix` for Linux hosts, cAdvisor + Docker discovery for containers, and CloudWatch / Azure Monitor / Google Cloud Monitoring datasource setup. Use when onboarding a new cluster or VM fleet to Grafana Cloud, picking the right Helm values for K8s scraping, wiring kube-state-metrics + node-exporter + cAdvisor, alerting on `PodCrashLooping` / node memory / PVC capacity, or pulling AWS / Azure / GCP cloud metrics — even when the user says "monitor my cluster", "send K8s metrics to Grafana", "scrape EC2 metrics", "cluster pod logs", or "install the monitoring helm chart" without naming `k8s-monitoring` or Alloy.
license
Apache-2.0

Grafana Cloud Infrastructure Monitoring

Docs: https://grafana.com/docs/grafana-cloud/monitor-infrastructure/

K8s + host + container + cloud-provider telemetry, mostly via the grafana/k8s-monitoring Helm chart or Alloy.

Prerequisites

  • Grafana Cloud stack with Prometheus / Loki / Tempo endpoints + API key (metrics:write, logs:write, traces:write)
  • For Kubernetes: a cluster + helm 3.x + kubectl context pointing at it
  • For hosts / Docker: Alloy installed on the node

Common Workflows

1. Onboard a Kubernetes cluster (k8s-monitoring chart)
bash
# 1. Create the namespace + secret
kubectl create namespace monitoring
kubectl create secret generic grafana-cloud-secret \
  -n monitoring --from-literal=api-key=<your-api-key>

# 2. Install — values.yaml in references/k8s-monitoring-values.md
helm repo add grafana https://grafana.github.io/helm-charts && helm repo update
helm install k8s-monitoring grafana/k8s-monitoring \
  --version 4.1.4 -n monitoring -f values.yaml

# 3. Verify every pod is Running
kubectl get pods -n monitoring
# Expect alloy-*, kube-state-metrics-*, node-exporter-*, etc. all Ready.

# 4. Verify no error logs in the metrics/logs/traces Alloys
kubectl -n monitoring logs deploy/k8s-monitoring-alloy-metrics --tail=50 | grep -iE 'error|level=err' || echo "clean"

# 5. Verify telemetry landed in Grafana Cloud
#    PromQL on the metrics datasource (should be > 0):
#      sum(up{cluster="production-us-east"})
#    LogQL on Loki:
#      sum(count_over_time({cluster="production-us-east"}[5m]))

Full values.yaml, key PromQL, dashboard IDs (15520, 1860, 14282…), and alert rules: references/k8s-monitoring-values.md.

2. Monitor a Linux host
alloy
# 1. /etc/alloy/config.alloy — see references/clouds-and-hosts.md for the full block
prometheus.exporter.unix "host"  { rootfs_path = "/" }
prometheus.scrape         "node" { targets = prometheus.exporter.unix.host.targets
                                   forward_to = [prometheus.remote_write.cloud.receiver] }
bash
# 2. Reload Alloy and verify the unix exporter is up
systemctl reload alloy
curl -s http://localhost:12345/api/v0/web/components | jq '.[] | select(.id|contains("prometheus.exporter.unix"))'

# 3. Verify in Grafana Cloud — open the "Node Exporter Full" dashboard (ID 1860)
#    and pick your host from the `instance` dropdown.
3. Pull AWS / Azure / GCP metrics

Provision the datasource (full YAML in references/clouds-and-hosts.md), then:

bash
# 1. After provisioning, restart Grafana to pick up the file
# 2. Verify the datasource — Grafana → Connections → Data sources → "Test"
#    Expect "Successfully queried the CloudWatch metrics API" (or equivalent).
# 3. Confirm a query — Explore → datasource → metric e.g.
#    CloudWatch namespace AWS/EC2 metric CPUUtilization, last 1h.

Troubleshooting

  • chart installed but no metrics in Cloud → check the grafana-cloud-secret api-key value; check Alloy logs for 401
  • kube-state-metrics pod Pending → likely RBAC; reapply the chart's CRDs/CRBs
  • Node-exporter pod CrashLoopBackOff → typically hostNetwork: true collision with the host's :9100; change the port
  • CloudWatch "Access denied" → IAM role missing cloudwatch:GetMetricData, cloudwatch:ListMetrics

Resources

© grafana, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/grafana-cloud/infrastructure of grafana/skills.

  • SKILL.md
  • references/clouds-and-hosts.md
  • references/k8s-monitoring-values.md

Open the folder on GitHubat commit 1ccacf2

Compare with similar skills

Infrastructure next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Infrastructure compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Infrastructure this skillgrafana/skills281—~1.2kAutomated safety check: PassApache-2.0
Cloud Devopsdavila7/claude-code-templates32k4 repos~1.4kAutomated safety check: PassMIT
Provider Bug Reviewmondoohq/mql412—~2.9kAutomated safety check: PassCustom licence
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
Extend Discovery Typerunwhen-contrib/runwhen-local163—~1.7kAutomated safety check: PassApache-2.0
Kclikarmab/kcli653—~2.6kAutomated safety check: WarnApache-2.0

Similar skills

  • Cloud Devops

    davila7/claude-code-templates

    Cloud infrastructure and DevOps workflow covering AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, monitoring, and cloud-native development.

    32k GitHub starsUsed in 4 repos~1.4k tokens
    DevOps & CloudAuto-check passed
  • Deep static code review of an mql provider for logic errors, nil-handling bugs, pagination truncation, caching/id collisions, and other defects that silently give users wrong data.

    412 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Extend Discovery Type

    runwhen-contrib/runwhen-local

    Add or enrich a resource type in an existing RunWhen Local discovery indexer (Azure azureapi, GCP gcpapi, AWS, or Kubernetes).

    163 GitHub stars~1.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Kcli

    karmab/kcli

    Comprehensive guide for kcli usage. An agent skill from karmab/kcli.

    653 GitHub stars~2.6k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • Azure Arc

    vinayaklatthe/microsoft-security-skills

    Guidance for Azure Arc — projecting on-premises, multicloud (AWS/GCP), and edge servers, Kubernetes, and data services into Azure Resource Manager for unified governance, security, and management.

    175 GitHub stars~1.9k tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed

More from grafana/skills

All 51 skills in this repo
  • K6 Docs

    grafana/skills

    Official

    Write or review k6 documentation across the three k6 repositories - k6-DefinitelyTyped (TypeScript types), k6-docs (user documentation), and k6 (release notes / changelog).

    281 GitHub stars~678 tokensUpdated today
    Auto-check passed
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    281 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Dashboarding

    grafana/skills

    Official

    Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units…

    281 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • K6 Perf Test Website

    grafana/skills

    Official

    A skill your agent uses when the user wants to performance-test, load-test, or stress-test a public website end-to-end with k6.

    281 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    281 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Adaptive Metrics

    grafana/skills

    Official

    Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config…

    281 GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Categories

Questions about Infrastructure

What does Infrastructure do?

Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — k8s-monitoring Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy…. Infrastructure is an agent skill from grafana/skills, published by the product's own GitHub organization.unix for Linux hosts, cAdvisor + Docker discovery for containers, and CloudWatch / Azure Monitor / Google Cloud Monitoring datasource setup.

When should I use Infrastructure?

Infrastructure fits situations like: onboarding a new cluster; VM fleet to Grafana Cloud; picking the right Helm values for K8s scraping; wiring kube-state-metrics + node-exporter + cAdvisor.

How do I install Infrastructure in Claude Code?

Run `npx skills add grafana/skills --skill infrastructure -a claude-code`. Or copy the skill folder (skills/grafana-cloud/infrastructure in grafana/skills) into .claude/skills/infrastructure in your project. Claude Code loads it when a task matches its description.

How do I install Infrastructure in Codex?

Run `npx skills add grafana/skills --skill infrastructure -a codex`. Or copy the skill folder (skills/grafana-cloud/infrastructure in grafana/skills) into .agents/skills/infrastructure in your project. Codex loads it when a task matches its description.

Can I use Infrastructure in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grafana/skills --skill infrastructure -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/infrastructure, .gemini/skills/infrastructure, .github/skills/infrastructure and .opencode/skills/infrastructure in your project.

What does Infrastructure need to run?

Going by SKILL.md and its folder, Infrastructure needs the command-line tools its instructions call (kubectl, helm, curl and jq). Our summary lists: Docker.

Does Infrastructure access the network?

SKILL.md names 3 domains. In commands or code: grafana.github.io; the agent is likely to contact it when it follows the instructions. As links in the text: grafana.com and github.com. This is read from the text; nothing was executed.

Is Infrastructure safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Infrastructure use?

Infrastructure is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Infrastructure use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Infrastructure?

Skills that share tags, products or a category with Infrastructure: Cloud Devops (davila7/claude-code-templates, 32k stars), Provider Bug Review (mondoohq/mql, 412 stars), Kcli Cluster Deployment (karmab/kcli, 653 stars) and Extend Discovery Type (runwhen-contrib/runwhen-local, 163 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Infrastructure?

grafana (a GitHub organization, an official publisher) maintains it in grafana/skills, which has 281 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on October 8, 2026.

Source: grafana/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.