Developer tool
Prometheus agent skills for Claude Code, Codex and other agents.
- skills
- 147
- official
- 39
- Type
- Developer tool
- Website
- prometheus.io
- Official GitHub
- prometheus
- Reviews
- See Prometheus on Enlisted
Prometheus skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
Official (39 skills)
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when changing metric CSV files, exporter-owned counters, Prometheus labels, or metric rendering. | NVIDIA/ | 1.9k | — | ~149 | Automated safety check: Pass | Apache-2.0 | 18 days ago |
| 2 | Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO… | redis/ | 165 | 2 repos | ~911 | Automated safety check: Pass | MIT | 29 days ago |
| 3 | Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)… | grafana/ | 278 | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech. | qdrant/ | 253 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 5 | Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units… | grafana/ | 278 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 278 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 7 | Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes. | google/ | 21k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 8 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 9 | Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously. | google/ | 21k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 10 | Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization. | qdrant/ | 253 | 2 repos | ~874 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 11 | Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /… | grafana/ | 278 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 12 | Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. | google/ | 21k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 13 | Generates Google Cloud Monitoring Server-Driven UI (SDUI) Widget and XyChart Protocol Buffer textprotos from resolved PromQL or ListTimeSeries queries. | google/ | 21k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 14 | Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus. | google/ | 21k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | Configures best-practice, high-signal alerting policies for Google Cloud Run resources (services, jobs, and worker pools) based on seasoned SRE practices. | google/ | 21k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 16 | Configures PromQL-based Service Level Objective (SLO) alerting policies for Google Cloud resources registered in App Hub or individually specified. | google/ | 21k | — | ~3k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Guides Qdrant monitoring and observability setup. An agent skill from github/awesome-copilot. | github/ | 40k | 1 repo | ~276 | Automated safety check: Pass | MIT | today |
| 18 | Manages scaling for GKE workloads using HPA and VPA. An agent skill from google/skills. | google/ | 21k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 19 | Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. | google/ | 21k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 20 | Configures GKE observability, including Cloud Logging, Cloud Monitoring, and managed Prometheus. | google/ | 21k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | today |
| 21 | Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart. | grafana/ | 278 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 22 | Find the Prometheus metrics that drive your Grafana Cloud bill. | grafana/ | 278 | — | ~966 | Automated safety check: Notes | Apache-2.0 | today |
| 23 | Manage a fleet of Grafana Alloy collectors with Fleet Management — author Alloy pipelines once, target them via attribute matchers (env="production", regex region=~"us-."), push remotely via OpAMP… | grafana/ | 278 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | Configure Grafana OSS — provisions dashboards from YAML, sets up data sources (Prometheus / Loki / Tempo / Pyroscope), writes dashboard JSON with template variables, builds panel queries, assigns… | grafana/ | 278 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | Stand up Grafana Mimir for horizontally scalable, multi-tenant, long-term Prometheus + OTLP metrics storage. | grafana/ | 278 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 26 | Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN… | grafana/ | 278 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | Route alerts, run on-call rotations, and drive incidents in Grafana IRM / OnCall — integrations (Alertmanager / Grafana Alerting / generic webhook / PagerDuty), Jinja2 routing + grouping templates… | grafana/ | 278 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 28 | Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for… | github/ | 40k | — | ~2.6k | Automated safety check: Pass | MIT | today |
| 29 | Prometheus and Grafana Cloud Metrics overview including PromQL query language, Metrics Drilldown, alerting, recording rules, and integration patterns. | grafana/ | 278 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 30 | Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid… | grafana/ | 278 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 31 | Sending telemetry data to Grafana Cloud — metrics via Prometheus remote write or OTLP, logs via Loki push or Alloy, traces via OTLP to Tempo, profiles via Pyroscope. | grafana/ | 278 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 32 | A skill your agent uses to deploy and operate a CollectX (clx) based DOCA telemetry collector on a host or BlueField — wiring providers / counters into the collector, running the collection daemon… | NVIDIA/ | 3.5k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 33 | Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. | aws/ | 2.8k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | today |
| 34 | Expert evaluator for Prometheus label strategy on Grafana Cloud. | grafana/ | 278 | — | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 35 | Analyze the experiment precompute result-consistency canary across prod-US and prod-EU, deep-dive any issues, and produce an actionable report. | PostHog/ | 721 | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 36 | Investigates server/infrastructure metric anomalies in PostHog Metrics — from "this metric is rising/dropping/spiking" or a fired alert to a probable cause with evidence. | PostHog/ | 721 | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 37 | Guides Qdrant monitoring and observability setup. An agent skill from qdrant/skills. | qdrant/ | 253 | — | ~368 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 38 | Expert knowledge for Azure Container Storage development including troubleshooting, decision making, limits & quotas, security, and configuration. | MicrosoftDocs/ | 775 | — | ~1.4k | Automated safety check: Pass | CC-BY-4.0 | yesterday |
| 39 | Expert knowledge for Azure Managed Grafana development including troubleshooting, decision making, limits & quotas, security, configuration, and integrations & coding patterns. | MicrosoftDocs/ | 775 | — | ~2k | Automated safety check: Pass | CC-BY-4.0 | yesterday |
Community
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 40 | Deploys KubeEye on KubeSphere and writes InspectRule and InspectPlan resources to inspect cluster health, then retrieves the inspection reports. | kubesphere/ | 17k | — | ~3.6k | Automated safety check: Pass | Unknown | 2 mo ago |
| 41 | Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API. | slopus/ | 24k | — | ~2k | Automated safety check: Notes | MIT | today |
| 42 | 42.Syncmeta Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta). | pawurb/ | 1.9k | — | ~1.2k | Automated safety check: Notes | MIT | today |
| 43 | Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions. | kubesphere/ | 17k | — | ~1.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 44 | Post a comment (or a formal PR review) to a GitHub issue or pull request in mariadb-operator/mariadb-operator. | mariadb-operator/ | 1k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 45 | Create, review or validate Netdata Prometheus chart profiles, exporter dashboard design, collection policy and stock semantic proofs. | netdata/ | 81k | — | ~4.7k | Automated safety check: Pass | GPL-3.0 | today |
| 46 | Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts. | NVlabs/ | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 19 days ago |
| 47 | Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments. | alibaba/ | 412 | — | ~1.9k | Automated safety check: Pass | Unknown | 13 days ago |
| 48 | Perform a structured maintainer-style PR review for the mariadb-operator repository. | mariadb-operator/ | 1k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
Questions, answered from the data.
What is the best Prometheus skill?
Metric Contract Changes (official) from NVIDIA/dcgm-exporter ranks first of the 147 Prometheus skills listed here, with the highest score: its repository has 1.9k GitHub stars, its SKILL.md loads about 149 tokens and it passes the automated safety check with no findings. Next come Redis Observability and Alerting Irm.
Is there an official Prometheus skill?
39 of the 147 Prometheus skills are official, published by the vendor's own GitHub organization: Metric Contract Changes, Redis Observability, Alerting Irm, Qdrant Advisor, Dashboarding and 34 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.