Search
Prometheus
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact. | prometheus/ | 121 | — | ~547 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 50 | LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs. | majiayu000/ | 118 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 51 | Author, modify, or review Netdata collectors across Go, IBM, C, Rust and external plugins. | netdata/ | 81k | — | ~1.9k | Automated safety check: Pass | GPL-3.0 | today |
| 52 | DevOps 工程师 Agent — CI/CD 流水线、容器化与 K8s、基础设施即代码、可观测性. An agent skill from peterfei/ai-agent-team. | peterfei/ | 441 | — | ~1k | Automated safety check: Pass | MIT | 3 mo ago |
| 53 | Improve and validate the Kaniop Grafana dashboard against repository metrics and the grigri live cluster. | pando85/ | 132 | — | ~987 | Automated safety check: Pass | AGPL-3.0 | today |
| 54 | 54.Expert Ops 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |
| 55 | Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes. | google/ | 21k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 56 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 57 | A skill your agent uses when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom… | coralogix/ | 121 | — | ~3k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 58 | Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery. | Jeffallan/ | 12k | — | ~1.6k | Automated safety check: Pass | MIT | 7 days ago |
| 59 | 59.SRE Engineer Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 7 days ago |
| 60 | Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML. | prometheus/ | 121 | — | ~765 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 61 | Installs and configures the WizTelemetry Notification extension for KubeSphere: channel setup, alert routing by tenant labels, silences and troubleshooting. | kubesphere/ | 17k | — | ~6.1k | Automated safety check: Pass | Unknown | 2 mo ago |
| 62 | 62.Live Debug Debug the running local stack with traces, logs, and a shared headless browser. | macro-inc/ | 4.6k | — | ~2.4k | Automated safety check: Notes | AGPL-3.0 | today |
| 63 | Find out what a running Mendix app actually does — logs, Prometheus metrics, OpenTelemetry traces and the model catalog, joined across sources. | mendixlabs/ | 129 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 64 | A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"… | coralogix/ | 121 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 65 | Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server. | prometheus/ | 121 | — | ~587 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 66 | 检查 Prometheus 数据源的连通性、数据延迟和指标采集健康度。 | kubehan/ | 125 | — | ~229 | Automated safety check: Pass | No licence | 23 days ago |
| 67 | Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp. | prometheus/ | 121 | — | ~724 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 68 | Build and deploy a Coralogix dashboard for a given service from its logs, spans, metrics, and service specs. | coralogix/ | 121 | — | ~4.7k | Automated safety check: Warn | Apache-2.0 | 4 days ago |
| 69 | Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /… | grafana/ | 282 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 70 | Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization. | qdrant/ | 254 | 2 repos | ~874 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 71 | 71.Prometheus Prometheus monitoring expert for PromQL, alerting rules, Grafana dashboards, and observability | RightNow-AI/ | 18k | — | ~738 | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 72 | Go observability — always-on production signals: slog logging, Prometheus metrics, OpenTelemetry tracing, pprof profiling, alerting, Grafana. | context-labs/ | 1.1k | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 6 days ago |
| 73 | Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status. | azrtydxb/ | 108 | — | ~916 | Automated safety check: Pass | Apache-2.0 | today |
| 74 | Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously. | google/ | 21k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 75 | 当需要为 funboost 任务添加监控、链路追踪或告警时使用。触发场景:Prometheus 指标、OpenTelemetry 链路追踪、异常告警通知、周期额度限制、函数结果持久化。关键词:Prometheus, OpenTelemetry, OTel, 告警, 监控, metrics, tracing, AlertNotifier, PeriodicQuota。 | ydf0509/ | 895 | — | ~4.9k | Automated safety check: Pass | No licence | 2 mo ago |
| 76 | Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~2k | Automated safety check: Notes | MIT | 2 days ago |
| 77 | Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions. | sickn33/ | 47k | 2 repos | ~3k | Automated safety check: Pass | MIT | 2 days ago |
| 78 | Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. | google/ | 21k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 79 | Set up alerting rules, configure on-call rotations, and manage incident response workflows. | sickn33/ | 47k | 1 repo | ~2.8k | Automated safety check: Pass | MIT | 2 days ago |
| 80 | 80.Nav Init Initialize Navigator documentation structure in a project. An agent skill from qf-studio/navigator. | qf-studio/ | 355 | — | ~3k | Automated safety check: Notes | MIT | 2 days ago |
| 81 | Monitoring, logging, and tracing implementation using OpenTelemetry as the unified standard. | ancoleman/ | 525 | — | ~3k | Automated safety check: Pass | MIT | 10 mo ago |
| 82 | Generates Cloud Monitoring Server-Driven UI (SDUI) Widget and XyChart Protocol Buffer textprotos on Google Cloud from resolved PromQL or ListTimeSeries queries. | google/ | 21k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 83 | Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus. | google/ | 21k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 84 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | 2 days ago |
| 85 | Set up metrics collection and visualization with Prometheus and Grafana. | sickn33/ | 47k | 1 repo | ~2.7k | Automated safety check: Pass | MIT | 2 days ago |
| 86 | Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices. | google/ | 21k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 87 | Configures Cloud Monitoring PromQL-based Service Level Objective (SLO) alerting policies on Google Cloud for resources registered in App Hub or individually specified. | google/ | 21k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 88 | Guides Qdrant monitoring and observability setup. An agent skill from github/awesome-copilot. | github/ | 40k | 1 repo | ~276 | Automated safety check: Pass | MIT | 2 days ago |
| 89 | Analyze the experiment precompute result-consistency canary across prod-US and prod-EU, deep-dive any issues, and produce an actionable report. | PostHog/ | 40k | — | ~3.5k | Automated safety check: Pass | Unknown | yesterday |
| 90 | Investigates server/infrastructure metric anomalies in PostHog Metrics — from "this metric is rising/dropping/spiking" or a fired alert to a probable cause with evidence. | PostHog/ | 40k | — | ~1.5k | Automated safety check: Pass | Unknown | yesterday |
| 91 | Manages scaling for GKE workloads using HPA and VPA. An agent skill from google/skills. | google/ | 21k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 92 | Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. | google/ | 21k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 93 | Configures GKE observability, including Cloud Logging, Cloud Monitoring, and managed Prometheus. | google/ | 21k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 94 | 94.Cost Export Export cost-tracking telemetry in Prometheus textfile or webhook JSON formats — for external observability (Grafana, Datadog, custom dashboards) | ruvnet/ | 74k | — | ~687 | Automated safety check: Notes | MIT | yesterday |
| 95 | Kubernetes deployment workflow for container orchestration, Helm charts, service mesh, and production-ready K8s configurations. | aiskillstore/ | 433 | 5 repos | ~839 | Automated safety check: Pass | No licence | yesterday |
| 96 | Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart. | grafana/ | 282 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |