Search

Prometheus · For devops and sre engineers

151 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML.

prometheus/prometheus-mcp121—~765Automated safety check: PassApache-2.0yesterday
50

Installs and configures the WizTelemetry Notification extension for KubeSphere: channel setup, alert routing by tenant labels, silences and troubleshooting.

kubesphere/kubesphere17k—~6.1kAutomated safety check: PassUnknown2 mo ago
51

Debug the running local stack with traces, logs, and a shared headless browser.

macro-inc/macro4.6k—~2.4kAutomated safety check: NotesAGPL-3.0today
52

Find out what a running Mendix app actually does — logs, Prometheus metrics, OpenTelemetry traces and the model catalog, joined across sources.

mendixlabs/mxcli129—~2.8kAutomated safety check: PassApache-2.0today
53

A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"…

coralogix/cx-cli121—~2.6kAutomated safety check: PassApache-2.04 days ago
54

Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server.

prometheus/prometheus-mcp121—~587Automated safety check: PassApache-2.0yesterday
55

检查 Prometheus 数据源的连通性、数据延迟和指标采集健康度。

kubehan/PromAI125—~229Automated safety check: PassNo licence23 days ago
56

Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp.

prometheus/prometheus-mcp121—~724Automated safety check: PassApache-2.0yesterday
57

Build and deploy a Coralogix dashboard for a given service from its logs, spans, metrics, and service specs.

coralogix/cx-cli121—~4.7kAutomated safety check: WarnApache-2.04 days ago
58
58.AlloyOfficial

Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.02 days ago
59

Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization.

qdrant/skills2542 repos~874Automated safety check: PassApache-2.0yesterday
60

Prometheus monitoring expert for PromQL, alerting rules, Grafana dashboards, and observability

RightNow-AI/openfang18k—~738Automated safety check: PassApache-2.03 mo ago
61

Go observability — always-on production signals: slog logging, Prometheus metrics, OpenTelemetry tracing, pprof profiling, alerting, Grafana.

context-labs/whip1.1k1 repo~3.3kAutomated safety check: PassMIT6 days ago
62

Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status.

azrtydxb/Fastllm-proxy108—~916Automated safety check: PassApache-2.0today
63

Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.

google/skills21k—~2.8kAutomated safety check: PassApache-2.0yesterday
64

当需要为 funboost 任务添加监控、链路追踪或告警时使用。触发场景:Prometheus 指标、OpenTelemetry 链路追踪、异常告警通知、周期额度限制、函数结果持久化。关键词:Prometheus, OpenTelemetry, OTel, 告警, 监控, metrics, tracing, AlertNotifier, PeriodicQuota。

ydf0509/funboost895—~4.9kAutomated safety check: PassNo licence2 mo ago
65

Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions.

sickn33/agentic-awesome-skills47k2 repos~3kAutomated safety check: PassMIT2 days ago
66

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.

google/skills21k—~1.9kAutomated safety check: PassApache-2.0yesterday
67

Set up alerting rules, configure on-call rotations, and manage incident response workflows.

sickn33/agentic-awesome-skills47k1 repo~2.8kAutomated safety check: PassMIT2 days ago
68

Monitoring, logging, and tracing implementation using OpenTelemetry as the unified standard.

ancoleman/ai-design-components525—~3kAutomated safety check: PassMIT10 mo ago
69

Generates Cloud Monitoring Server-Driven UI (SDUI) Widget and XyChart Protocol Buffer textprotos on Google Cloud from resolved PromQL or ListTimeSeries queries.

google/skills21k—~2.7kAutomated safety check: PassApache-2.0yesterday
70

Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus.

google/skills21k—~5.3kAutomated safety check: PassApache-2.0yesterday
71

Set up metrics collection and visualization with Prometheus and Grafana.

sickn33/agentic-awesome-skills47k1 repo~2.7kAutomated safety check: PassMIT2 days ago
72

Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices.

google/skills21k—~1.7kAutomated safety check: PassApache-2.0yesterday
73

Configures Cloud Monitoring PromQL-based Service Level Objective (SLO) alerting policies on Google Cloud for resources registered in App Hub or individually specified.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0yesterday
74

Guides Qdrant monitoring and observability setup. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~276Automated safety check: PassMIT2 days ago
75

Analyze the experiment precompute result-consistency canary across prod-US and prod-EU, deep-dive any issues, and produce an actionable report.

PostHog/posthog40k—~3.5kAutomated safety check: PassUnknownyesterday
76

Investigates server/infrastructure metric anomalies in PostHog Metrics — from "this metric is rising/dropping/spiking" or a fired alert to a probable cause with evidence.

PostHog/posthog40k—~1.5kAutomated safety check: PassUnknownyesterday
77

Manages scaling for GKE workloads using HPA and VPA. An agent skill from google/skills.

google/skills21k—~1.3kAutomated safety check: PassApache-2.0yesterday
78

Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.

google/skills21k—~1.6kAutomated safety check: PassApache-2.0yesterday
79

Configures GKE observability, including Cloud Logging, Cloud Monitoring, and managed Prometheus.

google/skills21k—~4.4kAutomated safety check: PassApache-2.0yesterday
80

Export cost-tracking telemetry in Prometheus textfile or webhook JSON formats — for external observability (Grafana, Datadog, custom dashboards)

ruvnet/ruflo74k—~687Automated safety check: NotesMITyesterday
81

Kubernetes deployment workflow for container orchestration, Helm charts, service mesh, and production-ready K8s configurations.

aiskillstore/marketplace4335 repos~839Automated safety check: PassNo licenceyesterday
82
82.BeylaOfficial

Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart.

grafana/skills282—~1.1kAutomated safety check: PassApache-2.02 days ago
83
83.Dpm FinderOfficial

Find the Prometheus metrics that drive your Grafana Cloud bill.

grafana/skills282—~966Automated safety check: NotesApache-2.02 days ago
84
84.Fleet ManagementOfficial

Manage a fleet of Grafana Alloy collectors with Fleet Management — author Alloy pipelines once, target them via attribute matchers (env="production", regex region=~"us-."), push remotely via OpAMP…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.02 days ago
85
85.Grafana OssOfficial

Configure Grafana OSS — provisions dashboards from YAML, sets up data sources (Prometheus / Loki / Tempo / Pyroscope), writes dashboard JSON with template variables, builds panel queries, assigns…

grafana/skills282—~1.5kAutomated safety check: PassApache-2.02 days ago
86
86.MimirOfficial

Stand up Grafana Mimir for horizontally scalable, multi-tenant, long-term Prometheus + OTLP metrics storage.

grafana/skills282—~1.2kAutomated safety check: PassApache-2.02 days ago
87
87.ML AIOfficial

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.02 days ago
88
88.Oncall IrmOfficial

Route alerts, run on-call rotations, and drive incidents in Grafana IRM / OnCall — integrations (Alertmanager / Grafana Alerting / generic webhook / PagerDuty), Jinja2 routing + grouping templates…

grafana/skills282—~1.4kAutomated safety check: PassApache-2.02 days ago
89

Validate, lint, audit, or fix PromQL queries and alerting rules; detects anti-patterns.

akin-ozer/cc-devops-skills320—~4kAutomated safety check: PassApache-2.02 mo ago
90

Canvas/A2UI inline network visualizations — topology maps, health dashboards, alert cards, change timelines, config diffs, path traces, and health scorecards rendered directly in the OpenClaw chat…

automateyournetwork/netclaw676—~941Automated safety check: PassApache-2.0yesterday
91

Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for…

github/awesome-copilot40k—~2.6kAutomated safety check: PassMIT2 days ago
92

Collect comprehensive infrastructure performance metrics across compute, storage, network, containers, load balancers, and databases.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.2kAutomated safety check: PassMITyesterday
93

Monitor use when deploying monitoring stacks including Prometheus, Grafana, and Datadog.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.2kAutomated safety check: PassMITyesterday
94
94.PrometheusOfficial

Prometheus and Grafana Cloud Metrics overview including PromQL query language, Metrics Drilldown, alerting, recording rules, and integration patterns.

grafana/skills282—~1.2kAutomated safety check: PassApache-2.02 days ago
95

Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid…

grafana/skills282—~4.6kAutomated safety check: PassApache-2.02 days ago
96
96.Send DataOfficial

Sending telemetry data to Grafana Cloud — metrics via Prometheus remote write or OTLP, logs via Loki push or Alloy, traces via OTLP to Tempo, profiles via Pyroscope.

grafana/skills282—~1.4kAutomated safety check: PassApache-2.02 days ago