Search

DevOps & Cloud · Prometheus

163 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex.

ReJeCtAll/ExpertTeam-Codex113—~625Automated safety check: PassMIT3 mo ago
50

Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes.

google/skills21k—~2.6kAutomated safety check: PassApache-2.0yesterday
51
51.Aqua MetricsOfficial

Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

oracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.01 mo ago
52

Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery.

Jeffallan/claude-skills12k—~1.6kAutomated safety check: PassMIT7 days ago
53

Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

Jeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT7 days ago
54

Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML.

prometheus/prometheus-mcp121—~765Automated safety check: PassApache-2.0today
55

Installs and configures the WizTelemetry Notification extension for KubeSphere: channel setup, alert routing by tenant labels, silences and troubleshooting.

kubesphere/kubesphere17k—~6.1kAutomated safety check: PassUnknown2 mo ago
56

Debug the running local stack with traces, logs, and a shared headless browser.

macro-inc/macro4.6k—~2.4kAutomated safety check: NotesAGPL-3.0yesterday
57

Find out what a running Mendix app actually does — logs, Prometheus metrics, OpenTelemetry traces and the model catalog, joined across sources.

mendixlabs/mxcli129—~2.8kAutomated safety check: PassApache-2.0today
58

A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"…

coralogix/cx-cli121—~2.6kAutomated safety check: PassApache-2.03 days ago
59

Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server.

prometheus/prometheus-mcp121—~587Automated safety check: PassApache-2.0today
60

检查 Prometheus 数据源的连通性、数据延迟和指标采集健康度。

kubehan/PromAI125—~229Automated safety check: PassNo licence23 days ago
61

Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp.

prometheus/prometheus-mcp121—~724Automated safety check: PassApache-2.0today
62

Build and deploy a Coralogix dashboard for a given service from its logs, spans, metrics, and service specs.

coralogix/cx-cli121—~4.7kAutomated safety check: WarnApache-2.03 days ago
63
63.AlloyOfficial

Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.02 days ago
64

Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization.

qdrant/skills2542 repos~874Automated safety check: PassApache-2.0yesterday
65

Prometheus monitoring expert for PromQL, alerting rules, Grafana dashboards, and observability

RightNow-AI/openfang18k—~738Automated safety check: PassApache-2.03 mo ago
66

Go observability — always-on production signals: slog logging, Prometheus metrics, OpenTelemetry tracing, pprof profiling, alerting, Grafana.

context-labs/whip1.1k1 repo~3.3kAutomated safety check: PassMIT5 days ago
67

Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status.

azrtydxb/Fastllm-proxy108—~916Automated safety check: PassApache-2.05 days ago
68

Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.

google/skills21k—~2.8kAutomated safety check: PassApache-2.0yesterday
69

当需要为 funboost 任务添加监控、链路追踪或告警时使用。触发场景:Prometheus 指标、OpenTelemetry 链路追踪、异常告警通知、周期额度限制、函数结果持久化。关键词:Prometheus, OpenTelemetry, OTel, 告警, 监控, metrics, tracing, AlertNotifier, PeriodicQuota。

ydf0509/funboost895—~4.9kAutomated safety check: PassNo licence2 mo ago
70

Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~2kAutomated safety check: NotesMITyesterday
71

Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions.

sickn33/agentic-awesome-skills47k2 repos~3kAutomated safety check: PassMITyesterday
72

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.

google/skills21k—~1.9kAutomated safety check: PassApache-2.0yesterday
73

Set up alerting rules, configure on-call rotations, and manage incident response workflows.

sickn33/agentic-awesome-skills47k1 repo~2.8kAutomated safety check: PassMITyesterday
74

Initialize Navigator documentation structure in a project. An agent skill from qf-studio/navigator.

qf-studio/navigator355—~3kAutomated safety check: NotesMIT2 days ago
75

Monitoring, logging, and tracing implementation using OpenTelemetry as the unified standard.

ancoleman/ai-design-components525—~3kAutomated safety check: PassMIT10 mo ago
76

Generates Cloud Monitoring Server-Driven UI (SDUI) Widget and XyChart Protocol Buffer textprotos on Google Cloud from resolved PromQL or ListTimeSeries queries.

google/skills21k—~2.7kAutomated safety check: PassApache-2.0yesterday
77

Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus.

google/skills21k—~5.3kAutomated safety check: PassApache-2.0yesterday
78

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMITyesterday
79

Set up metrics collection and visualization with Prometheus and Grafana.

sickn33/agentic-awesome-skills47k1 repo~2.7kAutomated safety check: PassMITyesterday
80

Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices.

google/skills21k—~1.7kAutomated safety check: PassApache-2.0yesterday
81

Configures Cloud Monitoring PromQL-based Service Level Objective (SLO) alerting policies on Google Cloud for resources registered in App Hub or individually specified.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0yesterday
82

Guides Qdrant monitoring and observability setup. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~276Automated safety check: PassMITyesterday
83

Analyze the experiment precompute result-consistency canary across prod-US and prod-EU, deep-dive any issues, and produce an actionable report.

PostHog/posthog40k—~3.5kAutomated safety check: PassUnknowntoday
84

Investigates server/infrastructure metric anomalies in PostHog Metrics — from "this metric is rising/dropping/spiking" or a fired alert to a probable cause with evidence.

PostHog/posthog40k—~1.5kAutomated safety check: PassUnknowntoday
85

Manages scaling for GKE workloads using HPA and VPA. An agent skill from google/skills.

google/skills21k—~1.3kAutomated safety check: PassApache-2.0yesterday
86

Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.

google/skills21k—~1.6kAutomated safety check: PassApache-2.0yesterday
87

Configures GKE observability, including Cloud Logging, Cloud Monitoring, and managed Prometheus.

google/skills21k—~4.4kAutomated safety check: PassApache-2.0yesterday
88

Export cost-tracking telemetry in Prometheus textfile or webhook JSON formats — for external observability (Grafana, Datadog, custom dashboards)

ruvnet/ruflo74k—~687Automated safety check: NotesMITtoday
89

Kubernetes deployment workflow for container orchestration, Helm charts, service mesh, and production-ready K8s configurations.

aiskillstore/marketplace4335 repos~839Automated safety check: PassNo licenceyesterday
90
90.BeylaOfficial

Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart.

grafana/skills282—~1.1kAutomated safety check: PassApache-2.02 days ago
91
91.Dpm FinderOfficial

Find the Prometheus metrics that drive your Grafana Cloud bill.

grafana/skills282—~966Automated safety check: NotesApache-2.02 days ago
92
92.Fleet ManagementOfficial

Manage a fleet of Grafana Alloy collectors with Fleet Management — author Alloy pipelines once, target them via attribute matchers (env="production", regex region=~"us-."), push remotely via OpAMP…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.02 days ago
93
93.Grafana OssOfficial

Configure Grafana OSS — provisions dashboards from YAML, sets up data sources (Prometheus / Loki / Tempo / Pyroscope), writes dashboard JSON with template variables, builds panel queries, assigns…

grafana/skills282—~1.5kAutomated safety check: PassApache-2.02 days ago
94
94.MimirOfficial

Stand up Grafana Mimir for horizontally scalable, multi-tenant, long-term Prometheus + OTLP metrics storage.

grafana/skills282—~1.2kAutomated safety check: PassApache-2.02 days ago
95
95.ML AIOfficial

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.02 days ago
96
96.Oncall IrmOfficial

Route alerts, run on-call rotations, and drive incidents in Grafana IRM / OnCall — integrations (Alertmanager / Grafana Alerting / generic webhook / PagerDuty), Jinja2 routing + grouping templates…

grafana/skills282—~1.4kAutomated safety check: PassApache-2.02 days ago