Search

Prometheus

169 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact.

prometheus/prometheus-mcp121—~547Automated safety check: PassApache-2.0yesterday
50

LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

majiayu000/litellm-rs118—~1.3kAutomated safety check: PassMITtoday
51

Author, modify, or review Netdata collectors across Go, IBM, C, Rust and external plugins.

netdata/netdata81k—~1.9kAutomated safety check: PassGPL-3.0today
52

DevOps 工程师 Agent — CI/CD 流水线、容器化与 K8s、基础设施即代码、可观测性. An agent skill from peterfei/ai-agent-team.

peterfei/ai-agent-team441—~1kAutomated safety check: PassMIT3 mo ago
53

Improve and validate the Kaniop Grafana dashboard against repository metrics and the grigri live cluster.

pando85/kaniop132—~987Automated safety check: PassAGPL-3.0today
54

基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex.

ReJeCtAll/ExpertTeam-Codex113—~625Automated safety check: PassMIT3 mo ago
55

Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes.

google/skills21k—~2.6kAutomated safety check: PassApache-2.0yesterday
56
56.Aqua MetricsOfficial

Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

oracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.01 mo ago
57

A skill your agent uses when the user asks to "set up parsing", "create parsing rule", "extract fields from logs", "regex extraction", "log parsing", "enrich logs", "add context to logs", "custom…

coralogix/cx-cli121—~3kAutomated safety check: PassApache-2.04 days ago
58

Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery.

Jeffallan/claude-skills12k—~1.6kAutomated safety check: PassMIT7 days ago
59

Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

Jeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT7 days ago
60

Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML.

prometheus/prometheus-mcp121—~765Automated safety check: PassApache-2.0yesterday
61

Installs and configures the WizTelemetry Notification extension for KubeSphere: channel setup, alert routing by tenant labels, silences and troubleshooting.

kubesphere/kubesphere17k—~6.1kAutomated safety check: PassUnknown2 mo ago
62

Debug the running local stack with traces, logs, and a shared headless browser.

macro-inc/macro4.6k—~2.4kAutomated safety check: NotesAGPL-3.0today
63

Find out what a running Mendix app actually does — logs, Prometheus metrics, OpenTelemetry traces and the model catalog, joined across sources.

mendixlabs/mxcli129—~2.8kAutomated safety check: PassApache-2.0today
64

A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"…

coralogix/cx-cli121—~2.6kAutomated safety check: PassApache-2.04 days ago
65

Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server.

prometheus/prometheus-mcp121—~587Automated safety check: PassApache-2.0yesterday
66

检查 Prometheus 数据源的连通性、数据延迟和指标采集健康度。

kubehan/PromAI125—~229Automated safety check: PassNo licence23 days ago
67

Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp.

prometheus/prometheus-mcp121—~724Automated safety check: PassApache-2.0yesterday
68

Build and deploy a Coralogix dashboard for a given service from its logs, spans, metrics, and service specs.

coralogix/cx-cli121—~4.7kAutomated safety check: WarnApache-2.04 days ago
69
69.AlloyOfficial

Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /…

grafana/skills282—~1.3kAutomated safety check: PassApache-2.02 days ago
70

Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization.

qdrant/skills2542 repos~874Automated safety check: PassApache-2.0yesterday
71

Prometheus monitoring expert for PromQL, alerting rules, Grafana dashboards, and observability

RightNow-AI/openfang18k—~738Automated safety check: PassApache-2.03 mo ago
72

Go observability — always-on production signals: slog logging, Prometheus metrics, OpenTelemetry tracing, pprof profiling, alerting, Grafana.

context-labs/whip1.1k1 repo~3.3kAutomated safety check: PassMIT6 days ago
73

Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status.

azrtydxb/Fastllm-proxy108—~916Automated safety check: PassApache-2.0today
74

Diagnoses GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads autonomously.

google/skills21k—~2.8kAutomated safety check: PassApache-2.0yesterday
75

当需要为 funboost 任务添加监控、链路追踪或告警时使用。触发场景:Prometheus 指标、OpenTelemetry 链路追踪、异常告警通知、周期额度限制、函数结果持久化。关键词:Prometheus, OpenTelemetry, OTel, 告警, 监控, metrics, tracing, AlertNotifier, PeriodicQuota。

ydf0509/funboost895—~4.9kAutomated safety check: PassNo licence2 mo ago
76

Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~2kAutomated safety check: NotesMIT2 days ago
77

Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions.

sickn33/agentic-awesome-skills47k2 repos~3kAutomated safety check: PassMIT2 days ago
78

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL.

google/skills21k—~1.9kAutomated safety check: PassApache-2.0yesterday
79

Set up alerting rules, configure on-call rotations, and manage incident response workflows.

sickn33/agentic-awesome-skills47k1 repo~2.8kAutomated safety check: PassMIT2 days ago
80

Initialize Navigator documentation structure in a project. An agent skill from qf-studio/navigator.

qf-studio/navigator355—~3kAutomated safety check: NotesMIT2 days ago
81

Monitoring, logging, and tracing implementation using OpenTelemetry as the unified standard.

ancoleman/ai-design-components525—~3kAutomated safety check: PassMIT10 mo ago
82

Generates Cloud Monitoring Server-Driven UI (SDUI) Widget and XyChart Protocol Buffer textprotos on Google Cloud from resolved PromQL or ListTimeSeries queries.

google/skills21k—~2.7kAutomated safety check: PassApache-2.0yesterday
83

Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus.

google/skills21k—~5.3kAutomated safety check: PassApache-2.0yesterday
84

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMIT2 days ago
85

Set up metrics collection and visualization with Prometheus and Grafana.

sickn33/agentic-awesome-skills47k1 repo~2.7kAutomated safety check: PassMIT2 days ago
86

Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices.

google/skills21k—~1.7kAutomated safety check: PassApache-2.0yesterday
87

Configures Cloud Monitoring PromQL-based Service Level Objective (SLO) alerting policies on Google Cloud for resources registered in App Hub or individually specified.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0yesterday
88

Guides Qdrant monitoring and observability setup. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~276Automated safety check: PassMIT2 days ago
89

Analyze the experiment precompute result-consistency canary across prod-US and prod-EU, deep-dive any issues, and produce an actionable report.

PostHog/posthog40k—~3.5kAutomated safety check: PassUnknownyesterday
90

Investigates server/infrastructure metric anomalies in PostHog Metrics — from "this metric is rising/dropping/spiking" or a fired alert to a probable cause with evidence.

PostHog/posthog40k—~1.5kAutomated safety check: PassUnknownyesterday
91

Manages scaling for GKE workloads using HPA and VPA. An agent skill from google/skills.

google/skills21k—~1.3kAutomated safety check: PassApache-2.0yesterday
92

Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE.

google/skills21k—~1.6kAutomated safety check: PassApache-2.0yesterday
93

Configures GKE observability, including Cloud Logging, Cloud Monitoring, and managed Prometheus.

google/skills21k—~4.4kAutomated safety check: PassApache-2.0yesterday
94

Export cost-tracking telemetry in Prometheus textfile or webhook JSON formats — for external observability (Grafana, Datadog, custom dashboards)

ruvnet/ruflo74k—~687Automated safety check: NotesMITyesterday
95

Kubernetes deployment workflow for container orchestration, Helm charts, service mesh, and production-ready K8s configurations.

aiskillstore/marketplace4335 repos~839Automated safety check: PassNo licenceyesterday
96
96.BeylaOfficial

Auto-instrument an application's HTTP / gRPC / DB traffic with Grafana Beyla eBPF — no code changes, no SDK, no restart.

grafana/skills282—~1.1kAutomated safety check: PassApache-2.02 days ago