Search
DevOps & Cloud · Prometheus
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Deploys KubeEye on KubeSphere and writes InspectRule and InspectPlan resources to inspect cluster health, then retrieves the inspection reports. | kubesphere/ | 17k | — | ~3.6k | Automated safety check: Pass | Unknown | 2 mo ago |
| 2 | Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API. | slopus/ | 24k | — | ~2k | Automated safety check: Notes | MIT | today |
| 3 | 3.Syncmeta Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta). | pawurb/ | 1.9k | — | ~1.2k | Automated safety check: Notes | MIT | today |
| 4 | Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions. | kubesphere/ | 17k | — | ~1.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 5 | Post a comment (or a formal PR review) to a GitHub issue or pull request in mariadb-operator/mariadb-operator. | mariadb-operator/ | 1k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 6 | Create, review or validate Netdata Prometheus chart profiles, exporter dashboard design, collection policy and stock semantic proofs. | netdata/ | 81k | — | ~4.7k | Automated safety check: Pass | GPL-3.0 | today |
| 7 | Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts. | NVlabs/ | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 21 days ago |
| 8 | Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments. | alibaba/ | 415 | — | ~1.9k | Automated safety check: Pass | Unknown | 15 days ago |
| 9 | Generate Perses dashboards or single panels for GreptimeDB. An agent skill from GreptimeTeam/dashboard. | GreptimeTeam/ | 111 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | today |
| 10 | 10.Code Review Reviews code for correctness and potential bugs, pinpoints bug locations by file and line, and suggests concrete fixes. | aide-family/ | 253 | — | ~815 | Automated safety check: Pass | No licence | 3 mo ago |
| 11 | 11.Opsany 通过 opsany-mcp-server 连接 OpsAny 运维平台,实现 CMDB 资源查询和操作、模型管理、工单管理、用户管理、作业执行及主机纳管等全栈运维操作。 | unixhot/ | 181 | — | ~2k | Automated safety check: Pass | Unknown | 2 mo ago |
| 12 | End-to-end docker-compose test harness for the minecraft-prometheus-exporter. | dirien/ | 142 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 13 | A skill your agent uses when changing metric CSV files, exporter-owned counters, Prometheus labels, or metric rendering. | NVIDIA/ | 1.9k | — | ~149 | Automated safety check: Pass | Apache-2.0 | 20 days ago |
| 14 | Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO… | redis/ | 165 | 2 repos | ~911 | Automated safety check: Pass | MIT | 1 mo ago |
| 15 | Visually verify Eneru browser-dashboard changes against a live daemon or audit an exact deployment. | m4r1k/ | 149 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 16 | Installs, uninstalls, checks and troubleshoots the KubeSphere Gateway extension built on ingress-nginx, including gateways stuck in bad states and Helm or pod failures. | kubesphere/ | 17k | — | ~2.9k | Automated safety check: Pass | Unknown | 2 mo ago |
| 17 | Develops Go microservices with Kratos v2 following official design philosophy, DDD/Clean Architecture layout, Protobuf API, error/config/middleware patterns, and observability. | aide-family/ | 253 | — | ~1.5k | Automated safety check: Pass | No licence | 3 mo ago |
| 18 | Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)… | grafana/ | 281 | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 19 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 20 | Bootstrap, create, connect to, operate, secure, scale, upgrade, troubleshoot, inspect, and tear down Alibaba Cloud Container Compute Service (ACS) Agent Sandbox environments. | cinience/ | 397 | — | ~2.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 21 | 21.Signoz Manage the self-hosted SigNoz observability stack in this GitOps repo. | qjoly/ | 112 | — | ~6.1k | Automated safety check: Pass | WTFPL | today |
| 22 | A skill your agent uses when publishing prometheus-proxy to Maven Central, cutting a release, running a snapshot publish, or bumping the project version — covers the Maven Central coordinates, GPG… | pambrose/ | 157 | — | ~506 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 23 | Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units… | grafana/ | 281 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech. | qdrant/ | 254 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | 当需要为 funboost 创建 Consumer 或 Publisher 的 Mixin 扩展类时使用。触发场景:添加监控、熔断、限流、链路追踪等横切关注点,编写自定义前置/后置处理钩子。关键词:mixin, consumeroverridecls, publisheroverridecls, ConsumerMixin, 自定义消费者, hook, 拦截器, 熔断器, 监控… | ydf0509/ | 893 | — | ~2.1k | Automated safety check: Pass | No licence | 1 mo ago |
| 26 | 26.Graft This repo is indexed by graft/. An agent skill from m4r1k/Eneru. | m4r1k/ | 149 | 1 repo | ~2.3k | Automated safety check: Pass | MIT | today |
| 27 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 281 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 28 | Mandatory pre-release deep review for minor/major releases (X.Y.0 / X.0.0). | m4r1k/ | 149 | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 29 | Implements backend modules from proto definitions for goddess, marksman, and rabbit apps. | aide-family/ | 253 | — | ~4.1k | Automated safety check: Pass | No licence | 3 mo ago |
| 30 | Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load. | prometheus/ | 120 | — | ~584 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 31 | A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup. | archestra-ai/ | 4.4k | — | ~1.2k | Automated safety check: Pass | Unknown | today |
| 32 | 32.Helm Chart A skill your agent uses for Helm chart work - creating charts, modifying existing charts, values design, testing. | astronomer/ | 491 | — | ~6.5k | Automated safety check: Pass | Unknown | today |
| 33 | A skill your agent uses when the user asks about AI Center Coding Agents data, wants to reproduce or extend the Coding Agents dashboards, or asks questions about usage, cost, tokens, sessions… | coralogix/ | 121 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 34 | 34.Docs Site A skill your agent uses when editing, building, previewing, or deploying the prometheus-proxy documentation site under website/prometheus-proxy — covers the Zensical config, code-snippet resolution… | pambrose/ | 157 | — | ~250 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 35 | A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server. | agentfront/ | 146 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 36 | Diagnose a live, running lean devnet from its Prometheus and container logs, then report findings with proposed fixes and stop for approval. | geanlabs/ | 177 | — | ~3.7k | Automated safety check: Warn | No licence | today |
| 37 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 38 | Set up metrics collection and visualization with Prometheus and Grafana. | BagelHole/ | 1.1k | — | ~2.5k | Automated safety check: Pass | MIT | 4 mo ago |
| 39 | A skill your agent uses when the user asks to "check data usage", "list TCO policies", "reduce Coralogix costs", "optimize observability spend", "lower our logging bill", "data budget exceeded"… | coralogix/ | 121 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 40 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 120 | — | ~592 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 41 | Aggregate and centralize performance metrics from applications, systems, databases, caches, and services. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 42 | Complete guide to Prometheus setup, metric collection, scrape configuration, and recording rules. | davila7/ | 32k | 12 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 43 | In-memory caching in Golang using samber/hot — eviction algorithms (LRU, LFU, TinyLFU, W-TinyLFU, S3FIFO, ARC, TwoQueue, SIEVE, FIFO), TTL, cache loaders, sharding, stale-while-revalidate, missing… | samber/ | 3.4k | — | ~2k | Automated safety check: Pass | MIT | 7 days ago |
| 44 | Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact. | prometheus/ | 120 | — | ~547 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 45 | LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs. | majiayu000/ | 117 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 46 | Author, modify, or review Netdata collectors across Go, IBM, C, Rust and external plugins. | netdata/ | 81k | — | ~1.9k | Automated safety check: Pass | GPL-3.0 | today |
| 47 | DevOps 工程师 Agent — CI/CD 流水线、容器化与 K8s、基础设施即代码、可观测性. An agent skill from peterfei/ai-agent-team. | peterfei/ | 441 | — | ~1k | Automated safety check: Pass | MIT | 3 mo ago |
| 48 | 48.Expert Ops 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |