Search
Prometheus · Monitoring and alerting
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Deploys KubeEye on KubeSphere and writes InspectRule and InspectPlan resources to inspect cluster health, then retrieves the inspection reports. | kubesphere/ | 17k | — | ~3.6k | Automated safety check: Pass | Unknown | 2 mo ago |
| 2 | Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API. | slopus/ | 24k | — | ~2k | Automated safety check: Notes | MIT | today |
| 3 | 3.Syncmeta Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta). | pawurb/ | 1.9k | — | ~1.2k | Automated safety check: Notes | MIT | yesterday |
| 4 | Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions. | kubesphere/ | 17k | — | ~1.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 5 | Create, review or validate Netdata Prometheus chart profiles, exporter dashboard design, collection policy and stock semantic proofs. | netdata/ | 81k | — | ~4.7k | Automated safety check: Pass | GPL-3.0 | today |
| 6 | Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts. | NVlabs/ | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 23 days ago |
| 7 | Generate Perses dashboards or single panels for GreptimeDB. An agent skill from GreptimeTeam/dashboard. | GreptimeTeam/ | 111 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | Reviews code for correctness and potential bugs, pinpoints bug locations by file and line, and suggests concrete fixes. | aide-family/ | 253 | — | ~815 | Automated safety check: Pass | No licence | 3 mo ago |
| 9 | End-to-end docker-compose test harness for the minecraft-prometheus-exporter. | dirien/ | 142 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 10 | A skill your agent uses when changing metric CSV files, exporter-owned counters, Prometheus labels, or metric rendering. | NVIDIA/ | 1.9k | — | ~149 | Automated safety check: Pass | Apache-2.0 | 22 days ago |
| 11 | Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO… | redis/ | 166 | 2 repos | ~911 | Automated safety check: Pass | MIT | 1 mo ago |
| 12 | Visually verify Eneru browser-dashboard changes against a live daemon or audit an exact deployment. | m4r1k/ | 149 | — | ~1.4k | Automated safety check: Pass | MIT | 3 days ago |
| 13 | Develops Go microservices with Kratos v2 following official design philosophy, DDD/Clean Architecture layout, Protobuf API, error/config/middleware patterns, and observability. | aide-family/ | 253 | — | ~1.5k | Automated safety check: Pass | No licence | 3 mo ago |
| 14 | Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)… | grafana/ | 282 | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 15 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 16 | Bootstrap, create, connect to, operate, secure, scale, upgrade, troubleshoot, inspect, and tear down Alibaba Cloud Container Compute Service (ACS) Agent Sandbox environments. | cinience/ | 397 | — | ~2.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 17 | 17.Signoz Manage the self-hosted SigNoz observability stack in this GitOps repo. | qjoly/ | 112 | — | ~6.1k | Automated safety check: Pass | WTFPL | yesterday |
| 18 | A skill your agent uses when publishing prometheus-proxy to Maven Central, cutting a release, running a snapshot publish, or bumping the project version — covers the Maven Central coordinates, GPG… | pambrose/ | 157 | — | ~506 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 19 | 对远程多实例MySQL数据库执行全方位深度巡检,覆盖基础健康、连接负载、性能慢查询、索引冗余、主从复制、容量空间、账号安全、配置风险八大维度,全自动完成巡检扫描、风险识别、问题定级、优化建议、报告归档与飞书推送,适用于生产/测试所有运行中MySQL实例常态化合规巡检。适用场景:用户要求进行 MySQL 全链路健康检查、MySQL 综合巡检、MySQL 风险扫描、MySQL 性能审计、MySQL… | openocta/ | 167 | — | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 20 | Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units… | grafana/ | 282 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 21 | Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech. | qdrant/ | 254 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 22 | 当需要为 funboost 创建 Consumer 或 Publisher 的 Mixin 扩展类时使用。触发场景:添加监控、熔断、限流、链路追踪等横切关注点,编写自定义前置/后置处理钩子。关键词:mixin, consumeroverridecls, publisheroverridecls, ConsumerMixin, 自定义消费者, hook, 拦截器, 熔断器, 监控… | ydf0509/ | 895 | — | ~2.1k | Automated safety check: Pass | No licence | 2 mo ago |
| 23 | 23.Oryxos Init 初始化 OryxOS(或同类 JDK 21 + Spring Boot 3.x 企业级单体)的工程地基:Maven 多模块骨架、 结构化日志、Actuator + Prometheus 监控、Spring MVC + 虚拟线程、springdoc OpenAPI、 统一响应体与全局异常/错误码、Google 格式 + 阿里编码规约(Spotless + 阿里 P3C +… | oryx-labs/ | 187 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | 24.Graft This repo is indexed by graft/. An agent skill from m4r1k/Eneru. | m4r1k/ | 149 | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 3 days ago |
| 25 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 282 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 26 | Mandatory pre-release deep review for minor/major releases (X.Y.0 / X.0.0). | m4r1k/ | 149 | — | ~1.9k | Automated safety check: Pass | MIT | 3 days ago |
| 27 | MUST USE when investigating performance issues on a ClickHouse-managed Postgres instance. | ClickHouse/ | 545 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 28 | Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load. | prometheus/ | 121 | — | ~584 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 29 | Implements backend modules from proto definitions for goddess, marksman, and rabbit apps. | aide-family/ | 253 | — | ~4.1k | Automated safety check: Pass | No licence | 3 mo ago |
| 30 | A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup. | archestra-ai/ | 4.4k | — | ~1.2k | Automated safety check: Pass | Unknown | today |
| 31 | 31.Docs Site A skill your agent uses when editing, building, previewing, or deploying the prometheus-proxy documentation site under website/prometheus-proxy — covers the Zensical config, code-snippet resolution… | pambrose/ | 157 | — | ~250 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 32 | A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server. | agentfront/ | 146 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 33 | Diagnose a live, running lean devnet from its Prometheus and container logs, then report findings with proposed fixes and stop for approval. | geanlabs/ | 177 | — | ~3.7k | Automated safety check: Warn | MIT | yesterday |
| 34 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 35 | Set up metrics collection and visualization with Prometheus and Grafana. | BagelHole/ | 1.2k | — | ~2.5k | Automated safety check: Pass | MIT | 4 mo ago |
| 36 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 121 | — | ~592 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 37 | Complete guide to Prometheus setup, metric collection, scrape configuration, and recording rules. | davila7/ | 33k | 12 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 38 | In-memory caching in Golang using samber/hot — eviction algorithms (LRU, LFU, TinyLFU, W-TinyLFU, S3FIFO, ARC, TwoQueue, SIEVE, FIFO), TTL, cache loaders, sharding, stale-while-revalidate, missing… | samber/ | 3.4k | — | ~2k | Automated safety check: Pass | MIT | 10 days ago |
| 39 | Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact. | prometheus/ | 121 | — | ~547 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 40 | LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs. | majiayu000/ | 118 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 41 | Author, modify, or review Netdata collectors across Go, IBM, C, Rust and external plugins. | netdata/ | 81k | — | ~1.9k | Automated safety check: Pass | GPL-3.0 | today |
| 42 | Improve and validate the Kaniop Grafana dashboard against repository metrics and the grigri live cluster. | pando85/ | 132 | — | ~987 | Automated safety check: Pass | AGPL-3.0 | today |
| 43 | 43.Expert Ops 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |
| 44 | Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes. | google/ | 21k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 45 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 46 | Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery. | Jeffallan/ | 12k | — | ~1.6k | Automated safety check: Pass | MIT | 8 days ago |
| 47 | 47.SRE Engineer Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 8 days ago |
| 48 | Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML. | prometheus/ | 121 | — | ~765 | Automated safety check: Pass | Apache-2.0 | yesterday |