Search
Monitoring and alerting
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Investigates distributed application performance using PostHog APM (OpenTelemetry span) data via MCP. | PostHog/ | 40k | — | ~3.5k | Automated safety check: Pass | Unknown | today |
| 50 | Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space. | huggingface/ | 11k | 2 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 51 | A skill your agent uses when publishing prometheus-proxy to Maven Central, cutting a release, running a snapshot publish, or bumping the project version — covers the Maven Central coordinates, GPG… | pambrose/ | 157 | — | ~506 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 52 | Adds Pydantic Logfire tracing, logging and metrics to Python, JavaScript or TypeScript and Rust projects, with the correct setup order and library extras. | basicmachines-co/ | 4.1k | — | ~2.3k | Automated safety check: Pass | AGPL-3.0 | today |
| 53 | 53.Loki Logs A skill your agent uses when investigating production behaviour from logs — a 500/error in prod, a failing or stuck Celery task, tracing one request/traceid/user across services, or confirming a… | letsrevel/ | 110 | — | ~918 | Automated safety check: Notes | MIT | 4 days ago |
| 54 | 对远程多实例MySQL数据库执行全方位深度巡检,覆盖基础健康、连接负载、性能慢查询、索引冗余、主从复制、容量空间、账号安全、配置风险八大维度,全自动完成巡检扫描、风险识别、问题定级、优化建议、报告归档与飞书推送,适用于生产/测试所有运行中MySQL实例常态化合规巡检。适用场景:用户要求进行 MySQL 全链路健康检查、MySQL 综合巡检、MySQL 风险扫描、MySQL 性能审计、MySQL… | openocta/ | 167 | — | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 55 | Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units… | grafana/ | 282 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 56 | Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech. | qdrant/ | 254 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 57 | AWS cost optimization, monitoring, and operational excellence expert. | zxkane/ | 367 | — | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 58 | 当需要为 funboost 创建 Consumer 或 Publisher 的 Mixin 扩展类时使用。触发场景:添加监控、熔断、限流、链路追踪等横切关注点,编写自定义前置/后置处理钩子。关键词:mixin, consumeroverridecls, publisheroverridecls, ConsumerMixin, 自定义消费者, hook, 拦截器, 熔断器, 监控… | ydf0509/ | 895 | — | ~2.1k | Automated safety check: Pass | No licence | 2 mo ago |
| 59 | 59.Oryxos Init 初始化 OryxOS(或同类 JDK 21 + Spring Boot 3.x 企业级单体)的工程地基:Maven 多模块骨架、 结构化日志、Actuator + Prometheus 监控、Spring MVC + 虚拟线程、springdoc OpenAPI、 统一响应体与全局异常/错误码、Google 格式 + 阿里编码规约(Spotless + 阿里 P3C +… | oryx-labs/ | 187 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 60 | A skill your agent uses to run a four-panel adversarial advisory review on any pull request that touches the OpenAPM specification artifact (docs/src/content/docs/specs/openapm-.md), its inline /… | microsoft/ | 4k | — | ~4.9k | Automated safety check: Pass | MIT | yesterday |
| 61 | Help choose, configure, and test local agento11y guard packs for coding-agent tool calls. | grafana/ | 128 | — | ~2.6k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 62 | 62.Graft This repo is indexed by graft/. An agent skill from m4r1k/Eneru. | m4r1k/ | 149 | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 3 days ago |
| 63 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 282 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 64 | Mandatory pre-release deep review for minor/major releases (X.Y.0 / X.0.0). | m4r1k/ | 149 | — | ~1.9k | Automated safety check: Pass | MIT | 3 days ago |
| 65 | MUST USE when investigating performance issues on a ClickHouse-managed Postgres instance. | ClickHouse/ | 545 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 66 | Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load. | prometheus/ | 121 | — | ~584 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 67 | Implements backend modules from proto definitions for goddess, marksman, and rabbit apps. | aide-family/ | 253 | — | ~4.1k | Automated safety check: Pass | No licence | 3 mo ago |
| 68 | Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality. | wshobson/ | 40k | 9 repos | ~607 | Automated safety check: Pass | MIT | 6 days ago |
| 69 | Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 70 | Provides patterns to configure Spring Boot Actuator for production-grade monitoring, health probes, secured management endpoints, and Micrometer metrics across JVM services. | giuseppe-trisciuoglio/ | 357 | — | ~2.2k | Automated safety check: Notes | MIT | 1 mo ago |
| 71 | A skill your agent uses to queue maintainer-accepted microsoft/apm issues (status/accepted) and fan them out through an isolated pool (default 2) of autopilot-issue-delivery-worker sessions. | microsoft/ | 4k | — | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 72 | A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup. | archestra-ai/ | 4.4k | — | ~1.2k | Automated safety check: Pass | Unknown | today |
| 73 | Find the top CPU, memory, or disk consumers. Use for capacity reviews, noisy-neighbor hunts, and top-N questions about which jobs, pods, or… | prometheus/ | 121 | — | ~569 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 74 | 74.Docs Site A skill your agent uses when editing, building, previewing, or deploying the prometheus-proxy documentation site under website/prometheus-proxy — covers the Zensical config, code-snippet resolution… | pambrose/ | 157 | — | ~250 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 75 | Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval. | superset-sh/ | 15k | — | ~1k | Automated safety check: Pass | Unknown | today |
| 76 | A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server. | agentfront/ | 146 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 77 | Diagnose a live, running lean devnet from its Prometheus and container logs, then report findings with proposed fixes and stop for approval. | geanlabs/ | 177 | — | ~3.7k | Automated safety check: Warn | MIT | yesterday |
| 78 | Overview of Neon, a complete set of cloud backend primitives around Lakebase Postgres: Auth, Object Storage, Functions, and the AI Gateway. | neondatabase/ | 100 | — | ~8.7k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 79 | Checks whether the Telegram daemon is alive, shows recent logs, finds the remote control URL and diagnoses MCP server problems. | k1p1l0/ | 132 | — | ~746 | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 80 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 81 | Queue open microsoft/apm issues and fan them out through an isolated pool (default 2) of autopilot-issue-triage-worker sessions. | microsoft/ | 4k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 82 | Create and manage production-ready Grafana dashboards for comprehensive system observability. | davila7/ | 33k | 14 repos | ~2.1k | Automated safety check: Pass | MIT | today |
| 83 | Set up metrics collection and visualization with Prometheus and Grafana. | BagelHole/ | 1.2k | — | ~2.5k | Automated safety check: Pass | MIT | 4 mo ago |
| 84 | APM - traces, services, dependencies, performance analysis. An agent skill from DataDog/pup. | DataDog/ | 1k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 85 | Build or refactor a Tool-channel provider (PagerDuty, Opsgenie, future incident/alerting tools) to be endpoint-routed: per-subscriber secrets encrypted on the channel endpoint resource, a stateless… | novuhq/ | 40k | — | ~2.8k | Automated safety check: Pass | Unknown | yesterday |
| 86 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 121 | — | ~592 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 87 | Complete guide to Prometheus setup, metric collection, scrape configuration, and recording rules. | davila7/ | 33k | 12 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 88 | In-memory caching in Golang using samber/hot — eviction algorithms (LRU, LFU, TinyLFU, W-TinyLFU, S3FIFO, ARC, TwoQueue, SIEVE, FIFO), TTL, cache loaders, sharding, stale-while-revalidate, missing… | samber/ | 3.4k | — | ~2k | Automated safety check: Pass | MIT | 10 days ago |
| 89 | Queue open microsoft/apm pull requests (community fixes, PRs without an issue, or PRs that close an issue) and fan them out through an isolated pool (default 2) of autopilot-pr-triage-worker sessions. | microsoft/ | 4k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 90 | Checks the health of an OmniRoute gateway: provider circuit breakers, latency percentiles, budget guard alerts, connection cooldowns and model lockouts. | diegosouzapw/ | 75k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 91 | 91.Ops Monitor OPS on-demand: This skill should be used when the user asks to "datadog", "APM alerts", or… | Lifecycle-Innovations-Limited/ | 542 | 1 repo | ~1.5k | Automated safety check: Notes | MIT | yesterday |
| 92 | Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record… | grafana/ | 128 | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 93 | Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact. | prometheus/ | 121 | — | ~547 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 94 | LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs. | majiayu000/ | 118 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 95 | Author, modify, or review Netdata collectors across Go, IBM, C, Rust and external plugins. | netdata/ | 81k | — | ~1.9k | Automated safety check: Pass | GPL-3.0 | today |
| 96 | Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config… | grafana/ | 282 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |