Search
Prometheus · Observability
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API. | slopus/ | 24k | — | ~2k | Automated safety check: Notes | MIT | today |
| 2 | Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions. | kubesphere/ | 17k | — | ~1.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 3 | Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments. | alibaba/ | 415 | — | ~1.9k | Automated safety check: Pass | Unknown | 17 days ago |
| 4 | Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO… | redis/ | 166 | 2 repos | ~911 | Automated safety check: Pass | MIT | 1 mo ago |
| 5 | Develops Go microservices with Kratos v2 following official design philosophy, DDD/Clean Architecture layout, Protobuf API, error/config/middleware patterns, and observability. | aide-family/ | 253 | — | ~1.5k | Automated safety check: Pass | No licence | 3 mo ago |
| 6 | 当需要为 funboost 创建 Consumer 或 Publisher 的 Mixin 扩展类时使用。触发场景:添加监控、熔断、限流、链路追踪等横切关注点,编写自定义前置/后置处理钩子。关键词:mixin, consumeroverridecls, publisheroverridecls, ConsumerMixin, 自定义消费者, hook, 拦截器, 熔断器, 监控… | ydf0509/ | 895 | — | ~2.1k | Automated safety check: Pass | No licence | 2 mo ago |
| 7 | Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load. | prometheus/ | 121 | — | ~584 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup. | archestra-ai/ | 4.4k | — | ~1.2k | Automated safety check: Pass | Unknown | today |
| 9 | A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server. | agentfront/ | 146 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 10 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 11 | A skill your agent uses when the user asks to "check data usage", "list TCO policies", "reduce Coralogix costs", "optimize observability spend", "lower our logging bill", "data budget exceeded"… | coralogix/ | 121 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 12 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 121 | — | ~592 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 13 | Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact. | prometheus/ | 121 | — | ~547 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 14 | LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs. | majiayu000/ | 118 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 15 | Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes. | google/ | 21k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 16 | Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery. | Jeffallan/ | 12k | — | ~1.6k | Automated safety check: Pass | MIT | 7 days ago |
| 17 | Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML. | prometheus/ | 121 | — | ~765 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 18 | Installs and configures the WizTelemetry Notification extension for KubeSphere: channel setup, alert routing by tenant labels, silences and troubleshooting. | kubesphere/ | 17k | — | ~6.1k | Automated safety check: Pass | Unknown | 2 mo ago |
| 19 | Find out what a running Mendix app actually does — logs, Prometheus metrics, OpenTelemetry traces and the model catalog, joined across sources. | mendixlabs/ | 129 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 20 | A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"… | coralogix/ | 121 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 21 | Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server. | prometheus/ | 121 | — | ~587 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 22 | Build and deploy a Coralogix dashboard for a given service from its logs, spans, metrics, and service specs. | coralogix/ | 121 | — | ~4.7k | Automated safety check: Warn | Apache-2.0 | 4 days ago |
| 23 | Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /… | grafana/ | 282 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 24 | Go observability — always-on production signals: slog logging, Prometheus metrics, OpenTelemetry tracing, pprof profiling, alerting, Grafana. | context-labs/ | 1.1k | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 6 days ago |
| 25 | Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status. | azrtydxb/ | 108 | — | ~916 | Automated safety check: Pass | Apache-2.0 | today |
| 26 | 当需要为 funboost 任务添加监控、链路追踪或告警时使用。触发场景:Prometheus 指标、OpenTelemetry 链路追踪、异常告警通知、周期额度限制、函数结果持久化。关键词:Prometheus, OpenTelemetry, OTel, 告警, 监控, metrics, tracing, AlertNotifier, PeriodicQuota。 | ydf0509/ | 895 | — | ~4.9k | Automated safety check: Pass | No licence | 2 mo ago |
| 27 | Monitoring, logging, and tracing implementation using OpenTelemetry as the unified standard. | ancoleman/ | 525 | — | ~3k | Automated safety check: Pass | MIT | 10 mo ago |
| 28 | Configures GKE observability, including Cloud Logging, Cloud Monitoring, and managed Prometheus. | google/ | 21k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 29 | 29.Cost Export Export cost-tracking telemetry in Prometheus textfile or webhook JSON formats — for external observability (Grafana, Datadog, custom dashboards) | ruvnet/ | 74k | — | ~687 | Automated safety check: Notes | MIT | yesterday |
| 30 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 31 | Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. | aws/ | 2.8k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 32 | A skill your agent uses when you need to implement or improve Java metrics observability with Micrometer — including meter design, naming/tag conventions, cardinality control… | jabrena/ | 447 | — | ~868 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 33 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 287 | — | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 34 | Observability patterns for Python applications. An agent skill from aiskillstore/marketplace. | aiskillstore/ | 433 | 1 repo | ~1.3k | Automated safety check: Pass | No licence | yesterday |
| 35 | This skill provides AWS cost optimization, monitoring, and operational best practices with integrated MCP servers for billing analysis, cost estimation, observability, and security assessment. | Microck/ | 404 | 1 repo | ~2.5k | Automated safety check: Pass | Unknown | 1 mo ago |
| 36 | Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection. | yonatangross/ | 292 | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 37 | OpenTelemetry, distributed tracing, structured logging, metrics (Prometheus, Grafana, Datadog). | TheBeardedBearSAS/ | 107 | — | ~547 | Automated safety check: Pass | MIT | 27 days ago |
| 38 | Monitoring and observability with OpenTelemetry, Prometheus, Grafana dashboards, and structured logging | rohitg00/ | 2.7k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 5 mo ago |
| 39 | Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI. | softspark/ | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 40 | Set up Apollo.io monitoring and observability. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 41 | Monitor ClickHouse with Prometheus metrics, Grafana dashboards, system table queries, and alerting for query performance, merge health, and resource usage. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 42 | Set up Customer.io monitoring and observability. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 43 | Set up comprehensive observability for Deepgram integrations. | jeremylongshore/ | 2.8k | — | ~3k | Automated safety check: Pass | MIT | yesterday |
| 44 | Implement observability for Evernote integrations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 45 | Set up monitoring, metrics, and alerting for Figma API integrations. | jeremylongshore/ | 2.8k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 46 | Implement comprehensive observability for Gamma integrations. | jeremylongshore/ | 2.8k | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 47 | Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts. | jeremylongshore/ | 2.8k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 48 | A skill your agent uses when you need production monitoring for an Intercom integration — instrumenting API calls with metrics and traces, standing up dashboards, or wiring alerts for error rate… | jeremylongshore/ | 2.8k | — | ~1.7k | Automated safety check: Pass | MIT | yesterday |