Search
OpenTelemetry · Site reliability engineering
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 2 | Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead… | EliasOulkadi/ | 114 | — | ~3.6k | Automated safety check: Notes | MIT | 6 days ago |
| 3 | Search and read official Coralogix platform documentation using cx docs search and cx docs fetch. | coralogix/ | 121 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 4 | Intent-based observability + traceability router across layers, boundaries, and signals. | first-fluke/ | 1.3k | — | ~4.9k | Automated safety check: Pass | MIT | yesterday |
| 5 | Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health… | elastic/ | 592 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 6 | Expert performance engineer specializing in modern observability. | Dokhacgiakhoa/ | 508 | — | ~743 | Automated safety check: Pass | Unknown | 4 mo ago |
| 7 | Design or troubleshoot telemetry, SLOs, and incident diagnostics. | first-fluke/ | 1.3k | — | ~3.7k | Automated safety check: Pass | MIT | yesterday |
| 8 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 9 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 287 | — | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 10 | Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI. | softspark/ | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 11 | Wire LangChain 1.0 / LangGraph 1.0 traces into an OpenTelemetry-native backend (Jaeger, Honeycomb, Grafana Tempo, Datadog) with LLM-specific SLOs, safe prompt-content policy, and subgraph-aware span… | jeremylongshore/ | 2.8k | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 12 | 12.Monitoring A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and… | ericrisco/ | 180 | — | ~3.1k | Automated safety check: Pass | MIT | yesterday |
| 13 | 13.Telemetry Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector… | magnus919/ | 115 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |