Search
Observability
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Audit, plan, and fix logging in any Python project. An agent skill from areed1192/interactive-brokers-api. | areed1192/ | 103 | — | ~1.9k | Automated safety check: Pass | MIT | 2 mo ago |
| 98 | Add lightweight application monitoring, health checks, and request metrics with minimal dependencies and operational overhead. | ejboy/ | 116 | — | ~1.9k | Automated safety check: Pass | MIT | 8 days ago |
| 99 | This skill helps an LLM generate correct AxAgent observability code using @ax-llm/ax. | dosco/ | 107 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 100 | 100.Inspect Inspect and debug live streaming agent sessions to understand what the agent did. | agentevals-dev/ | 163 | — | ~534 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 101 | Guides log level choices and when to raise a structured Sentry event instead of a plain log line in the Warp Rust codebase, keeping secrets out of logs. | warpdotdev/ | 65k | 1 repo | ~5.6k | Automated safety check: Pass | AGPL-3.0 | today |
| 102 | Plan what to measure in mobile apps. An agent skill from nexus-labs-automation/mobile-observability. | nexus-labs-automation/ | 116 | — | ~517 | Automated safety check: Pass | MIT | 1 mo ago |
| 103 | Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load. | prometheus/ | 121 | — | ~584 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 104 | 104.Otel Go OpenTelemetry in Go — SDK setup, API surface, breaking changes, contrib instrumentation libraries (otelhttp, otelgrpc, otelmongo), compile-time zero-code instrumentation (otelc), and performance… | ollygarden/ | 106 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 105 | Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality. | wshobson/ | 40k | 9 repos | ~607 | Automated safety check: Pass | MIT | 6 days ago |
| 106 | Correlates a Codex CLI session's local transcript with a model router's production logs to explain why a reply rendered the way it did. | weave-os/ | 5.6k | — | ~4.5k | Automated safety check: Warn | Apache-2.0 | today |
| 107 | 107.Clawmetry Real-time observability for OpenClaw agents — local dashboard + optional encrypted cloud sync. | vivekchand/ | 426 | — | ~992 | Automated safety check: Pass | MIT | 3 days ago |
| 108 | LLM observability platform for tracing, evaluation, and monitoring. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 109 | A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup. | archestra-ai/ | 4.4k | — | ~1.2k | Automated safety check: Pass | Unknown | today |
| 110 | 基于 LoongSuite Pilot / AI Coding Agent 日志生成事件洞察、组织洞察、数据质量、研发效能和 AI Native 使用类 SLS 报表时使用;包含 AI Coding 事件表语义,以及团队报表可选的部门维表、deptuser 组织关系、指标口径和公共 CTE,通常与 sls-dashboard-builder 一起使用。 | alibaba/ | 201 | — | ~944 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 111 | A skill your agent uses when an agent workflow needs production-like runtime controls for context, tools, permissions, observability, scheduling, evaluation, recovery, or maintenance. | Mark393295827/ | 141 | — | ~2.2k | Automated safety check: Pass | MIT | 22 days ago |
| 112 | 112.Observability Observability best practices. An agent skill from 2SSK/dot-files. | 2SSK/ | 249 | — | ~676 | Automated safety check: Pass | MIT | today |
| 113 | 113.Interlinked Overview and router for the Interlinked CLI — a local guard, quality-enforcement, simplification-review, semantic-code-search, and observability layer for AI coding agents. | QuentinCody/ | 178 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 114 | Vercel AI Gateway guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and… | vercel/ | 301 | — | ~4.4k | Automated safety check: Notes | Unknown | yesterday |
| 115 | 115.Network Tracing Instrument API requests with spans and distributed tracing. An agent skill from nexus-labs-automation/mobile-observability. | nexus-labs-automation/ | 116 | — | ~672 | Automated safety check: Pass | MIT | 1 mo ago |
| 116 | Find the top CPU, memory, or disk consumers. Use for capacity reviews, noisy-neighbor hunts, and top-N questions about which jobs, pods, or… | prometheus/ | 121 | — | ~569 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 117 | A skill your agent uses when building cloud-native apps. An agent skill from moeru-ai/auv. | moeru-ai/ | 100 | 1 repo | ~1k | Automated safety check: Pass | Apache-2.0 | today |
| 118 | 118.Setup Agent Tail Configure agent-tail log aggregation for the current project. | qdhenry/ | 1.3k | — | ~1.8k | Automated safety check: Pass | No licence | 7 mo ago |
| 119 | 119.Otel Ottl OpenTelemetry Transformation Language (OTTL) expert for writing and debugging telemetry transformations in the OpenTelemetry Collector. | ollygarden/ | 106 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 120 | Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов). | AnastasiyaW/ | 154 | — | ~3.8k | Automated safety check: Pass | MIT | today |
| 121 | Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside… | livekit-examples/ | 264 | 1 repo | ~2.4k | Automated safety check: Pass | MIT | yesterday |
| 122 | A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server. | agentfront/ | 146 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 123 | 123.Opik Instrument Add Opik tracing to an existing app and verify a real trace lands. | comet-ml/ | 220 | — | ~2.7k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 124 | 124.Backend Design Elite Tier Backend standards, including Vertical Slice Architecture, Zero Trust Security, and High-Performance API protocols. | xenitV1/ | 229 | — | ~2.1k | Automated safety check: Notes | MIT | 8 mo ago |
| 125 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 126 | 126.Agentsop Dify SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable. | agentsope/ | 436 | — | ~5.4k | Automated safety check: Notes | MIT | 2 days ago |
| 127 | 127.Mz Query Tracing Debug SQL execution time via distributed tracing (OpenTelemetry / Tempo). | MaterializeInc/ | 6.4k | — | ~1.8k | Automated safety check: Pass | Unknown | today |
| 128 | Guides building event-driven background work for LobeHub agents, with sources, signals, actions, policies, workflow handoff and deduplication. | lobehub/ | 83k | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 129 | 129.Agents SDK Build AI agents on Cloudflare Workers using the Agents SDK. An agent skill from hodgef/apiker. | hodgef/ | 127 | 3 repos | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 130 | INVOKE THIS SKILL when adding Arize AX tracing or observability to an app for the first time, or when the user wants to instrument their LLM app or get started with LLM observability. | boshi-xixixi/ | 276 | — | ~5.1k | Automated safety check: Notes | MIT | 5 mo ago |
| 131 | Monitor an AG2 beta agent's stream — log events, detect repeated tool calls, track token spend, build trigger-driven observers, route observer alerts to the model, and halt on FATAL conditions. | ag2ai/ | 252 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 132 | A skill your agent uses when the user asks to "check data usage", "list TCO policies", "reduce Coralogix costs", "optimize observability spend", "lower our logging bill", "data budget exceeded"… | coralogix/ | 121 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 133 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 121 | — | ~592 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 134 | Inspects LobeHub agent execution snapshots by operation ID: failed tool calls, arguments and results, available tools, LLM calls and the surrounding context. | lobehub/ | 83k | — | ~5.3k | Automated safety check: Notes | Unknown | today |
| 135 | Guides reading mecatl's perf MCP data to find why a running harness is slow, leaking goroutines or growing in memory, using cheap reads before any CPU capture. | stacklok/ | 254 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 136 | 136.Otel Profiles OpenTelemetry profiles signal and the eBPF profiler (otelcol-ebpf-profiler, profiling receiver). | ollygarden/ | 106 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 137 | 137.Add Metric TRIGGER when user asks to add or modify a metric, counter, gauge, or histogram, or to track/measure an operation. | microbus-io/ | 172 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 138 | Audit code for observability gaps — debug logs left in, errors caught without being logged, missing context on log entries, untracked slow operations. | markmdev/ | 187 | — | ~836 | Automated safety check: Pass | No licence | 7 mo ago |
| 139 | 139.Loop Architect Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them. | fabricioctelles/ | 106 | — | ~2.1k | Automated safety check: Notes | MIT | today |
| 140 | 140.Evidence Ledger Build the deterministic evidence ledger (artifactmanifest.json + claims.json) that every other Anti-Autoresearch auditor reads. | wanshuiyin/ | 160 | — | ~11k | Automated safety check: Notes | MIT | 4 days ago |
| 141 | Command reference for omniroute's audit, logs, policy and telemetry commands: search and export audit trails, manage access policies and review request history for compliance work. | diegosouzapw/ | 75k | — | ~733 | Automated safety check: Pass | MIT | today |
| 142 | Checks the health of an OmniRoute gateway: provider circuit breakers, latency percentiles, budget guard alerts, connection cooldowns and model lockouts. | diegosouzapw/ | 75k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 143 | Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs… | aws/ | 103 | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 144 | Product analytics with posthog. An agent skill from langfuse/langfuse. | langfuse/ | 36k | — | ~3.2k | Automated safety check: Pass | Unknown | today |