Topic · AI & LLM Engineering
Best LLM observability skills for Claude Code, Codex and other agents.
- skills
- 217
- official
- 17
LLM observability skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Finds every LLM workflow in a repository, proposes a labeling table and, once you agree, wires labels so Caveman Cloud groups spend per workflow. | JuliusBrussee/ | 110k | 1 repo | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Read-only review of Caveman Cloud data to explain where LLM spend goes: cost, score, workflows, traces, latency, errors, routing and verified savings. | JuliusBrussee/ | 110k | 1 repo | ~927 | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use. | ComposioHQ/ | 77k | 9 repos | ~2.7k | Automated safety check: Pass | No licence | 19 days ago |
| 4 | CodexBar read. Provider usage, limits, credits, config health. JSON. No writes. | steipete/ | 22k | — | ~320 | Automated safety check: Pass | MIT | today |
| 5 | Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs. | FailproofAI/ | 5.3k | — | ~6k | Automated safety check: Pass | Unknown | yesterday |
| 6 | Build or review Langfuse backend code. An agent skill from langfuse/langfuse. | langfuse/ | 35k | — | ~1.9k | Automated safety check: Pass | Unknown | today |
| 7 | Navigate Langfuse repositories, code areas, and agent skills. | langfuse/ | 35k | — | ~1.4k | Automated safety check: Pass | Unknown | today |
| 8 | Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior. | JuliusBrussee/ | 110k | 1 repo | ~2.6k | Automated safety check: Warn | Apache-2.0 | today |
| 9 | Refactor avoidable React useEffect usage in Langfuse frontend code. | langfuse/ | 35k | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 10 | Frontend development guidelines for the Phoenix AI observability platform. | Arize-ai/ | 12k | — | ~709 | Automated safety check: Pass | Unknown | today |
| 11 | 11.Langfuse Investigate AI traces, observations, exceptions, latency, sessions, prompts, datasets, annotation queues, and scores through Langfuse MCP. | avivsinai/ | 113 | 1 repo | ~580 | Automated safety check: Pass | MIT | 12 days ago |
| 12 | 12.Databuddy Integrate Databuddy analytics using the SDK, REST API, or MCP. | databuddy-analytics/ | 1.2k | — | ~2.1k | Automated safety check: Pass | AGPL-3.0 | today |
| 13 | Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix. | Arize-ai/ | 12k | — | ~2.2k | Automated safety check: Pass | Unknown | today |
| 14 | Runs a real Codex CLI session through claude-tap and produces trace evidence and viewer screenshots for pull requests that touch capture, proxying or the viewer. | liaohch3/ | 3.3k | — | ~3k | Automated safety check: Pass | MIT | 15 days ago |
| 15 | Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 16 | Run ServiceRadar web-ng locally against the live Kubernetes demo CNPG database for dashboard, SRQL, services, and UI testing. | carverauto/ | 921 | — | ~672 | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Create a new Langfuse integration page in the langfuse-docs repo. | langfuse/ | 246 | — | ~3.7k | Automated safety check: Pass | MIT | today |
| 18 | 18.Langfuse Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications. | langfuse/ | 299 | — | ~2.1k | Automated safety check: Notes | MIT | 6 days ago |
| 19 | Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI). | Arize-ai/ | 12k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 20 | Adds Olakai monitoring to an existing LLM application with minimal code changes, then configures custom KPIs so the dashboard tracks business outcomes instead of just token counts. | andrewyng/ | 14k | — | ~4.5k | Automated safety check: Pass | MIT | 4 mo ago |
| 21 | Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the… | amd/ | 395 | — | ~2.6k | Automated safety check: Pass | MIT | today |
| 22 | 22.Livetable A skill your agent uses when building, modifying, or reviewing Phoenix LiveView tables with LiveTable, including schema-backed tables, context-owned data providers, joined queries, filters… | gurujada/ | 211 | — | ~1.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 23 | Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads. | langchain-ai/ | 159 | — | ~1.4k | Automated safety check: Pass | MIT | 4 days ago |
| 24 | This skill should be used when the user wants to "set up tracing", "monitor my ADK agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring… | pifferologo/ | 129 | 1 repo | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 25 | Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook). | Arize-ai/ | 12k | — | ~1.9k | Automated safety check: Pass | Unknown | today |
| 26 | Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud. | google/ | 21k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | Use this skill working with Ash Framework or any of its extensions. | podlove/ | 136 | — | ~2.1k | Automated safety check: Pass | MIT | 12 days ago |
| 28 | Answers token usage and cost questions from the Claude Command Center throughput dashboard and shares its link, for the last 7 days or one session. | amirfish1/ | 177 | — | ~676 | Automated safety check: Pass | Unknown | today |
| 29 | Queries OmniRoute call logs, usage history and analytics, filters them by provider, model, status or cost, and reads or sets usage budgets. | diegosouzapw/ | 74k | 1 repo | ~2k | Automated safety check: Pass | MIT | today |
| 30 | A skill your agent uses when designing or architecting Elixir/Phoenix applications, creating comprehensive project documentation, planning OTP supervision trees, defining domain models with Ash… | maxim-ist/ | 145 | — | ~7.6k | Automated safety check: Pass | MIT | 10 mo ago |
| 31 | Use this skill working with Phoenix Framework. An agent skill from podlove/radiator. | podlove/ | 136 | — | ~562 | Automated safety check: Pass | MIT | 12 days ago |
| 32 | Full Sentry SDK setup for Elixir. An agent skill from getsentry/sentry-for-ai. | getsentry/ | 268 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 33 | Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs. | amd/ | 1.6k | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 34 | Manage GitHub issues, labels, project boards, sprint operations, and roadmap health for the Arize-ai/phoenix repository. | Arize-ai/ | 12k | — | ~6.5k | Automated safety check: Pass | Apache-2.0 | today |
| 35 | Builds a new AI agent with Olakai monitoring from the start: CLI login, SDK integration, per-agent KPI configuration and an end-to-end check that data flows. | andrewyng/ | 14k | — | ~5k | Automated safety check: Pass | MIT | 4 mo ago |
| 36 | Add or update a company in the Langfuse /users adopters table. | langfuse/ | 246 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 37 | Tune and review Langfuse autoscaling for web, web-iso, and web-ingestion. | langfuse/ | 35k | — | ~4.4k | Automated safety check: Pass | Unknown | today |
| 38 | 38.Elixir A skill your agent uses for Elixir/Phoenix development in this repo: implementing features, refactors, debugging, tests, Ecto changes, and production-safe fixes. | streamband/ | 146 | — | ~804 | Automated safety check: Pass | Apache-2.0 | 21 days ago |
| 39 | Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space. | huggingface/ | 11k | 2 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 40 | Conduz a jornada de onboarding 'do zero ao agente de atendimento' do fazer.ai agents num VPS, escolhendo o orquestrador de deploy (Tier A Coolify, B Portainer, C compose genérico para VM crua ou… | fazer-ai/ | 118 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | today |
| 41 | Migrate to Langfuse from another LLM observability/evals platform (LangSmith, Arize AX, Phoenix, Braintrust, Helicone, Promptfoo, ...). | langfuse/ | 299 | — | ~1.7k | Automated safety check: Notes | MIT | 6 days ago |
| 42 | Queries Langfuse traces, prompts, datasets and sessions, and analyzes local LLM gateway logs for requests, context growth, token use and cache hits. | KonghaYao/ | 223 | — | ~4.3k | Automated safety check: Notes | Apache-2.0 | today |
| 43 | A skill your agent uses when the user wants to analyze agent telemetry traces to find bugs and get fix recommendations — walks through exporting traces from a local or remote watsonx Orchestrate… | IBM/ | 178 | — | ~10k | Automated safety check: Notes | MIT | 4 days ago |
| 44 | Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK. | datadog-labs/ | 177 | — | ~2.3k | Automated safety check: Pass | MIT | 4 days ago |
| 45 | Add a new team member to Langfuse's canonical team data and shared team table. | langfuse/ | 246 | — | ~548 | Automated safety check: Pass | MIT | today |
| 46 | Use the writetodos tool effectively for task planning and decomposition in Deep Agents. | soba-labs/ | 107 | — | ~2.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 47 | Use Langfuse's disposable per-PR previews at pr-N.preview.langfuse.com (synthetic data only). | langfuse/ | 35k | — | ~2.8k | Automated safety check: Notes | Unknown | today |
| 48 | LLM observability platform for tracing, evaluation, and monitoring. | Orchestra-Research/ | 13k | 3 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
Questions, answered from the data.
What is the best LLM observability skill?
Caveman Workflow Labeler from JuliusBrussee/caveman ranks first of the 217 LLM observability skills listed here, with the highest score: its repository has 110k GitHub stars, 1 other GitHub owner carry a copy, its SKILL.md loads about 1.3k tokens and it passes the automated safety check with no findings. Next come Caveman Evidence Review and LangSmith Trace Debugging.
Which LLM observability skills are official?
17 of the 217 LLM observability skills are official, published by the vendor's own GitHub organization: Langsmith Online Eval Engineering, Agent Platform Alert Configuration, Sentry Elixir SDK, Trackio Experiment Tracking, Telemetry Analyzer and 12 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23