Topic · AI & LLM Engineering

Best LLM observability skills for Claude Code, Codex and other agents.

Skills that trace and monitor language-model calls, costs and quality in production.
skills
217
official
17

LLM observability skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM observability skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Finds every LLM workflow in a repository, proposes a labeling table and, once you agree, wires labels so Caveman Cloud groups spend per workflow.

JuliusBrussee/caveman110k1 repo~1.3kAutomated safety check: PassApache-2.0today
2

Read-only review of Caveman Cloud data to explain where LLM spend goes: cost, score, workflows, traces, latency, errors, routing and verified savings.

JuliusBrussee/caveman110k1 repo~927Automated safety check: PassApache-2.0today
3

Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.

ComposioHQ/awesome-claude-skills77k9 repos~2.7kAutomated safety check: PassNo licence19 days ago
4

CodexBar read. Provider usage, limits, credits, config health. JSON. No writes.

steipete/CodexBar22k—~320Automated safety check: PassMITtoday
5

Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

FailproofAI/failproofai5.3k—~6kAutomated safety check: PassUnknownyesterday
6

Build or review Langfuse backend code. An agent skill from langfuse/langfuse.

langfuse/langfuse35k—~1.9kAutomated safety check: PassUnknowntoday
7

Navigate Langfuse repositories, code areas, and agent skills.

langfuse/langfuse35k—~1.4kAutomated safety check: PassUnknowntoday
8

Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

JuliusBrussee/caveman110k1 repo~2.6kAutomated safety check: WarnApache-2.0today
9

Refactor avoidable React useEffect usage in Langfuse frontend code.

langfuse/langfuse35k—~1.7kAutomated safety check: PassUnknowntoday
10

Frontend development guidelines for the Phoenix AI observability platform.

Arize-ai/phoenix12k—~709Automated safety check: PassUnknowntoday
11

Investigate AI traces, observations, exceptions, latency, sessions, prompts, datasets, annotation queues, and scores through Langfuse MCP.

avivsinai/langfuse-mcp1131 repo~580Automated safety check: PassMIT12 days ago
12

Integrate Databuddy analytics using the SDK, REST API, or MCP.

databuddy-analytics/Databuddy1.2k—~2.1kAutomated safety check: PassAGPL-3.0today
13

Write efficient GraphQL queries against the Phoenix API. An agent skill from Arize-ai/phoenix.

Arize-ai/phoenix12k—~2.2kAutomated safety check: PassUnknowntoday
14

Runs a real Codex CLI session through claude-tap and produces trace evidence and viewer screenshots for pull requests that touch capture, proxying or the viewer.

liaohch3/claude-tap3.3k—~3kAutomated safety check: PassMIT15 days ago
15

Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

ai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.012 days ago
16

Run ServiceRadar web-ng locally against the live Kubernetes demo CNPG database for dashboard, SRQL, services, and UI testing.

carverauto/serviceradar921—~672Automated safety check: PassApache-2.0today
17

Create a new Langfuse integration page in the langfuse-docs repo.

langfuse/langfuse-docs246—~3.7kAutomated safety check: PassMITtoday
18

Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications.

langfuse/skills299—~2.1kAutomated safety check: NotesMIT6 days ago
19

Backend development guide for the Phoenix AI observability platform (Strawberry GraphQL, SQLAlchemy async, FastAPI).

Arize-ai/phoenix12k—~1.6kAutomated safety check: PassUnknowntoday
20

Adds Olakai monitoring to an existing LLM application with minimal code changes, then configures custom KPIs so the dashboard tracks business outcomes instead of just token counts.

andrewyng/context-hub14k—~4.5kAutomated safety check: PassMIT4 mo ago
21

Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…

amd/skills395—~2.6kAutomated safety check: PassMITtoday
22

A skill your agent uses when building, modifying, or reviewing Phoenix LiveView tables with LiveTable, including schema-backed tables, context-owned data providers, joined queries, filters…

gurujada/live_table211—~1.9kAutomated safety check: PassMIT3 mo ago
23

Design, test, create, and attach LangSmith online evaluators for production traces or conversation threads.

langchain-ai/langsmith-skills159—~1.4kAutomated safety check: PassMIT4 days ago
24

This skill should be used when the user wants to "set up tracing", "monitor my ADK agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring…

pifferologo/cloud-agents-cli1291 repo~2.5kAutomated safety check: PassApache-2.01 mo ago
25

Conventions for creating, modifying, and reviewing production-faithful Storybook stories in the Phoenix frontend (js/app/stories, js/app/.storybook).

Arize-ai/phoenix12k—~1.9kAutomated safety check: PassUnknowntoday
26

Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

google/skills21k—~4.2kAutomated safety check: PassApache-2.0today
27

Use this skill working with Ash Framework or any of its extensions.

podlove/radiator136—~2.1kAutomated safety check: PassMIT12 days ago
28

Answers token usage and cost questions from the Claude Command Center throughput dashboard and shares its link, for the last 7 days or one session.

amirfish1/claude-command-center177—~676Automated safety check: PassUnknowntoday
29

Queries OmniRoute call logs, usage history and analytics, filters them by provider, model, status or cost, and reads or sets usage budgets.

diegosouzapw/OmniRoute74k1 repo~2kAutomated safety check: PassMITtoday
30

A skill your agent uses when designing or architecting Elixir/Phoenix applications, creating comprehensive project documentation, planning OTP supervision trees, defining domain models with Ash…

maxim-ist/elixir-architect145—~7.6kAutomated safety check: PassMIT10 mo ago
31

Use this skill working with Phoenix Framework. An agent skill from podlove/radiator.

podlove/radiator136—~562Automated safety check: PassMIT12 days ago
32

Full Sentry SDK setup for Elixir. An agent skill from getsentry/sentry-for-ai.

getsentry/sentry-for-ai268—~3.5kAutomated safety check: PassApache-2.0today
33

Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

amd/gaia1.6k—~2.3kAutomated safety check: PassMITtoday
34

Manage GitHub issues, labels, project boards, sprint operations, and roadmap health for the Arize-ai/phoenix repository.

Arize-ai/phoenix12k—~6.5kAutomated safety check: PassApache-2.0today
35

Builds a new AI agent with Olakai monitoring from the start: CLI login, SDK integration, per-agent KPI configuration and an end-to-end check that data flows.

andrewyng/context-hub14k—~5kAutomated safety check: PassMIT4 mo ago
36

Add or update a company in the Langfuse /users adopters table.

langfuse/langfuse-docs246—~2.3kAutomated safety check: PassMITtoday
37

Tune and review Langfuse autoscaling for web, web-iso, and web-ingestion.

langfuse/langfuse35k—~4.4kAutomated safety check: PassUnknowntoday
38

A skill your agent uses for Elixir/Phoenix development in this repo: implementing features, refactors, debugging, tests, Ecto changes, and production-safe fixes.

streamband/hydra-srt146—~804Automated safety check: PassApache-2.021 days ago
39

Logs and visualizes ML training metrics with Trackio, firing alerts for issues like loss spikes, and syncing a live dashboard to a Hugging Face Space.

huggingface/skills11k2 repos~1.3kAutomated safety check: PassApache-2.06 days ago
40

Conduz a jornada de onboarding 'do zero ao agente de atendimento' do fazer.ai agents num VPS, escolhendo o orquestrador de deploy (Tier A Coolify, B Portainer, C compose genérico para VM crua ou…

fazer-ai/agents118—~4.4kAutomated safety check: PassApache-2.0today
41

Migrate to Langfuse from another LLM observability/evals platform (LangSmith, Arize AX, Phoenix, Braintrust, Helicone, Promptfoo, ...).

langfuse/skills299—~1.7kAutomated safety check: NotesMIT6 days ago
42

Queries Langfuse traces, prompts, datasets and sessions, and analyzes local LLM gateway logs for requests, context growth, token use and cache hits.

KonghaYao/peri223—~4.3kAutomated safety check: NotesApache-2.0today
43

A skill your agent uses when the user wants to analyze agent telemetry traces to find bugs and get fix recommendations — walks through exporting traces from a local or remote watsonx Orchestrate…

IBM/ibm-watsonx-orchestrate-adk178—~10kAutomated safety check: NotesMIT4 days ago
44

Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.

datadog-labs/agent-skills177—~2.3kAutomated safety check: PassMIT4 days ago
45

Add a new team member to Langfuse's canonical team data and shared team table.

langfuse/langfuse-docs246—~548Automated safety check: PassMITtoday
46

Use the writetodos tool effectively for task planning and decomposition in Deep Agents.

soba-labs/langchain-agent-skills107—~2.3kAutomated safety check: PassMIT1 mo ago
47

Use Langfuse's disposable per-PR previews at pr-N.preview.langfuse.com (synthetic data only).

langfuse/langfuse35k—~2.8kAutomated safety check: NotesUnknowntoday
48

LLM observability platform for tracing, evaluation, and monitoring.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.4kAutomated safety check: PassMIT3 mo ago

Questions, answered from the data.

What is the best LLM observability skill?

Caveman Workflow Labeler from JuliusBrussee/caveman ranks first of the 217 LLM observability skills listed here, with the highest score: its repository has 110k GitHub stars, 1 other GitHub owner carry a copy, its SKILL.md loads about 1.3k tokens and it passes the automated safety check with no findings. Next come Caveman Evidence Review and LangSmith Trace Debugging.

Which LLM observability skills are official?

17 of the 217 LLM observability skills are official, published by the vendor's own GitHub organization: Langsmith Online Eval Engineering, Agent Platform Alert Configuration, Sentry Elixir SDK, Trackio Experiment Tracking, Telemetry Analyzer and 12 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.