Agent skill

Monitoring Observability

by yonatangross in yonatangross/orchestkit

Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection.

MITAuto-check passedDevOps & Cloud

Install Monitoring Observability

skills CLI
$ npx skills add yonatangross/orchestkit --skill monitoring-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit monitoring-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/monitoring-observability .claude/skills/monitoring-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
monitoring-observability
GitHub stars
289
Token cost
~2.2k tokens
SKILL.md length
678 words
Files
28 (incl. scripts, references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection.

  • Distributed tracing
  • SKILL.md covers Upstream coverage (do not…, Quick Reference, Quick Start and Infrastructure Monitoring, plus 5 more sections
  • LLM cost tracking
  • Quality drift monitoring

What it does

Monitoring Observability is an agent skill from yonatangross/orchestkit. Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality drift monitoring.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 30 other files, including scripts and reference files (for example `examples/orchestkit-monitoring-dashboard.md`, `metadata.json` and `references/dashboards.md`). Compatibility notes: Claude Code 2.1.277+.

It sits in DevOps & Cloud, covering Monitoring and alerting, Observability and LLM observability. It works with Prometheus, Langfuse and Grafana. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • Distributed tracing
  • LLM cost tracking
  • Quality drift monitoring

Example prompts

  • “/monitoring-observability”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Claude Code 2.1.277+.
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, WebFetch, WebSearch

What it can do on your machine

Read from SKILL.md and the folder at commit 0ef71d2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • WebFetch
    • WebSearch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • langfuse.com
    • prometheus.io
    • grafana.com
    • opentelemetry.io
    • evidentlyai.com
    • itl.nist.gov
    • structlog.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+.

    From compatibility in the SKILL.md frontmatter.

Context cost

Monitoring Observability loads about 2.2k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 678 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~16k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit 0ef71d2, republished under its MIT licence (© yonatangross). 678 words, ~2,222 tokens.

Download SKILL.mdSave it as .claude/skills/monitoring-observability/SKILL.md (or your agent's skills folder). This skill also uses 27 other files; get the full folder from GitHub.
name
monitoring-observability
description
Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (as_type, score_current_span, should_export_span, LangfuseMedia), and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality drift monitoring.
allowed-tools
Read, Glob, Grep, WebFetch, WebSearch
compatibility
Claude Code 2.1.277+.
license
MIT
user-invocable
false
disable-model-invocation
true
upstream-version-tested
4.16.0
metadata.category
document-asset-creation
metadata.version
3.0.0
metadata.author
OrchestKit
metadata.complexity
medium
metadata.tags
monitoring, observability, prometheus, grafana, langfuse, tracing, metrics, drift-detection, logging
path_patterns
**/metrics/**, **/tracing/**, prometheus.*, grafana/**

Monitoring & Observability

A wrap around Prometheus, Grafana, OpenTelemetry and Langfuse, not a re-teaching of them. This skill carries OrchestKit's delta (version floors, house decisions, scars) and points at the vendor for everything else. Start at references/ork-delta.md.

Upstream coverage (do not restate)

These topics are fully covered first-party. Read the source, do not add a local copy.

TopicFirst-party source
Prometheus metric types, RED method, cardinality, PromQLhttps://prometheus.io/docs/practices/
Alertmanager grouping, inhibition, escalation, runbookshttps://prometheus.io/docs/alerting/latest/configuration/
Grafana dashboards, Loki and LogQL, Promtailhttps://grafana.com/docs/
OpenTelemetry spans, sampling, context propagationhttps://opentelemetry.io/docs/
Langfuse Python SDK (@observe, as_type, score_current_span, should_export_span, LangfuseMedia)https://langfuse.com/docs/sdk/python
Langfuse v2 to v4 Python and v3 to v5 JS migration pathshttps://langfuse.com/docs/sdk/python/v4-migration
Langfuse self-hosting (ClickHouse, Redis, S3, Helm)https://langfuse.com/docs/deployment/self-host
Langfuse cost tracking, model pricing, Metrics API v2https://langfuse.com/docs/model-usage-and-cost
Langfuse scores, online evaluators, annotation queues, prompt managementhttps://langfuse.com/docs/scores/overview
Langfuse framework integrations (LangChain, LangGraph, CrewAI, Pydantic AI, Bedrock, LiveKit)https://langfuse.com/docs/integrations
Agent Graphs, observation types, rendered tool callshttps://langfuse.com/docs/tracing-features/agent-graphs
PSI, KS test, KL and JS divergence, Wasserstein, embedding drifthttps://www.evidentlyai.com/blog/data-drift-detection-large-datasets
EWMA control chartshttps://www.itl.nist.gov/div898/handbook/pmc/section3/pmc324.htm
structlog, Winston, correlation IDs, log samplinghttps://www.structlog.org/en/stable/

Quick Reference

CategoryRulesImpactWhen to Use
Infrastructure Monitoring1CRITICALGrafana dashboards, Golden Signals, SLO/SLI
LLM Observability1HIGHLangfuse tracing, observation types, agent graphs
Silent Failures3HIGHTool skipping, quality degradation, loop/token spike alerting

Total: 5 rules across 3 categories. Drift detection, cost tracking, eval scoring, Prometheus instrumentation and alert-rule authoring moved to the upstream sources listed above.

Quick Start

python
# Langfuse v4 LLM tracing: semantic as_type plus inline scoring
from langfuse import observe, get_client

@observe(as_type="generation", name="analyze_content")
async def analyze_content(content: str):
    get_client().update_current_trace(
        user_id="user_123", session_id="session_abc",
        tags=["production", "orchestkit"],
    )
    result = await llm.generate(content)
    get_client().score_current_span(name="response_quality", value=0.85)
    return result
python
# Prometheus RED method, wired the way this repo expects (bounded labels only)
from prometheus_client import Counter, Histogram

http_requests = Counter('http_requests_total', 'Total requests', ['method', 'endpoint', 'status'])
http_duration = Histogram('http_request_duration_seconds', 'Request latency',
    buckets=[0.01, 0.05, 0.1, 0.5, 1, 2, 5])

Infrastructure Monitoring

Dashboard and health-check patterns. Metric instrumentation and alert-rule syntax are upstream.

RuleFileKey Pattern
Grafana Dashboardsrules/monitoring-grafana.mdGolden Signals, SLO/SLI, health checks

CC 2.1.161 — OTEL resource attributes as metric labels: OTEL_RESOURCE_ATTRIBUTES values are now attached as labels on metric datapoints, so usage metrics can be sliced by custom dimensions (team, repo, environment). Add label selectors to dashboards for multi-tenant / per-team cost and usage tracking.

LLM Observability

Langfuse-based tracing for LLM applications. Cost tracking, scoring and drift statistics are upstream; what stays here is how this repo wires traces.

RuleFileKey Pattern
Langfuse Tracesrules/llm-langfuse-traces.md@observe decorator, OTEL spans, agent graphs

Silent Failures

Detection and alerting for silent failures in LLM agents.

RuleFileKey Pattern
Tool Skippingrules/silent-tool-skipping.mdExpected vs actual tool calls, Langfuse traces
Quality Degradationrules/silent-degraded-quality.mdHeuristics + LLM-as-judge, z-score baselines
Silent Alertingrules/silent-alerting.mdLoop detection, token spikes, escalation workflow

CC 2.1.169 — OTEL client-cert paths require trust: untrusted project settings can no longer set OTEL client-certificate paths without a trust confirmation. If your OTEL exporter uses client certs configured in project .claude/settings.json, expect a one-time trust prompt on first use in an untrusted project — telemetry silently not flowing after 2.1.169 is usually this gate, not the collector.

Show full SKILL.md (238 more words)Show less

Key Decisions

DecisionRecommendationRationale
Metric methodologyRED method (Rate, Errors, Duration)Industry standard, covers essential service health
Log formatStructured JSONMachine-parseable, supports log aggregation
TracingOpenTelemetryVendor-neutral, auto-instrumentation, broad ecosystem
LLM observabilityLangfuse (not LangSmith)Open-source, self-hosted, built-in prompt management
LLM tracing API@observe(as_type=...) + score_current_span()v4: semantic types, inline scoring, span filtering
Langfuse APIsObservations API v2 + Metrics API v2v4 (Mar 2026): faster querying, aggregations at scale
Hook telemetry transportJSONL under ~/.claude/analytics/, never an SDK in-processHooks are per-event processes; SDK init would be paid on every spawn (references/ork-delta.md)

Detailed Documentation

ResourceDescription
references/ork-delta.mdStart here. Floors, house decisions and scars that upstream docs do not carry
references/langfuse-js-v5.mdJS/TS SDK v5 delta from Python 4.x: package map, phantom packages, SpanProcessor vs exporter. Read before writing any JS Langfuse code
references/experiments-api.mdLangfuse experiments and dataset runs as this repo uses them
references/evaluation-scores.mdScore shapes and scoring pipeline wiring
references/session-tracking.mdSession and user grouping across multi-step workflows
references/metrics-collection.mdClaude Code OTEL metric inventory and collector-side joins
references/dashboards.mdDashboard layout conventions
references/structured-logging.mdStructured log field conventions
references/dev-agent-lens.mdLiteLLM proxy layer for API-boundary observability
examples/orchestkit-monitoring-dashboard.mdWorked monitoring dashboard example
scripts/Templates: Prometheus, OpenTelemetry, health checks, Langfuse
  • defense-in-depth - Layer 8 observability as part of security architecture
  • devops-deployment - Observability integration with CI/CD and Kubernetes
  • resilience-patterns - Monitoring circuit breakers and failure scenarios
  • llm-evaluation - Evaluation patterns that integrate with Langfuse scoring
  • caching - Caching strategies that reduce costs tracked by Langfuse

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 27 other files (scripts, references) in src/skills/monitoring-observability of yonatangross/orchestkit.

  • SKILL.md
  • examples/orchestkit-monitoring-dashboard.md
  • metadata.json
  • references/dashboards.md
  • references/dev-agent-lens.md
  • references/evaluation-scores.md
  • references/experiments-api.md
  • references/langfuse-js-v5.md
  • references/metrics-collection.md
  • references/ork-delta.md
  • references/session-tracking.md
  • references/structured-logging.md
  • rules/_sections.md
  • rules/_template.md
  • rules/llm-langfuse-traces.md
  • rules/monitoring-grafana.md
  • rules/silent-alerting.md
  • rules/silent-degraded-quality.md
  • … and 10 more

Open the folder on GitHubat commit 0ef71d2

Compare with similar skills

Monitoring Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Monitoring Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Monitoring Observability this skillyonatangross/orchestkit289—~2.2kAutomated safety check: PassMIT
Langfuse Observabilityjeremylongshore/tons-of-skills-marketplace2.8k—~2.2kAutomated safety check: PassMIT
Ag2 Telemetryag2ai/build-with-ag2252—~1.9kAutomated safety check: PassApache-2.0
Cost Exportruvnet/ruflo74k1 repos~687Automated safety check: NotesMIT
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence
Frontmcp Observabilityagentfront/frontmcp146—~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Langfuse Observability

    jeremylongshore/tons-of-skills-marketplace

    Set up comprehensive observability for Langfuse with metrics, dashboards, and alerts.

    2.8k GitHub stars~2.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Ag2 Telemetry

    ag2ai/build-with-ag2

    Add OpenTelemetry traces to an AG2 beta Agent via TelemetryMiddleware (autogen.beta.middleware.builtin).

    252 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Cost Export

    ruvnet/ruflo

    Export cost-tracking telemetry in Prometheus textfile or webhook JSON formats — for external observability (Grafana, Datadog, custom dashboards)

    74k GitHub starsUsed in 1 repo~687 tokens
    DevOps & CloudAuto-check: notes
  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Monitoring Observability

    ahmedasmar/devops-claude-skills

    Monitoring and observability strategy, implementation, and troubleshooting.

    203 GitHub stars~3.9k tokensUpdated 6 mo ago
    DevOps & CloudAuto-check passed

More from yonatangross/orchestkit

All 108 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    289 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    289 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    289 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    289 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    289 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    289 GitHub stars~3.9k tokensUpdated yesterday
    Auto-check: notes

Categories

Questions about Monitoring Observability

What does Monitoring Observability do?

Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection. Monitoring Observability is an agent skill from yonatangross/orchestkit. Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection.

When should I use Monitoring Observability?

Monitoring Observability fits situations like: distributed tracing; LLM cost tracking; quality drift monitoring.

How do I install Monitoring Observability in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill monitoring-observability -a claude-code`. Or copy the skill folder (src/skills/monitoring-observability in yonatangross/orchestkit) into .claude/skills/monitoring-observability in your project. Claude Code loads it when a task matches its description.

How do I install Monitoring Observability in Codex?

Run `npx skills add yonatangross/orchestkit --skill monitoring-observability -a codex`. Or copy the skill folder (src/skills/monitoring-observability in yonatangross/orchestkit) into .agents/skills/monitoring-observability in your project. Codex loads it when a task matches its description.

Can I use Monitoring Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill monitoring-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/monitoring-observability, .gemini/skills/monitoring-observability, .github/skills/monitoring-observability and .opencode/skills/monitoring-observability in your project.

What does Monitoring Observability need to run?

SKILL.md names no scripts, command-line tools or credentials: Monitoring Observability is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..

Does Monitoring Observability access the network?

SKILL.md names 7 domains. As links in the text: langfuse.com, prometheus.io, grafana.com, opentelemetry.io, evidentlyai.com, itl.nist.gov and structlog.org. This is read from the text; nothing was executed.

Is Monitoring Observability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Monitoring Observability use?

Monitoring Observability is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Monitoring Observability use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Monitoring Observability?

Skills that share tags, products or a category with Monitoring Observability: Langfuse Observability (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Ag2 Telemetry (ag2ai/build-with-ag2, 252 stars), Cost Export (ruvnet/ruflo, 74k stars) and Archestra Dev Observability (archestra-ai/archestra, 4.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Monitoring Observability?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 289 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.