Agent skill

Dspy Debugging Observability

by OmidZamani in OmidZamani/dspy-skills

A skill your agent uses for debugging DSPy programs, inspecthistory, tracing LLM calls, custom callbacks, observability, monitoring, and cost tracking.

MITAuto-check: warningsDevOps & Cloud

Install Dspy Debugging Observability

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add OmidZamani/dspy-skills --skill dspy-debugging-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OmidZamani/dspy-skills dspy-debugging-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OmidZamani/dspy-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dspy-debugging-observability .claude/skills/dspy-debugging-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dspy-debugging-observability
GitHub stars
124
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
244 words
Files
2
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses for debugging DSPy programs, inspecthistory, tracing LLM calls, custom callbacks, observability, monitoring, and cost tracking.

  • Works in 4 steps: Basic Inspection with inspect_history() → MLflow Tracing → Custom Callbacks for Production → …
  • Debugging DSPy programs
  • SKILL.md covers Goal, When to Use, Related Skills and Inputs, plus 5 more sections
  • Runs Python scripts from its folder

What it does

Dspy Debugging Observability is an agent skill from OmidZamani/dspy-skills. Use for debugging DSPy programs, inspecthistory, tracing LLM calls, custom callbacks, observability, monitoring, and cost tracking.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `example.py`).

It sits in DevOps & Cloud, covering Observability, Debugging and LLM cost and token optimization. It works with MLflow. The repository describes itself as: Collection of Claude Skills for DSPy framework - program language models, optimize prompts, and build RAG pipelines systematically. The licence is MIT.

When your agent uses it

  • Debugging DSPy programs
  • Tracing LLM calls
  • Custom callbacks

Example prompts

  • “/dspy-debugging-observability”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Glob, Grep

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Basic Inspection with inspect_history()
  2. MLflow Tracing
  3. Custom Callbacks for Production
  4. Sampling for High-Volume Production

What it can do on your machine

Read from SKILL.md and the folder at commit f5db3b7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Glob
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • dspy.ai
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dspy Debugging Observability loads about 2.1k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 244 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningLinks to a raw public IP addressSKILL.md:100
    rm = dspy.ColBERTv2(url="http://20.102.90.50:2017/wiki17_abstracts")

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OmidZamani/dspy-skills at commit f5db3b7, republished under its MIT licence (© OmidZamani). 244 words, ~2,083 tokens.

Download SKILL.mdSave it as .claude/skills/dspy-debugging-observability/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
dspy-debugging-observability
description
Use for debugging DSPy programs, inspect_history, tracing LLM calls, custom callbacks, observability, monitoring, and cost tracking.
allowed-tools
Read, Write, Glob, Grep
version
1.0.0
dspy-compatibility
3.2.1
tags
production

DSPy Debugging & Observability

Goal

Debug, trace, and monitor DSPy programs using built-in inspection, MLflow tracing, and custom callbacks for production observability.

When to Use

  • Debugging unexpected outputs
  • Understanding multi-step program flow
  • Production monitoring (cost, latency, errors)
  • Analyzing optimizer behavior
  • Tracking LLM API usage

Inputs

InputTypeDescription
programdspy.ModuleProgram to debug/monitor
callbackBaseCallbackOptional custom callback (subclass of dspy.utils.callback.BaseCallback)

Outputs

OutputTypeDescription
GLOBAL_HISTORYlist[dict]Raw execution trace from dspy.clients.base_lm
metricsdictCost, latency, token counts from callbacks

Workflow

Phase 1: Basic Inspection with inspect_history()

The simplest debugging approach:

python
import dspy

dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))

# Run program
qa = dspy.ChainOfThought("question -> answer")
result = qa(question="What is the capital of France?")

# Inspect last execution (prints to console)
dspy.inspect_history(n=1)

# To access raw history programmatically:
from dspy.clients.base_lm import GLOBAL_HISTORY
for entry in GLOBAL_HISTORY[-1:]:
    print(f"Model: {entry['model']}")
    print(f"Usage: {entry.get('usage', {})}")
    print(f"Cost: {entry.get('cost', 0)}")
Phase 2: MLflow Tracing

MLflow integration requires explicit setup:

python
import dspy
import mlflow

# Setup MLflow (4 steps required)
# 1. Set tracking URI and experiment
mlflow.set_tracking_uri("http://localhost:5000")
mlflow.set_experiment("DSPy")

# 2. Enable DSPy autologging
mlflow.dspy.autolog(
    log_traces=True,              # Log traces during inference
    log_traces_from_compile=True, # Log traces when compiling/optimizing
    log_traces_from_eval=True,    # Log traces during evaluation
    log_compiles=True,            # Log optimization process info
    log_evals=True                # Log evaluation call info
)

dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))

# Configure retriever (required before using dspy.Retrieve)
rm = dspy.ColBERTv2(url="http://20.102.90.50:2017/wiki17_abstracts")
dspy.configure(rm=rm)

class RAGPipeline(dspy.Module):
    def __init__(self):
        self.retrieve = dspy.Retrieve(k=3)
        self.generate = dspy.ChainOfThought("context, question -> answer")

    def forward(self, question):
        context = self.retrieve(question).passages
        return self.generate(context=context, question=question)

pipeline = RAGPipeline()
result = pipeline(question="What is machine learning?")

# View traces in MLflow UI (run in terminal): mlflow ui --port 5000

MLflow captures LLM calls, token usage, costs, and execution times when autolog is enabled.

Phase 3: Custom Callbacks for Production

Build custom callbacks for specialized monitoring:

python
import dspy
from dspy.utils.callback import BaseCallback
import logging
import time
from typing import Any

logger = logging.getLogger(__name__)

class ProductionMonitoringCallback(BaseCallback):
    """Track cost, latency, and errors in production."""

    def __init__(self):
        super().__init__()
        self.total_cost = 0.0
        self.total_tokens = 0
        self.call_count = 0
        self.errors = []
        self.start_times = {}

    def on_lm_start(self, call_id: str, instance: Any, inputs: dict[str, Any]):
        """Called when LM is invoked."""
        self.start_times[call_id] = time.time()

    def on_lm_end(self, call_id: str, outputs: dict[str, Any] | None, exception: Exception | None = None):
        """Called after LM finishes."""
        if exception:
            self.errors.append(str(exception))
            logger.error(f"LLM error: {exception}")
            return

        # Calculate latency
        start = self.start_times.pop(call_id, time.time())
        latency = time.time() - start

        # Extract usage from outputs
        usage = outputs.get('usage', {}) if isinstance(outputs, dict) else {}
        tokens = usage.get('total_tokens', 0)
        model = outputs.get('model', 'unknown') if isinstance(outputs, dict) else 'unknown'
        cost = self._estimate_cost(model, usage)

        self.total_tokens += tokens
        self.total_cost += cost
        self.call_count += 1

        logger.info(f"LLM call: {latency:.2f}s, {tokens} tokens, ${cost:.4f}")

    def _estimate_cost(self, model: str, usage: dict[str, int]) -> float:
        """Estimate cost based on model pricing (update rates for 2026)."""
        pricing = {
            'gpt-4o-mini': {'input': 0.00015 / 1000, 'output': 0.0006 / 1000},
            'gpt-4o': {'input': 0.0025 / 1000, 'output': 0.01 / 1000},
        }
        model_key = next((k for k in pricing if k in model), 'gpt-4o-mini')
        input_cost = usage.get('prompt_tokens', 0) * pricing[model_key]['input']
        output_cost = usage.get('completion_tokens', 0) * pricing[model_key]['output']
        return input_cost + output_cost

    def get_metrics(self) -> dict[str, Any]:
        """Return aggregated metrics."""
        return {
            'total_cost': self.total_cost,
            'total_tokens': self.total_tokens,
            'call_count': self.call_count,
            'avg_cost_per_call': self.total_cost / max(self.call_count, 1),
            'error_count': len(self.errors)
        }

# Usage
monitor = ProductionMonitoringCallback()
dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"), callbacks=[monitor])

# Run your program
qa = dspy.ChainOfThought("question -> answer")
for question in questions:
    result = qa(question=question)

# Get metrics
metrics = monitor.get_metrics()
print(f"Total cost: ${metrics['total_cost']:.2f}")
print(f"Total calls: {metrics['call_count']}")
print(f"Errors: {metrics['error_count']}")
Phase 4: Sampling for High-Volume Production

For high-traffic applications, sample traces to reduce overhead:

python
import random
from dspy.utils.callback import BaseCallback
from typing import Any

class SamplingCallback(BaseCallback):
    """Sample 10% of traces."""

    def __init__(self, sample_rate: float = 0.1):
        super().__init__()
        self.sample_rate = sample_rate
        self.sampled_calls = []

    def on_lm_end(self, call_id: str, outputs: dict[str, Any] | None, exception: Exception | None = None):
        """Sample a subset of LM calls."""
        if random.random() < self.sample_rate:
            self.sampled_calls.append({
                'call_id': call_id,
                'outputs': outputs,
                'exception': exception
            })

# Use with high-volume apps
callback = SamplingCallback(sample_rate=0.1)
dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"), callbacks=[callback])

Best Practices

  1. Use inspect_history() for debugging - Quick inspection during development
  2. MLflow for comprehensive tracing - Automatic instrumentation in production
  3. Sample high-volume traces - Reduce overhead with 1-10% sampling
  4. Privacy-aware logging - Redact PII before logging
  5. Async callbacks - Non-blocking callbacks for production

Limitations

  • Callbacks are synchronous by default (can block LLM calls)
  • MLflow tracing adds ~5-10ms overhead per call
  • inspect_history() only stores recent calls (last 100 by default)
  • Custom callbacks don't capture internal optimizer steps
  • Cost estimation requires manual pricing table updates

Official Documentation

© OmidZamani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/dspy-debugging-observability of OmidZamani/dspy-skills.

  • SKILL.md
  • example.py

Open the folder on GitHubat commit f5db3b7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in OmidZamani/dspy-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Dspy Debugging Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dspy Debugging Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dspy Debugging Observability this skillOmidZamani/dspy-skills1241 repos~2.1kAutomated safety check: WarnMIT
Motel Debugkitlangton/motel298—~2.2kAutomated safety check: PassMIT
Codex Session Debuggingweave-os/router5.6k—~4.5kAutomated safety check: WarnApache-2.0
Log Aggregationaspectrr/deer405—~1.4kAutomated safety check: PassMIT
Gcloud Usagefcakyon/claude-codex-settings1.2k—~871Automated safety check: PassApache-2.0
Caveman Gateway SetupJuliusBrussee/caveman110k1 repos~2.6kAutomated safety check: WarnApache-2.0

Similar skills

  • Motel Debug

    kitlangton/motel

    Debug applications with motel, a local OpenTelemetry ingest and query server.

    298 GitHub stars~2.2k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Correlates a Codex CLI session's local transcript with a model router's production logs to explain why a reply rendered the way it did.

    5.6k GitHub stars~4.5k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • Log Aggregation

    aspectrr/deer

    ELK Stack deployment, Logstash pipeline building, Filebeat configuration, and Kibana dashboard setup.

    405 GitHub stars~1.4k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Gcloud Usage

    fcakyon/claude-codex-settings

    This skill should be used when user asks about "GCloud logs", "Cloud Logging queries", "Google Cloud metrics", "GCP observability", "trace analysis", or "debugging production issues on GCP".

    1.2k GitHub stars~871 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Caveman Gateway Setup

    JuliusBrussee/caveman

    Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

    110k GitHub starsUsed in 1 repo~2.6k tokens
    DevOps & CloudAuto-check: warnings
  • Agent Kill Switch

    vivekchand/clawmetry

    Give the human an off switch and a cost meter for the coding agents on this machine, using ClawMetry.

    425 GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed

More from OmidZamani/dspy-skills

All 17 skills in this repo
  • Dspy Haystack Integration

    OmidZamani/dspy-skills

    A skill your agent uses for integrating DSPy with Haystack, optimizing Haystack prompts, improving retrieval pipelines, and extracting DSPy prompts.

    124 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Skill Perfection

    OmidZamani/dspy-skills

    A skill your agent uses when you need to QA audit and fix a plugin skill file.

    124 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Dspy Advanced Module Composition

    OmidZamani/dspy-skills

    A skill your agent uses for composing DSPy modules with Ensemble, MultiChainComparison, ensemble voting, sequential pipelines, and multi-program workflows.

    124 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Dspy Better Together

    OmidZamani/dspy-skills

    A skill your agent uses for BetterTogether, prompt plus weight optimization, fine-tuning sequences, and strategy chains like p - w - p.

    124 GitHub starsUsed in 1 repo~756 tokens
    Auto-check passed
  • Dspy Bootstrap Fewshot

    OmidZamani/dspy-skills

    A skill your agent uses for BootstrapFewShot, bootstrapped demonstrations, teacher-model demos, and low-data DSPy prompt optimization.

    124 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Dspy Custom Module Design

    OmidZamani/dspy-skills

    A skill your agent uses for creating custom DSPy modules, extending dspy.Module, reusable components, stateful modules, serialization, and module testing.

    124 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed

Works with

Questions about Dspy Debugging Observability

What does Dspy Debugging Observability do?

A skill your agent uses for debugging DSPy programs, inspecthistory, tracing LLM calls, custom callbacks, observability, monitoring, and cost tracking. Dspy Debugging Observability is an agent skill from OmidZamani/dspy-skills. Use for debugging DSPy programs, inspecthistory, tracing LLM calls, custom callbacks, observability, monitoring, and cost tracking.

When should I use Dspy Debugging Observability?

Dspy Debugging Observability fits situations like: debugging DSPy programs; tracing LLM calls; custom callbacks.

How do I install Dspy Debugging Observability in Claude Code?

Run `npx skills add OmidZamani/dspy-skills --skill dspy-debugging-observability -a claude-code`. Or copy the skill folder (skills/dspy-debugging-observability in OmidZamani/dspy-skills) into .claude/skills/dspy-debugging-observability in your project. Claude Code loads it when a task matches its description.

How do I install Dspy Debugging Observability in Codex?

Run `npx skills add OmidZamani/dspy-skills --skill dspy-debugging-observability -a codex`. Or copy the skill folder (skills/dspy-debugging-observability in OmidZamani/dspy-skills) into .agents/skills/dspy-debugging-observability in your project. Codex loads it when a task matches its description.

Can I use Dspy Debugging Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OmidZamani/dspy-skills --skill dspy-debugging-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dspy-debugging-observability, .gemini/skills/dspy-debugging-observability, .github/skills/dspy-debugging-observability and .opencode/skills/dspy-debugging-observability in your project.

What does Dspy Debugging Observability need to run?

Going by SKILL.md and its folder, Dspy Debugging Observability needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Glob, Grep.

Does Dspy Debugging Observability access the network?

SKILL.md names 2 domains. As links in the text: dspy.ai and github.com. This is read from the text; nothing was executed.

Is Dspy Debugging Observability safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): links to a raw public ip address. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Dspy Debugging Observability use?

Dspy Debugging Observability is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dspy Debugging Observability use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dspy Debugging Observability?

Skills that share tags, products or a category with Dspy Debugging Observability: Motel Debug (kitlangton/motel, 298 stars), Codex Session Debugging (weave-os/router, 5.6k stars), Log Aggregation (aspectrr/deer, 405 stars) and Gcloud Usage (fcakyon/claude-codex-settings, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dspy Debugging Observability?

OmidZamani (a GitHub user) maintains it in OmidZamani/dspy-skills, which has 124 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on June 23, 2026.

Source: OmidZamani/dspy-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.