Agent skill

Dspy Evaluation Suite

by OmidZamani in OmidZamani/dspy-skills

A skill your agent uses for evaluating DSPy programs with Evaluate, answerexactmatch, SemanticF1, custom metrics, baselines, and program comparisons.

MITAuto-check passed

Install Dspy Evaluation Suite

skills CLI
$ npx skills add OmidZamani/dspy-skills --skill dspy-evaluation-suite -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OmidZamani/dspy-skills dspy-evaluation-suite --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OmidZamani/dspy-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dspy-evaluation-suite .claude/skills/dspy-evaluation-suite && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dspy-evaluation-suite
GitHub stars
123
Token cost
~2k tokens
SKILL.md length
178 words
Files
2
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses for evaluating DSPy programs with Evaluate, answerexactmatch, SemanticF1, custom metrics, baselines, and program comparisons.

  • Works in 2 steps: Setup Evaluator → Run Evaluation
  • Evaluating DSPy programs with Evaluate
  • SKILL.md covers Goal, When to Use, Related Skills and Inputs, plus 8 more sections
  • Runs Python scripts from its folder

What it does

Dspy Evaluation Suite is an agent skill from OmidZamani/dspy-skills. Use for evaluating DSPy programs with Evaluate, answerexactmatch, SemanticF1, custom metrics, baselines, and program comparisons.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `example.py`).

The repository describes itself as: Collection of Claude Skills for DSPy framework - program language models, optimize prompts, and build RAG pipelines systematically. The licence is MIT.

When your agent uses it

  • Evaluating DSPy programs with Evaluate
  • Answerexactmatch
  • Program comparisons

Example prompts

  • “/dspy-evaluation-suite”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Glob, Grep

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Setup Evaluator
  2. Run Evaluation

What it can do on your machine

Read from SKILL.md and the folder at commit f5db3b7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Glob
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • dspy.ai
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dspy Evaluation Suite loads about 2k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 178 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OmidZamani/dspy-skills at commit f5db3b7, republished under its MIT licence (© OmidZamani). 178 words, ~1,969 tokens.

Download SKILL.mdSave it as .claude/skills/dspy-evaluation-suite/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
dspy-evaluation-suite
description
Use for evaluating DSPy programs with Evaluate, answer_exact_match, SemanticF1, custom metrics, baselines, and program comparisons.
allowed-tools
Read, Write, Glob, Grep
version
1.0.0
dspy-compatibility
3.2.1
tags
evaluation

DSPy Evaluation Suite

Goal

Systematically evaluate DSPy programs using built-in and custom metrics with parallel execution.

When to Use

  • Measuring program performance before/after optimization
  • Comparing different program variants
  • Establishing baselines
  • Validating production readiness

Inputs

InputTypeDescription
programdspy.ModuleProgram to evaluate
devsetlist[dspy.Example]Evaluation examples
metriccallableScoring function
num_threadsintParallel threads

Outputs

OutputTypeDescription
scorefloatAverage metric score
resultslistPer-example results

Workflow

Phase 1: Setup Evaluator
python
from dspy.evaluate import Evaluate

evaluator = Evaluate(
    devset=devset,
    metric=my_metric,
    num_threads=8,
    display_progress=True
)
Phase 2: Run Evaluation
python
result = evaluator(my_program)
print(f"Score: {result.score:.2f}%")
# Access individual results: (example, prediction, score) tuples
for example, pred, score in result.results[:3]:
    print(f"Example: {example.question[:50]}... Score: {score}")

Built-in Metrics

answer_exact_match
python
import dspy

# Normalized, case-insensitive comparison
metric = dspy.evaluate.answer_exact_match
SemanticF1

LLM-based semantic evaluation:

python
from dspy.evaluate import SemanticF1

semantic = SemanticF1()
score = semantic(example, prediction)

Custom Metrics

Basic Metric
python
def exact_match(example, pred, trace=None):
    """Returns bool, int, or float."""
    return example.answer.lower().strip() == pred.answer.lower().strip()
Multi-Factor Metric
python
def quality_metric(example, pred, trace=None):
    """Score based on multiple factors."""
    score = 0.0
    
    # Correctness (50%)
    if example.answer.lower() in pred.answer.lower():
        score += 0.5
    
    # Conciseness (25%)
    if len(pred.answer.split()) <= 20:
        score += 0.25
    
    # Has reasoning (25%)
    if hasattr(pred, 'reasoning') and pred.reasoning:
        score += 0.25
    
    return score
GEPA-Compatible Metric
python
def feedback_metric(example, pred, trace=None, pred_name=None, pred_trace=None):
    """Return a GEPA-compatible score and textual feedback."""
    correct = example.answer.lower() in pred.answer.lower()
    
    if correct:
        return dspy.Prediction(score=1.0, feedback="Correct answer provided.")
    else:
        return dspy.Prediction(
            score=0.0,
            feedback=f"Expected '{example.answer}', got '{pred.answer}'"
        )

Production Example

python
import dspy
from dspy.evaluate import Evaluate, SemanticF1
import json
import logging
from typing import Optional
from dataclasses import dataclass

logger = logging.getLogger(__name__)

@dataclass
class EvaluationResult:
    score: float
    num_examples: int
    correct: int
    incorrect: int
    errors: int

def comprehensive_metric(example, pred, trace=None) -> float:
    """Multi-dimensional evaluation metric."""
    scores = []
    
    # 1. Correctness
    if hasattr(example, 'answer') and hasattr(pred, 'answer'):
        correct = example.answer.lower().strip() in pred.answer.lower().strip()
        scores.append(1.0 if correct else 0.0)
    
    # 2. Completeness (answer not empty or error)
    if hasattr(pred, 'answer'):
        complete = len(pred.answer.strip()) > 0 and "error" not in pred.answer.lower()
        scores.append(1.0 if complete else 0.0)
    
    # 3. Reasoning quality (if available)
    if hasattr(pred, 'reasoning'):
        has_reasoning = len(str(pred.reasoning)) > 20
        scores.append(1.0 if has_reasoning else 0.5)
    
    return sum(scores) / len(scores) if scores else 0.0

class EvaluationSuite:
    def __init__(self, devset, num_threads=8):
        self.devset = devset
        self.num_threads = num_threads
    
    def evaluate(self, program, metric=None) -> EvaluationResult:
        """Run full evaluation with detailed results."""
        metric = metric or comprehensive_metric

        evaluator = Evaluate(
            devset=self.devset,
            metric=metric,
            num_threads=self.num_threads,
            display_progress=True
        )

        eval_result = evaluator(program)

        # Extract individual scores from results
        scores = [score for example, pred, score in eval_result.results]
        correct = sum(1 for s in scores if s >= 0.5)
        errors = sum(1 for s in scores if s == 0)

        return EvaluationResult(
            score=eval_result.score,
            num_examples=len(self.devset),
            correct=correct,
            incorrect=len(self.devset) - correct - errors,
            errors=errors
        )
    
    def compare(self, programs: dict, metric=None) -> dict:
        """Compare multiple programs."""
        results = {}
        
        for name, program in programs.items():
            logger.info(f"Evaluating: {name}")
            results[name] = self.evaluate(program, metric)
        
        # Rank by score
        ranked = sorted(results.items(), key=lambda x: x[1].score, reverse=True)
        
        print("\n=== Comparison Results ===")
        for rank, (name, result) in enumerate(ranked, 1):
            print(f"{rank}. {name}: {result.score:.2%}")
        
        return results
    
    def export_report(self, program, output_path: str, metric=None):
        """Export detailed evaluation report."""
        result = self.evaluate(program, metric)
        
        report = {
            "summary": {
                "score": result.score,
                "total": result.num_examples,
                "correct": result.correct,
                "accuracy": result.correct / result.num_examples
            },
            "config": {
                "num_threads": self.num_threads,
                "num_examples": len(self.devset)
            }
        }
        
        with open(output_path, 'w') as f:
            json.dump(report, f, indent=2)
        
        logger.info(f"Report saved to {output_path}")
        return report

# Usage
suite = EvaluationSuite(devset, num_threads=8)

# Single evaluation
result = suite.evaluate(my_program)
print(f"Score: {result.score:.2%}")

# Compare variants
results = suite.compare({
    "baseline": baseline_program,
    "optimized": optimized_program,
    "finetuned": finetuned_program
})

Best Practices

  1. Hold out test data - Never optimize on evaluation set
  2. Multiple metrics - Combine correctness, quality, efficiency
  3. Statistical significance - Use enough examples (100+)
  4. Track over time - Version control evaluation results

Limitations

  • Metrics are task-specific; no universal measure
  • SemanticF1 requires LLM calls (cost)
  • Parallel evaluation can hit rate limits
  • Edge cases may not be captured

Official Documentation

© OmidZamani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/dspy-evaluation-suite of OmidZamani/dspy-skills.

  • SKILL.md
  • example.py

Open the folder on GitHubat commit f5db3b7

Compare with similar skills

Dspy Evaluation Suite next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dspy Evaluation Suite compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dspy Evaluation Suite this skillOmidZamani/dspy-skills123—~2kAutomated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
DSPy Language Model ProgrammingOrchestra-Research/AI-Research-SKILLs13k9 repos~3.8kAutomated safety check: PassMIT
Agent Benchmark Suiteruvnet/ruflo74k2 repos~4.9kAutomated safety check: PassMIT
LLM Evaluationdavila7/claude-code-templates33k12 repos~3.5kAutomated safety check: PassMIT
Evals Create Suiteelastic/kibana21k—~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • DSPy Language Model Programming

    Orchestra-Research/AI-Research-SKILLs

    Teaches an agent to build LM pipelines, RAG systems and agents in DSPy using signatures, modules and optimizers instead of hand-tuned prompts.

    13k GitHub starsUsed in 9 repos~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent skill for benchmark-suite - invoke with $agent-benchmark-suite

    74k GitHub starsUsed in 2 repos~4.9k tokens
    Agent WorkflowsAuto-check passed
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    33k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Evals Create Suite

    elastic/kibana

    Official

    Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files.

    21k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed

More from OmidZamani/dspy-skills

All 17 skills in this repo
  • Skill Perfection

    OmidZamani/dspy-skills

    A skill your agent uses when you need to QA audit and fix a plugin skill file.

    123 GitHub stars~1.6k tokensUpdated 3 mo ago
    Auto-check passed
  • Dspy Haystack Integration

    OmidZamani/dspy-skills

    A skill your agent uses for integrating DSPy with Haystack, optimizing Haystack prompts, improving retrieval pipelines, and extracting DSPy prompts.

    123 GitHub stars~1.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Dspy Adapters Multimodal

    OmidZamani/dspy-skills

    A skill your agent uses for DSPy adapter selection, JSONAdapter, XMLAdapter, ChatAdapter, native function calling, structured outputs, and multimodal inputs like dspy.Image or dspy.Audio.

    123 GitHub stars~864 tokensUpdated 3 mo ago
    Auto-check passed
  • Dspy Advanced Module Composition

    OmidZamani/dspy-skills

    A skill your agent uses for composing DSPy modules with Ensemble, MultiChainComparison, ensemble voting, sequential pipelines, and multi-program workflows.

    123 GitHub stars~2.2k tokensUpdated 3 mo ago
    Auto-check passed
  • Dspy Better Together

    OmidZamani/dspy-skills

    A skill your agent uses for BetterTogether, prompt plus weight optimization, fine-tuning sequences, and strategy chains like p - w - p.

    123 GitHub stars~756 tokensUpdated 3 mo ago
    Auto-check passed
  • Dspy Bootstrap Fewshot

    OmidZamani/dspy-skills

    A skill your agent uses for BootstrapFewShot, bootstrapped demonstrations, teacher-model demos, and low-data DSPy prompt optimization.

    123 GitHub stars~1.3k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Dspy Evaluation Suite

What does Dspy Evaluation Suite do?

A skill your agent uses for evaluating DSPy programs with Evaluate, answerexactmatch, SemanticF1, custom metrics, baselines, and program comparisons. Dspy Evaluation Suite is an agent skill from OmidZamani/dspy-skills. Use for evaluating DSPy programs with Evaluate, answerexactmatch, SemanticF1, custom metrics, baselines, and program comparisons.

When should I use Dspy Evaluation Suite?

Dspy Evaluation Suite fits situations like: evaluating DSPy programs with Evaluate; answerexactmatch; program comparisons.

How do I install Dspy Evaluation Suite in Claude Code?

Run `npx skills add OmidZamani/dspy-skills --skill dspy-evaluation-suite -a claude-code`. Or copy the skill folder (skills/dspy-evaluation-suite in OmidZamani/dspy-skills) into .claude/skills/dspy-evaluation-suite in your project. Claude Code loads it when a task matches its description.

How do I install Dspy Evaluation Suite in Codex?

Run `npx skills add OmidZamani/dspy-skills --skill dspy-evaluation-suite -a codex`. Or copy the skill folder (skills/dspy-evaluation-suite in OmidZamani/dspy-skills) into .agents/skills/dspy-evaluation-suite in your project. Codex loads it when a task matches its description.

Can I use Dspy Evaluation Suite in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OmidZamani/dspy-skills --skill dspy-evaluation-suite -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dspy-evaluation-suite, .gemini/skills/dspy-evaluation-suite, .github/skills/dspy-evaluation-suite and .opencode/skills/dspy-evaluation-suite in your project.

What does Dspy Evaluation Suite need to run?

Going by SKILL.md and its folder, Dspy Evaluation Suite needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Glob, Grep.

Does Dspy Evaluation Suite access the network?

SKILL.md names 2 domains. As links in the text: dspy.ai and github.com. This is read from the text; nothing was executed.

Is Dspy Evaluation Suite safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dspy Evaluation Suite use?

Dspy Evaluation Suite is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dspy Evaluation Suite use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dspy Evaluation Suite?

Skills that share tags, products or a category with Dspy Evaluation Suite: Arize Evaluator (github/awesome-copilot, 40k stars), DSPy Language Model Programming (Orchestra-Research/AI-Research-SKILLs, 13k stars), Agent Benchmark Suite (ruvnet/ruflo, 74k stars) and LLM Evaluation (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dspy Evaluation Suite?

OmidZamani (a GitHub user) maintains it in OmidZamani/dspy-skills, which has 123 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on June 23, 2026.

Source: OmidZamani/dspy-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.