Agent skill

Dspy Advanced Workflow

by intertwine in intertwine/dspy-agent-skills

Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget.

MITAuto-check passedAI & LLM Engineering

Install Dspy Advanced Workflow

skills CLI
$ npx skills add intertwine/dspy-agent-skills --skill dspy-advanced-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install intertwine/dspy-agent-skills dspy-advanced-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/intertwine/dspy-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dspy-advanced-workflow .claude/skills/dspy-advanced-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dspy-advanced-workflow
GitHub stars
278
Token cost
~1.7k tokens
SKILL.md length
385 words
Files
3
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget.

  • Works in 7 steps: Spec → Program → Data → …
  • Greenfield DSPy builds
  • SKILL.md covers The seven steps, Full orchestration template, Guardrails and Runnable scaffold →…
  • Runs Python scripts from its folder

What it does

Dspy Advanced Workflow is an agent skill from intertwine/dspy-agent-skills. Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget. Orchestrates the other four DSPy skills (dspy-fundamentals, dspy-evaluation-harness, dspy-gepa-optimizer, dspy-rlm-module) in the correct order. Use for greenfield DSPy builds; prototypes may stop at a validated baseline.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `example_pipeline.py` and `reference.md`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Production-grade DSPy 3.2.x agent skills + validated end-to-end examples for Claude Code and Codex CLI — fundamentals, evaluation, GEPA, BetterTogether, and RLM. The licence is MIT.

When your agent uses it

  • Greenfield DSPy builds
  • Prototypes may stop at a validated baseline

Example prompts

  • “/dspy-advanced-workflow”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Spec
  2. Program
  3. Data
  4. Rich metric
  5. Baseline
  6. GEPA optimize
  7. Export & deploy

What it can do on your machine

Read from SKILL.md and the folder at commit 623dca0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dspy Advanced Workflow loads about 1.7k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 385 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from intertwine/dspy-agent-skills at commit 623dca0, republished under its MIT licence (© intertwine). 385 words, ~1,672 tokens.

Download SKILL.mdSave it as .claude/skills/dspy-advanced-workflow/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
dspy-advanced-workflow
description
Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget. Orchestrates the other four DSPy skills (dspy-fundamentals, dspy-evaluation-harness, dspy-gepa-optimizer, dspy-rlm-module) in the correct order. Use for greenfield DSPy builds; prototypes may stop at a validated baseline.
when_to_use
User wants to build, optimize, and ship a new DSPy pipeline; says "full workflow" / "end to end" / "from scratch"; or needs the standard loop applied to a…

DSPy Advanced Workflow (2026)

This skill runs the seven-step loop that turns a natural-language task description into an optimized, saved, deployable DSPy program. Use the relevant steps in order. Stop at a validated baseline for a prototype; optimizer runs require an appropriate authorized budget and evidence of need. Exporting a local artifact does not authorize deployment.

The seven steps

1. Spec

Rephrase the user's task in one sentence. Identify inputs, outputs, the quality axis that matters, and any constraints (latency, cost, tool access, context size). Pick predictor shape:

Task shapePredictor
Single-step structured I/Odspy.Predict / dspy.ChainOfThought
Tool use / multi-stepdspy.ReAct
Code executiondspy.ProgramOfThought
Long context / codebasedspy.RLM → dspy-rlm-module
2. Program

Write the typed dspy.Signature + dspy.Module subclass per dspy-fundamentals. No hard-coded prompts. Keep predictors named so GEPA can target them.

3. Data

Build trainset and separate valset as dspy.Example(...).with_inputs(...). For GEPA, maximize trainset size and keep validation just large enough to represent downstream behavior; held-out testset is reported on at the end only. See dspy-evaluation-harness.

4. Rich metric

Write rich_metric(gold, pred, trace=None, pred_name=None, pred_trace=None) returning dspy.Prediction(score=0..1, feedback="natural-language critique"). The feedback is load-bearing — it's what GEPA's reflection LM learns from. A dict with the same fields crashes dspy.Evaluate; only dspy.Prediction aggregates correctly. See dspy-evaluation-harness.

5. Baseline
python
evaluator = dspy.Evaluate(devset=valset, metric=rich_metric,
                          num_threads=8, display_progress=True,
                          provide_traceback=True,
                          save_as_json="runs/baseline.json")
baseline = evaluator(program)
print("Baseline:", baseline.score)
6. GEPA optimize
python
reflection_lm = dspy.LM("openai/gpt-5", temperature=1.0, max_tokens=32000)
optimizer = dspy.GEPA(
    metric=rich_metric,
    auto="medium",
    reflection_lm=reflection_lm,
    candidate_selection_strategy="pareto",
    track_stats=True,
    track_best_outputs=True,
    log_dir="./gepa_logs",
    num_threads=8,
    seed=0,
)
optimized = optimizer.compile(student=program, trainset=trainset, valset=valset)
print("Optimized:", evaluator(optimized).score)

Run auto="light" first as a sanity check; move to auto="medium"/"heavy" for the final run. See dspy-gepa-optimizer.

If you need a deliberate multi-stage compile loop, DSPy 3.2.x also exposes dspy.BetterTogether(metric=..., bootstrap=..., gepa=...) for chaining named optimizers after you have a clean baseline GEPA setup.

Show full SKILL.md (135 more words)Show less
7. Export & deploy
python
optimized.save("artifacts/program.json", save_program=False)     # state, portable
# or for full deployment artifact:
optimized.save("artifacts/program_dir/", save_program=True)

Deploy:

  • Load with dspy.load("artifacts/program_dir/") or reconstruct + .load("program.json").
  • Wrap in FastAPI/CLI.
  • Enable track_usage=True for cost/latency observability.
  • Log with MLflow (mlflow.dspy.autolog()) or W&B in CI.
  • Keep an offline regression test that runs the evaluator against the saved program and fails CI below a threshold.

Full orchestration template

python
"""DSPy end-to-end pipeline — spec → optimize → deploy."""

import dspy
from pathlib import Path

# ----- 1–2. Spec & program (dspy-fundamentals) -----
class MyTask(dspy.Signature):
    """<one-line instruction from the spec>."""
    input_field: str = dspy.InputField()
    output_field: str = dspy.OutputField()

class MyProgram(dspy.Module):
    def __init__(self):
        super().__init__()
        self.step = dspy.ChainOfThought(MyTask)
    def forward(self, **kw):
        return self.step(**kw)

# ----- 3. Data (dspy-evaluation-harness) -----
trainset = [...]   # list[dspy.Example(...).with_inputs(...)]
valset   = [...]

# ----- 4. Rich metric (dspy-evaluation-harness) -----
def rich_metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
    score = ...          # compute 0..1
    feedback = ...       # detailed critique
    return dspy.Prediction(score=score, feedback=feedback)  # NOT a dict

# ----- 5. Baseline -----
dspy.configure(lm=dspy.LM("openai/gpt-4o"), track_usage=True)
evaluator = dspy.Evaluate(devset=valset, metric=rich_metric, num_threads=8,
                          display_progress=True, provide_traceback=True,
                          save_as_json="runs/baseline.json")
program = MyProgram()
print("Baseline:", evaluator(program).score)

# ----- 6. GEPA optimize (dspy-gepa-optimizer) -----
optimizer = dspy.GEPA(
    metric=rich_metric,
    auto="medium",
    reflection_lm=dspy.LM("openai/gpt-5", temperature=1.0, max_tokens=32000),
    candidate_selection_strategy="pareto",
    track_stats=True, track_best_outputs=True,
    log_dir="./gepa_logs", num_threads=8, seed=0,
)
optimized = optimizer.compile(student=program, trainset=trainset, valset=valset)
print("Optimized:", evaluator(optimized).score)

# ----- 7. Export (dspy-fundamentals) -----
Path("artifacts").mkdir(exist_ok=True)
optimized.save("artifacts/program.json", save_program=False)

Guardrails

  • Define the metric in step 4 before optimization; use informative feedback appropriate to the task.
  • Always baseline before optimizing — no baseline, no claim.
  • Save both pre- and post-optimization metrics to JSON for auditability.
  • If held-out test score drops, preserve that result and diagnose using training/validation evidence rather than assuming a cause. After test-informed changes, use a new untouched final holdout or label subsequent results exploratory; do not repeatedly tune against the original test set.
  • Freeze optimized program with module._compiled = True before multi-stage re-compilation.

Runnable scaffold → example_pipeline.py

© intertwine, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/dspy-advanced-workflow of intertwine/dspy-agent-skills.

  • SKILL.md
  • example_pipeline.py
  • reference.md

Open the folder on GitHubat commit 623dca0

Compare with similar skills

Dspy Advanced Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dspy Advanced Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dspy Advanced Workflow this skillintertwine/dspy-agent-skills278—~1.7kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from intertwine/dspy-agent-skills

  • Dspy Evaluation Harness

    intertwine/dspy-agent-skills

    Build DSPy evaluation harnesses with rich-feedback metrics that are essential for GEPA optimization.

    278 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Dspy Fundamentals

    intertwine/dspy-agent-skills

    Write idiomatic DSPy 3.2.x programs — typed Signatures, dspy.Module subclasses, Predict/ChainOfThought/ReAct/ProgramOfThought, and save/load.

    278 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Dspy Gepa Optimizer

    intertwine/dspy-agent-skills

    Optimize DSPy programs with dspy.GEPA — a reflective/evolutionary optimizer to consider against task-specific baselines within an authorized evaluation budget.

    278 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dspy Rlm Module

    intertwine/dspy-agent-skills

    Use dspy.RLM (Recursive Language Model) for reasoning over contexts too large to fit in an LLM's working window — entire codebases, long logs, massive documents, or multi-step data exploration that…

    278 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Dspy Advanced Workflow

What does Dspy Advanced Workflow do?

Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget. Dspy Advanced Workflow is an agent skill from intertwine/dspy-agent-skills.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget.

When should I use Dspy Advanced Workflow?

Dspy Advanced Workflow fits situations like: greenfield DSPy builds; prototypes may stop at a validated baseline.

How do I install Dspy Advanced Workflow in Claude Code?

Run `npx skills add intertwine/dspy-agent-skills --skill dspy-advanced-workflow -a claude-code`. Or copy the skill folder (skills/dspy-advanced-workflow in intertwine/dspy-agent-skills) into .claude/skills/dspy-advanced-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Dspy Advanced Workflow in Codex?

Run `npx skills add intertwine/dspy-agent-skills --skill dspy-advanced-workflow -a codex`. Or copy the skill folder (skills/dspy-advanced-workflow in intertwine/dspy-agent-skills) into .agents/skills/dspy-advanced-workflow in your project. Codex loads it when a task matches its description.

Can I use Dspy Advanced Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intertwine/dspy-agent-skills --skill dspy-advanced-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dspy-advanced-workflow, .gemini/skills/dspy-advanced-workflow, .github/skills/dspy-advanced-workflow and .opencode/skills/dspy-advanced-workflow in your project.

What does Dspy Advanced Workflow need to run?

Going by SKILL.md and its folder, Dspy Advanced Workflow needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Dspy Advanced Workflow access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dspy Advanced Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dspy Advanced Workflow use?

Dspy Advanced Workflow is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dspy Advanced Workflow use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dspy Advanced Workflow?

Skills that share tags, products or a category with Dspy Advanced Workflow: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dspy Advanced Workflow?

intertwine (a GitHub user) maintains it in intertwine/dspy-agent-skills, which has 278 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 6, 2026.

Source: intertwine/dspy-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.