Agent skill

Dspy Gepa Optimizer

by intertwine in intertwine/dspy-agent-skills

Optimize DSPy programs with dspy.GEPA — a reflective/evolutionary optimizer to consider against task-specific baselines within an authorized evaluation budget.

MITAuto-check passed

Install Dspy Gepa Optimizer

skills CLI
$ npx skills add intertwine/dspy-agent-skills --skill dspy-gepa-optimizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install intertwine/dspy-agent-skills dspy-gepa-optimizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/intertwine/dspy-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dspy-gepa-optimizer .claude/skills/dspy-gepa-optimizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dspy-gepa-optimizer
GitHub stars
278
Token cost
~2.6k tokens
SKILL.md length
914 words
Files
4
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Optimize DSPy programs with dspy.GEPA — a reflective/evolutionary optimizer to consider against task-specific baselines within an authorized evaluation budget.

  • Works in 4 steps: A dspy.Module that runs end-to-end (see… → A rich-feedback metric returning… → trainset and a separate valset. For… → …
  • The user says optimize
  • SKILL.md covers Prerequisites — do these first…, Canonical call, Import paths and Metric contract (precise), plus 12 more sections
  • Runs Python scripts from its folder; needs WANDB_API_KEY

What it does

Dspy Gepa Optimizer is an agent skill from intertwine/dspy-agent-skills. Optimize DSPy programs with dspy.GEPA — a reflective/evolutionary optimizer to consider against task-specific baselines within an authorized evaluation budget. Use when the user says optimize, compile, GEPA, reflective optimization, or "make this program better" and a DSPy program + metric + trainset exist.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `example_bettertogether.py`, `example_gepa.py` and `reference.md`).

The repository describes itself as: Production-grade DSPy 3.2.x agent skills + validated end-to-end examples for Claude Code and Codex CLI — fundamentals, evaluation, GEPA, BetterTogether, and RLM. The licence is MIT.

When your agent uses it

  • The user says optimize
  • Reflective optimization
  • Make this program better and a DSPy program + metric + trainset exist

Example prompts

  • “make this program better”
  • “/dspy-gepa-optimizer”

Requirements

  • Python 3
  • A credential in WANDB_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. A dspy.Module that runs end-to-end (see dspy-fundamentals).
  2. A rich-feedback metric returning dspy.Prediction(score=float, feedback=str) (see dspy-evaluation-harness). Informative feedback can…
  3. trainset and a separate valset. For GEPA, maximize training examples and keep validation just large enough to represent the downstream…
  4. A reflection_lm — a strong LM (often the same or stronger than the task LM) set to temperature=1.0 for creative proposals. Current DSPy…

What it can do on your machine

Read from SKILL.md and the folder at commit 623dca0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • WANDB_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dspy Gepa Optimizer loads about 2.6k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 914 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from intertwine/dspy-agent-skills at commit 623dca0, republished under its MIT licence (© intertwine). 914 words, ~2,612 tokens.

Download SKILL.mdSave it as .claude/skills/dspy-gepa-optimizer/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
dspy-gepa-optimizer
description
Optimize DSPy programs with dspy.GEPA — a reflective/evolutionary optimizer to consider against task-specific baselines within an authorized evaluation budget. Use when the user says optimize, compile, GEPA, reflective optimization, or "make this program better" and a DSPy program + metric + trainset exist.
when_to_use
User asks to optimize/compile/tune a DSPy program, mentions GEPA or reflective optimization, or has a working program with a non-trivial metric and wants to…

DSPy GEPA Optimizer (3.2.x)

GEPA (Genetic-Pareto) is a reflective optimizer: it mutates a program's instructions and few-shots using an LM that reads your metric's textual feedback and proposes improvements. It maintains a Pareto frontier across validation tasks and is the default recommendation for complex DSPy workloads in 2026.

The expansion "Genetic-Evolutionary Prompt Adaptation" that appears in some AI-generated summaries is an LLM-hallucinated backronym. The paper defines GEPA as Genetic-Pareto; the "Pareto" is load-bearing (GEPA keeps a frontier of candidates rather than collapsing to one).

Prerequisites — do these first or GEPA wastes rollouts

  1. A dspy.Module that runs end-to-end (see dspy-fundamentals).
  2. A rich-feedback metric returning dspy.Prediction(score=float, feedback=str) (see dspy-evaluation-harness). Informative feedback can support reflection; evaluate optimizer benefit on the task rather than assuming superiority. A dict with the same fields still crashes dspy.Evaluate under DSPy 3.2.1 — use dspy.Prediction.
  3. trainset and a separate valset. For GEPA, maximize training examples and keep validation just large enough to represent the downstream distribution; do not reuse the same examples for both.
  4. A reflection_lm — a strong LM (often the same or stronger than the task LM) set to temperature=1.0 for creative proposals. Current DSPy docs use a GPT-5-class reflection model with a large output budget.

Canonical call

python
import dspy

dspy.configure(lm=dspy.LM("openai/gpt-5-mini"))
reflection_lm = dspy.LM("openai/gpt-5", temperature=1.0, max_tokens=32000)

optimizer = dspy.GEPA(
    metric=rich_metric,
    auto="medium",                       # "light" / "medium" / "heavy"
    reflection_lm=reflection_lm,
    reflection_minibatch_size=3,
    candidate_selection_strategy="pareto",  # or "current_best"
    skip_perfect_score=True,
    use_merge=True,
    num_threads=8,
    track_stats=True,
    track_best_outputs=True,             # enables inference-time best-of selection
    log_dir="./gepa_logs",               # resume/checkpoint
    seed=0,
)

optimized = optimizer.compile(
    student=program,
    trainset=trainset,
    valset=valset,
)

# Pareto inspection
pareto = optimized.detailed_results.val_aggregate_scores
print("Pareto frontier:", sorted(pareto, reverse=True)[:5])

optimized.save("optimized_program.json", save_program=False)

Import paths

Either works; use the top-level in new code:

python
import dspy
dspy.GEPA(...)                              # preferred
# equivalently:
from dspy.teleprompt import GEPA

Metric contract (precise)

python
import dspy

def rich_metric(gold, pred, trace=None, pred_name=None, pred_trace=None):
    score = ...      # 0.0..1.0
    feedback = ...   # detailed natural-language critique
    return dspy.Prediction(score=score, feedback=feedback)

Return dspy.Prediction, not a dict. Some upstream GEPA prose describes score/feedback as a dict-like shape, but dspy.Evaluate in DSPy 3.2.1 still crashes on a literal dict metric (TypeError: unsupported operand type(s) for +: 'int' and 'dict'). GEPA uses dspy.Evaluate internally for candidate scoring, so a dict return can fail inside GEPA too, not just in your explicit Evaluate(...) calls.

  • pred_name / pred_trace are set during reflection on a specific predictor inside your module — write per-predictor feedback when possible (credit assignment). If you cannot localize feedback, return program-level feedback rather than a vague score-only critique.
  • Feedback quality is the load-bearing part: specifics about why it failed and what good looks like are what the reflection LM acts on.

Budget knobs

Use either auto=... or explicit budget — not both.

ModeRough rolloutsWhen to use
auto="light"~20–40 full evalsSanity-check GEPA works on your metric
auto="medium"~80–150 full evalsEveryday optimization
auto="heavy"~300–600 full evalsFinal run before ship
max_full_evals=NExplicitDeterministic budget
max_metric_calls=NExplicitHard cap on metric invocations (more predictable cost)

Each "full eval" ≈ len(valset) metric calls. Budget accordingly for cost.

Constructor parameters (every one, DSPy 3.2.x)

python
dspy.GEPA(
    metric,                                  # required
    auto=None,                               # Literal["light","medium","heavy"] | None
    max_full_evals=None,
    max_metric_calls=None,
    reflection_minibatch_size=3,
    candidate_selection_strategy="pareto",   # or "current_best"
    reflection_lm=None,                      # required in practice
    skip_perfect_score=True,
    add_format_failure_as_feedback=False,
    instruction_proposer=None,               # custom ProposalFn
    component_selector="round_robin",        # or a callable
    use_merge=True,
    max_merge_invocations=5,
    num_threads=None,
    failure_score=0.0,
    perfect_score=1.0,
    log_dir=None,
    track_stats=False,
    use_wandb=False,
    wandb_api_key=None,                      # overrides WANDB_API_KEY env var
    wandb_init_kwargs=None,                  # dict forwarded to wandb.init(...)
    track_best_outputs=False,
    warn_on_score_mismatch=True,
    use_mlflow=False,
    seed=0,
    gepa_kwargs=None,                        # e.g. {"use_cloudpickle": True} for dynamic signatures
)

.compile(student, *, trainset, valset=None, teacher=None) — teacher is not currently used.

Data split guidance

DSPy's general prompt-optimizer docs often recommend a validation-heavy split, such as 20% train / 80% validation, because small prompt optimizers can overfit tiny trainsets. GEPA is different: maximize the training set and reserve only enough validation examples to represent downstream behavior. The Pareto frontier still needs a real valset, but GEPA learns from traces and textual feedback on training examples, so starving trainset hurts.

BetterTogether in DSPy 3.2.x

If you want a multi-stage optimizer loop, DSPy 3.2.0's BetterTogether now accepts arbitrary named optimizers instead of the older fixed prompt_optimizer / weight_optimizer pair:

python
optimizer = dspy.BetterTogether(
    metric=rich_metric,
    bootstrap=dspy.BootstrapFewShotWithRandomSearch(metric=rich_metric),
    gepa=dspy.GEPA(metric=rich_metric, auto="light", reflection_lm=reflection_lm),
)

optimized = optimizer.compile(
    student=program,
    trainset=trainset,
    valset=valset,
    strategy="bootstrap -> gepa",
)

Pass strategy= explicitly when you use named stages like bootstrap=... and gepa=.... DSPy 3.2.0's default strategy is still "p -> w -> p", which only works if your optimizer keys are literally p and w.

Keep plain GEPA as the default first pass. Reach for BetterTogether only when you have a specific reason to chain optimizers and want the valset to pick the best intermediate program.

Show full SKILL.md (346 more words)Show less

When GEPA > MIPROv2

  • Your metric can produce specific, teachable critiques (GEPA's superpower).
  • The program has multiple predictors that need targeted improvements (GEPA gives per-predictor feedback; MIPRO doesn't).
  • Rollout budget is small (GEPA converges faster with rich feedback).

When MIPROv2 > GEPA

  • Metric is scalar-only (no signal to reflect on) — use dspy.MIPROv2.
  • You want pure few-shot bootstrapping with no instruction mutation.
  • Very large trainset (500+) where Bayesian search over demos pays off.

When SIMBA is worth trying

dspy.SIMBA is a lighter reflective optimizer. Try it when you want a cheaper reflective pass than GEPA, your program is simple, or you need quick exploration before committing to a full GEPA run. Keep GEPA as the default for multi-predictor programs where per-predictor feedback and Pareto candidate selection matter.

Resume & checkpointing

log_dir writes candidate programs + scores per round. To resume an interrupted run, point log_dir at the same directory — GEPA picks up from the last checkpoint. Inspect <log_dir>/candidates/ to see every proposed program.

Inference-time best-of with track_best_outputs

With track_best_outputs=True, GEPA records, per task, the best prediction seen across all candidates. At inference time on held-out data, you can ensemble or select among the top-Pareto programs for robustness. Access via optimized.detailed_results.best_outputs_valset.

Anti-patterns

  • Float-only metric ("score is 0.7") with no feedback — GEPA collapses to random search.
  • Same set used for train and val — Pareto selection overfits.
  • reflection_lm = small model — it can't critique; use the strongest LM you can afford for this role.
  • Running auto="heavy" on an untested metric — burn money to learn the metric was bugged. Run auto="light" first.
  • Ignoring log_dir — losing a 4-hour run to a disconnect is very painful.

Gotcha: reflection_lm is required at construction, not compile

dspy.GEPA(...) asserts reflection_lm is not None (or a custom instruction_proposer) at init time — you cannot defer it to .compile(). If you see

AssertionError: GEPA requires a reflection language model...

add reflection_lm=dspy.LM("openai/gpt-5", temperature=1.0, max_tokens=32000) to the constructor, or substitute the strongest instruction-following model available on your provider. dspy.LM(...) is a cheap stub until you actually call it, so constructing one doesn't hit the network.

Next

© intertwine, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in skills/dspy-gepa-optimizer of intertwine/dspy-agent-skills.

  • SKILL.md
  • example_bettertogether.py
  • example_gepa.py
  • reference.md

Open the folder on GitHubat commit 623dca0

Compare with similar skills

Dspy Gepa Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dspy Gepa Optimizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dspy Gepa Optimizer this skillintertwine/dspy-agent-skills278—~2.6kAutomated safety check: PassMIT
DSPy Language Model ProgrammingOrchestra-Research/AI-Research-SKILLs13k9 repos~3.8kAutomated safety check: PassMIT
SQL Optimizationgithub/awesome-copilot40k2 repos~2.3kAutomated safety check: PassMIT
Reflectalirezarezvani/claude-skills28k1 repos~2.4kAutomated safety check: PassMIT
Agent Performance Optimizerruvnet/ruflo74k2 repos~3.6kAutomated safety check: PassMIT
Database Optimizerdavila7/claude-code-templates33k8 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • DSPy Language Model Programming

    Orchestra-Research/AI-Research-SKILLs

    Teaches an agent to build LM pipelines, RAG systems and agents in DSPy using signatures, modules and optimizers instead of hand-tuned prompts.

    13k GitHub starsUsed in 9 repos~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • SQL Optimization

    github/awesome-copilot

    Official

    Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…

    40k GitHub starsUsed in 2 repos~2.3k tokens
    DatabasesAuto-check passed
  • Reflect

    alirezarezvani/claude-skills

    Mid-conversation reflection skill that pauses execution and zooms out from detail-mode to honestly reassess direction, assumptions, and bias.

    28k GitHub starsUsed in 1 repo~2.4k tokens
    Agent WorkflowsAuto-check passed
  • Agent skill for performance-optimizer - invoke with $agent-performance-optimizer

    74k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Database Optimizer

    davila7/claude-code-templates

    Expert database optimizer specializing in modern performance tuning, query optimization, and scalable architectures.

    33k GitHub starsUsed in 8 repos~2.5k tokens
    DatabasesAuto-check passed
  • Prompt Optimizer

    affaan-m/ECC

    分析原始提示,识别意图和差距,匹配ECC组件(技能/命令/代理/钩子),并输出一个可直接粘贴的优化提示。仅提供咨询角色——绝不自行执行任务。触发时机:当用户说“优化提示”、“改进我的提示”、“如何编写提示”、“帮我优化这个指令”或明确要求提高提示质量时。中文等效表达同样触发:“优化prompt”、“改进prompt”、“怎么写prompt”、“帮我优化这个指令”。不触发时机:当用户希望直接执行任…

    276k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed

More from intertwine/dspy-agent-skills

  • Dspy Advanced Workflow

    intertwine/dspy-agent-skills

    Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget.

    278 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Dspy Evaluation Harness

    intertwine/dspy-agent-skills

    Build DSPy evaluation harnesses with rich-feedback metrics that are essential for GEPA optimization.

    278 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Dspy Fundamentals

    intertwine/dspy-agent-skills

    Write idiomatic DSPy 3.2.x programs — typed Signatures, dspy.Module subclasses, Predict/ChainOfThought/ReAct/ProgramOfThought, and save/load.

    278 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Dspy Rlm Module

    intertwine/dspy-agent-skills

    Use dspy.RLM (Recursive Language Model) for reasoning over contexts too large to fit in an LLM's working window — entire codebases, long logs, massive documents, or multi-step data exploration that…

    278 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Dspy Gepa Optimizer

What does Dspy Gepa Optimizer do?

Optimize DSPy programs with dspy.GEPA — a reflective/evolutionary optimizer to consider against task-specific baselines within an authorized evaluation budget. Dspy Gepa Optimizer is an agent skill from intertwine/dspy-agent-skills.GEPA — a reflective/evolutionary optimizer to consider against task-specific baselines within an authorized evaluation budget.

When should I use Dspy Gepa Optimizer?

Dspy Gepa Optimizer fits situations like: the user says optimize; reflective optimization; make this program better and a DSPy program + metric + trainset exist.

How do I install Dspy Gepa Optimizer in Claude Code?

Run `npx skills add intertwine/dspy-agent-skills --skill dspy-gepa-optimizer -a claude-code`. Or copy the skill folder (skills/dspy-gepa-optimizer in intertwine/dspy-agent-skills) into .claude/skills/dspy-gepa-optimizer in your project. Claude Code loads it when a task matches its description.

How do I install Dspy Gepa Optimizer in Codex?

Run `npx skills add intertwine/dspy-agent-skills --skill dspy-gepa-optimizer -a codex`. Or copy the skill folder (skills/dspy-gepa-optimizer in intertwine/dspy-agent-skills) into .agents/skills/dspy-gepa-optimizer in your project. Codex loads it when a task matches its description.

Can I use Dspy Gepa Optimizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add intertwine/dspy-agent-skills --skill dspy-gepa-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dspy-gepa-optimizer, .gemini/skills/dspy-gepa-optimizer, .github/skills/dspy-gepa-optimizer and .opencode/skills/dspy-gepa-optimizer in your project.

What does Dspy Gepa Optimizer need to run?

Going by SKILL.md and its folder, Dspy Gepa Optimizer needs Python for the scripts in its folder and credentials named WANDB_API_KEY. Our summary lists: Python 3; A credential in WANDB_API_KEY.

Does Dspy Gepa Optimizer access the network?

SKILL.md names 1 domain. As links in the text: arxiv.org. This is read from the text; nothing was executed.

Is Dspy Gepa Optimizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dspy Gepa Optimizer use?

Dspy Gepa Optimizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dspy Gepa Optimizer use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dspy Gepa Optimizer?

Skills that share tags, products or a category with Dspy Gepa Optimizer: DSPy Language Model Programming (Orchestra-Research/AI-Research-SKILLs, 13k stars), SQL Optimization (github/awesome-copilot, 40k stars), Reflect (alirezarezvani/claude-skills, 28k stars) and Agent Performance Optimizer (ruvnet/ruflo, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dspy Gepa Optimizer?

intertwine (a GitHub user) maintains it in intertwine/dspy-agent-skills, which has 278 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 6, 2026.

Source: intertwine/dspy-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.