Bridging Presidio And Spacy
maziyarpanahi/openmed
Combine OpenMed clinical NLP with Microsoft Presidio, spaCy, or LangChain through OpenMed's built-in interop adapter registry (openmed.interop).
Operating SOP for DSPy (Stanford NLP) — the declarative framework for "programming, not prompting" language models.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-dspy --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-dspy .claude/skills/agentsop-dspy && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-dspy" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-dspy into .claude/skills/agentsop-dspy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-dspy", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-dspyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-dspy --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-dspy .agents/skills/agentsop-dspy && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-dspy" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-dspy into .agents/skills/agentsop-dspy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-dspy", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-dspy --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-dspy .cursor/skills/agentsop-dspy && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-dspy" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-dspy into .cursor/skills/agentsop-dspy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-dspy", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-dspy--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-dspy --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-dspy .gemini/skills/agentsop-dspy && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-dspy" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-dspy into .gemini/skills/agentsop-dspy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-dspy", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-dspyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-dspy .github/skills/agentsop-dspy && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-dspy" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-dspy into .github/skills/agentsop-dspy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-dspy", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-dspy --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-dspy .opencode/skills/agentsop-dspy && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-dspy" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-dspy into .opencode/skills/agentsop-dspy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-dspy", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-dspyOperating SOP for DSPy (Stanford NLP) — the declarative framework for "programming, not prompting" language models.
Agentsop Dspy is an agent skill from agentsope/SkillAlchemy. Operating SOP for DSPy (Stanford NLP) — the declarative framework for "programming, not prompting" language models. Activate when the user says any of: "use DSPy", "compile a prompt", "optimize prompts/programs", "MIPRO/MIPROv2", "BootstrapFewShot", "GEPA", "Signatures + Modules", "teleprompter", "auto-tune prompts for a different LM", or whenever a brittle hand-crafted prompt pipeline needs to be turned into a compiled, measurable, swappable program. Do NOT activate for one-shot prompt tweaks, no-metric…
Its SKILL.md is about 7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-architecture.md`).
It sits in AI & LLM Engineering, covering Operations and SOPs, Building AI agents and Prompt engineering. It works with LangChain. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Dspy loads about 7k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 165 tokens; SKILL.md has 2,930 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 2,930 words, ~7,037 tokens.
.claude/skills/agentsop-dspy/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub."DSPy isn't a prompt-optimization agent framework. It's the LLM compiler for the shortest, cleanest code." — Eito Miyamura [eito.substack.com/p/dspy-the-most-misunderstood-agent]
"Prompts are effectively the weights of an LLM application." — Core philosophy [arxiv.org/abs/2310.03714]
Activate this skill when any of the following triggers are present in the user's intent or codebase:
| Trigger | Signal |
|---|---|
| Imports / mentions | import dspy, dspy.Signature, dspy.ChainOfThought, dspy.ReAct, Predict, MIPROv2, BootstrapFewShot, GEPA, teleprompter, compile( on an LM program |
| Tasks | "auto-tune this prompt", "I want to swap GPT-4 for a smaller model without re-engineering prompts", "I have 50/200/1000 labeled examples — optimize this", "compile a pipeline for our metric", "distill GPT-4 into Llama-3-8B" |
| Symptoms | Hand-written prompts grow past ~50 lines; brittleness on model swap; the team manually tunes few-shot examples; a metric exists but isn't being used to drive prompt design |
| Cross-skill bridges | LangGraph node calls an LLM and needs better prompts → wrap the node body in a DSPy module. LlamaIndex retriever feeds a reranker → DSPy-compile the reranker against a labeled set |
Do NOT activate when:
client.messages.create.DSPy's full name is Declarative Self-improving Python. The three primitives form a PyTorch-like compile chain [arxiv.org/abs/2310.03714]:
┌─────────────┐ ┌──────────┐ ┌──────────────┐ ┌─────────┐
│ Signature │ → │ Module │ → │ Teleprompter │ → │ Compile │
│ (what) │ │ (how) │ │ (optimizer) │ │ (tune) │
└─────────────┘ └──────────┘ └──────────────┘ └─────────┘
I/O spec Predict/CoT/ MIPROv2/GEPA/ Bake demos
field names ReAct/PoT BootstrapFewShot + instructions
= semantic = strategy = search algorithm into JSONThree mental shifts the agent must internalize:
Prompts are weights. The prompt string is not the artifact you ship — the compiled program (a JSON of demonstrations + instructions + structural choices) is. You ship program.json, not a .txt prompt [dspy.ai/tutorials/saving/].
Signatures carry semantic load. question -> answer is not the same as query -> response. DSPy uses the field names as the only natural-language hint the optimizer has about intent before it sees data. Name them like you'd name function parameters in well-written code [dspy.ai/learn/programming/signatures/].
Compile is a hyperparameter search, not a one-shot call. Compilation runs hundreds-to-thousands of LM calls (typically $2–$3 USD, 6–20 minutes, 3.2k API calls in the reference run). Costs scale with num_trials × |trainset| × |program LM calls| [dspy.ai/faqs/].
The PyTorch analogy is load-bearing. Signatures ≈ nn.Module.forward() shape contract. Modules ≈ nn.Linear / nn.Transformer. Teleprompters ≈ torch.optim.Adam. compile() ≈ training loop. save()/load() ≈ checkpoint.
The DSPy team is explicit about a three-stage gate [dspy.ai/learn/]:
"It's unproductive to launch optimization runs using a poorly designed program or a bad metric."
Do not skip stages. Each stage has an exit criterion.
"question -> answer"); upgrade to a class-based dspy.Signature with InputField(desc=...) / OutputField(desc=...) when types matter or fields need disambiguation.dspy.ChainOfThought. Use dspy.Predict for trivial classification, dspy.ReAct only when tools are needed, dspy.ProgramOfThought for arithmetic-heavy tasks [dspy.ai/learn/programming/modules/].dspy.Module, instantiate sub-modules in __init__, call them in forward(). No special DSL.dspy.inspect_history(n=3).Exit criterion: the un-optimized program produces plausible outputs on 5+ examples. Not great — plausible.
def metric(example, pred, trace=None) -> float|bool. Start with exact-match; only escalate to LLM-as-judge when the task demands it (open-ended generation, multi-criteria).dspy.Evaluate(devset=dev, metric=metric, num_threads=16) and record a baseline score.Exit criterion: baseline score is stable across two runs (cache-free) AND the metric agrees with human judgment on 10 spot-checks.
auto="light". Only escalate to "medium"/"heavy" if dev-set gains flatten and budget allows.compiled.save("v1.json") for state, or compiled.save("./v1/", save_program=True) for whole-program (preferred for production with metadata) [dspy.ai/tutorials/saving/].dspy.asyncify) or MLflow (mlflow.dspy.log_model) [dspy.ai/tutorials/deployment/].Exit criterion: compiled program beats baseline on a held-out test set (not the val set used in optimization) by ≥ task-relevant delta.
Loop to Stage 1 if optimization plateaus. Per the docs: "Is your task well-defined? Do you need more data? Should your evaluation metric change?" — these are the questions to re-ask, not "should I try a different optimizer?" [dspy.ai/learn/optimization/overview/].
| Trigger | Action | Output | Evidence |
|---|---|---|---|
| ≤10 labeled examples | BootstrapFewShot(metric=m, max_bootstrapped_demos=4, max_rounds=1) | Compiled program with self-generated demos | [dspy.ai/learn/optimization/optimizers/] |
| 30–50 examples | BootstrapFewShotWithRandomSearch | Best-of-N candidate programs | [dspy.ai/learn/optimization/optimizers/] |
| 200+ examples, willing to spend compute | MIPROv2(metric=m, auto="light") then escalate | Jointly-tuned instructions + few-shot demos via Bayesian optimization | [dspy.ai/api/optimizers/MIPROv2/] |
| Need zero-shot prompts (no demos in final) | MIPROv2(..., max_bootstrapped_demos=0, max_labeled_demos=0) | Instruction-only optimization | [dspy.ai/learn/optimization/optimizers/] |
| Have textual error feedback (test diffs, schema violations, judge rationales) | dspy.GEPA(metric=m_with_feedback) | Reflection-evolved prompts; sample-efficient | [dspy.ai/tutorials/gepa_ai_program/], [arxiv.org/abs/2507.19457] |
| Already optimized with MIPROv2 / want to ship a smaller model | Chain into BootstrapFinetune(student=small_lm, teacher=optimized) | Finetuned weights (not just prompts) | [dspy.ai/api/optimizers/BootstrapFinetune/] |
| Just want labeled demos in prompt (no search) | LabeledFewShot(k=8) | Trivial — fastest, cheapest, weakest | [dspy.ai/cheatsheet/] |
| Trigger | Action | Why |
|---|---|---|
| Simple input → output | dspy.Predict(Sig) | Lowest overhead |
| Reasoning helps | dspy.ChainOfThought(Sig) | Default choice per docs |
| Math / counting / parsing | dspy.ProgramOfThought(Sig) | Code execution grounds the answer |
| Tools (search, calc, API) | dspy.ReAct(Sig, tools=[...]) | Built-in tool loop |
| Ensemble for hard cases | dspy.MultiChainComparison or dspy.majority | Vote across N CoT samples |
| Trigger | Action | Caveat |
|---|---|---|
| Exact answer expected | lambda ex, pred: ex.answer.lower() == pred.answer.lower() | Cheap, deterministic |
| Open-ended generation | LLM-as-judge with dspy.ChainOfThought(JudgeSig) | Watch for self-preference bias, recency bias, score-ID bias [arxiv.org/pdf/2509.26072] |
| Multi-criteria (factuality + tone + length) | Sub-judge each dim, return bool during optimization (trace is not None) and float during evaluation | Documented pattern [dspy.ai/learn/evaluation/metrics/] |
| Have rich error context | Return dspy.Prediction(score=..., feedback="missing field X") and use GEPA | Textual feedback is GEPA's superpower [dspy.ai/api/optimizers/GEPA/overview/] |
| Trigger | Action | Reference |
|---|---|---|
Before any MIPROv2 call | Estimate: auto="light" ≈ a few $; auto="heavy" on 1000+ examples can hit tens of $ | [dspy.ai/faqs/] |
| Budget tight | Use a cheap optimizer LM (e.g. gpt-4o-mini) to optimize prompts for a more expensive task LM — community-reported parity [github.com/stanfordnlp/dspy/issues/1596] | |
| Compile stuck mid-trial | Check issue #1970 pattern; reduce minibatch_size or kill and restart with smaller num_trials | |
| Need reproducibility | dspy.configure(track_usage=True) + log program.get_lm_usage() |
困境 (Dilemma): User has a 3-stage RAG pipeline. Hand-tuned prompts already hit 72% on dev. MIPROv2 auto="heavy" would cost ~$40 and 4 hours. Worth it?
约束 (Constraints):
决策步骤 (Decision steps):
auto="light" (~$2) typically yields 10–30%+ on hand-tuned baselines per the paper's GPT-3.5/Llama2 results (25%/65% lift over standard few-shot) [arxiv.org/abs/2310.03714].auto="light" first as a cheap signal. The docs explicitly recommend "start with moderate values, observe behavior, and scale up only if you see clear gains" [github.com/stanfordnlp/dspy issue #1596].light gives <2% lift, do not escalate to heavy. Instead, revisit Stage 1: is the signature ambiguous? Is the program structure (3 stages) actually right?light gives 5–10% lift, run medium. Only escalate to heavy if data ≥ 300 and you have a held-out test set distinct from val.结果 (Outcome): Typical: light exposes whether more compute helps. Often the answer is "no — fix the program/metric first."
可提取的操作 (Extractable operation): Never start compilation at auto="heavy". Always probe with light and use a cheap optimizer LM.
困境: Compiled program for GPT-4o works at 85%. Need to switch to Llama-3-8B for cost. Re-use the GPT-4o-compiled program.json or recompile?
约束:
决策步骤:
BootstrapFinetune as a follow-on: optimize prompts on the big model, then distill into a 1B–7B student. Typical setup: student=Llama-3.2-1B-Instruct, teacher=gpt-4o-mini [dspy.ai/api/optimizers/BootstrapFinetune/].program.gpt4o.json and program.llama8b.json checked in; A/B in production.结果: Recompiled programs typically recover 70–90% of the larger-model performance at 1/10–1/50 the per-call cost. The "transfer without recompile" path is reliably worse.
可提取的操作: Treat the compiled program as a (program × LM) pair. Changing the LM invalidates the artifact — recompile.
困境: Open-ended customer-support response task. No exact-match metric possible. LLM-as-judge "feels right" but the team worries the judge will be biased toward verbose, hedged outputs.
约束:
决策步骤:
dspy.Predict(Assess) call with a single yes/no question (factual? on-topic? concise? non-hedging?). Documented pattern [dspy.ai/learn/evaluation/metrics/].trace is not None to return bool during compile, float during eval — same metric function, two modes. Avoids the optimizer overfitting to score noise.dspy.GEPA instead of MIPROv2 — GEPA leverages text feedback for faster, more sample-efficient convergence [dspy.ai/api/optimizers/GEPA/overview/, arxiv.org/abs/2507.19457].结果: Multi-dimension metric with explicit length penalty + GEPA's textual feedback typically beats single-judge + MIPROv2 by 10–13% on AIME-style benchmarks [arxiv.org/abs/2507.19457] and is the empirically robust path.
可提取的操作: Never compile against a metric you haven't human-validated on ≥ 20 spot-checks. Decompose multi-criteria metrics. Prefer GEPA when you can express textual feedback.
困境: MIPROv2 compile stuck mid-trial (no progress logs for 30 min). Reported pattern in issue #1970 [github.com/stanfordnlp/dspy/issues/1970]. Abort and restart, or wait?
约束:
决策步骤:
dspy.inspect_history(n=3) — does the last LM call show truncation or rate-limit error?max_bootstrapped_demos and max_labeled_demos (default 4 each); the docs explicitly cite this as the #1 context-length fix [dspy.ai/faqs/].num_threads in the underlying Evaluate; add retry/backoff in the LM client.minibatch_size (default 35; try 16) and smaller num_trials. Hanging is a known failure mode without graceful resume.student retains best demo candidates — check compiled._predictors state.可提取的操作: Compile is not atomic. Treat long hangs as failure. The cost of restart < cost of indefinite wait.
auto="heavy". Always probe with light first [Case A].cache=False in Lambda / stateless deploys. Caches default to a writable dir and break in serverless [dspy.ai/faqs/].dspy.streamify from 2.6.0+ but it's newer than the rest of the stack — verify your version [dspy.ai/tutorials/deployment/].DSPy is not the same layer as LangChain / LlamaIndex / LangGraph. It sits underneath them as a compiler for the individual LM calls inside those orchestration layers [langwatch.ai/blog/best-ai-agent-frameworks-in-2025-...].
┌──────────────────────────────────────────────┐
│ Orchestration: LangGraph, CrewAI │ ← graphs, agents, state
├──────────────────────────────────────────────┤
│ Retrieval: LlamaIndex │ ← ingestion, indexing
├──────────────────────────────────────────────┤
│ Compiler: DSPy │ ← signatures, modules, compile
├──────────────────────────────────────────────┤
│ Generation: Guidance, LMQL, Outlines │ ← single-call grammar control
├──────────────────────────────────────────────┤
│ Inference: vLLM, llama.cpp, Anthropic │ ← serving
└──────────────────────────────────────────────┘Predict(context, question -> answer) synthesizes. JetBlue's chatbot uses exactly this split — retrieval quality + answer quality as separate metrics, DSPy optimizes both [databricks.com/blog/optimizing-databricks-llm-pipelines-dspy].OutputField already pushes the LM toward structure but doesn't guarantee grammar conformance [dspy.ai/faqs/].import dspy
# 1. Signature
class BasicQA(dspy.Signature):
"""Answer questions with short factoid answers."""
question: str = dspy.InputField()
answer: str = dspy.OutputField(desc="often between 1 and 5 words")
# 2. Module
qa = dspy.ChainOfThought(BasicQA)
# 3. Metric
def metric(ex, pred, trace=None):
return ex.answer.lower() in pred.answer.lower()
# 4. Compile
from dspy.teleprompt import MIPROv2
optimizer = MIPROv2(metric=metric, auto="light")
compiled = optimizer.compile(qa, trainset=trainset)
# 5. Save / load
compiled.save("v1.json")Have a metric? ─── No ──► Stop. Build a metric first. (Or skip DSPy.)
│
Yes
│
Have ≥ 30 examples? ─── No ──► Stop. Collect more data, or use LabeledFewShot(k=8) as floor.
│
Yes
│
Have textual error feedback? ─── Yes ──► dspy.GEPA
│
No
│
≤ 10 examples? ──► BootstrapFewShot
30–50? ──► BootstrapFewShotWithRandomSearch
50–200? ──► MIPROv2(auto="light", max_bootstrapped_demos=4)
200+? ──► MIPROv2(auto="light" → "medium" if gains; "heavy" only if 300+ and budget)
Need to ship small model? ──► chain BootstrapFinetune after MIPROv2program.jsonAfter compiled.save("v1.json"), the file is plain JSON. Per-predictor it contains [dspy.ai/tutorials/saving/]:
{
"predictor_name": {
"signature_instructions": "Given the context, answer the question with a short factoid...",
"signature_prefix": "Answer:",
"extended_signature_instructions": "...",
"demos": [
{"question": "...", "reasoning": "...", "answer": "..."},
...
],
"signature": {
"instructions": "...",
"fields": [{"prefix": "Question:", "description": "..."}, ...]
}
}
}What changes when you compile:
What does NOT change between LMs (so you can read across artifacts):
What DOES change between LMs (so you can't reuse):
dspy.Assert vs dspy.SuggestFor self-refining pipelines [dspy.ai/learn/programming/7-assertions/, arxiv.org/pdf/2312.13382]:
# Hard: halts after max retries with dspy.AssertionError
dspy.Assert(len(pred.answer) < 100, "Answer must be < 100 chars")
# Soft: retries with feedback in prompt, logs failure, continues
dspy.Suggest(is_valid_json(pred.output), "Output must be valid JSON")When a constraint fails, DSPy backtracks to the previous module and re-runs with the error message injected into the prompt. This is self-refinement at inference time — distinct from compile-time optimization.
Use Assert during development (catch bugs hard). Use Suggest in production (degrade gracefully).
© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (references) in skills/agentsop-dspy of agentsope/SkillAlchemy.
Open the folder on GitHubat commit d0f0355
Agentsop Dspy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Dspy this skillagentsope/SkillAlchemy | 466 | — | ~7k | Automated safety check: Pass | MIT | |
| Bridging Presidio And Spacymaziyarpanahi/openmed | 5.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Agent Prompt Engineeringagentailor/fullstack-langgraph-nextjs-agent | 132 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Awesome Chatgpt Searchtaishi-i/awesome-ChatGPT-repositories | 3.3k | — | ~3.8k | Automated safety check: Pass | CC0-1.0 | |
| Module 1brevdev/workshop-build-an-agent | 146 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Module 1brevdev/workshop-build-an-agent | 146 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 |
maziyarpanahi/openmed
Combine OpenMed clinical NLP with Microsoft Presidio, spaCy, or LangChain through OpenMed's built-in interop adapter registry (openmed.interop).
agentailor/fullstack-langgraph-nextjs-agent
Comprehensive guide for designing, refining, and auditing system prompts for autonomous AI agents based on Anthropic's production practices.
taishi-i/awesome-ChatGPT-repositories
Search 2500+ curated ChatGPT and LLM open-source repositories.
brevdev/workshop-build-an-agent
This skill should be used when a learner is working through Module 1 ("Build an Agent") of the Build-an-Agent workshop and wants help understanding the concepts, notebooks, or code — e.g.
brevdev/workshop-build-an-agent
This skill should be used when a learner is working through Module 1 ("Build an Agent") of the Build-an-Agent workshop and wants help understanding the concepts, notebooks, or code — e.g.
jeremylongshore/tons-of-skills-marketplace
Manage LangChain 1.0 prompts like code — LangSmith prompt hub versioning, XML-tag conventions for Claude, few-shot example selection, discriminated-union extraction schemas, and A/B test wiring.
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Works with
Categories
Operating SOP for DSPy (Stanford NLP) — the declarative framework for "programming, not prompting" language models. Agentsop Dspy is an agent skill from agentsope/SkillAlchemy. Operating SOP for DSPy (Stanford NLP) — the declarative framework for "programming, not prompting" language models.
Agentsop Dspy fits situations like: says any of: use DSPy; compile a prompt; optimize prompts/programs; bootstrapFewShot.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a claude-code`. Or copy the skill folder (skills/agentsop-dspy in agentsope/SkillAlchemy) into .claude/skills/agentsop-dspy in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a codex`. Or copy the skill folder (skills/agentsop-dspy in agentsope/SkillAlchemy) into .agents/skills/agentsop-dspy in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-dspy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-dspy, .gemini/skills/agentsop-dspy, .github/skills/agentsop-dspy and .opencode/skills/agentsop-dspy in your project.
SKILL.md names no scripts, command-line tools or credentials: Agentsop Dspy is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Dspy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Dspy: Bridging Presidio And Spacy (maziyarpanahi/openmed, 5.5k stars), Agent Prompt Engineering (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Awesome Chatgpt Search (taishi-i/awesome-ChatGPT-repositories, 3.3k stars) and Module 1 (brevdev/workshop-build-an-agent, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 466 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.