Agent skill

Agentsop Metric Design

by agentsope in agentsope/SkillAlchemy

Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.

MITAuto-check passedAI & LLM Engineering

Install Agentsop Metric Design

skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentsope/SkillAlchemy agentsop-metric-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-metric-design .claude/skills/agentsop-metric-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentsop-metric-design
GitHub stars
436
Token cost
~6.4k tokens
SKILL.md length
2,651 words
Files
5 (incl. references)
Skills in repo
46
Repo updated
First seen
Licence
MIT

At a glance

Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.

  • Works in 7 steps: 何时激活 (When to Activate) → 核心心智模型 (Core Mental Model) → SOP (Standard Operating Procedure) → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers 1. 何时激活 (When to Activate), 2. 核心心智模型 (Core Mental Model), 3. SOP (Standard Operating… and 4. 操作模型 (Operations), plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agentsop Metric Design is an agent skill from agentsope/SkillAlchemy. Decomposed, multi-criteria metric design for LLM pipelines. The metric IS the model — change the metric and the optimizer changes behavior. Decompose by default; bool during compile, float during eval; calibrate against human; mitigate judge bias. Search keywords: LLM-as-judge, llm as judge, eval metric, evaluation score, scoring function, rubric, RAGAS, G-Eval, judge bias, verbosity bias, how to evaluate LLM output.

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).

It sits in AI & LLM Engineering, covering LLM evaluation, Quizzes and assessments and Operations and SOPs. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve Quizzes and assessments
  • Tasks that involve Operations and SOPs

Example prompts

  • “/agentsop-metric-design”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 何时激活 (When to Activate)
  2. 核心心智模型 (Core Mental Model)
  3. SOP (Standard Operating Procedure)
  4. 操作模型 (Operations)
  5. 困境决策案例 (Dilemma Cases)
  6. 反模式与边界 (Anti-Patterns & Boundaries)
  7. 跨框架对照 (Cross-Framework Mapping)

What it can do on your machine

Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentsop Metric Design loads about 6.4k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 2,651 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 2,651 words, ~6,414 tokens.

Download SKILL.mdSave it as .claude/skills/agentsop-metric-design/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
agentsop-metric-design
description
Decomposed, multi-criteria metric design for LLM pipelines. The metric IS the model — change the metric and the optimizer changes behavior. Decompose by default; bool during compile, float during eval; calibrate against human; mitigate judge bias. Search keywords: LLM-as-judge, llm as judge, eval metric, evaluation score, scoring function, rubric, RAGAS, G-Eval, judge bias, verbosity bias, how to evaluate LLM output.
version
0.1.0
phase
D
tier
core
frequency
high
status
opinionated

metric-design — Decomposed, Multi-Criteria Metrics for LLM Pipelines

"It's unproductive to launch optimization runs using a poorly designed program or a bad metric." — DSPy core team [dspy.ai/learn/optimization/overview/]

"LLM judges exhibit self-preference, recency, rubric-order, score-ID, and length biases." — Synthesized from [arxiv.org/pdf/2506.02592, arxiv.org/pdf/2509.26072]

This is a tool skill. It produces a metric function (and a calibration receipt) that other skills consume — DSPy compilers (MIPROv2 / GEPA / BootstrapFewShot), LlamaIndex FaithfulnessEvaluator/RelevancyEvaluator/RetrieverEvaluator, LangGraph eval judges, RAGAS, TruLens. The metric is the optimization target. Get it wrong and every downstream optimizer is theatre.


1. 何时激活 (When to Activate)

Activate when any of these is true:

  • Optimization runs: a DSPy / OpenAI-Evals / RAGAS / TruLens job is about to consume a metric(example, pred) -> bool|float. The metric drives gradient-free search; bias propagates into the artifact.
  • RAG evaluation: deciding chunk size, reranker, hybrid alpha, retriever-k. A bad metric here picks the wrong chunking strategy and you ship it.
  • Agent benchmarks: tool-use, multi-step, planning. A single holistic LLM judge cannot distinguish "wrong tool" from "right tool, wrong args".
  • Prompt tuning that has gone past 2 manual iterations: if you're tuning a prompt and have no quantitative metric, you are guessing. Stop the prompt edits, write the metric.
  • Production regression test: every chunking / embedding / retriever / prompt PR change needs a metric gate (LlamaIndex OP-10 EvalLoop).

Do NOT activate for:

  • One-off exploratory prompt tests where no decision rides on the output.
  • Tasks where exact-match / unit-test / schema-validity already gives ≥95% of signal — don't over-engineer.
  • Where the user explicitly refuses to commit to any evaluation criteria (then dspy-sop will refuse to compile anyway; this skill cannot help).

2. 核心心智模型 (Core Mental Model)

"The metric IS the model. Change the metric, change the behavior."

A DSPy/GEPA/MIPRO optimizer is a black-box search that maximizes metric(pred, example). Whatever the metric rewards, the compiled prompt will produce. If the metric prefers verbose, hedged answers (which an LLM judge will, by default — judges over-prefer length [arxiv.org/pdf/2506.02592]), the optimizer will produce verbose, hedged answers. A bad metric beats a good optimizer every time.

Three corollaries:

  1. Decompose by default. A single holistic LLM-as-judge call ("is this answer good? rate 1-5") collapses orthogonal axes (factuality, tone, length, relevance) into one noisy scalar. Decompose into N orthogonal yes/no sub-judges, then aggregate. Same number of LM calls in the limit, vastly less noise.
  2. Bool during compile, float during eval. Same metric function, two return types based on the trace argument. Compile-time bool prevents the optimizer from chasing noise in the middle of the distribution; eval-time float gives gradient for reporting and debugging. [dspy.ai/learn/evaluation/metrics/]
  3. The metric must be calibrated against humans. ≥20 spot-checks where you (the human) rate the same examples the metric does. If metric disagrees with human on >20% of cases, fix the metric before any compile. Otherwise the optimizer just learns the metric's bias.
Why holistic judges fail (the catalog)
BiasWhat it doesSource
Length biasJudges prefer longer answers regardless of quality[arxiv.org/pdf/2506.02592]
Self-preferenceA judge from family X prefers outputs from family X[arxiv.org/pdf/2506.02592]
Recency / positionLast option in a pairwise rated higher[arxiv.org/pdf/2509.26072]
Rubric-orderCriteria listed first weighted more[arxiv.org/pdf/2509.26072]
Score-ID"5" and "10" anchor differently across rubrics[arxiv.org/pdf/2509.26072]
ProvenanceKnowing the source model biases the rating[arxiv.org/pdf/2506.02592]

Full catalog in references/R2-judge-bias-catalog.md.


3. SOP (Standard Operating Procedure)

0. Confirm activation criteria (§1)
1. DECOMPOSE: list orthogonal criteria
2. CHOICE: bool vs float per criterion
3. WRITE: implement sub-judges + aggregator
4. BIAS-TEST: probe for length, position, self-preference
5. CALIBRATE: ≥20 human spot-checks; agreement ≥80%
6. SHIP: hand metric to optimizer / eval loop
Stage 1 — Decompose

Write down the criteria the user actually cares about. For an open-ended Q&A:

  • Factual? (yes/no) — answer makes no unsupported claims
  • On-topic? (yes/no) — addresses the question
  • Concise? (length penalty, scalar) — token count vs budget
  • Non-hedging? (yes/no) — no "as an AI language model..." or "it depends" cop-outs
  • Cites source? (yes/no) — if grounded retrieval is required

For RAG specifically, copy LlamaIndex's triad: Faithfulness (answer entailed by context), Relevancy (answer addresses query), Retriever quality (MRR / hit-rate on labeled QA pairs). See references/R1-source-evidence.md.

Output of this stage: a checklist of 3–6 sub-criteria, each with type (bool / scalar) and aggregation rule (AND for hard gates, weighted sum for soft).

Stage 2 — Bool vs Float Choice (per criterion)

Per [dspy.ai/learn/evaluation/metrics/]:

Criterion shapeReturnWhy
Hard requirement (factual, schema-valid, no PII)boolOptimizer should reject anything that fails, not partially credit
Soft preference (concise, fluent, on-tone)float ∈ [0,1]Optimizer benefits from gradient
Lengthexplicit scalar penalty, not LLM-ratedLLM judges over-prefer length; bake the penalty in deterministically

Universal rule (DSPy OP-030): wrap the whole metric so it returns bool when trace is not None (compile mode) and float otherwise (eval mode). Same function, two semantics.

Stage 3 — Write the Judge
python
import dspy

class Assess(dspy.Signature):
    """Assess a single binary criterion. Reply yes/no."""
    text_to_assess = dspy.InputField()
    assessment_question = dspy.InputField()
    assessment_answer: bool = dspy.OutputField()

def metric(example, pred, trace=None):
    # Hard gates (bool sub-judges)
    factual = dspy.Predict(Assess)(text_to_assess=pred.answer,
        assessment_question="Is every factual claim supported by the context?").assessment_answer
    on_topic = dspy.Predict(Assess)(text_to_assess=pred.answer,
        assessment_question=f"Does this address the question: '{example.question}'?").assessment_answer
    non_hedging = dspy.Predict(Assess)(text_to_assess=pred.answer,
        assessment_question="Does this answer commit to a position (no 'it depends' / 'as an AI' hedging)?").assessment_answer

    # Length penalty (deterministic — DO NOT delegate to judge)
    over_budget = len(pred.answer.split()) > 150
    length_score = 1.0 if not over_budget else max(0.0, 1.0 - (len(pred.answer.split()) - 150) / 150)

    if trace is not None:  # compile mode → strict bool
        return factual and on_topic and non_hedging and not over_budget

    # eval mode → float for reporting
    return (factual + on_topic + non_hedging) / 3.0 * length_score

Rules:

  • One yes/no question per sub-judge. Combining ("is it factual AND concise?") re-introduces holistic confusion.
  • Length is a deterministic scalar, never an LLM call. Judges have systematic length bias [arxiv.org/pdf/2506.02592].
  • Use a cheaper / different model family for the judge than the task model (mitigates self-preference). E.g. task = GPT-4o, judge = Claude-Haiku.
Stage 4 — Bias Test

Before calibration, run these probes on a 20-example dev set:

  1. Length probe: take 10 good answers, append a redundant paragraph. Does the metric score them higher? If yes, your length penalty is too weak.
  2. Self-preference probe: generate the same answer from 2 model families. Does the judge prefer its own family by >10pp? If yes, swap judge family.
  3. Position probe (pairwise only): swap A/B order. Does winner flip >10% of the time? If yes, randomize order or average both orderings.
  4. Rubric-order probe: list sub-judges in 2 different orders. Does aggregated score shift >5pp? If yes, randomize sub-judge order per call.

Document failures in references/R2-judge-bias-catalog.md for this project.

Stage 5 — Human Calibration
  • Sample 20 (or 30 if open-ended) examples spanning the task distribution.
  • A human rates each on the same sub-criteria.
  • Compute per-sub-judge agreement (Cohen's κ or simple accuracy).
  • Threshold: ≥80% agreement on each sub-judge, ≥80% on aggregate.
  • If below: fix the metric, do not compile. Common fixes: rephrase the assessment question, switch judge model, narrow the question scope.

This is non-negotiable. Per DSPy Case C: "Never compile against a metric you haven't human-validated on ≥ 20 spot-checks." Garbage metric → garbage compiled program.

Stage 6 — Ship

Hand off:

  • The metric function (callable).
  • A calibration receipt: {n_spot_checks, per_judge_agreement, bias_probe_results, judge_model_id, task_model_id, date}. Stored next to the compiled artifact. Required for any later audit.
  • Recommended optimizer pairing: if sub-judges produce textual feedback (not just bool), pipe into dspy.GEPA for sample efficiency; otherwise MIPROv2.

4. 操作模型 (Operations)

OP-M01 — DecomposeMultiCriteria
  • Trigger: Open-ended generation; multiple correctness axes (factuality + tone + length + relevance).
  • Action: List 3–6 orthogonal yes/no sub-criteria. One dspy.Predict(Assess) call per criterion. Aggregate with AND (hard) + weighted sum (soft).
  • Output: Multi-judge metric callable; one criterion per call.
  • Evidence: [dspy.ai/learn/evaluation/metrics/]; DSPy Case C; LlamaIndex Faithfulness+Relevancy+Retriever triad.
OP-M02 — BoolDuringCompileFloatDuringEval
  • Trigger: Writing any metric used for both compile (teleprompt.compile) and dspy.Evaluate.
  • Action: Inside the metric, if trace is not None: return bool_aggregate; else: return float_aggregate.
  • Output: One metric function, two modes; optimizer avoids chasing float noise.
  • Evidence: [dspy.ai/learn/evaluation/metrics/]; [dspy.ai/cheatsheet/]; DSPy OP-030.
OP-M03 — ExplicitLengthPenalty
  • Trigger: Any LLM-as-judge metric (always).
  • Action: Add a deterministic scalar length_score = clip(1 - max(0, len - budget) / budget, 0, 1) multiplied into the final float. Do NOT delegate length judgment to the LLM.
  • Output: Length-controlled metric immune to the judge's length bias.
  • Evidence: [arxiv.org/pdf/2506.02592]; DSPy SOP step "Add length penalty as separate scalar".
OP-M04 — CrossFamilyJudge
  • Trigger: Judge model is in same family as task model (e.g., both GPT-4).
  • Action: Swap judge to a different family (Claude, Llama, Gemini, Mistral). Document choice.
  • Output: Self-preference mitigated.
  • Evidence: [arxiv.org/pdf/2506.02592].
OP-M05 — HumanSpotCheck20
  • Trigger: Before any compile / production rollout of a new or modified metric.
  • Action: 20 (open-ended: 30) human ratings on the same sub-criteria. Compute agreement. Reject if <80% per-sub or aggregate.
  • Output: Calibration receipt; metric is human-validated.
  • Evidence: DSPy Case C step 4; [dspy.ai/learn/evaluation/metrics/].
OP-M06 — RAGFaithfulnessRelevancyContext
  • Trigger: RAG pipeline; tuning chunk size / reranker / retriever-k / hybrid alpha.
  • Action: Instantiate three evaluators: FaithfulnessEvaluator (answer entailed by retrieved context), RelevancyEvaluator (answer addresses query), RetrieverEvaluator(["mrr","hit_rate"]) (does retrieval pull the gold passage). Gate every PR.
  • Output: 3-axis RAG quality vector; never a single number.
  • Evidence: LlamaIndex OP-10 EvalLoop; [developers.llamaindex.ai/python/framework-api-reference/evaluation/].
OP-M07 — BiasProbeSuite
  • Trigger: Before calibration.
  • Action: Run length-probe, self-preference probe, position probe, rubric-order probe (§3 Stage 4) on 20 examples. Log deltas.
  • Output: Bias-probe report; rejected if length probe ≥+5pp on padded answers or position flips >10%.
  • Evidence: [arxiv.org/pdf/2509.26072]; [arxiv.org/pdf/2506.02592].
OP-M08 — TextualFeedbackForGEPA
  • Trigger: Sub-judge can articulate why it failed (e.g., "answer was verbose", "missed cited source").
  • Action: Return dspy.Prediction(score=float, feedback=str) from the metric. Pair with dspy.GEPA optimizer (not MIPROv2).
  • Output: Sample-efficient, reflection-evolved prompts.
  • Evidence: [dspy.ai/api/optimizers/GEPA/overview/]; [arxiv.org/abs/2507.19457]; DSPy OP-004.
OP-M09 — RandomizeJudgeOrder
  • Trigger: Pairwise judge or rubric with N>3 criteria.
  • Action: Per-call: randomize A/B order in pairwise; randomize sub-judge ordering. Or run both and average.
  • Output: Position and rubric-order bias mitigated.
  • Evidence: [arxiv.org/pdf/2509.26072].
OP-M10 — CalibrationReceipt
  • Trigger: Ship a compiled artifact.
  • Action: Save JSON next to artifact: {metric_version, n_spot_checks, per_judge_agreement, bias_probe_results, judge_model_id, task_model_id, date, criteria_list}.
  • Output: Audit-ready provenance for the metric used to compile.
  • Evidence: General audit best-practice; DSPy save(save_program=True) analog.

5. 困境决策案例 (Dilemma Cases)

Dilemma 1 — MIPRO inherited the judge's verbosity bias

困境: Team compiled an open-ended customer-support response program with MIPROv2(metric=llm_judge, auto="medium"). Auto-eval score climbed from 0.62 to 0.84. Production rolled out. Users complained answers were too long. Auditor traced it: the LLM judge gave +0.18 to answers >120 words across the eval set, including some objectively wrong ones. The optimizer faithfully chased that signal.

约束: Re-compile costs ~$30. Cannot re-label dataset. Cannot change judge model (vendor approval cycle).

决策步骤:

  1. Diagnose: Confirm length bias on a held-out set: bin answers by length, plot judge-score vs length. Look for monotonic positive slope. (It was there: +0.21 per 50 words.)
  2. Patch the metric, not the artifact: Add explicit length penalty (OP-M03) and decompose the holistic judge into factual / on-topic / non-hedging sub-judges (OP-M01). Length is now deterministic, not judge-rated.
  3. Re-calibrate on 30 human spot-checks (OP-M05). Agreement: was 64%, now 87%.
  4. Re-compile: Same MIPROv2, new metric. Score on new metric: starts at 0.55 (because old prompts over-verbose under new metric), recovers to 0.79 after compile.
  5. Production A/B: New artifact has 38% shorter answers, equal factuality, +12% user thumbs-up.

结果: The "score regression" (0.84 → 0.79) was misleading — the old metric was the bug. Real quality improved because the metric finally matched the user.

可提取的操作: OP-M01 DecomposeMultiCriteria, OP-M03 ExplicitLengthPenalty, OP-M05 HumanSpotCheck20. Lesson: when a compiled program is "good on metric, bad in production", the metric is wrong. Don't tune harder — fix the metric.

Show full SKILL.md (957 more words)Show less
Dilemma 2 — RAG faithfulness = 100% but answers are wrong

困境: RAG pipeline reports Faithfulness = 1.0 across 200 eval QA pairs. Users complain answers don't address their questions. Team is confused: "but we're 100% faithful…"

约束: Cannot change evaluator vendor. Existing eval suite has only FaithfulnessEvaluator.

决策步骤:

  1. Realize the missing axis: FaithfulnessEvaluator only asks "is the answer entailed by retrieved context?" An answer of "The context discusses Q3 financials" is faithful to the context but does not answer "What was Q3 revenue?".
  2. Add RelevancyEvaluator (does the answer address the question?) and RetrieverEvaluator(["mrr", "hit_rate"]) (is the gold passage even retrieved?). [LlamaIndex OP-10]. This is OP-M06.
  3. Re-evaluate: Faithfulness still 1.00. Relevancy: 0.41. MRR: 0.62. Now the picture is clear — retrieval pulls some relevant passages, but the synthesizer hedges with topic-level statements rather than answering the question.
  4. Fix the synthesizer prompt, gated on the 3-axis metric. Add a sub-judge: "Does the answer give a direct response to the question (not just describe the context)?" — bool, AND-gated.
  5. Result: Faithfulness 0.97, Relevancy 0.79, MRR 0.74. Users stop complaining.

结果: A single metric (faithfulness) created a blind spot. Decomposition exposed the actual failure (synthesizer hedging) which no single number could surface.

可提取的操作: OP-M01, OP-M06 RAGFaithfulnessRelevancyContext. Lesson: a single RAG metric, however precise, is structurally incomplete. Triad is the minimum.

Dilemma 3 — Vendor metric ships in 11 installs but covers one axis

困境: A vendor-tied cekura-metric-design skill (11 installs) wraps a single proprietary judge. Convenient, but: (a) cannot inspect the rubric, (b) cannot swap judge family, (c) cannot add length penalty, (d) calibration receipt does not include bias probes. Optimizer is chasing the vendor judge's blind spots.

决策步骤:

  1. Diagnose dependency: Identify which axes the vendor judge covers (often: a holistic "quality" score). Identify what is missing (length, hedging, position bias mitigation).
  2. Wrap, don't replace: Use the vendor score as one sub-judge in a composed metric (OP-M01). Add length penalty (OP-M03), non-hedging sub-judge, cross-family fact-check (OP-M04).
  3. Calibrate the composed metric (OP-M05), not the vendor's alone.
  4. If vendor score correlates <0.7 with human aggregate after composition: drop the vendor; you are paying for noise.

结果: Vendor metrics buy convenience but lose audit and bias control. Open, decomposed metrics dominate for any pipeline that will be optimized against.

可提取的操作: OP-M01, OP-M03, OP-M04, OP-M10. Lesson: a vendor-tied holistic metric is a managed optimizer target you cannot inspect. Decompose around it.


6. 反模式与边界 (Anti-Patterns & Boundaries)

Anti-patterns
#Anti-patternWhy it's wrongFix
AP-1Single holistic LLM-as-judge ("rate 1-5") as optimizer metricConflates orthogonal axes; inherits all judge biases; optimizer chases noiseDecompose into yes/no sub-judges (OP-M01)
AP-2No length penalty in a generation metricJudges over-prefer length; optimizer learns to be verboseDeterministic length term (OP-M03)
AP-3Same family for judge and task modelSelf-preference bias inflates score 5–15ppCross-family judge (OP-M04)
AP-4Skipping human calibrationOptimizer learns the metric's bias, not the task≥20 spot-checks (OP-M05)
AP-5Throwing optimizers at a stalled compileIf metric is bad, no optimizer can save itReturn to metric design
AP-6Single-axis RAG metric (faithfulness only)Misses query-relevance and retrieval-qualityTriad (OP-M06)
AP-7LLM-rated length ("is this concise?" as a judge call)Inherits length bias; deterministic is freeLength is len(tokens), not a judge call
AP-8Hand-edited metric mid-experiment without re-calibrationPast compile artifacts now have invalid calibration receiptsBump metric version; re-calibrate
AP-9Float metric in compile modeOptimizer chases sub-percent noise; overfitbool in compile, float in eval (OP-M02)
AP-10Compiling without a metric at allWithout a metric DSPy degenerates to verbose promptingRefuse the compile (DSPy OP-024)
Boundaries (when this skill cannot help)
  • No willingness to define success criteria: if the user cannot articulate any criterion for "good", no metric exists. Route to scientific-critical-thinking for problem definition first.
  • Tasks where exact-match dominates (classification, structured extraction, code-passes-tests): use the obvious metric; this skill's decomposition machinery is overkill. Don't gold-plate.
  • Pure preference-pair tasks (RLHF data): preference judges have their own bias profile (Bradley-Terry-style); see simpo / trl-fine-tuning skills, not this one.
  • Aesthetic / subjective tasks with no human ground truth (poetry, art): metric design here is essentially arbitrary. Be explicit that you are encoding one aesthetic, not measuring quality.

7. 跨框架对照 (Cross-Framework Mapping)

ConceptDSPyLlamaIndexRAGASTruLensThis skill
Metric typedef metric(ex, pred, trace=None) -> float|boolEvaluationResult from BaseEvaluatorMetric class (Faithfulness, AnswerRelevancy, etc.)Feedback functionmetric function + calibration receipt
Multi-criteriasub-judges via dspy.Predict(Assess)stack of FaithfulnessEvaluator + RelevancyEvaluator + RetrieverEvaluatormetric ensembleFeedback per criterionOP-M01 DecomposeMultiCriteria
Mode switch (compile vs eval)trace is not None returns boolN/A (eval only)N/AN/AOP-M02
Faithfulnesssub-judge "every claim supported?"FaithfulnessEvaluator (entailment-based)Faithfulness metric (claim-by-claim NLI)groundedness feedbacksub-judge or pre-built
Relevancysub-judge "addresses question?"RelevancyEvaluatorAnswerRelevancy (embedding-based)answer-relevance feedbacksub-judge
Retrieval qualityLlamaIndex RetrieverEvaluator MRR/hit-rateRetrieverEvaluator(["mrr","hit_rate"])ContextPrecision, ContextRecallcontext-relevance feedbackpre-built (use LlamaIndex / RAGAS)
Length penaltymanual scalar in metricmanualmanualmanualOP-M03 (deterministic)
Textual feedbackdspy.Prediction(score, feedback) → GEPAN/Arationale stringsfeedback rationaleOP-M08
Cross-family judgeswap judge_lm in metricswap service_context.llm of evaluatorswap evaluator LLMswap feedback LLMOP-M04 (always)

Combination patterns:

  • DSPy + LlamaIndex: use LlamaIndex FaithfulnessEvaluator + RelevancyEvaluator as the sub-judges inside a DSPy metric; aggregate to bool/float per OP-M02. (JetBlue/Databricks pattern.)
  • DSPy + RAGAS: use Faithfulness / AnswerRelevancy / ContextPrecision as sub-scalars; pass into a DSPy metric for compile.
  • TruLens for production monitoring: same sub-judges, different runtime (live, sampled). Calibration receipts should match.

Opinionated default: build the metric in pure Python with dspy.Predict(Assess) calls (transparent, version-controllable) and import LlamaIndex evaluators only for the RAG triad where the LlamaIndex evaluators are battle-tested. Avoid vendor SDKs whose rubrics you cannot inspect.


Appendix — Quickstart Skeleton

python
import dspy

class Assess(dspy.Signature):
    """One yes/no assessment."""
    text = dspy.InputField()
    question = dspy.InputField()
    answer: bool = dspy.OutputField()

# Use a DIFFERENT family from the task model
judge_lm = dspy.LM("anthropic/claude-haiku-4")  # task model is openai/gpt-4o

def make_metric(budget_words=150):
    def metric(example, pred, trace=None):
        with dspy.context(lm=judge_lm):
            factual = dspy.Predict(Assess)(text=pred.answer,
                question="Is every claim supported by the provided context?").answer
            on_topic = dspy.Predict(Assess)(text=pred.answer,
                question=f"Does this directly address: {example.question}?").answer
            non_hedging = dspy.Predict(Assess)(text=pred.answer,
                question="Does this commit to a position (no 'as an AI', no 'it depends')?").answer
        n_words = len(pred.answer.split())
        length_score = 1.0 if n_words <= budget_words else max(0.0, 1 - (n_words - budget_words)/budget_words)
        if trace is not None:
            return bool(factual and on_topic and non_hedging and n_words <= budget_words)
        return (int(factual) + int(on_topic) + int(non_hedging)) / 3.0 * length_score
    return metric

Then: bias-probe → human-calibrate → save receipt → ship.


References

  • references/R1-source-evidence.md — DSPy / LlamaIndex / LangGraph source quotes
  • references/R2-judge-bias-catalog.md — Judge bias taxonomy + mitigations
  • intermediate/operation_candidates.json — Operation registry

Citations: [dspy.ai/learn/evaluation/metrics/], [dspy.ai/learn/optimization/overview/], [dspy.ai/cheatsheet/], [dspy.ai/api/optimizers/GEPA/overview/], [arxiv.org/abs/2507.19457], [arxiv.org/pdf/2506.02592], [arxiv.org/pdf/2509.26072], [developers.llamaindex.ai/python/framework-api-reference/evaluation/], [cookbook.openai.com/examples/evaluation/evaluate_rag_with_llamaindex].

© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/agentsop-metric-design of agentsope/SkillAlchemy.

  • SKILL.md
  • README.md
  • intermediate/operation_candidates.json
  • references/R1-source-evidence.md
  • references/R2-judge-bias-catalog.md

Open the folder on GitHubat commit d0f0355

Compare with similar skills

Agentsop Metric Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentsop Metric Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentsop Metric Design this skillagentsope/SkillAlchemy436—~6.4kAutomated safety check: PassMIT
Advanced Evaluationguanyang/open-agent-hub9772 repos~4.2kAutomated safety check: PassMIT
Agentic Evalgithub/awesome-copilot40k3 repos~1.5kAutomated safety check: PassMIT
Clawpathy AutoresearchClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT
Suede AI EvalJasonColapietro/suede-creator-skills127—~3.3kAutomated safety check: PassMIT
Commerce Evalsanthropics/commerce-agents3.2k—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Advanced Evaluation

    guanyang/open-agent-hub

    This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Agentic Eval

    github/awesome-copilot

    Official

    Patterns and techniques for evaluating and improving AI agent outputs.

    40k GitHub starsUsed in 3 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Suede AI Eval

    JasonColapietro/suede-creator-skills

    Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates.

    127 GitHub stars~3.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Commerce Evals

    anthropics/commerce-agents

    Official

    Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.

    3.2k GitHub stars~1.8k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Advanced Evaluation

    aiskillstore/marketplace

    This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…

    433 GitHub starsUsed in 3 repos~4.2k tokens
    EducationAuto-check passed

More from agentsope/SkillAlchemy

All 46 skills in this repo
  • Agentsop Aider

    agentsope/SkillAlchemy

    SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).

    436 GitHub stars~3.5k tokensUpdated 2 days ago
    Auto-check passed
  • Agentsop Context Scope Discipline

    agentsope/SkillAlchemy

    Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…

    436 GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed
  • Agentsop Cost Tiered Models

    agentsope/SkillAlchemy

    Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

    436 GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed
  • Agentsop Crewai

    agentsope/SkillAlchemy

    SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.

    436 GitHub stars~4.8k tokensUpdated 2 days ago
    Auto-check passed
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    436 GitHub stars~5.4k tokensUpdated 2 days ago
    Auto-check: notes
  • Agentsop Multiscale Chunking

    agentsope/SkillAlchemy

    Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

    436 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed

Questions about Agentsop Metric Design

What does Agentsop Metric Design do?

Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy. Agentsop Metric Design is an agent skill from agentsope/SkillAlchemy. Decomposed, multi-criteria metric design for LLM pipelines.

When should I use Agentsop Metric Design?

Agentsop Metric Design fits situations like: tasks that involve LLM evaluation; tasks that involve Quizzes and assessments; tasks that involve Operations and SOPs.

How do I install Agentsop Metric Design in Claude Code?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a claude-code`. Or copy the skill folder (skills/agentsop-metric-design in agentsope/SkillAlchemy) into .claude/skills/agentsop-metric-design in your project. Claude Code loads it when a task matches its description.

How do I install Agentsop Metric Design in Codex?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a codex`. Or copy the skill folder (skills/agentsop-metric-design in agentsope/SkillAlchemy) into .agents/skills/agentsop-metric-design in your project. Codex loads it when a task matches its description.

Can I use Agentsop Metric Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-metric-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-metric-design, .gemini/skills/agentsop-metric-design, .github/skills/agentsop-metric-design and .opencode/skills/agentsop-metric-design in your project.

What does Agentsop Metric Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentsop Metric Design is instructions for the agent only. Our summary lists: Python 3.

Does Agentsop Metric Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentsop Metric Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentsop Metric Design use?

Agentsop Metric Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentsop Metric Design use?

About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.6k tokens, read only when the agent opens those files.

What are the alternatives to Agentsop Metric Design?

Skills that share tags, products or a category with Agentsop Metric Design: Advanced Evaluation (guanyang/open-agent-hub, 977 stars), Agentic Eval (github/awesome-copilot, 40k stars), Clawpathy Autoresearch (ClawBio/ClawBio, 1.2k stars) and Suede AI Eval (JasonColapietro/suede-creator-skills, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentsop Metric Design?

agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 436 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.

Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.